ducoterra ad7585d474
Build and Push Containers / build-and-push-app (push) Successful in 2m11s
Build and Push Containers / build-and-push-db (push) Successful in 11s
phase: 97_kb_tree_catalog
All completion criteria verified — everything is green, no defects found. Final report:

## Phase 97 final verification pass — ALL GREEN

**Verified (no code changes needed):**
- `GET /api/docs/tree` (admin), `build_kb_tree` pure builder, `PATCH /api/folders/summary`, migration 0018 (`manually_edited`, head confirmed), generator skip/keep + `kept_manual` stat, RAG tree UI + edit affordance in `sources.js`/`index.html`/`styles.css`
- `tests/e2e/test_kb_tree.py`: 8 passed — top level, drill source/folder, edit round-trip, clear, manual-desc-survives-sync, reload fallback, anonymous gate
- Integration: tree shape/order/403/empty/indexed-only + PATCH update/create/root/clear/404/403/no-LLM + stat-walk equivalence (in `test_docs_api.py`); 3-field `folder_summaries=` import token preserved

**Gates (exact commands):**
- `uv run pytest --cov=app --cov-report=term-missing` → **2053 passed**, TOTAL coverage **99%** (>90% ✓)
- `uv run ruff check . && uv run pyright` → **All checks passed / 0 errors**
- `uv run pytest tests/e2e/test_kb_tree.py -v --no-cov` → **8 passed** in isolation
- 30 story/RAG-view E2E suites run **one per process**: all passed, incl. `test_ls_tree_drilldown` (agent `ls` byte-identical ✓), `test_import_documents`, `test_edit_summaries`, `test_admin_auth`, `test_kb_overview`

**Completion criteria:** tree view ✓ · edit round-trip + clear ✓ · manual persists/clear resets ✓ · `ls` unchanged ✓ · pytest/coverage/lint ✓ · E2E isolation ✓ · commit — left to harness per protocol (working tree untouched, `git add/commit` not run)

**Deviations:** none. **Next pending phase:** none — `todo/` contains only 97 (96 already committed).
2026-09-11 22:48:02 -04:00
2026-09-11 22:48:02 -04:00
2026-08-27 10:00:21 -04:00
2026-09-11 22:48:02 -04:00
2026-09-11 22:48:02 -04:00
2026-09-11 22:48:02 -04:00
2026-09-11 13:16:20 -04:00
2026-09-11 22:48:02 -04:00
2026-09-11 13:16:20 -04:00
2026-09-09 17:22:24 -04:00
2026-09-09 17:22:24 -04:00

🧠 Brain of Reese

A chippy, honest RAG chatbot over your ~/Homelab and ~/Deployments projects. Point it at your notes — markdown, YAML, JSON, Python, plain text — ask it anything, and it retrieves the relevant chunks with hybrid search (pgvector cosine ∪ Postgres full-text search, fused with Reciprocal Rank Fusion), feeds the whole relevant document to a self-hosted LLM (turbo via https://aipi.reeseapps.com/v1), and streams a grounded answer back.

If it doesn't have notes for your question, it admits it: "I haven't done anything like that" — plus suggestions for what it does know.

Updated your notes? Re-run the import — it's idempotent and only re-embeds what changed:

uv run python -m scripts.import_docs --prune

See Updating the documents for details.


Quick Start

# 1. Dependencies
uv sync

# 2. Configure (copy and edit)
cp .env.example .env
# Set BOR_LLM_API_KEY — your aipi key (falls back to $AIPI_KEY)

# 3. Start the database
podman compose up -d db
podman compose ps          # wait until "healthy"

# 4. Apply migrations & import
uv run alembic upgrade head
uv run python -m scripts.import_docs

# 5. Run the app
uv run uvicorn app.main:app --reload
# → http://localhost:8000  (chat)

Stack: FastAPI · Pydantic v2 · SQLAlchemy 2 · Alembic · pgvector · vanilla HTML/CSS/JS (no CDN) · Playwright E2E


Pages at a glance

The app is a single-page application — all views load from / with a client-side router. Direct URLs deep-link to the matching view.

Page URL Who can see it
Chat / Everyone (token gate for non-shared)
RAG / Knowledge base /sources.html Admin only
Git sources /git-sources.html Admin only
Tuning /tuning.html Admin only
Saved chats /history.html Admin only
Access tokens /tokens.html Admin only
Document viewer /document.html?source=X&path=Y Token users & admins
Login /login.html Everyone
Shared chat /shared/<token> Everyone (anonymous)
Edit doc /doc-edit.html Admin only

Chat (/)

Ask questions; answers stream in with source chips that cite the exact documents used. Clicking a chip opens the document in an almost-fullscreen modal on the same page (no new tab). The onboarding suggestion chips follow the last 3 questions asked — on a fresh deployment they seed from BOR_SUGGESTIONS.

RAG / Knowledge base (/sources.html)

The indexed document list. The Path column opens each document in the same almost-fullscreen modal. Admin-only — anonymous visitors see a sign-in gate instead.

Git sources (/git-sources.html)

The admin-managed source registry: add or remove git repositories, upload source archives (.tar, .tar.gz, .tgz, .zip), and register local directories — all in one table with a kind discriminator. No .env editing, no restart. The Sync sources button clones/pulls/walks every source and re-imports in one click.

Tuning (/tuning.html)

Set global instructions that are injected into the system prompt of every chat turn. Add notes like "be more concise" or "assume I'm on NixOS". List, edit, or remove them here — no conversation required.

Saved chats (/history.html)

Every conversation is saved automatically. Click a title to return to that chat. Admin-only view.

Access tokens (/tokens.html)

Generate tokens and hand them out — a token opens chat, the answers, and the documents they cite. Shared chats stay open to everyone.

Shared chat (/shared/<token>)

A read-only snapshot of any conversation, shareable via a link. Open to everyone — no login or token required.

Edit doc (/doc-edit.html)

Review and adjust an AI-generated answer before committing it to the docs repository. Admin-only flow page.


Admin & Sign-in

Brain of Reese has exactly one account: the admin (you). Signing in unlocks the full Sources catalog, the answer-tuning controls, and token management. Shared chats are the only content that stays open to anonymous visitors.

Setup (one-time)

python -c 'import secrets;print(secrets.token_hex(32))'   # → paste into .env
BOR_ADMIN_PASSWORD=your-password        # plaintext — homelab scope, by design
BOR_SESSION_SECRET=<the hex from above> # signs the session cookie

Fail-loud: while either variable is empty the app refuses to start, naming the missing one(s).

How authentication works

Endpoint Purpose
POST /api/login Password → signed bor_session cookie
POST /api/logout Clears session cookie
GET /api/whoami `{"authenticated": bool, "role": "admin"
POST /api/token-auth Token (bor_…) → same cookie, role user

The cookie uses same_site="lax", https_only off — no HTTPS enforcement (homelab HTTP; the cookie is single-admin convenience, not a cloud boundary). Max age is BOR_SESSION_MAX_AGE (default 43200 = 12 h).

Access tokens

The admin can hand out access without sharing the admin password.

  1. Generate: sign in → Tokens in the navbar. Type a label and hit Generate — a bor_ + 32-hex token appears in the shown once block. Copy it now: only its SHA-256 hash is stored.
  2. Use: the holder enters the token at the in-app gate. The browser caches it in localStorage["bor.token"] for silent re-auth on reloads.
  3. Revoke: click Revoke on the token's row. Revocation is immediate.

A token user can chat and open documents — nothing else. The Sources catalog, git sources, tuning, and history stay admin-only.


Using the App

Updating the documents

The knowledge base is refreshed by re-running the import. It is idempotent and delta-based (sha256 per file), so a refresh after a normal editing session takes seconds:

# After editing/adding/removing notes:
uv run python -m scripts.import_docs                 # re-index what changed
uv run python -m scripts.import_docs --prune         # also drop deleted files

# Point it at extra directories (repeatable):
uv run python -m scripts.import_docs --source ~/SomeOtherDocs

With git-based sources each run first pulls the latest commits of your repos, so this same command is the whole update loop: commit → re-run.

Then check the Sources page (http://localhost:8000/sources.html): the documents / chunks counters and last indexed timestamp should reflect the new files.

  • The import prints one line per file and ends with a greppable summary, so it is safe to run from cron or after every commit.
  • Indexed formats: md, markdown, txt, yaml, yml, json, py, the Podman quadlet family (container, network, volume, image, pod, kube, swap, os, endpoint), and j2 Jinja templates (case-insensitive).
  • Hidden files and common cache dirs (.venv, node_modules, .git, __pycache__, etc.) are skipped.
  • Non-markdown files get format-aware chunking (YAML keys, JSON keys, Python classes/defs).

Git-based sources

Rather than pointing the import at local folders, point it at git repositories. The admin Git sources page (/git-sources.html) is the primary management surface — add or remove repositories there and the list is stored in Postgres.

# .env — the fallback list (fresh setups, or until the admin page
# stores a source; the page becomes the source of truth)
BOR_GIT_SOURCES=https://git.reeseapps.com/reese/homelab.git,git@github.com:reese/deployments.git
BOR_SOURCES_DIR=~/bor-sources   # default; each repo lands in <dir>/<repo-name>/
  • Auth is whatever the machine supplies — HTTPS via the OS credential helper, or SSH via your key; no credentials are stored in the app.
  • Every run clones (first time, shallow --depth 1) or pulls (git pull --ff-only) each repo, then indexes the checkout.
  • A failed sync aborts the run — no partial junk.

Archive upload sources

Upload a .tar, .tar.gz, .tgz, or .zip archive to make it a source. The form on the Git sources page — and the POST /api/git-sources/upload route behind it — accepts archives and unpacks + scans them immediately.

  • Re-uploading the same filename replaces the source in place (no second folder, no duplicate row).
  • Zip-bomb guard: absolute member paths, .. traversal, and symlinks escaping the unpack folder are rejected.
  • Archives unpack under BOR_UPLOAD_DIR (default ~/bor-sources/uploads).
  • BOR_UPLOAD_MAX_MB (default 512) caps both compressed and extracted size.

Local directory sources

A plain directory can be a first-class source too. It shares the git sources' one table (the git_sources registry with a kind discriminator: git | local), one admin page, and one Sync button.

  • Register via API: POST /api/git-sources with {"kind": "local", "path": …}.
  • Sync walks it directly — no clone, no checkout copy.
  • The directory is re-verified to exist at sync time; a missing directory fails the run loudly.

Sync from the UI

The Sync sources button on the Sources page (admin only) runs the whole git-source refresh in one click:

  1. clone/pull + walk every configured source — git repos and local directories, including uploaded archives.
  2. re-import with prune — files deleted upstream leave the index.
  3. regenerate the KB overview (the <knowledge_base> outline every chat turn injects) — only when the import changed the knowledge base.
  • States: the button shows Syncing… with a spinner while polling GET /api/sync/status every 2 s. No client-side timeout — a clone + embed can legitimately take minutes.
  • One sync at a time: a second trigger while a run is in flight gets a 409.

Thinking

The self-hosted turbo model reasons before it answers. That reasoning is streamed with the turn as thinking SSE events and shown in a collapsible "Thinking" block above the answer bubble: it opens and fills in live while the model thinks, tucks itself away the moment the first answer token lands, and stays click-toggleable afterwards.

To hide it, set BOR_STREAM_THINKING=0.


Agent document tools (ls + read + grep)

Retrieval only puts the top documents in context. When an answer depends on a file a note references ("the exact JSON shape is in example-record-file.json"), the model can extend its own context with three server-side tools — on grounded (high-relevance) turns only:

Tool Description
ls Lists every indexed document (source: X | path: Y | title: Z); pass a source name as path to list one source
read(path) Appends the full text of one indexed document to context (never truncated); path is the combined source/path string
grep(pattern, path?) Searches indexed documents for an exact string (case-insensitive fixed substring); returns up to 20 source/path:line: text matches; an optional path limits to one document

Each call is executed against Postgres only (no extra LLM round trip) and streamed as an SSE tool frame. In the chat, each call shows a transient calling-tool status alongside "thinking" (the send button keeps its busy state — "Stop" — for the whole turn).

The model may call tools as many times as needed, bounded by a round cap:

Env Default Meaning
BOR_AGENT_MAX_ROUNDS 10 hard cap on agent tool rounds per grounded turn (0 = no tools, the kill switch)

Tuning your answers

Admin-only — sign in first. If an answer isn't quite right — too chatty, wrong assumption, missing context — tune Brain right there:

  1. Press "Tune" in the meta row under any completed answer (deflected ones included).
  2. Type a short instruction (1–2000 chars), e.g. "be more concise" or "assume I'm on NixOS", and Save.

The note is stored in Postgres (steering_notes) and read into the system prompt of every subsequent chat turn as a <tuning> section (numbered, oldest first, capped at BOR_STEERING_MAX_CHARS chars — default 8000, overflow marked […truncated…]).

List or remove notes from the "Tuning" button in the chat header (count badge, newest-first, per-note delete). The API is stateless JSON if you prefer curl:

curl -s localhost:8000/api/steering                                  # list
curl -s -X POST localhost:8000/api/steering \
     -H 'Content-Type: application/json' -d '{"note": "be more concise"}'
curl -s -X DELETE localhost:8000/api/steering/<note-id>              # remove

How retrieval works (hybrid)

Every question is embedded and also lexically tokenized (OR-joined, English stemming) and searched twice against Postgres:

  1. Vector — pgvector cosine top-N (default BOR_HYBRID_VECTOR_CANDIDATES=100)
  2. Lexical — a stored tsvector (GIN-indexed) matched with to_tsquery, top-N by ts_rank (default BOR_HYBRID_LEXICAL_CANDIDATES=30)

The two ranked lists are fused with Reciprocal Rank Fusion (score = Σ 1/(k + rank), BOR_RRF_K=60) — a chunk in both lists scores nearly double, which lets a name-your-tool question ("gitlab") find its own document even when the question embeds close to generic templates.

The honesty gate (A8) then answers (HIGH) when the best cosine is ≥ BOR_RELEVANCE_THRESHOLD (default 0.62) or at least one chunk matched lexically (fts_hits > 0) — it deflects (LOW) only when both signals are absent. The top BOR_TOP_N_DOCS full documents are what the LLM sees.


Document summaries (non-markdown)

Raw yaml/json/py/txt embeds badly — flags and keys are not language, so retrieval can miss exactly the documents that are all configuration. At import time, every non-markdown A9 document is summarized by the aipi lite model (BOR_LLM_SUMMARY_MODEL, default lite):

  • The summary is stored on documents.summary and indexed as one extra embedded chunk (chunks.is_summary, position −1), so hybrid search has a natural-language target to hit instead of the raw text.
  • The last line is a code-deterministic pointer — Source: <source>/<path> — appended by the app, never model-generated.
  • The model only sees the first BOR_SUMMARY_MAX_CHARS (default 12000) characters; overflow is cut and marked with […truncated…].

Summary generation is best-effort: if lite fails, the document is still indexed (without a summary), the failure is logged, and counted in the import summary line (summaries=N summary_errors=N).


Caching & deploys

A deploy is a commit — and the browser must see it without a hard refresh. One Starlette middleware (app/core/caching.py) applies the rule at the transport layer:

  • HTML pages ship Cache-Control: no-cache, no etag, no last-modified — each visit always gets a fresh 200 body.
  • Assets are versioned and cached for a year. Pages reference their CSS/JS with a token (/assets/styles.css?v=<token>), and every /assets/* response ships Cache-Control: public, max-age=31536000, immutable.
  • The token is the deploy. In a git checkout it is the short SHA of HEAD — so every commit/deploy flips the token and the versioned asset URLs change with it. A checkout without .git falls back to a stable content hash of the frontend/ tree.

No CDN, no new services, no build-step change: the middleware rewrites the asset references of the known pages in flight.

Deploy note: the very first deploy onto this scheme needs one normal page visit, so the browser revalidates the HTML once and starts requesting the versioned assets; every commit after that is picked up automatically.


Configuration reference

Env Default Meaning
BOR_APP_NAME Brain of Reese Display name everywhere (page titles, header brand, status labels)
BOR_INPUT_PLACEHOLDER Ask me anything… Chat composer placeholder
BOR_FOOTER_TEXT Powered by self-hosted models Footer line on every page
BOR_DATABASE_URL local compose URL SQLAlchemy URL (psycopg)
BOR_LLM_BASE_URL https://aipi.reeseapps.com/v1 OpenAI-compatible endpoint
BOR_LLM_API_KEY — (falls back to $AIPI_KEY) aipi API key
BOR_LLM_CHAT_MODEL turbo Chat model
BOR_LLM_EMBED_MODEL embed Embedding model
BOR_LLM_SUMMARY_MODEL lite One-shot completions: document summaries at import, KB overview
BOR_EMBEDDING_DIM 768 Vector dimension (fixed at table creation)
BOR_TOP_N_DOCS 2 Full documents fed to the LLM
BOR_RELEVANCE_THRESHOLD 0.62 Answer when best cosine ≥ this or an FTS hit; below + no FTS ⇒ honest deflection
BOR_HYBRID_VECTOR_CANDIDATES 100 Cosine list width for the RRF fusion
BOR_HYBRID_LEXICAL_CANDIDATES 30 FTS list width for the RRF fusion
BOR_RRF_K 60 RRF damping constant (1/(k + rank))
BOR_AGENT_MAX_ROUNDS 10 Hard cap on agent tool rounds per grounded turn (0 = no tools)
BOR_IMPORT_EXTENSIONS csv (see below) Importable formats (may only narrow the A9 set)
BOR_GIT_SOURCES — (empty) CSV of git repo URLs — fallback while the admin Git sources page's list is empty
BOR_SOURCES_DIR ~/bor-sources Where git repos are cloned/pulled
BOR_UPLOAD_DIR ~/bor-sources/uploads Where uploaded source archives are unpacked
BOR_UPLOAD_MAX_MB 512 Cap (MiB) for uploaded source archives (zip-bomb guard)
BOR_STEERING_MAX_CHARS 8000 Char budget for the <tuning> prompt section
BOR_SUMMARY_MAX_CHARS 12000 Cap on document content sent to the lite summary model
BOR_KB_OVERVIEW_MAX_CHARS 4000 Char budget for the <knowledge_base> prompt section
BOR_OVERVIEW_INPUT_MAX_CHARS 40000 Cap on the document list sent to lite for KB overview generation
BOR_SUGGESTIONS built-in list JSON seed for onboarding chips (shown before the first saved question)
BOR_ADMIN_PASSWORD (required) The single admin's password — app refuses to start when empty
BOR_SESSION_SECRET (required) Signing key for the bor_session cookie
BOR_SESSION_MAX_AGE 43200 Session-cookie lifetime in seconds (12 h, sliding)
DEBUGPY 0 1 ⇒ attach-on-demand debugpy on DEBUGPY_PORT (default 5678)
BOR_LOG_LEVEL INFO App log level

Customizing the look

Every identity string is an env var: BOR_APP_NAME (display name), BOR_INPUT_PLACEHOLDER (chat composer placeholder), and BOR_FOOTER_TEXT (footer line) — all served by GET /api/config and applied by assets/brand.js at boot.

The colors are set from the admin Theme tab (/theme.html, admin-only — phase 91): the 8 identity colors plus the three strings above are edited with text fields and color pickers, persisted in the ui_settings table, and injected into every served page as an inline <style> BEFORE first paint (no red flash, no pop-in). The old CSS-file theming is retired — a leftover BOR_THEME line in a deployment's .env is simply ignored. Leave everything unset and the app renders the built-in dark-tech palette byte-identically.


Checking retrieval quality

Ask the real pipeline (live aipi embeddings + the current KB) whether a question lands on the right document, with gate verdict and per-document cosine / FTS / fused scores:

uv run python -m scripts.eval_retrieval "How did I install gitlab?"
uv run python -m scripts.eval_retrieval --from-file questions.txt --top 8

Requires AIPI_KEY in the environment and an imported knowledge base.


Debugging

debugpy is off by default and never imported unless you opt in:

DEBUGPY=1 uv run uvicorn app.main:app
# → log line: debugpy: remote debugging ENABLED, listening on 0.0.0.0:5678

Attach from VS Code (.vscode/launch.json):

{
  "name": "Attach to Brain of Reese",
  "type": "debugpy",
  "request": "attach",
  "connect": { "host": "localhost", "port": 5678 },
  "pathMappings": [
    { "localRoot": "${workspaceFolder}", "remoteRoot": "/app" }
  ]
}

The port is non-blocking and attach-on-demand: the app keeps running normally until you attach. Override the port with DEBUGPY_PORT.


QA / Testing

Three layers — the project rule is one story, one phase, one Playwright suite (see AGENTS.md):

# Unit + integration (FastAPI TestClient)
uv run pytest

# Coverage gate (phases require >90% on app/)
uv run pytest --cov=app --cov-report=term-missing

# Lint + static types
uv run ruff check .
uv run pyright

# Playwright E2E — install the browser once:
uv run playwright install chromium

# Each story's E2E runs IN ISOLATION:
uv run pytest tests/e2e/test_import_documents.py -v --no-cov
uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
# ...one file per story in .agents/user_stories/ (see .agents/phases/todo/)

Deterministic E2E: by default the E2E app talks to a local mock aipi (tests/e2e/mock_llm.py) whose embeddings are real token-overlap vectors — so the cosine relevance threshold behaves like production. To run E2E against the live self-hosted models instead:

E2E_REAL_LLM=1 uv run pytest tests/e2e/test_chat_rag.py -v --no-cov

Production Deployment

Build the multi-stage image (frontend minified by esbuild in the builder stage, deps installed by uv, non-root runtime):

podman build -t brain-of-reese/app:latest .

Run standalone (bring your own Postgres + pgvector):

podman run -d --name brain-of-reese \
  -p 8000:8000 \
  -e BOR_DATABASE_URL=postgresql+psycopg://reese:SECRETPASSWORD@dbhost:5432/brain_of_reese \
  -e BOR_LLM_BASE_URL=https://aipi.reeseapps.com/v1 \
  -e BOR_LLM_API_KEY=$AIPI_KEY \
  brain-of-reese/app:latest

The entrypoint runs alembic upgrade head automatically on start.

Or run the whole stack from compose (app + db):

podman compose --profile prod up -d --build

Production hardening: app runs as non-root (uid 10001), slim image, healthcheck on /api/health, debugpy off unless DEBUGPY=1, all assets served locally (no CDN), BOR_ENVIRONMENT=production.


Troubleshooting

  • 401 from aipi — set BOR_LLM_API_KEY (or $AIPI_KEY).
  • litellm.UnsupportedParamsError … encoding_format — the aipi proxy (litellm openai_like) rejects the encoding_format parameter. The app already works around this by POSTing a minimal {model, input} payload. If you see this, you are likely calling the endpoint with a different client — drop the parameter (or set litellm.drop_params = True on the proxy).
  • Embedding dimension mismatch — aipi changed models; run uv run python -m scripts.llm_probe, update BOR_EMBEDDING_DIM, then drop + recreate the chunks table.
  • Honest deflection (the amber "I haven't done anything like that" bubble) — every question passes the honesty gate: deflection happens only when the best cosine similarity is below BOR_RELEVANCE_THRESHOLD (default 0.62) and no chunk matched lexically (fts_hits = 0). A weak cosine with a lexical hit (name-your-tool questions) still gets a grounded answer. When it does deflect, the LLM prompt carries weak-hit titles only (no document content), the reply opens with "I haven't done anything like that", the bubble renders amber with "Maybe try" chips derived from the closest indexed titles, and the query_log row records deflected=true. This is a feature, not a bug.
  • Answers deflect too often / too rarely — tune BOR_RELEVANCE_THRESHOLD (lower = answers more, higher = more honest deflection): 0.0 ⇒ the gate leans entirely on FTS hits; 1.0 ⇒ everything deflects unless a chunk matches lexically. The embed model's cosines cluster in a ~0.6–0.85 band on the live KB, so the default is 0.62. Check real scores:
    SELECT question, top_score, fts_hits, deflected
    FROM query_log ORDER BY created_at DESC LIMIT 20;
    
  • KB offline banner in the chat — Postgres isn't running: podman compose up -d db.
  • Stuck "Thinking…" — the LLM is slow or down; a 120s client timeout turns it into an error banner automatically.

Planning & Architecture

The architecture, LOCKED decisions, and the phase roadmap live in .agents/PLAN.md; per-story specs in .agents/user_stories/.

S
Description
No description provided
Readme
16 MiB
Languages
Python 87.9%
JavaScript 7.6%
CSS 2.6%
HTML 1.8%