🧠 Brain of Reese

A chippy, honest RAG chatbot over the ~/Homelab and ~/Deployments projects. Point it at your notes — markdown, YAML, JSON, Python, plain text — ask it anything, and it retrieves the relevant chunks with hybrid search (pgvector cosine ∪ Postgres full-text search, fused with Reciprocal Rank Fusion), feeds the whole relevant document to a self-hosted LLM (turbo via https://aipi.reeseapps.com/v1), and streams a grounded answer back.

If it doesn't have notes for your question, it admits it: "I haven't done anything like that" — plus suggestions for what it does know.

Updated your notes? The knowledge base is refreshed by re-running the import — it's idempotent and only re-embeds what changed:

uv run python -m scripts.import_docs --prune

Details in Updating the documents.

  • Stack: FastAPI · Pydantic v2 · SQLAlchemy 2 · Alembic · pgvector · vanilla HTML/CSS/JS (no CDN) · Playwright E2E
  • Planning: architecture, LOCKED decisions and the phase roadmap live in .agent/PLAN.md; per-story specs in .agent/user_stories/.

Development Setup

Prerequisites

  • uv
  • Podman (with the podman compose provider)
  • Node.js is not needed locally (asset minification happens in the container build only)

1. Install dependencies

uv sync

2. Configure

cp .env.example .env
# edit .env — the defaults already match the local compose setup.
# BOR_LLM_API_KEY: your aipi key (falls back to $AIPI_KEY if unset)

3. Start the database (Postgres 17 + pgvector)

podman compose up -d db
podman compose ps          # wait until "healthy"

4. Apply migrations

uv run alembic upgrade head

5. Import your knowledge base

uv run python -m scripts.llm_probe      # sanity: models + 768-dim check
uv run python -m scripts.import_docs    # defaults: ~/Homelab + ~/Deployments

6. Run the app

uv run uvicorn app.main:app --reload
# → http://localhost:8000  (chat)   http://localhost:8000/sources.html (KB)

📝 After this, day-to-day is just: edit markdown → re-run the import. See Updating the documents below.

Using the UI

  • Chat (/) — ask questions; answers stream in with source chips that cite the exact documents used. Clicking a chip opens that document in a new tab.
  • Document viewer (/document.html?source=…&path=…) — the full text of any indexed document, served from the database (no filesystem access): markdown is rendered, every other format (yaml, json, py, txt, …) is shown as escaped monospace text. Unknown documents get a designed not-found state with a link back to the index.
  • Sources (/sources.html) — the indexed document list; the Path column links each document to the viewer in a new tab. Admin-only — anonymous visitors see a sign-in gate instead (the catalog is what the login locks; the document viewer itself stays open to everyone).

Thinking

The self-hosted turbo model reasons before it answers. That reasoning is streamed with the turn as thinking SSE events and shown in a collapsible "Thinking" block above the answer bubble: it opens and fills in live while the model thinks, tucks itself away the moment the first answer token lands, and stays click-toggleable afterwards. Thinking persists with the message, so a reloaded conversation restores the block (collapsed) alongside the answer. How much the model thinks — or whether it thinks at all — is the model's call: turns without reasoning render exactly as before.

To hide it, set BOR_STREAM_THINKING=0 — the thinking events stop (the per-turn log line still counts thinking_chars).

Admin & sign-in

Brain of Reese has exactly one account: the admin (you). Signing in unlocks the full Sources catalog and the answer-tuning controls; everyone else stays anonymous and keeps chat and the document viewer (any document an answer cites can be opened by its direct URL — the catalog is gated, not the viewer).

Setup (one-time)

python -c 'import secrets;print(secrets.token_hex(32))'   # → paste into .env
BOR_ADMIN_PASSWORD=your-password        # plaintext — homelab scope, by design
BOR_SESSION_SECRET=<the hex from above> # signs the session cookie

Fail-loud: while either variable is empty the app refuses to start, naming the missing one(s):

RuntimeError: Brain of Reese cannot start: admin auth is not configured.
Set the missing variable(s): BOR_ADMIN_PASSWORD, BOR_SESSION_SECRET …

How it works

  • POST /api/login {"password": …} → 204 + signed bor_session cookie (Starlette SessionMiddleware — an itsdangerous-signed cookie, no server-side store, no new service, no DB table); any mismatch → 401 {"detail": "invalid password"} (constant-time compare, one generic message — no user enumeration, there is only one user).
  • POST /api/logout → 204 (session cleared and cookie expired; idempotent for anonymous callers).
  • GET /api/whoami → {"authenticated": bool, "role": "admin"|"anonymous"} — the single source of truth for every UI gating decision.
  • Cookie flags: same_site="lax", https_only off — no HTTPS enforcement on purpose (homelab HTTP; the cookie is single-admin convenience, not a cloud boundary). Max age BOR_SESSION_MAX_AGE (default 43200 = 12 h, refreshed while active).
  • Sign in from the chat header (Sign in) or /login.html directly; the header then offers Sign out (logout + reload).

Who can do what

Capability Anonymous Admin (signed in)
Chat (/) + suggestion chips yes yes
Document viewer (/document.html?source=…&path=…) yes — any indexed doc by direct URL yes
Sources catalog (/sources.html, GET /api/docs) sign-in gate full catalog
Tuning (Tune button, Tuning panel, /api/steering) UI hidden full

The public API endpoints stay stateless — the signed cookie is the only session state in the system.

Tuning your answers

Admin-only — sign in first (see Admin & sign-in above); anonymous visitors never see the Tune button or the Tuning panel.

If an answer isn't quite right — too chatty, wrong assumption, missing context — tune Brain right there:

  1. Press “Tune” in the meta row under any completed answer (deflected ones included).
  2. Type a short instruction (1–2000 chars), e.g. “be more concise” or “assume I'm on NixOS”, and Save.

The note is stored in Postgres (steering_notes) and read into the system prompt of every subsequent chat turn as a <tuning> section (numbered, oldest first, capped at BOR_STEERING_MAX_CHARS chars — default 8000, overflow marked […truncated…]). With no stored notes the prompt is byte-identical to the un-tuned one, so tuning is opt-in per note.

List or remove notes at any time from the “Tuning” button in the chat header (count badge, newest-first, per-note delete). The API is stateless JSON if you prefer curl:

curl -s localhost:8000/api/steering                                  # list (newest first)
curl -s -X POST localhost:8000/api/steering \
     -H 'Content-Type: application/json' -d '{"note": "be more concise"}'
curl -s -X DELETE localhost:8000/api/steering/<note-id>              # remove

Updating the documents

This is the workflow you'll use most. The knowledge base is refreshed by re-running the import. It is idempotent and delta-based (sha256 per file), so a refresh after a normal editing session takes seconds:

# After editing/adding/removing notes in your projects:
uv run python -m scripts.import_docs                 # re-index what changed
uv run python -m scripts.import_docs --prune         # also drop deleted/out-of-scope files

# Point it at extra directories (repeatable):
uv run python -m scripts.import_docs --source ~/SomeOtherDocs

Then check the Sources page (http://localhost:8000/sources.html): the documents / chunks counters and last indexed timestamp should reflect the new files, and each document row shows when it was last embedded.

  • The import prints one line per file (import: added|updated|unchanged| pruned …) and ends with a greppable summary (import: summary files=… added=… updated=… unchanged=… pruned=… chunks=… embed_batches=… formats=md:203,yaml:267,…), so it is safe to run from a cron job or after every commit.
  • Indexed formats (A9): md, markdown, txt, yaml, yml, json, py (case-insensitive; narrow with BOR_IMPORT_EXTENSIONS). Any path with a dot-prefixed component — hidden files or vendored caches like .esphome/.espressif/** — is skipped, along with .venv, node_modules, .git, __pycache__, .pytest_cache, dist, build. --prune also drops documents whose files no longer match the filter — that's how previously imported junk leaves the index.
  • Non-markdown files get format-aware chunking (YAML top-level keys / --- docs, JSON top-level keys, Python top-level defs/classes via stdlib ast) and their title comes from the file stem.
  • Unchanged files are not re-embedded — only new/changed ones, so refreshes are cheap.
  • To sanity-check the LLM backend (models + embedding dimension) after any aipi change: uv run python -m scripts.llm_probe.

Checking retrieval quality

Ask the real pipeline (live aipi embeddings + the current KB) whether a question lands on the right document, with the gate verdict and per-document cosine / FTS / fused scores:

uv run python -m scripts.eval_retrieval "How did I install gitlab?"
uv run python -m scripts.eval_retrieval --from-file questions.txt --top 8

Requires AIPI_KEY in the environment (same convention as scripts/llm_probe.py) and an imported knowledge base.

How retrieval works (hybrid)

Every question is embedded and also lexically tokenized (OR-joined, English stemming) and searched twice against Postgres:

  1. Vector — pgvector cosine top-N (default BOR_HYBRID_VECTOR_CANDIDATES=100)
  2. Lexical — a stored tsvector (GIN-indexed) matched with to_tsquery, top-N by ts_rank (default BOR_HYBRID_LEXICAL_CANDIDATES=30)

The two ranked lists are fused with Reciprocal Rank Fusion (score = Σ 1/(k + rank), BOR_RRF_K=60) — a chunk in both lists scores nearly double, which is what lets a name-your-tool question ("gitlab") find its own document even when the question embeds close to generic templates.

The honesty gate (A8) then answers (HIGH) when the best cosine is ≥ BOR_RELEVANCE_THRESHOLD (default 0.62) or at least one chunk matched lexically (fts_hits > 0) — it deflects (LOW) only when both signals are absent. The top BOR_TOP_N_DOCS full documents are still what the LLM sees.

query_log records every turn (top_score = best cosine, fts_hits, chunk_hits, deflected, sources, latency_ms) — the raw material for tuning: psql … -c 'SELECT question, top_score, fts_hits, deflected FROM query_log ORDER BY created_at DESC LIMIT 20'.

Debugging

debugpy is off by default and never imported unless you opt in — zero overhead in normal runs.

DEBUGPY=1 uv run uvicorn app.main:app
# → log line: debugpy: remote debugging ENABLED, listening on 0.0.0.0:5678

Then attach from VS Code (.vscode/launch.json):

{
  "name": "Attach to Brain of Reese",
  "type": "debugpy",
  "request": "attach",
  "connect": { "host": "localhost", "port": 5678 },
  "pathMappings": [
    { "localRoot": "${workspaceFolder}", "remoteRoot": "/app" }
  ]
}

The port is non-blocking and attach-on-demand: the app keeps running normally until you attach. Override the port with DEBUGPY_PORT.

QA / Testing Environment

Three layers — the project rule is one story, one phase, one Playwright suite (see AGENTS.md):

# Unit + integration (FastAPI TestClient)
uv run pytest

# Same, with the coverage gate (phases require >90% on app/)
uv run pytest --cov=app --cov-report=term-missing

# Lint + static types
uv run ruff check .
uv run pyright

# Playwright E2E — install the browser once:
uv run playwright install chromium

# Each story's E2E runs IN ISOLATION (DB must be up):
podman compose up -d db
uv run pytest tests/e2e/test_import_documents.py -v --no-cov
uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
# ...one file per story in .agent/user_stories/ (see .agent/phases/todo/)

Deterministic E2E: by default the E2E app talks to a local mock aipi (tests/e2e/mock_llm.py) whose embeddings are real token-overlap vectors — so the cosine relevance threshold behaves like production (on-topic questions answer, off-topic ones deflect). To run E2E against the live self-hosted models instead:

E2E_REAL_LLM=1 uv run pytest tests/e2e/test_chat_rag.py -v --no-cov

(requires a real import of your docs first).

Production Deployment

Build the multi-stage image (frontend minified by esbuild in the builder stage, deps installed by uv, non-root runtime):

podman build -t brain-of-reese/app:latest .

Run standalone (bring your own Postgres + pgvector):

podman run -d --name brain-of-reese \
  -p 8000:8000 \
  -e BOR_DATABASE_URL=postgresql+psycopg://reese:SECRETPASSWORD@dbhost:5432/brain_of_reese \
  -e BOR_LLM_BASE_URL=https://aipi.reeseapps.com/v1 \
  -e BOR_LLM_API_KEY=$AIPI_KEY \
  brain-of-reese/app:latest

The entrypoint runs alembic upgrade head automatically on start.

Or run the whole stack from compose (app + db):

podman compose --profile prod up -d --build

Production hardening notes: app runs as non-root (uid 10001), slim image, healthcheck on /api/health, debugpy off unless DEBUGPY=1, all assets served locally (no CDN), BOR_ENVIRONMENT=production.

Configuration reference

Env Default Meaning
BOR_DATABASE_URL local compose URL SQLAlchemy URL (psycopg)
BOR_LLM_BASE_URL https://aipi.reeseapps.com/v1 OpenAI-compatible endpoint
BOR_LLM_API_KEY — (falls back to $AIPI_KEY) aipi API key
BOR_LLM_CHAT_MODEL turbo chat model
BOR_LLM_EMBED_MODEL embed embedding model
BOR_EMBEDDING_DIM 768 vector dimension (fixed at table creation)
BOR_TOP_N_DOCS 2 full documents fed to the LLM
BOR_RELEVANCE_THRESHOLD 0.62 answer when best cosine ≥ this or an FTS hit; below + no FTS ⇒ honest deflection
BOR_HYBRID_VECTOR_CANDIDATES 100 cosine list width for the RRF fusion
BOR_HYBRID_LEXICAL_CANDIDATES 30 FTS list width for the RRF fusion
BOR_RRF_K 60 RRF damping constant (1/(k + rank))
BOR_IMPORT_EXTENSIONS md,markdown,txt,yaml,yml,json,py csv of importable formats (may only narrow the A9 set)
BOR_MAX_CONTEXT_CHARS 24000 cap on total document text sent to the LLM
BOR_STEERING_MAX_CHARS 8000 char budget for the <tuning> (steering notes) prompt section
BOR_SUGGESTIONS built-in list JSON list of onboarding chips
BOR_ADMIN_PASSWORD (required) the single admin's password (plaintext, .env); app refuses to start when empty
BOR_SESSION_SECRET (required) signing key for the bor_session cookie; python -c 'import secrets;print(secrets.token_hex(32))'
BOR_SESSION_MAX_AGE 43200 session-cookie lifetime in seconds (12 h, sliding)
DEBUGPY 0 1 ⇒ attach-on-demand debugpy on DEBUGPY_PORT (default 5678)
BOR_LOG_LEVEL INFO app log level

Troubleshooting

  • 401 from aipi — set BOR_LLM_API_KEY (or $AIPI_KEY).
  • litellm.UnsupportedParamsError … encoding_format from aipi — the aipi proxy (litellm openai_like) rejects the encoding_format parameter that the openai SDK injects into every embeddings request. The app already works around this by POSTing a minimal {model, input} payload through the openai client's own httpx transport (app/rag/llm.py → LLMClient._embed_batch). If you see this, you are likely calling the endpoint with a different client — drop the parameter (or set litellm.drop_params = True on the proxy).
  • Embedding dimension mismatch — aipi changed models; run uv run python -m scripts.llm_probe, update BOR_EMBEDDING_DIM, then drop + recreate the chunks table (new migration or manual TRUNCATE chunks, documents).
  • Honest deflection (the amber “I haven't done anything like that” bubble) — every question passes the honesty gate: deflection happens only when the best cosine similarity is below BOR_RELEVANCE_THRESHOLD (default 0.62) and no chunk matched the question lexically (fts_hits = 0). A weak cosine with a lexical hit (name-your-tool questions) still gets a grounded answer. When it does deflect, the LLM prompt carries weak-hit titles only (no document content), the reply opens with “I haven't done anything like that”, the bubble renders amber with “Maybe try” chips derived from the closest indexed titles, the SSE done event carries deflected: true + suggestions[], and the query_log row records deflected=true + the weak top_score + fts_hits. This is a feature, not a bug — the KB simply has no notes that close; the chips always point at topics Brain really covers.
  • Answers deflect too often / too rarely — tune BOR_RELEVANCE_THRESHOLD (lower = answers more, higher = more honest deflection): 0.0 ⇒ the gate leans entirely on FTS hits; 1.0 ⇒ everything deflects unless a chunk matches lexically. The embed model's cosines cluster in a ~0.6–0.85 band on the live KB, so the default is 0.62; after changing it, check the real scores: psql … -c 'SELECT question, top_score, fts_hits, deflected FROM query_log ORDER BY created_at DESC LIMIT 20'
  • KB offline banner in the chat — Postgres isn't running: podman compose up -d db.
  • Stuck "Thinking…" — the LLM is slow or down; a 120s client timeout turns it into an error banner automatically.
S
Description
No description provided
Readme
16 MiB
Languages
Python 87.9%
JavaScript 7.6%
CSS 2.6%
HTML 1.8%