8.9 KiB
🧠 Brain of Reese
A chippy, honest RAG chatbot over the ~/Homelab and ~/Deployments
projects. Point it at your markdown docs, ask it anything — it retrieves
the relevant notes with Postgres 17 + pgvector cosine search, feeds the
whole relevant document to a self-hosted LLM (turbo via
https://aipi.reeseapps.com/v1), and streams a grounded answer back.
If it doesn't have notes for your question, it admits it: "I haven't done anything like that" — plus suggestions for what it does know.
- Stack: FastAPI · Pydantic v2 · SQLAlchemy 2 · Alembic · pgvector · vanilla HTML/CSS/JS (no CDN) · Playwright E2E
- Planning: architecture, LOCKED decisions and the phase roadmap live
in
.agent/PLAN.md; per-story specs in.agent/user_stories/.
Development Setup
Prerequisites
- uv
- Podman (with the
podman composeprovider) - Node.js is not needed locally (asset minification happens in the container build only)
1. Install dependencies
uv sync
2. Configure
cp .env.example .env
# edit .env — the defaults already match the local compose setup.
# BOR_LLM_API_KEY: your aipi key (falls back to $AIPI_KEY if unset)
3. Start the database (Postgres 17 + pgvector)
podman compose up -d db
podman compose ps # wait until "healthy"
4. Apply migrations
uv run alembic upgrade head
5. Import your knowledge base
uv run python -m scripts.llm_probe # sanity: models + 768-dim check
uv run python -m scripts.import_docs # defaults: ~/Homelab + ~/Deployments
6. Run the app
uv run uvicorn app.main:app --reload
# → http://localhost:8000 (chat) http://localhost:8000/sources.html (KB)
Updating the documents
The knowledge base is refreshed by re-running the import. It is idempotent and delta-based (sha256 per file):
# After editing/adding/removing markdown in your projects:
uv run python -m scripts.import_docs # re-index what changed
uv run python -m scripts.import_docs --prune # also drop deleted files
# Point it at extra directories (repeatable):
uv run python -m scripts.import_docs --source ~/SomeOtherDocs
- Only
*.mdfiles are indexed. Directories like.venv,node_modules,.git,__pycache__,.pytest_cache,dist,buildare skipped (see.agent/PLAN.mdanchor A9). - Every file is logged on its own line (
import: added|updated|unchanged| pruned …), and the run ends with a one-line summary (import: summary files=… added=… updated=… unchanged=… pruned=… chunks=… embed_batches=…) so the counts are greppable in logs. - Unchanged files are not re-embedded — only new/changed ones, so refreshes are cheap.
- To sanity-check the LLM backend (models + embedding dimension) after any
aipi change:
uv run python -m scripts.llm_probe.
Debugging
debugpy is off by default and never imported unless you opt in —
zero overhead in normal runs.
DEBUGPY=1 uv run uvicorn app.main:app
# → log line: debugpy: remote debugging ENABLED, listening on 0.0.0.0:5678
Then attach from VS Code (.vscode/launch.json):
{
"name": "Attach to Brain of Reese",
"type": "debugpy",
"request": "attach",
"connect": { "host": "localhost", "port": 5678 },
"pathMappings": [
{ "localRoot": "${workspaceFolder}", "remoteRoot": "/app" }
]
}
The port is non-blocking and attach-on-demand: the app keeps running
normally until you attach. Override the port with DEBUGPY_PORT.
QA / Testing Environment
Three layers — the project rule is one story, one phase, one Playwright
suite (see AGENTS.md):
# Unit + integration (FastAPI TestClient)
uv run pytest
# Same, with the coverage gate (phases require >90% on app/)
uv run pytest --cov=app --cov-report=term-missing
# Lint + static types
uv run ruff check .
uv run pyright
# Playwright E2E — install the browser once:
uv run playwright install chromium
# Each story's E2E runs IN ISOLATION (DB must be up):
podman compose up -d db
uv run pytest tests/e2e/test_import_documents.py -v --no-cov
uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
# ...one file per story in .agent/user_stories/ (see .agent/phases/todo/)
Deterministic E2E: by default the E2E app talks to a local mock
aipi (tests/e2e/mock_llm.py) whose embeddings are real
token-overlap vectors — so the cosine relevance threshold behaves like
production (on-topic questions answer, off-topic ones deflect).
To run E2E against the live self-hosted models instead:
E2E_REAL_LLM=1 uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
(requires a real import of your docs first).
Production Deployment
Build the multi-stage image (frontend minified by esbuild in the builder
stage, deps installed by uv, non-root runtime):
podman build -t brain-of-reese/app:latest .
Run standalone (bring your own Postgres + pgvector):
podman run -d --name brain-of-reese \
-p 8000:8000 \
-e BOR_DATABASE_URL=postgresql+psycopg://reese:SECRETPASSWORD@dbhost:5432/brain_of_reese \
-e BOR_LLM_BASE_URL=https://aipi.reeseapps.com/v1 \
-e BOR_LLM_API_KEY=$AIPI_KEY \
brain-of-reese/app:latest
The entrypoint runs alembic upgrade head automatically on start.
Or run the whole stack from compose (app + db):
podman compose --profile prod up -d --build
Production hardening notes: app runs as non-root (uid 10001), slim image,
healthcheck on /api/health, debugpy off unless DEBUGPY=1, all assets
served locally (no CDN), BOR_ENVIRONMENT=production.
Configuration reference
| Env | Default | Meaning |
|---|---|---|
BOR_DATABASE_URL |
local compose URL | SQLAlchemy URL (psycopg) |
BOR_LLM_BASE_URL |
https://aipi.reeseapps.com/v1 |
OpenAI-compatible endpoint |
BOR_LLM_API_KEY |
— (falls back to $AIPI_KEY) |
aipi API key |
BOR_LLM_CHAT_MODEL |
turbo |
chat model |
BOR_LLM_EMBED_MODEL |
embed |
embedding model |
BOR_EMBEDDING_DIM |
768 |
vector dimension (fixed at table creation) |
BOR_TOP_K_CHUNKS |
4 |
chunks retrieved per question |
BOR_TOP_N_DOCS |
2 |
full documents fed to the LLM |
BOR_RELEVANCE_THRESHOLD |
0.30 |
best cosine similarity required to answer; below ⇒ honest deflection |
BOR_MAX_CONTEXT_CHARS |
24000 |
cap on total document text sent to the LLM |
BOR_SUGGESTIONS |
built-in list | JSON list of onboarding chips |
DEBUGPY |
0 |
1 ⇒ attach-on-demand debugpy on DEBUGPY_PORT (default 5678) |
BOR_LOG_LEVEL |
INFO |
app log level |
Troubleshooting
401from aipi — setBOR_LLM_API_KEY(or$AIPI_KEY).litellm.UnsupportedParamsError … encoding_formatfrom aipi — the aipi proxy (litellmopenai_like) rejects theencoding_formatparameter that theopenaiSDK injects into every embeddings request. The app already works around this by POSTing a minimal{model, input}payload through the openai client's own httpx transport (app/rag/llm.py→LLMClient._embed_batch). If you see this, you are likely calling the endpoint with a different client — drop the parameter (or setlitellm.drop_params = Trueon the proxy).- Embedding dimension mismatch — aipi changed models; run
uv run python -m scripts.llm_probe, updateBOR_EMBEDDING_DIM, then drop + recreate the chunks table (new migration or manualTRUNCATE chunks, documents). - Honest deflection (the amber “I haven't done anything like that”
bubble) — every question passes the honesty gate: when the best
cosine similarity is below
BOR_RELEVANCE_THRESHOLD(default0.30), Brain switches to deflection mode instead of guessing. The LLM prompt then carries weak-hit titles only (no document content), the reply opens with “I haven't done anything like that”, the bubble renders amber with “Maybe try” chips derived from the closest indexed titles, the SSEdoneevent carriesdeflected: true+suggestions[], and thequery_logrow recordsdeflected=true+ the weaktop_score. This is a feature, not a bug — the KB simply has no notes that close; the chips always point at topics Brain really covers. - Answers deflect too often / too rarely — tune
BOR_RELEVANCE_THRESHOLD(lower = answers more, higher = more honest deflection):0.0⇒ every question gets answered, even unknown topics (expect confident-sounding guesses);1.0⇒ everything deflects (nothing but a perfect 1.0 score counts as relevant). After changing it, check the real scores:psql … -c 'SELECT question, top_score, deflected FROM query_log ORDER BY created_at DESC LIMIT 20' - KB offline banner in the chat — Postgres isn't running:
podman compose up -d db. - Stuck "Thinking…" — the LLM is slow or down; a 120s client timeout turns it into an error banner automatically.