Phase 02 (story: import documents):
- fence-aware markdown chunker (heading sections, 200-char overlap,
heading anchor on every chunk, 1200-char hard cap, fence blocks
kept atomic and split under the cap)
- LLMClient over aipi (LiteLLM) reusing the openai client's httpx
transport to send a clean {model, input} payload — the openai SDK
injects encoding_format, which aipi's openai_like group rejects;
token-budget batching + halving retry for the endpoint's
~1024-token per-request input cap
- two-phase per-file upsert importer: sha256 delta (unchanged skip),
atomic commit, A9 exclusion walk, per-source prune, per-file error
tolerance (rollback + log + continue, non-zero CLI exit), adaptive
re-chunk at half target for URL-dense files the endpoint rejects
- scripts/import_docs CLI (repeatable --source, --prune, --limit,
defaults ~/Homelab + ~/Deployments)
- GET /api/docs with per-doc chunk counts; Sources page wired to the
real endpoint (stat cards, full-width a11y table, designed empty
state, DOM-built rows — no innerHTML)
- tests: 63 passed (chunker/llm/importer units, docs API + importer
integration), story E2E 3/3 (real endpoints, in-thread import);
app/ coverage 98%
- real KB imported: 672 docs / 8969 chunks in ~3m, idempotent
re-run (672 unchanged, 0 batches)
- harness: .agent/validate.sh now gates through uv (pytest +
coverage >90% + ruff + pyright) instead of system python3
🧠 Brain of Reese
A chippy, honest RAG chatbot over the ~/Homelab and ~/Deployments
projects. Point it at your markdown docs, ask it anything — it retrieves
the relevant notes with Postgres 17 + pgvector cosine search, feeds the
whole relevant document to a self-hosted LLM (turbo via
https://aipi.reeseapps.com/v1), and streams a grounded answer back.
If it doesn't have notes for your question, it admits it: "I haven't done anything like that" — plus suggestions for what it does know.
- Stack: FastAPI · Pydantic v2 · SQLAlchemy 2 · Alembic · pgvector · vanilla HTML/CSS/JS (no CDN) · Playwright E2E
- Planning: architecture, LOCKED decisions and the phase roadmap live
in
.agent/PLAN.md; per-story specs in.agent/user_stories/.
Development Setup
Prerequisites
- uv
- Podman (with the
podman composeprovider) - Node.js is not needed locally (asset minification happens in the container build only)
1. Install dependencies
uv sync
2. Configure
cp .env.example .env
# edit .env — the defaults already match the local compose setup.
# BOR_LLM_API_KEY: your aipi key (falls back to $AIPI_KEY if unset)
3. Start the database (Postgres 17 + pgvector)
podman compose up -d db
podman compose ps # wait until "healthy"
4. Apply migrations
uv run alembic upgrade head
5. Import your knowledge base
uv run python -m scripts.llm_probe # sanity: models + 768-dim check
uv run python -m scripts.import_docs # defaults: ~/Homelab + ~/Deployments
6. Run the app
uv run uvicorn app.main:app --reload
# → http://localhost:8000 (chat) http://localhost:8000/sources.html (KB)
Updating the documents
The knowledge base is refreshed by re-running the import. It is idempotent and delta-based (sha256 per file):
# After editing/adding/removing markdown in your projects:
uv run python -m scripts.import_docs # re-index what changed
uv run python -m scripts.import_docs --prune # also drop deleted files
# Point it at extra directories (repeatable):
uv run python -m scripts.import_docs --source ~/SomeOtherDocs
- Only
*.mdfiles are indexed. Directories like.venv,node_modules,.git,__pycache__,.pytest_cache,dist,buildare skipped (see.agent/PLAN.mdanchor A9). - Every file is logged on its own line (
import: added|updated|unchanged| pruned …), and the run ends with a one-line summary (import: summary files=… added=… updated=… unchanged=… pruned=… chunks=… embed_batches=…) so the counts are greppable in logs. - Unchanged files are not re-embedded — only new/changed ones, so refreshes are cheap.
- To sanity-check the LLM backend (models + embedding dimension) after any
aipi change:
uv run python -m scripts.llm_probe.
Debugging
debugpy is off by default and never imported unless you opt in —
zero overhead in normal runs.
DEBUGPY=1 uv run uvicorn app.main:app
# → log line: debugpy: remote debugging ENABLED, listening on 0.0.0.0:5678
Then attach from VS Code (.vscode/launch.json):
{
"name": "Attach to Brain of Reese",
"type": "debugpy",
"request": "attach",
"connect": { "host": "localhost", "port": 5678 },
"pathMappings": [
{ "localRoot": "${workspaceFolder}", "remoteRoot": "/app" }
]
}
The port is non-blocking and attach-on-demand: the app keeps running
normally until you attach. Override the port with DEBUGPY_PORT.
QA / Testing Environment
Three layers — the project rule is one story, one phase, one Playwright
suite (see AGENTS.md):
# Unit + integration (FastAPI TestClient)
uv run pytest
# Same, with the coverage gate (phases require >90% on app/)
uv run pytest --cov=app --cov-report=term-missing
# Lint + static types
uv run ruff check .
uv run pyright
# Playwright E2E — install the browser once:
uv run playwright install chromium
# Each story's E2E runs IN ISOLATION (DB must be up):
podman compose up -d db
uv run pytest tests/e2e/test_import_documents.py -v --no-cov
uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
# ...one file per story in .agent/user_stories/ (see .agent/phases/todo/)
Deterministic E2E: by default the E2E app talks to a local mock
aipi (tests/e2e/mock_llm.py) whose embeddings are real
token-overlap vectors — so the cosine relevance threshold behaves like
production (on-topic questions answer, off-topic ones deflect).
To run E2E against the live self-hosted models instead:
E2E_REAL_LLM=1 uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
(requires a real import of your docs first).
Production Deployment
Build the multi-stage image (frontend minified by esbuild in the builder
stage, deps installed by uv, non-root runtime):
podman build -t brain-of-reese/app:latest .
Run standalone (bring your own Postgres + pgvector):
podman run -d --name brain-of-reese \
-p 8000:8000 \
-e BOR_DATABASE_URL=postgresql+psycopg://reese:SECRETPASSWORD@dbhost:5432/brain_of_reese \
-e BOR_LLM_BASE_URL=https://aipi.reeseapps.com/v1 \
-e BOR_LLM_API_KEY=$AIPI_KEY \
brain-of-reese/app:latest
The entrypoint runs alembic upgrade head automatically on start.
Or run the whole stack from compose (app + db):
podman compose --profile prod up -d --build
Production hardening notes: app runs as non-root (uid 10001), slim image,
healthcheck on /api/health, debugpy off unless DEBUGPY=1, all assets
served locally (no CDN), BOR_ENVIRONMENT=production.
Configuration reference
| Env | Default | Meaning |
|---|---|---|
BOR_DATABASE_URL |
local compose URL | SQLAlchemy URL (psycopg) |
BOR_LLM_BASE_URL |
https://aipi.reeseapps.com/v1 |
OpenAI-compatible endpoint |
BOR_LLM_API_KEY |
— (falls back to $AIPI_KEY) |
aipi API key |
BOR_LLM_CHAT_MODEL |
turbo |
chat model |
BOR_LLM_EMBED_MODEL |
embed |
embedding model |
BOR_EMBEDDING_DIM |
768 |
vector dimension (fixed at table creation) |
BOR_TOP_K_CHUNKS |
4 |
chunks retrieved per question |
BOR_TOP_N_DOCS |
2 |
full documents fed to the LLM |
BOR_RELEVANCE_THRESHOLD |
0.30 |
best cosine similarity required to answer; below ⇒ honest deflection |
BOR_MAX_CONTEXT_CHARS |
24000 |
cap on total document text sent to the LLM |
BOR_SUGGESTIONS |
built-in list | JSON list of onboarding chips |
DEBUGPY |
0 |
1 ⇒ attach-on-demand debugpy on DEBUGPY_PORT (default 5678) |
BOR_LOG_LEVEL |
INFO |
app log level |
Troubleshooting
401from aipi — setBOR_LLM_API_KEY(or$AIPI_KEY).litellm.UnsupportedParamsError … encoding_formatfrom aipi — the aipi proxy (litellmopenai_like) rejects theencoding_formatparameter that theopenaiSDK injects into every embeddings request. The app already works around this by POSTing a minimal{model, input}payload through the openai client's own httpx transport (app/rag/llm.py→LLMClient._embed_batch). If you see this, you are likely calling the endpoint with a different client — drop the parameter (or setlitellm.drop_params = Trueon the proxy).- Embedding dimension mismatch — aipi changed models; run
uv run python -m scripts.llm_probe, updateBOR_EMBEDDING_DIM, then drop + recreate the chunks table (new migration or manualTRUNCATE chunks, documents). - Answers deflect too often / too rarely — tune
BOR_RELEVANCE_THRESHOLD(lower = answers more, higher = more honest deflection). Checkquery_logfor the actual scores:psql … -c 'SELECT question, top_score, deflected FROM query_log ORDER BY created_at DESC LIMIT 20' - KB offline banner in the chat — Postgres isn't running:
podman compose up -d db. - Stuck "Thinking…" — the LLM is slow or down; a 120s client timeout turns it into an error banner automatically.