Files
brain-of-reese/README.md
T
ducoterra 022da8e2bc feat: scaffold Brain of Reese — FastAPI RAG chat over Postgres 17 + pgvector
Foundation (phase 01, verified):
- FastAPI app: /api/health, /api/suggestions, /api/chat (placeholder),
  static frontend served locally (no CDN)
- Postgres 17 + pgvector via db/Containerfile + compose.yaml
  (podman compose up -d db), Alembic initial migration (documents,
  chunks with vector(768), query_log)
- LLM client targeting https://aipi.reeseapps.com/v1 (turbo/embed);
  scripts/llm_probe.py verified models + 768-dim embeddings live
- Conditional debugpy: imported only when DEBUGPY=1 (attach on demand,
  :5678); logging config for clean single-line logs
- Frontend shell: mobile-first chat + Sources pages, tokens, a11y baselines
- Tests: 24 unit+integration (99% coverage on app/), ruff + pyright clean,
  Playwright smoke E2E (3 tests) against a deterministic mock LLM
- Planning: .agent/PLAN.md (architecture + LOCKED decisions), AGENTS.md,
  6 user stories, 7 phase files (one story / one phase / one Playwright
  suite each)
2026-08-21 13:42:21 -04:00

7.2 KiB

🧠 Brain of Reese

A chippy, honest RAG chatbot over the ~/Homelab and ~/Deployments projects. Point it at your markdown docs, ask it anything — it retrieves the relevant notes with Postgres 17 + pgvector cosine search, feeds the whole relevant document to a self-hosted LLM (turbo via https://aipi.reeseapps.com/v1), and streams a grounded answer back.

If it doesn't have notes for your question, it admits it: "I haven't done anything like that" — plus suggestions for what it does know.

  • Stack: FastAPI · Pydantic v2 · SQLAlchemy 2 · Alembic · pgvector · vanilla HTML/CSS/JS (no CDN) · Playwright E2E
  • Planning: architecture, LOCKED decisions and the phase roadmap live in .agent/PLAN.md; per-story specs in .agent/user_stories/.

Development Setup

Prerequisites

  • uv
  • Podman (with the podman compose provider)
  • Node.js is not needed locally (asset minification happens in the container build only)

1. Install dependencies

uv sync

2. Configure

cp .env.example .env
# edit .env — the defaults already match the local compose setup.
# BOR_LLM_API_KEY: your aipi key (falls back to $AIPI_KEY if unset)

3. Start the database (Postgres 17 + pgvector)

podman compose up -d db
podman compose ps          # wait until "healthy"

4. Apply migrations

uv run alembic upgrade head

5. Import your knowledge base

uv run python -m scripts.llm_probe      # sanity: models + 768-dim check
uv run python -m scripts.import_docs    # defaults: ~/Homelab + ~/Deployments

6. Run the app

uv run uvicorn app.main:app --reload
# → http://localhost:8000  (chat)   http://localhost:8000/sources.html (KB)

Updating the documents

The knowledge base is refreshed by re-running the import. It is idempotent and delta-based (sha256 per file):

# After editing/adding/removing markdown in your projects:
uv run python -m scripts.import_docs                 # re-index what changed
uv run python -m scripts.import_docs --prune         # also drop deleted files

# Point it at extra directories (repeatable):
uv run python -m scripts.import_docs --source ~/SomeOtherDocs
  • Only *.md files are indexed. Directories like .venv, node_modules, .git, __pycache__, .pytest_cache, dist, build are skipped (see .agent/PLAN.md anchor A9).
  • Unchanged files are not re-embedded — only new/changed ones, so refreshes are cheap.
  • To sanity-check the LLM backend (models + embedding dimension) after any aipi change: uv run python -m scripts.llm_probe.

Debugging

debugpy is off by default and never imported unless you opt in — zero overhead in normal runs.

DEBUGPY=1 uv run uvicorn app.main:app
# → log line: debugpy: remote debugging ENABLED, listening on 0.0.0.0:5678

Then attach from VS Code (.vscode/launch.json):

{
  "name": "Attach to Brain of Reese",
  "type": "debugpy",
  "request": "attach",
  "connect": { "host": "localhost", "port": 5678 },
  "pathMappings": [
    { "localRoot": "${workspaceFolder}", "remoteRoot": "/app" }
  ]
}

The port is non-blocking and attach-on-demand: the app keeps running normally until you attach. Override the port with DEBUGPY_PORT.

QA / Testing Environment

Three layers — the project rule is one story, one phase, one Playwright suite (see AGENTS.md):

# Unit + integration (FastAPI TestClient)
uv run pytest

# Same, with the coverage gate (phases require >90% on app/)
uv run pytest --cov=app --cov-report=term-missing

# Lint + static types
uv run ruff check .
uv run pyright

# Playwright E2E — install the browser once:
uv run playwright install chromium

# Each story's E2E runs IN ISOLATION (DB must be up):
podman compose up -d db
uv run pytest tests/e2e/test_import_documents.py -v --no-cov
uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
# ...one file per story in .agent/user_stories/ (see .agent/phases/todo/)

Deterministic E2E: by default the E2E app talks to a local mock aipi (tests/e2e/mock_llm.py) whose embeddings are real token-overlap vectors — so the cosine relevance threshold behaves like production (on-topic questions answer, off-topic ones deflect). To run E2E against the live self-hosted models instead:

E2E_REAL_LLM=1 uv run pytest tests/e2e/test_chat_rag.py -v --no-cov

(requires a real import of your docs first).

Production Deployment

Build the multi-stage image (frontend minified by esbuild in the builder stage, deps installed by uv, non-root runtime):

podman build -t brain-of-reese/app:latest .

Run standalone (bring your own Postgres + pgvector):

podman run -d --name brain-of-reese \
  -p 8000:8000 \
  -e BOR_DATABASE_URL=postgresql+psycopg://reese:SECRETPASSWORD@dbhost:5432/brain_of_reese \
  -e BOR_LLM_BASE_URL=https://aipi.reeseapps.com/v1 \
  -e BOR_LLM_API_KEY=$AIPI_KEY \
  brain-of-reese/app:latest

The entrypoint runs alembic upgrade head automatically on start.

Or run the whole stack from compose (app + db):

podman compose --profile prod up -d --build

Production hardening notes: app runs as non-root (uid 10001), slim image, healthcheck on /api/health, debugpy off unless DEBUGPY=1, all assets served locally (no CDN), BOR_ENVIRONMENT=production.

Configuration reference

Env Default Meaning
BOR_DATABASE_URL local compose URL SQLAlchemy URL (psycopg)
BOR_LLM_BASE_URL https://aipi.reeseapps.com/v1 OpenAI-compatible endpoint
BOR_LLM_API_KEY — (falls back to $AIPI_KEY) aipi API key
BOR_LLM_CHAT_MODEL turbo chat model
BOR_LLM_EMBED_MODEL embed embedding model
BOR_EMBEDDING_DIM 768 vector dimension (fixed at table creation)
BOR_TOP_K_CHUNKS 4 chunks retrieved per question
BOR_TOP_N_DOCS 2 full documents fed to the LLM
BOR_RELEVANCE_THRESHOLD 0.30 best cosine similarity required to answer; below ⇒ honest deflection
BOR_MAX_CONTEXT_CHARS 24000 cap on total document text sent to the LLM
BOR_SUGGESTIONS built-in list JSON list of onboarding chips
DEBUGPY 0 1 ⇒ attach-on-demand debugpy on DEBUGPY_PORT (default 5678)
BOR_LOG_LEVEL INFO app log level

Troubleshooting

  • 401 from aipi — set BOR_LLM_API_KEY (or $AIPI_KEY).
  • Embedding dimension mismatch — aipi changed models; run uv run python -m scripts.llm_probe, update BOR_EMBEDDING_DIM, then drop + recreate the chunks table (new migration or manual TRUNCATE chunks, documents).
  • Answers deflect too often / too rarely — tune BOR_RELEVANCE_THRESHOLD (lower = answers more, higher = more honest deflection). Check query_log for the actual scores: psql … -c 'SELECT question, top_score, deflected FROM query_log ORDER BY created_at DESC LIMIT 20'
  • KB offline banner in the chat — Postgres isn't running: podman compose up -d db.
  • Stuck "Thinking…" — the LLM is slow or down; a 120s client timeout turns it into an error banner automatically.