# 🧠 Brain of Reese A chippy, honest **RAG chatbot** over the `~/Homelab` and `~/Deployments` projects. Point it at your markdown docs, ask it anything β€” it retrieves the relevant notes with **Postgres 17 + pgvector** cosine search, feeds the **whole relevant document** to a **self-hosted LLM** (`turbo` via `https://aipi.reeseapps.com/v1`), and streams a grounded answer back. If it doesn't have notes for your question, it admits it: *"I haven't done anything like that"* β€” plus suggestions for what it **does** know. - **Stack:** FastAPI Β· Pydantic v2 Β· SQLAlchemy 2 Β· Alembic Β· pgvector Β· vanilla HTML/CSS/JS (no CDN) Β· Playwright E2E - **Planning:** architecture, LOCKED decisions and the phase roadmap live in [`.agent/PLAN.md`](.agent/PLAN.md); per-story specs in [`.agent/user_stories/`](.agent/user_stories/). --- ## Development Setup ### Prerequisites - [uv](https://docs.astral.sh/uv/) - [Podman](https://podman.io/) (with the `podman compose` provider) - Node.js is **not** needed locally (asset minification happens in the container build only) ### 1. Install dependencies ```bash uv sync ``` ### 2. Configure ```bash cp .env.example .env # edit .env β€” the defaults already match the local compose setup. # BOR_LLM_API_KEY: your aipi key (falls back to $AIPI_KEY if unset) ``` ### 3. Start the database (Postgres 17 + pgvector) ```bash podman compose up -d db podman compose ps # wait until "healthy" ``` ### 4. Apply migrations ```bash uv run alembic upgrade head ``` ### 5. Import your knowledge base ```bash uv run python -m scripts.llm_probe # sanity: models + 768-dim check uv run python -m scripts.import_docs # defaults: ~/Homelab + ~/Deployments ``` ### 6. Run the app ```bash uv run uvicorn app.main:app --reload # β†’ http://localhost:8000 (chat) http://localhost:8000/sources.html (KB) ``` ## Updating the documents The knowledge base is refreshed by **re-running the import**. It is idempotent and delta-based (sha256 per file): ```bash # After editing/adding/removing markdown in your projects: uv run python -m scripts.import_docs # re-index what changed uv run python -m scripts.import_docs --prune # also drop deleted files # Point it at extra directories (repeatable): uv run python -m scripts.import_docs --source ~/SomeOtherDocs ``` - Only **`*.md`** files are indexed. Directories like `.venv`, `node_modules`, `.git`, `__pycache__`, `.pytest_cache`, `dist`, `build` are skipped (see `.agent/PLAN.md` anchor A9). - Unchanged files are **not re-embedded** β€” only new/changed ones, so refreshes are cheap. - To sanity-check the LLM backend (models + embedding dimension) after any aipi change: `uv run python -m scripts.llm_probe`. ## Debugging `debugpy` is **off by default** and *never imported* unless you opt in β€” zero overhead in normal runs. ```bash DEBUGPY=1 uv run uvicorn app.main:app # β†’ log line: debugpy: remote debugging ENABLED, listening on 0.0.0.0:5678 ``` Then attach from VS Code (`.vscode/launch.json`): ```json { "name": "Attach to Brain of Reese", "type": "debugpy", "request": "attach", "connect": { "host": "localhost", "port": 5678 }, "pathMappings": [ { "localRoot": "${workspaceFolder}", "remoteRoot": "/app" } ] } ``` The port is non-blocking and attach-on-demand: the app keeps running normally until you attach. Override the port with `DEBUGPY_PORT`. ## QA / Testing Environment Three layers β€” the project rule is **one story, one phase, one Playwright suite** (see `AGENTS.md`): ```bash # Unit + integration (FastAPI TestClient) uv run pytest # Same, with the coverage gate (phases require >90% on app/) uv run pytest --cov=app --cov-report=term-missing # Lint + static types uv run ruff check . uv run pyright # Playwright E2E β€” install the browser once: uv run playwright install chromium # Each story's E2E runs IN ISOLATION (DB must be up): podman compose up -d db uv run pytest tests/e2e/test_import_documents.py -v --no-cov uv run pytest tests/e2e/test_chat_rag.py -v --no-cov # ...one file per story in .agent/user_stories/ (see .agent/phases/todo/) ``` **Deterministic E2E:** by default the E2E app talks to a local **mock aipi** (`tests/e2e/mock_llm.py`) whose embeddings are real token-overlap vectors β€” so the cosine relevance threshold behaves like production (on-topic questions answer, off-topic ones deflect). To run E2E against the **live** self-hosted models instead: ```bash E2E_REAL_LLM=1 uv run pytest tests/e2e/test_chat_rag.py -v --no-cov ``` (requires a real import of your docs first). ## Production Deployment Build the multi-stage image (frontend minified by esbuild in the builder stage, deps installed by `uv`, non-root runtime): ```bash podman build -t brain-of-reese/app:latest . ``` Run standalone (bring your own Postgres + pgvector): ```bash podman run -d --name brain-of-reese \ -p 8000:8000 \ -e BOR_DATABASE_URL=postgresql+psycopg://reese:SECRETPASSWORD@dbhost:5432/brain_of_reese \ -e BOR_LLM_BASE_URL=https://aipi.reeseapps.com/v1 \ -e BOR_LLM_API_KEY=$AIPI_KEY \ brain-of-reese/app:latest ``` The entrypoint runs `alembic upgrade head` automatically on start. Or run the whole stack from compose (app + db): ```bash podman compose --profile prod up -d --build ``` Production hardening notes: app runs as non-root (uid 10001), slim image, healthcheck on `/api/health`, debugpy off unless `DEBUGPY=1`, all assets served locally (no CDN), `BOR_ENVIRONMENT=production`. ## Configuration reference | Env | Default | Meaning | |-----|---------|---------| | `BOR_DATABASE_URL` | local compose URL | SQLAlchemy URL (psycopg) | | `BOR_LLM_BASE_URL` | `https://aipi.reeseapps.com/v1` | OpenAI-compatible endpoint | | `BOR_LLM_API_KEY` | β€” (falls back to `$AIPI_KEY`) | aipi API key | | `BOR_LLM_CHAT_MODEL` | `turbo` | chat model | | `BOR_LLM_EMBED_MODEL` | `embed` | embedding model | | `BOR_EMBEDDING_DIM` | `768` | vector dimension (fixed at table creation) | | `BOR_TOP_K_CHUNKS` | `4` | chunks retrieved per question | | `BOR_TOP_N_DOCS` | `2` | full documents fed to the LLM | | `BOR_RELEVANCE_THRESHOLD` | `0.30` | best cosine similarity required to answer; below β‡’ honest deflection | | `BOR_MAX_CONTEXT_CHARS` | `24000` | cap on total document text sent to the LLM | | `BOR_SUGGESTIONS` | built-in list | JSON list of onboarding chips | | `DEBUGPY` | `0` | `1` β‡’ attach-on-demand debugpy on `DEBUGPY_PORT` (default 5678) | | `BOR_LOG_LEVEL` | `INFO` | app log level | ## Troubleshooting - **`401` from aipi** β€” set `BOR_LLM_API_KEY` (or `$AIPI_KEY`). - **Embedding dimension mismatch** β€” aipi changed models; run `uv run python -m scripts.llm_probe`, update `BOR_EMBEDDING_DIM`, then drop + recreate the chunks table (new migration or manual `TRUNCATE chunks, documents`). - **Answers deflect too often / too rarely** β€” tune `BOR_RELEVANCE_THRESHOLD` (lower = answers more, higher = more honest deflection). Check `query_log` for the actual scores: `psql … -c 'SELECT question, top_score, deflected FROM query_log ORDER BY created_at DESC LIMIT 20'` - **KB offline banner in the chat** β€” Postgres isn't running: `podman compose up -d db`. - **Stuck "Thinking…"** β€” the LLM is slow or down; a 120s client timeout turns it into an error banner automatically.