Files
brain-of-reese/README.md
T

317 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 🧠 Brain of Reese
A chippy, honest **RAG chatbot** over the `~/Homelab` and `~/Deployments`
projects. Point it at your notes — markdown, YAML, JSON, Python, plain
text — ask it anything, and it retrieves the relevant chunks with
**hybrid search** (pgvector cosine ∪ Postgres full-text search, fused with
Reciprocal Rank Fusion), feeds the **whole relevant document** to a
**self-hosted LLM** (`turbo` via `https://aipi.reeseapps.com/v1`), and
streams a grounded answer back.
If it doesn't have notes for your question, it admits it:
*"I haven't done anything like that"* — plus suggestions for what it **does** know.
> **Updated your notes?** The knowledge base is refreshed by re-running the
> import — it's idempotent and only re-embeds what changed:
> ```bash
> uv run python -m scripts.import_docs --prune
> ```
> Details in [Updating the documents](#updating-the-documents).
- **Stack:** FastAPI · Pydantic v2 · SQLAlchemy 2 · Alembic · pgvector ·
vanilla HTML/CSS/JS (no CDN) · Playwright E2E
- **Planning:** architecture, LOCKED decisions and the phase roadmap live
in [`.agent/PLAN.md`](.agent/PLAN.md); per-story specs in
[`.agent/user_stories/`](.agent/user_stories/).
---
## Development Setup
### Prerequisites
- [uv](https://docs.astral.sh/uv/)
- [Podman](https://podman.io/) (with the `podman compose` provider)
- Node.js is **not** needed locally (asset minification happens in the
container build only)
### 1. Install dependencies
```bash
uv sync
```
### 2. Configure
```bash
cp .env.example .env
# edit .env — the defaults already match the local compose setup.
# BOR_LLM_API_KEY: your aipi key (falls back to $AIPI_KEY if unset)
```
### 3. Start the database (Postgres 17 + pgvector)
```bash
podman compose up -d db
podman compose ps # wait until "healthy"
```
### 4. Apply migrations
```bash
uv run alembic upgrade head
```
### 5. Import your knowledge base
```bash
uv run python -m scripts.llm_probe # sanity: models + 768-dim check
uv run python -m scripts.import_docs # defaults: ~/Homelab + ~/Deployments
```
### 6. Run the app
```bash
uv run uvicorn app.main:app --reload
# → http://localhost:8000 (chat) http://localhost:8000/sources.html (KB)
```
> 📝 **After this, day-to-day is just: edit markdown → re-run the import.**
> See [Updating the documents](#updating-the-documents) below.
## Using the UI
- **Chat** (`/`) — ask questions; answers stream in with **source chips**
that cite the exact documents used. Clicking a chip opens that document
**in a new tab**.
- **Document viewer** (`/document.html?source=…&path=…`) — the full text of
any indexed document, served from the database (no filesystem access):
markdown is rendered, every other format (`yaml`, `json`, `py`, `txt`, …)
is shown as escaped monospace text. Unknown documents get a designed
not-found state with a link back to the index.
- **Sources** (`/sources.html`) — the indexed document list; the *Path*
column links each document to the viewer in a new tab.
## Updating the documents
**This is the workflow you'll use most.** The knowledge base is refreshed by
**re-running the import**. It is idempotent and delta-based (sha256 per
file), so a refresh after a normal editing session takes seconds:
```bash
# After editing/adding/removing notes in your projects:
uv run python -m scripts.import_docs # re-index what changed
uv run python -m scripts.import_docs --prune # also drop deleted/out-of-scope files
# Point it at extra directories (repeatable):
uv run python -m scripts.import_docs --source ~/SomeOtherDocs
```
Then check the **Sources** page (`http://localhost:8000/sources.html`):
the *documents* / *chunks* counters and *last indexed* timestamp should
reflect the new files, and each document row shows when it was last
embedded.
- The import prints one line per file (`import: added|updated|unchanged|
pruned …`) and ends with a greppable summary (`import: summary files=…
added=… updated=… unchanged=… pruned=… chunks=… embed_batches=…
formats=md:203,yaml:267,…`), so it is safe to run from a cron job or
after every commit.
- Indexed formats (A9): **`md, markdown, txt, yaml, yml, json, py`**
(case-insensitive; narrow with `BOR_IMPORT_EXTENSIONS`). Any path with a
**dot-prefixed component** — hidden files or vendored caches like
`.esphome/.espressif/**` — is skipped, along with `.venv`,
`node_modules`, `.git`, `__pycache__`, `.pytest_cache`, `dist`, `build`.
`--prune` also drops documents whose files no longer match the filter —
that's how previously imported junk leaves the index.
- Non-markdown files get format-aware chunking (YAML top-level keys /
`---` docs, JSON top-level keys, Python top-level defs/classes via
stdlib `ast`) and their title comes from the file stem.
- Unchanged files are **not re-embedded** — only new/changed ones, so
refreshes are cheap.
- To sanity-check the LLM backend (models + embedding dimension) after any
aipi change: `uv run python -m scripts.llm_probe`.
## Checking retrieval quality
Ask the *real* pipeline (live aipi embeddings + the current KB) whether a
question lands on the right document, with the gate verdict and per-document
cosine / FTS / fused scores:
```bash
uv run python -m scripts.eval_retrieval "How did I install gitlab?"
uv run python -m scripts.eval_retrieval --from-file questions.txt --top 8
```
Requires `AIPI_KEY` in the environment (same convention as
`scripts/llm_probe.py`) and an imported knowledge base.
## How retrieval works (hybrid)
Every question is embedded and also lexically tokenized (OR-joined, English
stemming) and searched **twice** against Postgres:
1. **Vector** — pgvector cosine top-N (default `BOR_HYBRID_VECTOR_CANDIDATES=100`)
2. **Lexical** — a stored `tsvector` (GIN-indexed) matched with `to_tsquery`,
top-N by `ts_rank` (default `BOR_HYBRID_LEXICAL_CANDIDATES=30`)
The two ranked lists are fused with **Reciprocal Rank Fusion**
(`score = Σ 1/(k + rank)`, `BOR_RRF_K=60`) — a chunk in both lists scores
nearly double, which is what lets a name-your-tool question ("gitlab") find
its own document even when the question embeds close to generic templates.
The **honesty gate** (A8) then answers (HIGH) when the best cosine is ≥
`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **or** at least one chunk matched
lexically (`fts_hits > 0`) — it deflects (LOW) only when *both* signals are
absent. The top `BOR_TOP_N_DOCS` full documents are still what the LLM sees.
`query_log` records every turn (`top_score` = best cosine, `fts_hits`,
`chunk_hits`, `deflected`, `sources`, `latency_ms`) — the raw material for
tuning: `psql … -c 'SELECT question, top_score, fts_hits, deflected FROM query_log ORDER BY created_at DESC LIMIT 20'`.
## Debugging
`debugpy` is **off by default** and *never imported* unless you opt in —
zero overhead in normal runs.
```bash
DEBUGPY=1 uv run uvicorn app.main:app
# → log line: debugpy: remote debugging ENABLED, listening on 0.0.0.0:5678
```
Then attach from VS Code (`.vscode/launch.json`):
```json
{
"name": "Attach to Brain of Reese",
"type": "debugpy",
"request": "attach",
"connect": { "host": "localhost", "port": 5678 },
"pathMappings": [
{ "localRoot": "${workspaceFolder}", "remoteRoot": "/app" }
]
}
```
The port is non-blocking and attach-on-demand: the app keeps running
normally until you attach. Override the port with `DEBUGPY_PORT`.
## QA / Testing Environment
Three layers — the project rule is **one story, one phase, one Playwright
suite** (see `AGENTS.md`):
```bash
# Unit + integration (FastAPI TestClient)
uv run pytest
# Same, with the coverage gate (phases require >90% on app/)
uv run pytest --cov=app --cov-report=term-missing
# Lint + static types
uv run ruff check .
uv run pyright
# Playwright E2E — install the browser once:
uv run playwright install chromium
# Each story's E2E runs IN ISOLATION (DB must be up):
podman compose up -d db
uv run pytest tests/e2e/test_import_documents.py -v --no-cov
uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
# ...one file per story in .agent/user_stories/ (see .agent/phases/todo/)
```
**Deterministic E2E:** by default the E2E app talks to a local **mock
aipi** (`tests/e2e/mock_llm.py`) whose embeddings are real
token-overlap vectors — so the cosine relevance threshold behaves like
production (on-topic questions answer, off-topic ones deflect).
To run E2E against the **live** self-hosted models instead:
```bash
E2E_REAL_LLM=1 uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
```
(requires a real import of your docs first).
## Production Deployment
Build the multi-stage image (frontend minified by esbuild in the builder
stage, deps installed by `uv`, non-root runtime):
```bash
podman build -t brain-of-reese/app:latest .
```
Run standalone (bring your own Postgres + pgvector):
```bash
podman run -d --name brain-of-reese \
-p 8000:8000 \
-e BOR_DATABASE_URL=postgresql+psycopg://reese:SECRETPASSWORD@dbhost:5432/brain_of_reese \
-e BOR_LLM_BASE_URL=https://aipi.reeseapps.com/v1 \
-e BOR_LLM_API_KEY=$AIPI_KEY \
brain-of-reese/app:latest
```
The entrypoint runs `alembic upgrade head` automatically on start.
Or run the whole stack from compose (app + db):
```bash
podman compose --profile prod up -d --build
```
Production hardening notes: app runs as non-root (uid 10001), slim image,
healthcheck on `/api/health`, debugpy off unless `DEBUGPY=1`, all assets
served locally (no CDN), `BOR_ENVIRONMENT=production`.
## Configuration reference
| Env | Default | Meaning |
|-----|---------|---------|
| `BOR_DATABASE_URL` | local compose URL | SQLAlchemy URL (psycopg) |
| `BOR_LLM_BASE_URL` | `https://aipi.reeseapps.com/v1` | OpenAI-compatible endpoint |
| `BOR_LLM_API_KEY` | — (falls back to `$AIPI_KEY`) | aipi API key |
| `BOR_LLM_CHAT_MODEL` | `turbo` | chat model |
| `BOR_LLM_EMBED_MODEL` | `embed` | embedding model |
| `BOR_EMBEDDING_DIM` | `768` | vector dimension (fixed at table creation) |
| `BOR_TOP_N_DOCS` | `2` | full documents fed to the LLM |
| `BOR_RELEVANCE_THRESHOLD` | `0.62` | answer when best cosine ≥ this **or** an FTS hit; below + no FTS ⇒ honest deflection |
| `BOR_HYBRID_VECTOR_CANDIDATES` | `100` | cosine list width for the RRF fusion |
| `BOR_HYBRID_LEXICAL_CANDIDATES` | `30` | FTS list width for the RRF fusion |
| `BOR_RRF_K` | `60` | RRF damping constant (`1/(k + rank)`) |
| `BOR_IMPORT_EXTENSIONS` | `md,markdown,txt,yaml,yml,json,py` | csv of importable formats (may only narrow the A9 set) |
| `BOR_MAX_CONTEXT_CHARS` | `24000` | cap on total document text sent to the LLM |
| `BOR_SUGGESTIONS` | built-in list | JSON list of onboarding chips |
| `DEBUGPY` | `0` | `1` ⇒ attach-on-demand debugpy on `DEBUGPY_PORT` (default 5678) |
| `BOR_LOG_LEVEL` | `INFO` | app log level |
## Troubleshooting
- **`401` from aipi** — set `BOR_LLM_API_KEY` (or `$AIPI_KEY`).
- **`litellm.UnsupportedParamsError … encoding_format` from aipi** — the
aipi proxy (litellm `openai_like`) rejects the `encoding_format` parameter
that the `openai` SDK injects into every embeddings request. The app
already works around this by POSTing a minimal `{model, input}` payload
through the openai client's own httpx transport (`app/rag/llm.py` →
`LLMClient._embed_batch`). If you see this, you are likely calling the
endpoint with a different client — drop the parameter (or set
`litellm.drop_params = True` on the proxy).
- **Embedding dimension mismatch** — aipi changed models; run
`uv run python -m scripts.llm_probe`, update `BOR_EMBEDDING_DIM`, then
drop + recreate the chunks table (new migration or manual `TRUNCATE
chunks, documents`).
- **Honest deflection (the amber “I haven't done anything like that”
bubble)** — every question passes the honesty gate: deflection happens
only when the best cosine similarity is below
`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **and** no chunk matched the
question lexically (`fts_hits = 0`). A weak cosine with a lexical hit
(name-your-tool questions) still gets a grounded answer. When it does
deflect, the LLM prompt carries weak-hit *titles only* (no document
content), the reply opens with “I haven't done anything like that”, the
bubble renders amber with “Maybe try” chips derived from the closest
indexed titles, the SSE `done` event carries `deflected: true` +
`suggestions[]`, and the `query_log` row records `deflected=true` + the
weak `top_score` + `fts_hits`. This is a feature, not a bug — the KB
simply has no notes that close; the chips always point at topics Brain
really covers.
- **Answers deflect too often / too rarely** — tune
`BOR_RELEVANCE_THRESHOLD` (lower = answers more, higher = more honest
deflection): `0.0` ⇒ the gate leans entirely on FTS hits; `1.0` ⇒
everything deflects unless a chunk matches lexically. The `embed` model's
cosines cluster in a ~0.6–0.85 band on the live KB, so the default is
`0.62`; after changing it, check the real scores:
`psql … -c 'SELECT question, top_score, fts_hits, deflected FROM query_log ORDER BY created_at DESC LIMIT 20'`
- **KB offline banner in the chat** — Postgres isn't running:
`podman compose up -d db`.
- **Stuck "Thinking…"** — the LLM is slow or down; a 120s client timeout
turns it into an error banner automatically.