Phase 103 final verification pass — all green.
**Verified (all 3 tasks already in `complete/`; no code changes needed):**
- `opening_questions` in `app/api/suggestions.py` — forward walk, one opener per chat (first non-blank user msg, A3), reads raw `messages` not `title` (A4), phase-80 order/dedup/cap/seed contracts; `last_questions` name gone from `app/`+`tests/`
- Docs updated: `app/config.py` seed docstring, `.env.example` `BOR_SUGGESTIONS`, `README.md` — "session openers" wording
- Diff scope correct: only the 6 expected files + phase-file moves; `app/rag/suggestions.py` and `frontend/` untouched
**Test / lint / coverage results:**
- `uv run pytest tests/integration/test_suggestions_api.py -v` → 12 passed
- `uv run pytest tests/e2e/test_suggestion_chips.py -v --no-cov` → 8 passed in isolation (opener-only core pin included)
- `test_responsive_polish.py` → 7 passed; `test_chat_persistence.py` → 4 passed (both isolated, no edits)
- `uv run pytest --cov=app --cov-report=term-missing` → 2086 passed, TOTAL 99% (>90%); `app/api/suggestions.py` 100%
- `uv run ruff check .` → clean; `uv run pyright` → 0 errors, 0 warnings
**Completion criteria:** all 7 ✅ (follow-ups-never-surface pin; cap-across-chats pin; seed/dedup/case/partial/A3/401 pins; E2E suites isolated; deflection chips unchanged; full suite + lint; commit + dir move left to harness per executor rules).
**Deviations:** none — no defects found; nothing changed in this pass.
**Next pending phase:** `98_sync_summary_visibility` (numeric order in `todo/`).
607 lines
24 KiB
Markdown
607 lines
24 KiB
Markdown
# 🧠 Brain of Reese
|
||
|
||
A chippy, honest **RAG chatbot** over your `~/Homelab` and `~/Deployments`
|
||
projects. Point it at your notes — markdown, YAML, JSON, Python, plain
|
||
text — ask it anything, and it retrieves the relevant chunks with
|
||
**hybrid search** (pgvector cosine ∪ Postgres full-text search, fused with
|
||
Reciprocal Rank Fusion), feeds the **whole relevant document** to a
|
||
**self-hosted LLM** (`turbo` via `https://aipi.reeseapps.com/v1`), and
|
||
streams a grounded answer back.
|
||
|
||
If it doesn't have notes for your question, it admits it: *"I haven't done
|
||
anything like that"* — plus suggestions for what it **does** know.
|
||
|
||
> **Updated your notes?** Re-run the import — it's idempotent and only
|
||
> re-embeds what changed:
|
||
> ```bash
|
||
> uv run python -m scripts.import_docs --prune
|
||
> ```
|
||
> See [Updating the documents](#updating-the-documents) for details.
|
||
|
||
---
|
||
|
||
## Quick Start
|
||
|
||
```bash
|
||
# 1. Dependencies
|
||
uv sync
|
||
|
||
# 2. Configure (copy and edit)
|
||
cp .env.example .env
|
||
# Set BOR_LLM_API_KEY — your aipi key (falls back to $AIPI_KEY)
|
||
|
||
# 3. Start the database
|
||
podman compose up -d db
|
||
podman compose ps # wait until "healthy"
|
||
|
||
# 4. Apply migrations & import
|
||
uv run alembic upgrade head
|
||
uv run python -m scripts.import_docs
|
||
|
||
# 5. Run the app
|
||
uv run uvicorn app.main:app --reload
|
||
# → http://localhost:8000 (chat)
|
||
```
|
||
|
||
**Stack:** FastAPI · Pydantic v2 · SQLAlchemy 2 · Alembic · pgvector ·
|
||
vanilla HTML/CSS/JS (no CDN) · Playwright E2E
|
||
|
||
---
|
||
|
||
## Pages at a glance
|
||
|
||
The app is a single-page application — all views load from `/` with a
|
||
client-side router. Direct URLs deep-link to the matching view.
|
||
|
||
| Page | URL | Who can see it |
|
||
|------|-----|----------------|
|
||
| **Chat** | `/` | Everyone (token gate for non-shared) |
|
||
| **RAG / Knowledge base** | `/sources.html` | Admin only |
|
||
| **Git sources** | `/git-sources.html` | Admin only |
|
||
| **Tuning** | `/tuning.html` | Admin only |
|
||
| **Saved chats** | `/history.html` | Admin only |
|
||
| **Access tokens** | `/tokens.html` | Admin only |
|
||
| **Document viewer** | `/document.html?source=X&path=Y` | Token users & admins |
|
||
| **Login** | `/login.html` | Everyone |
|
||
| **Shared chat** | `/shared/<token>` | Everyone (anonymous) |
|
||
| **Edit doc** | `/doc-edit.html` | Admin only |
|
||
|
||
### Chat (`/`)
|
||
|
||
Ask questions; answers stream in with **source chips** that cite the exact
|
||
documents used. Clicking a chip opens the document in an **almost-fullscreen
|
||
modal** on the same page (no new tab). The onboarding suggestion chips
|
||
follow the opening questions of the 3 most recent saved chats (each
|
||
chat's first question — a follow-up never chips, phase 103) — on a
|
||
fresh deployment they seed from `BOR_SUGGESTIONS`.
|
||
|
||
### RAG / Knowledge base (`/sources.html`)
|
||
|
||
The indexed document list. The *Path* column opens each document in the same
|
||
**almost-fullscreen modal**. Admin-only — anonymous visitors see a sign-in
|
||
gate instead.
|
||
|
||
### Git sources (`/git-sources.html`)
|
||
|
||
The admin-managed source registry: add or remove git repositories, upload
|
||
source archives (`.tar`, `.tar.gz`, `.tgz`, `.zip`), and register local
|
||
directories — all in one table with a `kind` discriminator. No `.env`
|
||
editing, no restart. The **Sync sources** button clones/pulls/walks every
|
||
source and re-imports in one click.
|
||
|
||
### Tuning (`/tuning.html`)
|
||
|
||
Set global instructions that are injected into the system prompt of every
|
||
chat turn. Add notes like *"be more concise"* or *"assume I'm on NixOS"*.
|
||
List, edit, or remove them here — no conversation required.
|
||
|
||
### Saved chats (`/history.html`)
|
||
|
||
Every conversation is saved automatically. Click a title to return to that
|
||
chat. Admin-only view.
|
||
|
||
### Access tokens (`/tokens.html`)
|
||
|
||
Generate tokens and hand them out — a token opens chat, the answers, and
|
||
the documents they cite. Shared chats stay open to everyone.
|
||
|
||
### Shared chat (`/shared/<token>`)
|
||
|
||
A read-only snapshot of any conversation, shareable via a link. Open to
|
||
everyone — no login or token required.
|
||
|
||
### Edit doc (`/doc-edit.html`)
|
||
|
||
Review and adjust an AI-generated answer before committing it to the docs
|
||
repository. Admin-only flow page.
|
||
|
||
---
|
||
|
||
## Admin & Sign-in
|
||
|
||
Brain of Reese has exactly **one account: the admin (you)**. Signing in
|
||
unlocks the **full Sources catalog**, the **answer-tuning** controls, and
|
||
**token management**. **Shared chats** are the only content that stays open
|
||
to anonymous visitors.
|
||
|
||
### Setup (one-time)
|
||
|
||
```bash
|
||
python -c 'import secrets;print(secrets.token_hex(32))' # → paste into .env
|
||
```
|
||
|
||
```env
|
||
BOR_ADMIN_PASSWORD=your-password # plaintext — homelab scope, by design
|
||
BOR_SESSION_SECRET=<the hex from above> # signs the session cookie
|
||
```
|
||
|
||
**Fail-loud:** while either variable is empty the app refuses to start,
|
||
naming the missing one(s).
|
||
|
||
### How authentication works
|
||
|
||
| Endpoint | Purpose |
|
||
|----------|---------|
|
||
| `POST /api/login` | Password → signed `bor_session` cookie |
|
||
| `POST /api/logout` | Clears session cookie |
|
||
| `GET /api/whoami` | `{"authenticated": bool, "role": "admin"|"user"|"anonymous"}` |
|
||
| `POST /api/token-auth` | Token (`bor_…`) → same cookie, role `user` |
|
||
|
||
The cookie uses `same_site="lax"`, `https_only` off — **no HTTPS
|
||
enforcement** (homelab HTTP; the cookie is single-admin convenience, not a
|
||
cloud boundary). Max age is `BOR_SESSION_MAX_AGE` (default `43200` = 12 h).
|
||
|
||
### Access tokens
|
||
|
||
The admin can hand out access without sharing the admin password.
|
||
|
||
1. **Generate:** sign in → **Tokens** in the navbar. Type a label and hit
|
||
**Generate** — a `bor_` + 32-hex token appears in the *shown once* block.
|
||
Copy it now: only its SHA-256 hash is stored.
|
||
2. **Use:** the holder enters the token at the in-app gate. The browser
|
||
caches it in `localStorage["bor.token"]` for silent re-auth on reloads.
|
||
3. **Revoke:** click **Revoke** on the token's row. Revocation is immediate.
|
||
|
||
A token user can chat and open documents — nothing else. The Sources
|
||
catalog, git sources, tuning, and history stay admin-only.
|
||
|
||
---
|
||
|
||
## Using the App
|
||
|
||
### Updating the documents
|
||
|
||
The knowledge base is refreshed by **re-running the import**. It is
|
||
idempotent and delta-based (sha256 per file), so a refresh after a normal
|
||
editing session takes seconds:
|
||
|
||
```bash
|
||
# After editing/adding/removing notes:
|
||
uv run python -m scripts.import_docs # re-index what changed
|
||
uv run python -m scripts.import_docs --prune # also drop deleted files
|
||
|
||
# Point it at extra directories (repeatable):
|
||
uv run python -m scripts.import_docs --source ~/SomeOtherDocs
|
||
```
|
||
|
||
With **git-based sources** each run first pulls the latest commits of your
|
||
repos, so this same command is the whole update loop: *commit → re-run*.
|
||
|
||
Then check the **Sources** page (`http://localhost:8000/sources.html`):
|
||
the *documents* / *chunks* counters and *last indexed* timestamp should
|
||
reflect the new files.
|
||
|
||
- The import prints one line per file and ends with a greppable summary,
|
||
so it is safe to run from cron or after every commit.
|
||
- Indexed formats: **md, markdown, txt, yaml, yml, json, py**, the Podman
|
||
quadlet family (**container, network, volume, image, pod, kube, swap,
|
||
os, endpoint**), and **j2** Jinja templates (case-insensitive).
|
||
- Hidden files and common cache dirs (`.venv`, `node_modules`, `.git`,
|
||
`__pycache__`, etc.) are skipped.
|
||
- Non-markdown files get format-aware chunking (YAML keys, JSON keys,
|
||
Python classes/defs).
|
||
|
||
### Git-based sources
|
||
|
||
Rather than pointing the import at local folders, point it at
|
||
**git repositories**. The admin **Git sources page** (`/git-sources.html`)
|
||
is the primary management surface — add or remove repositories there and
|
||
the list is stored in Postgres.
|
||
|
||
```env
|
||
# .env — the fallback list (fresh setups, or until the admin page
|
||
# stores a source; the page becomes the source of truth)
|
||
BOR_GIT_SOURCES=https://git.reeseapps.com/reese/homelab.git,git@github.com:reese/deployments.git
|
||
BOR_SOURCES_DIR=~/bor-sources # default; each repo lands in <dir>/<repo-name>/
|
||
```
|
||
|
||
- Auth is whatever the machine supplies — HTTPS via the OS credential
|
||
helper, or SSH via your key; no credentials are stored in the app.
|
||
- Every run **clones** (first time, shallow `--depth 1`) or **pulls**
|
||
(`git pull --ff-only`) each repo, then indexes the checkout.
|
||
- A failed sync aborts the run — no partial junk.
|
||
|
||
### Archive upload sources
|
||
|
||
Upload a `.tar`, `.tar.gz`, `.tgz`, or `.zip` archive to make it a source.
|
||
The form on the Git sources page — and the `POST /api/git-sources/upload`
|
||
route behind it — accepts archives and unpacks + scans them immediately.
|
||
|
||
- Re-uploading the same filename replaces the source **in place** (no
|
||
second folder, no duplicate row).
|
||
- Zip-bomb guard: absolute member paths, `..` traversal, and symlinks
|
||
escaping the unpack folder are rejected.
|
||
- Archives unpack under `BOR_UPLOAD_DIR` (default `~/bor-sources/uploads`).
|
||
- `BOR_UPLOAD_MAX_MB` (default 512) caps both compressed and extracted size.
|
||
|
||
### Local directory sources
|
||
|
||
A plain directory can be a first-class source too. It shares the git
|
||
sources' **one table** (the `git_sources` registry with a `kind`
|
||
discriminator: `git` | `local`), **one admin page**, and **one Sync button**.
|
||
|
||
- Register via API: `POST /api/git-sources` with `{"kind": "local", "path": …}`.
|
||
- Sync walks it directly — no clone, no checkout copy.
|
||
- The directory is re-verified to exist at sync time; a missing directory
|
||
fails the run loudly.
|
||
|
||
### Sync from the UI
|
||
|
||
The **Sync sources** button on the **Sources** page (admin only) runs the
|
||
whole git-source refresh in one click:
|
||
|
||
1. **clone/pull + walk** every configured source — git repos and local
|
||
directories, including uploaded archives.
|
||
2. **re-import with prune** — files deleted upstream leave the index.
|
||
3. **regenerate the KB overview** (the `<knowledge_base>` outline every
|
||
chat turn injects) — only when the import changed the knowledge base.
|
||
|
||
- States: the button shows **Syncing…** with a spinner while polling
|
||
`GET /api/sync/status` every 2 s. No client-side timeout — a clone +
|
||
embed can legitimately take minutes.
|
||
- One sync at a time: a second trigger while a run is in flight gets a
|
||
`409`.
|
||
|
||
---
|
||
|
||
## Thinking
|
||
|
||
The self-hosted `turbo` model reasons before it answers. That reasoning is
|
||
streamed with the turn as `thinking` SSE events and shown in a
|
||
**collapsible "Thinking" block** above the answer bubble: it opens and
|
||
fills in live while the model thinks, tucks itself away the moment the
|
||
first answer token lands, and stays click-toggleable afterwards.
|
||
|
||
To hide it, set `BOR_STREAM_THINKING=0`.
|
||
|
||
---
|
||
|
||
## Agent document tools (ls + read + grep)
|
||
|
||
Retrieval only puts the top documents in context. When an answer depends on
|
||
a file a note *references* ("the exact JSON shape is in
|
||
`example-record-file.json`"), the model can extend its own context with
|
||
three server-side tools — on **grounded** (high-relevance) turns only:
|
||
|
||
| Tool | Description |
|
||
|------|-------------|
|
||
| **`ls`** | Lists every indexed document (`source: X \| path: Y \| title: Z`); pass a source name as `path` to list one source |
|
||
| **`read(path)`** | Appends the **full** text of one indexed document to context (never truncated); `path` is the combined `source/path` string |
|
||
| **`grep(pattern, path?)`** | Searches indexed documents for an exact string (case-insensitive fixed substring); returns up to 20 `source/path:line: text` matches; an optional `path` limits to one document |
|
||
|
||
Each call is executed against Postgres only (no extra LLM round trip) and
|
||
streamed as an SSE `tool` frame. In the chat, each call shows a transient
|
||
**calling-tool status** alongside "thinking" (the send button keeps its
|
||
busy state — "Stop" — for the whole turn).
|
||
|
||
The model may call tools as many times as needed, bounded by a round cap:
|
||
|
||
| Env | Default | Meaning |
|
||
|-----|---------|---------|
|
||
| `BOR_AGENT_MAX_ROUNDS` | `10` | hard cap on agent tool rounds per grounded turn (0 = no tools, the kill switch) |
|
||
|
||
---
|
||
|
||
## Tuning your answers
|
||
|
||
*Admin-only* — sign in first. If an answer isn't quite right — too chatty,
|
||
wrong assumption, missing context — **tune** Brain right there:
|
||
|
||
1. Press **"Tune"** in the meta row under any completed answer (deflected
|
||
ones included).
|
||
2. Type a short instruction (1–2000 chars), e.g. *"be more concise"* or
|
||
*"assume I'm on NixOS"*, and **Save**.
|
||
|
||
The note is stored in Postgres (`steering_notes`) and read into the
|
||
**system prompt of every subsequent chat turn** as a `<tuning>` section
|
||
(numbered, oldest first, capped at `BOR_STEERING_MAX_CHARS` chars — default
|
||
8000, overflow marked `[…truncated…]`).
|
||
|
||
List or remove notes from the **"Tuning"** button in the chat header (count
|
||
badge, newest-first, per-note delete). The API is stateless JSON if you
|
||
prefer curl:
|
||
|
||
```bash
|
||
curl -s localhost:8000/api/steering # list
|
||
curl -s -X POST localhost:8000/api/steering \
|
||
-H 'Content-Type: application/json' -d '{"note": "be more concise"}'
|
||
curl -s -X DELETE localhost:8000/api/steering/<note-id> # remove
|
||
```
|
||
|
||
---
|
||
|
||
## How retrieval works (hybrid)
|
||
|
||
Every question is embedded and also lexically tokenized (OR-joined, English
|
||
stemming) and searched **twice** against Postgres:
|
||
|
||
1. **Vector** — pgvector cosine top-N (default `BOR_HYBRID_VECTOR_CANDIDATES=100`)
|
||
2. **Lexical** — a stored `tsvector` (GIN-indexed) matched with `to_tsquery`,
|
||
top-N by `ts_rank` (default `BOR_HYBRID_LEXICAL_CANDIDATES=30`)
|
||
|
||
The two ranked lists are fused with **Reciprocal Rank Fusion**
|
||
(`score = Σ 1/(k + rank)`, `BOR_RRF_K=60`) — a chunk in both lists scores
|
||
nearly double, which lets a name-your-tool question ("gitlab") find its own
|
||
document even when the question embeds close to generic templates.
|
||
|
||
The **honesty gate** (A8) then answers (HIGH) when the best cosine is ≥
|
||
`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **or** at least one chunk matched
|
||
lexically (`fts_hits > 0`) — it deflects (LOW) only when *both* signals are
|
||
absent. The top `BOR_TOP_N_DOCS` full documents are what the LLM sees.
|
||
|
||
---
|
||
|
||
## Document summaries (non-markdown)
|
||
|
||
Raw yaml/json/py/txt embeds badly — flags and keys are not language, so
|
||
retrieval can miss exactly the documents that are all configuration. At
|
||
import time, every **non-markdown** A9 document is summarized by the aipi
|
||
`lite` model (`BOR_LLM_SUMMARY_MODEL`, default `lite`):
|
||
|
||
* The summary is stored on `documents.summary` **and** indexed as one extra
|
||
embedded chunk (`chunks.is_summary`, position −1), so hybrid search has a
|
||
natural-language target to hit instead of the raw text.
|
||
* The last line is a **code-deterministic** pointer — `Source: <source>/<path>`
|
||
— appended by the app, never model-generated.
|
||
* The model only sees the first `BOR_SUMMARY_MAX_CHARS` (default 12000)
|
||
characters; overflow is cut and marked with `[…truncated…]`.
|
||
|
||
Summary generation is best-effort: if `lite` fails, the document is still
|
||
indexed (without a summary), the failure is logged, and counted in the
|
||
import summary line (`summaries=N summary_errors=N`).
|
||
|
||
---
|
||
|
||
## Caching & deploys
|
||
|
||
A deploy is a commit — and the browser must see it **without a hard refresh**.
|
||
One Starlette middleware (`app/core/caching.py`) applies the rule at the
|
||
transport layer:
|
||
|
||
- **HTML pages** ship `Cache-Control: no-cache`, **no `etag`**, **no
|
||
`last-modified`** — each visit always gets a fresh 200 body.
|
||
- **Assets are versioned and cached for a year.** Pages reference their
|
||
CSS/JS with a token (`/assets/styles.css?v=<token>`), and every
|
||
`/assets/*` response ships `Cache-Control: public, max-age=31536000, immutable`.
|
||
- **The token is the deploy.** In a git checkout it is the short SHA of
|
||
`HEAD` — so every commit/deploy flips the token and the versioned asset
|
||
URLs change with it. A checkout without `.git` falls back to a stable
|
||
content hash of the `frontend/` tree.
|
||
|
||
No CDN, no new services, no build-step change: the middleware rewrites the
|
||
asset references of the known pages in flight.
|
||
|
||
> **Deploy note:** the very first deploy onto this scheme needs one normal
|
||
> page visit, so the browser revalidates the HTML once and starts requesting
|
||
> the versioned assets; every commit after that is picked up automatically.
|
||
|
||
---
|
||
|
||
## Configuration reference
|
||
|
||
| Env | Default | Meaning |
|
||
|-----|---------|---------|
|
||
| `BOR_APP_NAME` | `Brain of Reese` | Display name everywhere (page titles, header brand, status labels) |
|
||
| `BOR_INPUT_PLACEHOLDER` | `Ask me anything…` | Chat composer placeholder |
|
||
| `BOR_FOOTER_TEXT` | `Powered by self-hosted models` | Footer line on every page |
|
||
| `BOR_DATABASE_URL` | local compose URL | SQLAlchemy URL (psycopg) |
|
||
| `BOR_LLM_BASE_URL` | `https://aipi.reeseapps.com/v1` | OpenAI-compatible endpoint |
|
||
| `BOR_LLM_API_KEY` | — (falls back to `$AIPI_KEY`) | aipi API key |
|
||
| `BOR_LLM_CHAT_MODEL` | `turbo` | Chat model |
|
||
| `BOR_LLM_EMBED_MODEL` | `embed` | Embedding model |
|
||
| `BOR_LLM_SUMMARY_MODEL` | `lite` | One-shot completions: document summaries at import, KB overview |
|
||
| `BOR_EMBEDDING_DIM` | `768` | Vector dimension (fixed at table creation) |
|
||
| `BOR_TOP_N_DOCS` | `2` | Full documents fed to the LLM |
|
||
| `BOR_RELEVANCE_THRESHOLD` | `0.62` | Answer when best cosine ≥ this **or** an FTS hit; below + no FTS ⇒ honest deflection |
|
||
| `BOR_HYBRID_VECTOR_CANDIDATES` | `100` | Cosine list width for the RRF fusion |
|
||
| `BOR_HYBRID_LEXICAL_CANDIDATES` | `30` | FTS list width for the RRF fusion |
|
||
| `BOR_RRF_K` | `60` | RRF damping constant (`1/(k + rank)`) |
|
||
| `BOR_AGENT_MAX_ROUNDS` | `10` | Hard cap on agent tool rounds per grounded turn (0 = no tools) |
|
||
| `BOR_IMPORT_EXTENSIONS` | csv (see below) | Importable formats (may only narrow the A9 set) |
|
||
| `BOR_GIT_SOURCES` | — (empty) | CSV of git repo URLs — **fallback while the admin Git sources page's list is empty** |
|
||
| `BOR_SOURCES_DIR` | `~/bor-sources` | Where git repos are cloned/pulled |
|
||
| `BOR_UPLOAD_DIR` | `~/bor-sources/uploads` | Where uploaded source archives are unpacked |
|
||
| `BOR_UPLOAD_MAX_MB` | `512` | Cap (MiB) for uploaded source archives (zip-bomb guard) |
|
||
| `BOR_STEERING_MAX_CHARS` | `8000` | Char budget for the `<tuning>` prompt section |
|
||
| `BOR_SUMMARY_MAX_CHARS` | `12000` | Cap on document content sent to the `lite` summary model |
|
||
| `BOR_KB_OVERVIEW_MAX_CHARS` | `4000` | Char budget for the `<knowledge_base>` prompt section |
|
||
| `BOR_OVERVIEW_INPUT_MAX_CHARS` | `40000` | Cap on the document list sent to `lite` for KB overview generation |
|
||
| `BOR_SUGGESTIONS` | built-in list | JSON seed for onboarding chips (shown before the first saved question) |
|
||
| `BOR_ADMIN_PASSWORD` | *(required)* | The single admin's password — app refuses to start when empty |
|
||
| `BOR_SESSION_SECRET` | *(required)* | Signing key for the `bor_session` cookie |
|
||
| `BOR_SESSION_MAX_AGE` | `43200` | Session-cookie lifetime in seconds (12 h, sliding) |
|
||
| `DEBUGPY` | `0` | `1` ⇒ attach-on-demand debugpy on `DEBUGPY_PORT` (default 5678) |
|
||
| `BOR_LOG_LEVEL` | `INFO` | App log level |
|
||
|
||
### Customizing the look
|
||
|
||
Every identity string is an env var: `BOR_APP_NAME` (display name),
|
||
`BOR_INPUT_PLACEHOLDER` (chat composer placeholder), and
|
||
`BOR_FOOTER_TEXT` (footer line) — all served by `GET /api/config` and
|
||
applied by `assets/brand.js` at boot.
|
||
|
||
The colors are set from the admin **Theme tab** (`/theme.html`,
|
||
admin-only — phase 91): the 8 identity colors plus the three strings
|
||
above are edited with text fields and color pickers, persisted in the
|
||
`ui_settings` table, and injected into every served page as an inline
|
||
`<style>` BEFORE first paint (no red flash, no pop-in). The old
|
||
CSS-file theming is retired — a leftover `BOR_THEME` line in a
|
||
deployment's `.env` is simply ignored. Leave everything unset and the
|
||
app renders the built-in dark-tech palette byte-identically.
|
||
|
||
---
|
||
|
||
## Checking retrieval quality
|
||
|
||
Ask the *real* pipeline (live aipi embeddings + the current KB) whether a
|
||
question lands on the right document, with gate verdict and per-document
|
||
cosine / FTS / fused scores:
|
||
|
||
```bash
|
||
uv run python -m scripts.eval_retrieval "How did I install gitlab?"
|
||
uv run python -m scripts.eval_retrieval --from-file questions.txt --top 8
|
||
```
|
||
|
||
Requires `AIPI_KEY` in the environment and an imported knowledge base.
|
||
|
||
---
|
||
|
||
## Debugging
|
||
|
||
`debugpy` is **off by default** and *never imported* unless you opt in:
|
||
|
||
```bash
|
||
DEBUGPY=1 uv run uvicorn app.main:app
|
||
# → log line: debugpy: remote debugging ENABLED, listening on 0.0.0.0:5678
|
||
```
|
||
|
||
Attach from VS Code (`.vscode/launch.json`):
|
||
```json
|
||
{
|
||
"name": "Attach to Brain of Reese",
|
||
"type": "debugpy",
|
||
"request": "attach",
|
||
"connect": { "host": "localhost", "port": 5678 },
|
||
"pathMappings": [
|
||
{ "localRoot": "${workspaceFolder}", "remoteRoot": "/app" }
|
||
]
|
||
}
|
||
```
|
||
|
||
The port is non-blocking and attach-on-demand: the app keeps running
|
||
normally until you attach. Override the port with `DEBUGPY_PORT`.
|
||
|
||
---
|
||
|
||
## QA / Testing
|
||
|
||
Three layers — the project rule is **one story, one phase, one Playwright
|
||
suite** (see `AGENTS.md`):
|
||
|
||
```bash
|
||
# Unit + integration (FastAPI TestClient)
|
||
uv run pytest
|
||
|
||
# Coverage gate (phases require >90% on app/)
|
||
uv run pytest --cov=app --cov-report=term-missing
|
||
|
||
# Lint + static types
|
||
uv run ruff check .
|
||
uv run pyright
|
||
|
||
# Playwright E2E — install the browser once:
|
||
uv run playwright install chromium
|
||
|
||
# Each story's E2E runs IN ISOLATION:
|
||
uv run pytest tests/e2e/test_import_documents.py -v --no-cov
|
||
uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
|
||
# ...one file per story in .agents/user_stories/ (see .agents/phases/todo/)
|
||
```
|
||
|
||
**Deterministic E2E:** by default the E2E app talks to a local **mock
|
||
aipi** (`tests/e2e/mock_llm.py`) whose embeddings are real token-overlap
|
||
vectors — so the cosine relevance threshold behaves like production. To
|
||
run E2E against the **live** self-hosted models instead:
|
||
|
||
```bash
|
||
E2E_REAL_LLM=1 uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
|
||
```
|
||
|
||
---
|
||
|
||
## Production Deployment
|
||
|
||
Build the multi-stage image (frontend minified by esbuild in the builder
|
||
stage, deps installed by `uv`, non-root runtime):
|
||
|
||
```bash
|
||
podman build -t brain-of-reese/app:latest .
|
||
```
|
||
|
||
Run standalone (bring your own Postgres + pgvector):
|
||
```bash
|
||
podman run -d --name brain-of-reese \
|
||
-p 8000:8000 \
|
||
-e BOR_DATABASE_URL=postgresql+psycopg://reese:SECRETPASSWORD@dbhost:5432/brain_of_reese \
|
||
-e BOR_LLM_BASE_URL=https://aipi.reeseapps.com/v1 \
|
||
-e BOR_LLM_API_KEY=$AIPI_KEY \
|
||
brain-of-reese/app:latest
|
||
```
|
||
|
||
The entrypoint runs `alembic upgrade head` automatically on start.
|
||
|
||
Or run the whole stack from compose (app + db):
|
||
```bash
|
||
podman compose --profile prod up -d --build
|
||
```
|
||
|
||
Production hardening: app runs as non-root (uid 10001), slim image,
|
||
healthcheck on `/api/health`, debugpy off unless `DEBUGPY=1`, all assets
|
||
served locally (no CDN), `BOR_ENVIRONMENT=production`.
|
||
|
||
---
|
||
|
||
## Troubleshooting
|
||
|
||
- **`401` from aipi** — set `BOR_LLM_API_KEY` (or `$AIPI_KEY`).
|
||
- **`litellm.UnsupportedParamsError … encoding_format`** — the aipi proxy
|
||
(litellm `openai_like`) rejects the `encoding_format` parameter. The app
|
||
already works around this by POSTing a minimal `{model, input}` payload.
|
||
If you see this, you are likely calling the endpoint with a different
|
||
client — drop the parameter (or set `litellm.drop_params = True` on the proxy).
|
||
- **Embedding dimension mismatch** — aipi changed models; run
|
||
`uv run python -m scripts.llm_probe`, update `BOR_EMBEDDING_DIM`, then
|
||
drop + recreate the chunks table.
|
||
- **Honest deflection** (the amber *"I haven't done anything like that"*
|
||
bubble) — every question passes the honesty gate: deflection happens only
|
||
when the best cosine similarity is below `BOR_RELEVANCE_THRESHOLD`
|
||
(default `0.62`) **and** no chunk matched lexically (`fts_hits = 0`).
|
||
A weak cosine with a lexical hit (name-your-tool questions) still gets a
|
||
grounded answer. When it does deflect, the LLM prompt carries weak-hit
|
||
*titles only* (no document content), the reply opens with *"I haven't
|
||
done anything like that"*, the bubble renders amber with *"Maybe try"*
|
||
chips derived from the closest indexed titles, and the `query_log` row
|
||
records `deflected=true`. This is a feature, not a bug.
|
||
- **Answers deflect too often / too rarely** — tune
|
||
`BOR_RELEVANCE_THRESHOLD` (lower = answers more, higher = more honest
|
||
deflection): `0.0` ⇒ the gate leans entirely on FTS hits; `1.0` ⇒
|
||
everything deflects unless a chunk matches lexically. The `embed` model's
|
||
cosines cluster in a ~0.6–0.85 band on the live KB, so the default is
|
||
`0.62`. Check real scores:
|
||
```sql
|
||
SELECT question, top_score, fts_hits, deflected
|
||
FROM query_log ORDER BY created_at DESC LIMIT 20;
|
||
```
|
||
- **KB offline banner in the chat** — Postgres isn't running:
|
||
`podman compose up -d db`.
|
||
- **Stuck "Thinking…"** — the LLM is slow or down; a 120s client timeout
|
||
turns it into an error banner automatically.
|
||
|
||
---
|
||
|
||
## Planning & Architecture
|
||
|
||
The architecture, LOCKED decisions, and the phase roadmap live in
|
||
[`.agents/PLAN.md`](.agents/PLAN.md); per-story specs in
|
||
[`.agents/user_stories/`](.agents/user_stories/).
|