Phase 47 (owner permission 2026-08-27, TODO.md L10–11, roadmap R1): the
full Podman quadlet family (.container, .network, .volume, .image,
.pod, .kube, .swap, .os, .endpoint) and .j2 Jinja templates join the
allowed + default A9 import formats, chunked as plain text (owner
decision — no TOML/Jinja-aware splitter). No env configuration needed:
a default import now indexes them.
- app/config.py: _ALLOWED_IMPORT_EXTENSIONS + the default
import_extensions CSV gain the ten names (the original seven first);
the never-widen BOR_IMPORT_EXTENSIONS validator is untouched and
still rejects truly unknown extensions.
- app/rag/chunker.py: ten _FORMAT_CHUNKERS entries -> chunk_text
(HARD_MAX_CHARS 1200 honored, unknown-suffix fallback unchanged);
docstring/comments cite the A9 revision 2026-08-27.
- tests/fixtures/docs/homelab/: quadlet/compose.container (realistic
quadlet TOML, >1500 chars, [Unit]/[Service]/[Container] sections,
RESE-QUADLET-SENTINEL-77aa), quadlet/lan.network,
quadlet/cache.volume, templates/deploy.j2 (for/set/if Jinja
constructs + RESE-JINJA-SENTINEL-33dd). Every suite that seeds the
fixture tree updates its 9 -> 13 document-count constants.
- tests/unit/test_config.py: allowed set carries all seventeen formats,
default CSV + dotted import_extension_set include the ten, the
validator accepts the new names and still rejects unknowns.
- tests/unit/test_chunker.py: dispatch parity with chunk_text for every
new suffix (parametrized), the .container fixture chunks >=2 under
the cap with the sentinel surviving, the .j2 fixture keeps {{ }}
verbatim, the unknown-suffix fallback is unchanged.
- tests/unit/test_importer.py: a default-extensions walk over a temp
tree indexes exactly the ten new files (unknown/hidden/excluded
filtered), the original seven still walk, stem-title fallback holds.
- tests/integration/test_import_quadlet_jinja.py (new): import_sources
over a temp tree with .container/.volume/.j2 -> documents + chunks
rows with stem titles; delta re-import updates only the changed .j2
doc; prune drops the deleted .volume doc with cascade.
- tests/e2e/test_quadlet_jinja_import.py (new, story suite, mock-only,
isolation): GET /api/docs (admin session) lists the four new-format
docs with non-zero chunk counts and stem titles; the Sources table
renders a row + .doc-link per file; the phase-26 modal shows the
.container TOML ([Container] section + sentinel) with stem title and
the container format badge; a RESE-JINJA-SENTINEL-33dd question
FTS-matches the .j2 chunk -> honest-positive (A8: LOW requires zero
FTS hits) — the bubble is not .is-deflected and a source chip names
templates/deploy.j2.
- README.md + .env.example: the extended default format set (A9
revised 2026-08-27, plain-text chunking, narrow-only rule intact).
- .agent/PLAN.md: the A9 revision (owner-locked R1) — A9 row status,
the revision note under the anchors table, and the §5 chunking-policy
+ §11 workflow lines. The only PLAN edit this phase.
Gates: uv run pytest 795 passed; app/ coverage TOTAL 99% (>90%);
ruff check + pyright clean; story E2E 4/4 in isolation (DB up);
regression E2E suites test_import_documents (3) / test_sync_button
(3) / test_git_sources_admin (6) green in isolation.
Also records the 47_quadlet_jinja_import task-file moves (01–03)
todo/ -> complete/.
705 lines
35 KiB
Markdown
705 lines
35 KiB
Markdown
# 🧠 Brain of Reese
|
||
|
||
A chippy, honest **RAG chatbot** over the `~/Homelab` and `~/Deployments`
|
||
projects. Point it at your notes — markdown, YAML, JSON, Python, plain
|
||
text — ask it anything, and it retrieves the relevant chunks with
|
||
**hybrid search** (pgvector cosine ∪ Postgres full-text search, fused with
|
||
Reciprocal Rank Fusion), feeds the **whole relevant document** to a
|
||
**self-hosted LLM** (`turbo` via `https://aipi.reeseapps.com/v1`), and
|
||
streams a grounded answer back.
|
||
|
||
If it doesn't have notes for your question, it admits it:
|
||
*"I haven't done anything like that"* — plus suggestions for what it **does** know.
|
||
|
||
> **Updated your notes?** The knowledge base is refreshed by re-running the
|
||
> import — it's idempotent and only re-embeds what changed:
|
||
> ```bash
|
||
> uv run python -m scripts.import_docs --prune
|
||
> ```
|
||
> Details in [Updating the documents](#updating-the-documents).
|
||
|
||
- **Stack:** FastAPI · Pydantic v2 · SQLAlchemy 2 · Alembic · pgvector ·
|
||
vanilla HTML/CSS/JS (no CDN) · Playwright E2E
|
||
- **Planning:** architecture, LOCKED decisions and the phase roadmap live
|
||
in [`.agent/PLAN.md`](.agent/PLAN.md); per-story specs in
|
||
[`.agent/user_stories/`](.agent/user_stories/).
|
||
|
||
---
|
||
|
||
## Development Setup
|
||
|
||
### Prerequisites
|
||
- [uv](https://docs.astral.sh/uv/)
|
||
- [Podman](https://podman.io/) (with the `podman compose` provider)
|
||
- Node.js is **not** needed locally (asset minification happens in the
|
||
container build only)
|
||
|
||
### 1. Install dependencies
|
||
```bash
|
||
uv sync
|
||
```
|
||
|
||
### 2. Configure
|
||
```bash
|
||
cp .env.example .env
|
||
# edit .env — the defaults already match the local compose setup.
|
||
# BOR_LLM_API_KEY: your aipi key (falls back to $AIPI_KEY if unset)
|
||
```
|
||
|
||
### 3. Start the database (Postgres 17 + pgvector)
|
||
```bash
|
||
podman compose up -d db
|
||
podman compose ps # wait until "healthy"
|
||
```
|
||
|
||
### 4. Apply migrations
|
||
```bash
|
||
uv run alembic upgrade head
|
||
```
|
||
|
||
### 5. Import your knowledge base
|
||
```bash
|
||
uv run python -m scripts.llm_probe # sanity: models + 768-dim check
|
||
uv run python -m scripts.import_docs # import the configured sources (below)
|
||
```
|
||
|
||
Two managed source kinds (one page, one registry) plus a manual override:
|
||
|
||
- **Git sources** — managed on the **admin Git sources page**
|
||
(`/git-sources.html`) and stored in Postgres (see
|
||
[Git-based sources](#git-based-sources)). `import_docs` clones each repo
|
||
(first run) or pulls it (subsequent runs) into
|
||
`BOR_SOURCES_DIR/<repo-name>/` (default `~/bor-sources`) and indexes the
|
||
checkouts. While the stored list is empty, the `BOR_GIT_SOURCES` variable
|
||
in `.env` is the fallback — the moment the page stores a source, the
|
||
variable is ignored.
|
||
- **Local directory sources** (phase 38) — an existing, non-git directory
|
||
on the server, registered on the *same* admin page (see
|
||
[Local directory sources](#local-directory-sources)). No clone, no
|
||
checkout copy: the directory is walked in place. There is **no env var
|
||
for local paths** — the DB is the registry.
|
||
- **Manual directories** — `--source <path>` (repeatable) imports local
|
||
directories directly and *always wins* over the stored sources (git and
|
||
local) and the env fallback.
|
||
- If neither is set (stored list, `--source`, and `BOR_GIT_SOURCES` all
|
||
empty), `import_docs` falls back to the **previous** default,
|
||
`~/Homelab` + `~/Deployments` — kept only for backwards compatibility,
|
||
now replaced by the managed sources; the UI Sync button instead fails
|
||
loudly ("no sources configured (git or local)").
|
||
|
||
### 6. Run the app
|
||
```bash
|
||
uv run uvicorn app.main:app --reload
|
||
# → http://localhost:8000 (chat) http://localhost:8000/sources.html (KB)
|
||
```
|
||
|
||
> 📝 **After this, day-to-day is just: edit markdown → re-run the import.**
|
||
> See [Updating the documents](#updating-the-documents) below.
|
||
|
||
## Using the UI
|
||
|
||
- **Chat** (`/`) — ask questions; answers stream in with **source chips**
|
||
that cite the exact documents used. Clicking a chip opens that document
|
||
in an **almost-fullscreen modal** on the same page (no new tab).
|
||
- **Document viewer** — the modal above *is* the viewer; the full text of
|
||
any indexed document is served from the database (no filesystem access):
|
||
markdown is rendered, every other format (`yaml`, `json`, `py`, `txt`, …)
|
||
is shown as escaped monospace text. `/document.html?source=…&path=…`
|
||
stays as the **full-page / direct-link** form (the modal's “Full page”
|
||
button and the URL to share — it works without JS). Unknown documents
|
||
get a designed not-found state with a link back to the index.
|
||
- **Sources** (`/sources.html`) — the indexed document list; the *Path*
|
||
column opens each document in the same **almost-fullscreen modal** (no
|
||
new tab). **Admin-only** — anonymous visitors see a sign-in gate instead
|
||
(the catalog is what the login locks; the document viewer itself stays
|
||
open to everyone).
|
||
- **Git sources** (`/git-sources.html`) — the admin-managed source
|
||
registry: the git repositories the **Sync sources** button clones and
|
||
indexes, **and** existing local directories it imports directly
|
||
(phase 38 — one table with a `kind` discriminator, one page); **admin-only**
|
||
(the same sign-in gate as Sources). Add or remove sources here — no
|
||
`.env` editing, no restart. A local directory must be an absolute,
|
||
existing directory at add-time (a missing/relative path is rejected
|
||
inline, naming the path; so are duplicates); list rows carry a **Git**
|
||
or **Local** badge. Adding/removing does not clone or prune on its
|
||
own: the Sync button performs that (git + local together, one run,
|
||
prune over the union), and a removed source's documents leave the index
|
||
on the next sync.
|
||
|
||
## Thinking
|
||
|
||
The self-hosted `turbo` model reasons before it answers. That reasoning is
|
||
streamed with the turn as `thinking` SSE events and shown in a
|
||
**collapsible "Thinking" block** above the answer bubble: it opens and
|
||
fills in live while the model thinks, tucks itself away the moment the
|
||
first answer token lands, and stays click-toggleable afterwards. Thinking
|
||
persists with the message, so a reloaded conversation restores the block
|
||
(collapsed) alongside the answer. How much the model thinks — or whether
|
||
it thinks at all — is the model's call: turns without reasoning render
|
||
exactly as before.
|
||
|
||
To hide it, set `BOR_STREAM_THINKING=0` — the `thinking` events stop
|
||
(the per-turn log line still counts `thinking_chars`).
|
||
|
||
## Agent document tools (list + read)
|
||
|
||
Retrieval only puts the top documents in context. When an answer depends
|
||
on a file a note *references* ("the exact JSON shape is in
|
||
example-record-file.json"), the model can extend its own context with two
|
||
server-side tools — on **grounded** (high-relevance) turns only:
|
||
|
||
* **`list_documents`** — lists every indexed document, one
|
||
`source/path — title` line each (the same order as the Sources page);
|
||
* **`read_document(source, path)`** — appends the **full** text of
|
||
one more indexed document to the context (never truncated).
|
||
|
||
Each call the model requests is executed against Postgres only (no extra
|
||
LLM round trip) and streamed as an SSE `tool` frame ahead of the answer —
|
||
`{"type": "tool", "name": …, "argument": "source/path" | null}`. In the
|
||
chat, each call shows a **"calling tool" state** in addition to
|
||
"thinking": the send button keeps its busy state ("Calling tool…") and a
|
||
visible tool line (`🔎 Listing documents` / `📄 Reading source/path`) lands
|
||
above the answer, one per call, in order. The tool lines persist with the
|
||
message, so a reloaded conversation re-renders them. The read document is
|
||
reflected in the answer's **source chips** and in the `query_log` row.
|
||
|
||
The tools stay offered for the whole turn — the model may call them as
|
||
many times as it needs (re-lists included), bounded only by a round cap
|
||
that stops a pathological infinite loop:
|
||
|
||
| Env | Default | Meaning |
|
||
|---|---|---|
|
||
| `BOR_AGENT_MAX_ROUNDS` | `10` | hard cap on agent tool rounds per grounded turn — every call the model emits consumes a round; at the cap the loop forces one final no-tools answer |
|
||
|
||
`BOR_AGENT_MAX_ROUNDS=0` reproduces the pre-agent chat behavior exactly
|
||
(no `tools` in the request, no `tool` frames) — the kill switch.
|
||
Deflected turns run no tools at all — the low-relevance path is
|
||
unchanged.
|
||
|
||
## Admin & sign-in
|
||
|
||
Brain of Reese has exactly **one account: the admin (you)**. Signing in
|
||
unlocks the **full Sources catalog** and the **answer-tuning** controls;
|
||
everyone else stays anonymous and keeps **chat** and the **document
|
||
viewer** (any document an answer cites can be opened by its direct URL —
|
||
the catalog is gated, not the viewer).
|
||
|
||
### Setup (one-time)
|
||
|
||
```bash
|
||
python -c 'import secrets;print(secrets.token_hex(32))' # → paste into .env
|
||
```
|
||
|
||
```env
|
||
BOR_ADMIN_PASSWORD=your-password # plaintext — homelab scope, by design
|
||
BOR_SESSION_SECRET=<the hex from above> # signs the session cookie
|
||
```
|
||
|
||
**Fail-loud:** while either variable is empty the app refuses to start,
|
||
naming the missing one(s):
|
||
|
||
```
|
||
RuntimeError: Brain of Reese cannot start: admin auth is not configured.
|
||
Set the missing variable(s): BOR_ADMIN_PASSWORD, BOR_SESSION_SECRET …
|
||
```
|
||
|
||
### How it works
|
||
|
||
- `POST /api/login {"password": …}` → `204` + signed `bor_session` cookie
|
||
(Starlette `SessionMiddleware` — an itsdangerous-signed cookie, no
|
||
server-side store, no new service, no DB table); any mismatch → `401`
|
||
`{"detail": "invalid password"}` (constant-time compare, one generic
|
||
message — no user enumeration, there is only one user).
|
||
- `POST /api/logout` → `204` (session cleared and cookie expired;
|
||
idempotent for anonymous callers).
|
||
- `GET /api/whoami` → `{"authenticated": bool, "role": "admin"|"anonymous"}`
|
||
— the single source of truth for every UI gating decision.
|
||
- Cookie flags: `same_site="lax"`, `https_only` off — **no HTTPS
|
||
enforcement on purpose** (homelab HTTP; the cookie is single-admin
|
||
convenience, not a cloud boundary). Max age `BOR_SESSION_MAX_AGE`
|
||
(default `43200` = 12 h, refreshed while active).
|
||
- Sign in from the chat header (**Sign in**) or `/login.html` directly;
|
||
the header then offers **Sign out** (logout + reload).
|
||
|
||
### Who can do what
|
||
|
||
| Capability | Anonymous | Admin (signed in) |
|
||
|---|---|---|
|
||
| Chat (`/`) + suggestion chips | yes | yes |
|
||
| Document viewer (`/document.html?source=…&path=…`) | yes — any indexed doc by direct URL | yes |
|
||
| Sources catalog (`/sources.html`, `GET /api/docs`) | sign-in gate | full catalog |
|
||
| Tuning (Tune button, Tuning panel, `/api/steering`) | UI hidden | full |
|
||
|
||
The public API endpoints stay stateless — the signed cookie is the only
|
||
session state in the system.
|
||
|
||
## Tuning your answers
|
||
|
||
*Admin-only* — sign in first (see **Admin & sign-in** above); anonymous
|
||
visitors never see the Tune button or the Tuning panel.
|
||
|
||
If an answer isn't quite right — too chatty, wrong assumption, missing
|
||
context — **tune** Brain right there:
|
||
|
||
1. Press **“Tune”** in the meta row under any completed answer (deflected
|
||
ones included).
|
||
2. Type a short instruction (1–2000 chars), e.g. *“be more concise”* or
|
||
*“assume I'm on NixOS”*, and **Save**.
|
||
|
||
The note is stored in Postgres (`steering_notes`) and read into the
|
||
**system prompt of every subsequent chat turn** as a `<tuning>` section
|
||
(numbered, oldest first, capped at `BOR_STEERING_MAX_CHARS` chars —
|
||
default 8000, overflow marked `[…truncated…]`). With no stored notes the
|
||
prompt is byte-identical to the un-tuned one, so tuning is opt-in per
|
||
note.
|
||
|
||
List or remove notes at any time from the **“Tuning”** button in the chat
|
||
header (count badge, newest-first, per-note delete). The API is stateless
|
||
JSON if you prefer curl:
|
||
|
||
```bash
|
||
curl -s localhost:8000/api/steering # list (newest first)
|
||
curl -s -X POST localhost:8000/api/steering \
|
||
-H 'Content-Type: application/json' -d '{"note": "be more concise"}'
|
||
curl -s -X DELETE localhost:8000/api/steering/<note-id> # remove
|
||
```
|
||
|
||
## Updating the documents
|
||
|
||
**This is the workflow you'll use most.** The knowledge base is refreshed by
|
||
**re-running the import**. It is idempotent and delta-based (sha256 per
|
||
file), so a refresh after a normal editing session takes seconds:
|
||
|
||
```bash
|
||
# After editing/adding/removing notes in your projects:
|
||
uv run python -m scripts.import_docs # re-index what changed
|
||
uv run python -m scripts.import_docs --prune # also drop deleted/out-of-scope files
|
||
|
||
# Point it at extra directories (repeatable):
|
||
uv run python -m scripts.import_docs --source ~/SomeOtherDocs
|
||
```
|
||
|
||
With **git-based sources** (below) each run first pulls the latest commits
|
||
of your repos, so this same command is the whole update loop: commit in the
|
||
repo → re-run the import.
|
||
|
||
Then check the **Sources** page (`http://localhost:8000/sources.html`):
|
||
the *documents* / *chunks* counters and *last indexed* timestamp should
|
||
reflect the new files, and each document row shows when it was last
|
||
embedded.
|
||
|
||
- The import prints one line per file (`import: added|updated|unchanged|
|
||
pruned …`) and ends with a greppable summary (`import: summary files=…
|
||
added=… updated=… unchanged=… pruned=… chunks=… embed_batches=…
|
||
formats=md:203,yaml:267,…`), so it is safe to run from a cron job or
|
||
after every commit.
|
||
- Indexed formats (A9, revised 2026-08-27): **`md, markdown, txt,
|
||
yaml, yml, json, py`**, the Podman quadlet family (**`container,
|
||
network, volume, image, pod, kube, swap, os, endpoint`**), and **`j2`**
|
||
Jinja templates (case-insensitive; narrow with `BOR_IMPORT_EXTENSIONS`).
|
||
Any path with a
|
||
**dot-prefixed component** — hidden files or vendored caches like
|
||
`.esphome/.espressif/**` — is skipped, along with `.venv`,
|
||
`node_modules`, `.git`, `__pycache__`, `.pytest_cache`, `dist`, `build`.
|
||
`--prune` also drops documents whose files no longer match the filter —
|
||
that's how previously imported junk leaves the index.
|
||
- Non-markdown files get format-aware chunking (YAML top-level keys /
|
||
`---` docs, JSON top-level keys, Python top-level defs/classes via
|
||
stdlib `ast`; quadlet unit files and `j2` templates are paragraph-
|
||
packed as plain text) and their title comes from the file stem.
|
||
- Unchanged files are **not re-embedded** — only new/changed ones, so
|
||
refreshes are cheap.
|
||
- After a run that **changed** the knowledge base (at least one document
|
||
added or updated), the import also regenerates the stored **KB
|
||
overview** — a plain-text outline of the KB's basic categories that
|
||
every chat turn injects into the system prompt as `<knowledge_base>`
|
||
(phase 31). The regeneration is best-effort and change-gated: unchanged
|
||
re-imports and `--limit` debug runs skip it (no `lite` call), and a
|
||
`lite`-model failure leaves the previous outline intact without failing
|
||
the import. The summary line ends `overview=updated|skipped|failed`.
|
||
- To sanity-check the LLM backend (models + embedding dimension) after any
|
||
aipi change: `uv run python -m scripts.llm_probe`.
|
||
|
||
### Git-based sources
|
||
|
||
Rather than pointing the import at local folders, point it at **git
|
||
repositories** — the notes live in the repos and `import_docs` keeps local
|
||
checkouts of them up to date for you.
|
||
|
||
**Where the list lives (phase 35):** the primary management surface is
|
||
the **admin Git sources page** (`/git-sources.html`) — add or remove
|
||
repositories there and the list is stored in Postgres (the `git_sources`
|
||
table). `BOR_GIT_SOURCES` in `.env` is the **empty-table fallback**: it
|
||
only applies while the stored list is empty, and is ignored once the page
|
||
has any row (the page becomes the source of truth — no `.env` editing, no
|
||
restart needed afterwards).
|
||
|
||
```env
|
||
# .env — the fallback list (fresh setups, or until the admin page
|
||
# stores a source; phase 35 demotes this variable, it does not remove it)
|
||
BOR_GIT_SOURCES=https://git.reeseapps.com/reese/homelab.git,git@github.com:reese/deployments.git
|
||
BOR_SOURCES_DIR=~/bor-sources # default; each repo lands in <dir>/<repo-name>/
|
||
```
|
||
|
||
- The effective list (stored rows, else `BOR_GIT_SOURCES` while the stored
|
||
list is empty) is a set of git repo URLs. Auth is whatever the machine
|
||
supplies — `https://…` via the OS credential helper, or `git@host:repo.git`
|
||
via your SSH key; no credentials are stored in the app or `.env`.
|
||
Stored URLs are shape-validated on the page (`https://`, `ssh://`,
|
||
`git@…` — scp-style `host:repo` is rejected).
|
||
- Every run **clones** each repo (first time, shallow `--depth 1`) or
|
||
**pulls** it (`git pull --ff-only` — fast-forward only, so a diverged or
|
||
broken checkout fails loudly instead of merging) into
|
||
`BOR_SOURCES_DIR/<repo-name>/`, then indexes the checkouts exactly like
|
||
any local directory (A9 format filter, hidden-dir skip, sha256 delta).
|
||
`documents.source` is the repo directory name (e.g. `homelab`).
|
||
- **`--source <path>` overrides**: when the flag is given, the git
|
||
sources (stored list *and* `BOR_GIT_SOURCES`) are ignored and the
|
||
manual directory(ies) are imported.
|
||
- **A failed sync aborts the run**: if any repo cannot be cloned/pulled,
|
||
`import_docs` exits non-zero naming the failing repo and imports
|
||
**nothing** (no partial junk). Fix the URL/connectivity and re-run — the
|
||
other checkouts stay on disk and are pulled as usual.
|
||
|
||
### Local directory sources
|
||
|
||
Not every set of notes lives in a git repo — a plain directory can be a
|
||
first-class source too (phase 38). It shares the git sources' **one
|
||
table** (the `git_sources` registry with a `kind` discriminator: `git` |
|
||
`local`, migration 0007), **one admin page**, and **one Sync button**:
|
||
|
||
- **Add it on the Git sources page** — the “Add a local directory” form
|
||
next to the git form. Add-time validation fails loud: the path is
|
||
trimmed, `~` is expanded, and must be an **absolute, existing directory
|
||
on the server** — anything else (missing, relative, a file) is rejected
|
||
with the path named inline; a duplicate path is rejected the same way.
|
||
There is **no env var for local paths** — the DB is the local-source
|
||
registry (`BOR_GIT_SOURCES` stays a git-only fallback).
|
||
- **Sync walks it directly** — no clone, no checkout copy: each run
|
||
indexes the directory in place (A9 format filter, hidden-dir skip,
|
||
sha256 delta), together with the git checkouts in the **same run**.
|
||
`documents.source` is the directory's name. The directory is
|
||
re-verified to exist **at sync time** (it may have moved or been
|
||
deleted since add-time): a missing directory fails the run loudly,
|
||
naming the path, and imports **nothing** (the same pre-import fail-loud
|
||
as a failing git clone).
|
||
- **Pruning is over the union** — git checkouts and local directories are
|
||
imported together with `prune=True`, so a file removed from a local
|
||
directory, a repo, or a removed source leaves the index on that run.
|
||
Removing the row on the page stops the directory being a source; its
|
||
documents leave the index on the next sync (exactly like git sources).
|
||
- **`import_docs`** (no `--source`) resolves the stored git **and** local
|
||
rows — git cloned/pulled as above, local walked directly — in one run;
|
||
`--source` still wins over everything; while the table is empty,
|
||
`BOR_GIT_SOURCES` is the git-only fallback; no git rows, no local rows,
|
||
and no env URLs fails loudly ("no sources configured (git or local)").
|
||
|
||
### Sync from the UI
|
||
|
||
The **Sync sources** button on the **Sources** page — visible to the
|
||
**admin only** (anonymous visitors never see it) — runs the whole
|
||
git-source refresh in one click, in-process:
|
||
|
||
1. **clone/pull + walk** every configured source — the git sources (the
|
||
admin-managed `git_sources` table; `BOR_GIT_SOURCES` only while that
|
||
list is empty) through the same `clone_or_pull` the CLI uses (shallow
|
||
clone on first run, `git pull --ff-only` afterwards), **and** the
|
||
local directories registered on the same page, walked directly
|
||
(re-verified to exist at sync time — a missing directory fails the run
|
||
loudly, naming the path);
|
||
2. **re-import with prune** — the `--prune` equivalent, so files deleted
|
||
upstream leave the index (the button is the canonical "mirror the
|
||
repos" action); the sha256 delta still skips unchanged files, so an
|
||
unchanged re-sync re-embeds nothing;
|
||
3. **regenerate the KB overview** (the `<knowledge_base>` outline every
|
||
chat turn injects) — but only when the import actually changed the
|
||
knowledge base.
|
||
|
||
- **Prerequisites:** at least one source must be configured — a git or
|
||
local row on the admin Git sources page, or `BOR_GIT_SOURCES` in `.env`
|
||
while the stored list is empty (git-only); **all** empty fails the sync
|
||
loudly ("no sources configured (git or local)"), because the button
|
||
targets the admin-managed registry (manual `--source` directories have
|
||
no place in it) — and `git` must be on the app's `PATH` for git sources.
|
||
- **States:** clicking starts the run (`202`) and the button goes
|
||
disabled with **Syncing…** (spinning icon) while the page polls
|
||
`GET /api/sync/status` every 2 s. There is deliberately **no
|
||
client-side timeout** — a clone + embed can legitimately take minutes,
|
||
so the poll is the feedback loop and the server state is authoritative.
|
||
On success the button settles to **Synced HH:MM** with the last result
|
||
in a live region (`1 added`, `0 added · 1 unchanged`, …); on failure it
|
||
re-enables (retry-ready) and a red error banner names the failure (git's
|
||
stderr, with any embedded credentials masked).
|
||
- **One sync at a time:** a second trigger while a run is in flight gets
|
||
`409` ("a sync is already running"); the UI adopts the in-flight run
|
||
instead of starting a second one, and a page reload mid-sync re-attaches
|
||
to it the same way.
|
||
- **Idempotent:** re-syncing unchanged repos is a no-op — fast-forward
|
||
pull, hash skip, and the overview is left alone (its regeneration is
|
||
change-gated).
|
||
|
||
## Caching / deploys
|
||
|
||
A deploy is a commit — and the browser must see it **without a hard
|
||
refresh** (the "the pages are too sticky" problem, phase 33). One
|
||
Starlette middleware (`app/core/caching.py`) applies the rule at the
|
||
transport layer:
|
||
|
||
- **HTML pages always revalidate.** Every page (`/`, `/sources.html`,
|
||
`/document.html`, `/login.html`, `/tuning.html`) ships
|
||
`Cache-Control: no-cache`, so each visit re-checks the page with the
|
||
server — a page never lingers in the browser's cache unchecked.
|
||
- **Assets are versioned and cached for a year.** The pages reference
|
||
their CSS/JS with a token (`/assets/styles.css?v=<token>`), and every
|
||
`/assets/*` response ships
|
||
`Cache-Control: public, max-age=31536000, immutable`. The token is what
|
||
identifies the content, so long-term caching is safe: a new token means
|
||
a new URL, which the browser fetches fresh.
|
||
- **The token is the deploy.** In a git checkout (the normal case) it is
|
||
the short SHA of `HEAD` (`git rev-parse --short HEAD`), computed once
|
||
per process start — so **every commit/deploy flips the token** and the
|
||
versioned asset URLs change with it. A checkout without `.git` (or a git
|
||
failure) falls back to a stable content hash of the `frontend/` tree
|
||
(sorted path + mtime + size), so dev checkouts still bust; a missing
|
||
static dir gets the placeholder token `dev`.
|
||
- **The API is untouched.** Nothing under `/api/*` — the SSE chat stream
|
||
in particular — gains or loses a header or has its body read; the SSE
|
||
endpoint's own `Cache-Control: no-cache` is set by the endpoint itself.
|
||
|
||
No CDN, no new services, no build-step change: the middleware rewrites
|
||
the asset references of the five known pages in flight. The unversioned
|
||
asset paths keep working too (the static mount ignores the query string),
|
||
so old tabs and direct links to `/assets/…` still resolve.
|
||
|
||
> **Deploy note:** the very first deploy onto this scheme needs one
|
||
> normal page visit, so the browser revalidates the HTML once and starts
|
||
> requesting the versioned assets; every commit after that is picked up
|
||
> automatically.
|
||
|
||
## Checking retrieval quality
|
||
|
||
Ask the *real* pipeline (live aipi embeddings + the current KB) whether a
|
||
question lands on the right document, with the gate verdict and per-document
|
||
cosine / FTS / fused scores:
|
||
|
||
```bash
|
||
uv run python -m scripts.eval_retrieval "How did I install gitlab?"
|
||
uv run python -m scripts.eval_retrieval --from-file questions.txt --top 8
|
||
```
|
||
|
||
Requires `AIPI_KEY` in the environment (same convention as
|
||
`scripts/llm_probe.py`) and an imported knowledge base.
|
||
|
||
## How retrieval works (hybrid)
|
||
|
||
Every question is embedded and also lexically tokenized (OR-joined, English
|
||
stemming) and searched **twice** against Postgres:
|
||
|
||
1. **Vector** — pgvector cosine top-N (default `BOR_HYBRID_VECTOR_CANDIDATES=100`)
|
||
2. **Lexical** — a stored `tsvector` (GIN-indexed) matched with `to_tsquery`,
|
||
top-N by `ts_rank` (default `BOR_HYBRID_LEXICAL_CANDIDATES=30`)
|
||
|
||
The two ranked lists are fused with **Reciprocal Rank Fusion**
|
||
(`score = Σ 1/(k + rank)`, `BOR_RRF_K=60`) — a chunk in both lists scores
|
||
nearly double, which is what lets a name-your-tool question ("gitlab") find
|
||
its own document even when the question embeds close to generic templates.
|
||
|
||
The **honesty gate** (A8) then answers (HIGH) when the best cosine is ≥
|
||
`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **or** at least one chunk matched
|
||
lexically (`fts_hits > 0`) — it deflects (LOW) only when *both* signals are
|
||
absent. The top `BOR_TOP_N_DOCS` full documents are still what the LLM sees.
|
||
|
||
`query_log` records every turn (`top_score` = best cosine, `fts_hits`,
|
||
`chunk_hits`, `deflected`, `sources`, `latency_ms`) — the raw material for
|
||
tuning: `psql … -c 'SELECT question, top_score, fts_hits, deflected FROM query_log ORDER BY created_at DESC LIMIT 20'`.
|
||
|
||
## Document summaries (non-markdown)
|
||
|
||
Raw yaml/json/py/txt embeds badly — flags and keys are not language, so
|
||
retrieval can miss exactly the documents that are all configuration. At
|
||
import time, every **non-markdown** A9 document is summarized by the aipi
|
||
`lite` model (one-shot `LLMClient.chat`, `BOR_LLM_SUMMARY_MODEL`, default
|
||
`lite`):
|
||
|
||
* the summary is stored on `documents.summary` **and** indexed as one
|
||
extra embedded chunk (`chunks.is_summary`, position −1), so hybrid
|
||
search has a natural-language target to hit instead of the raw text;
|
||
* the last line is a **code-deterministic** pointer — `Source:
|
||
<source>/<path>` — appended by the app, never model-generated;
|
||
* the model only sees the first `BOR_SUMMARY_MAX_CHARS` (default 12000)
|
||
characters of the document; overflow is cut and marked with the shared
|
||
`[…truncated…]` marker.
|
||
|
||
A summary hit resolves through the normal chunk→document mapping: the
|
||
LLM receives the **full source document** (never the summary alone,
|
||
never truncated). Markdown files are natural language already and get no
|
||
summary. Summary generation is best-effort: if `lite` fails for a file,
|
||
the document is still indexed (without a summary), the failure is logged
|
||
and counted in the import summary line (`summaries=N summary_errors=N`).
|
||
On a chat turn, the per-turn log line records `summary_hits=N` — how many
|
||
summary chunks of the selected context the question landed on.
|
||
|
||
## Debugging
|
||
|
||
`debugpy` is **off by default** and *never imported* unless you opt in —
|
||
zero overhead in normal runs.
|
||
|
||
```bash
|
||
DEBUGPY=1 uv run uvicorn app.main:app
|
||
# → log line: debugpy: remote debugging ENABLED, listening on 0.0.0.0:5678
|
||
```
|
||
|
||
Then attach from VS Code (`.vscode/launch.json`):
|
||
```json
|
||
{
|
||
"name": "Attach to Brain of Reese",
|
||
"type": "debugpy",
|
||
"request": "attach",
|
||
"connect": { "host": "localhost", "port": 5678 },
|
||
"pathMappings": [
|
||
{ "localRoot": "${workspaceFolder}", "remoteRoot": "/app" }
|
||
]
|
||
}
|
||
```
|
||
The port is non-blocking and attach-on-demand: the app keeps running
|
||
normally until you attach. Override the port with `DEBUGPY_PORT`.
|
||
|
||
## QA / Testing Environment
|
||
|
||
Three layers — the project rule is **one story, one phase, one Playwright
|
||
suite** (see `AGENTS.md`):
|
||
|
||
```bash
|
||
# Unit + integration (FastAPI TestClient)
|
||
uv run pytest
|
||
|
||
# Same, with the coverage gate (phases require >90% on app/)
|
||
uv run pytest --cov=app --cov-report=term-missing
|
||
|
||
# Lint + static types
|
||
uv run ruff check .
|
||
uv run pyright
|
||
|
||
# Playwright E2E — install the browser once:
|
||
uv run playwright install chromium
|
||
|
||
# Each story's E2E runs IN ISOLATION (DB must be up):
|
||
podman compose up -d db
|
||
uv run pytest tests/e2e/test_import_documents.py -v --no-cov
|
||
uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
|
||
# ...one file per story in .agent/user_stories/ (see .agent/phases/todo/)
|
||
```
|
||
|
||
**Deterministic E2E:** by default the E2E app talks to a local **mock
|
||
aipi** (`tests/e2e/mock_llm.py`) whose embeddings are real
|
||
token-overlap vectors — so the cosine relevance threshold behaves like
|
||
production (on-topic questions answer, off-topic ones deflect).
|
||
To run E2E against the **live** self-hosted models instead:
|
||
|
||
```bash
|
||
E2E_REAL_LLM=1 uv run pytest tests/e2e/test_chat_rag.py -v --no-cov
|
||
```
|
||
(requires a real import of your docs first).
|
||
|
||
## Production Deployment
|
||
|
||
Build the multi-stage image (frontend minified by esbuild in the builder
|
||
stage, deps installed by `uv`, non-root runtime):
|
||
|
||
```bash
|
||
podman build -t brain-of-reese/app:latest .
|
||
```
|
||
|
||
Run standalone (bring your own Postgres + pgvector):
|
||
```bash
|
||
podman run -d --name brain-of-reese \
|
||
-p 8000:8000 \
|
||
-e BOR_DATABASE_URL=postgresql+psycopg://reese:SECRETPASSWORD@dbhost:5432/brain_of_reese \
|
||
-e BOR_LLM_BASE_URL=https://aipi.reeseapps.com/v1 \
|
||
-e BOR_LLM_API_KEY=$AIPI_KEY \
|
||
brain-of-reese/app:latest
|
||
```
|
||
The entrypoint runs `alembic upgrade head` automatically on start.
|
||
|
||
Or run the whole stack from compose (app + db):
|
||
```bash
|
||
podman compose --profile prod up -d --build
|
||
```
|
||
|
||
Production hardening notes: app runs as non-root (uid 10001), slim image,
|
||
healthcheck on `/api/health`, debugpy off unless `DEBUGPY=1`, all assets
|
||
served locally (no CDN), `BOR_ENVIRONMENT=production`.
|
||
|
||
## Configuration reference
|
||
|
||
| Env | Default | Meaning |
|
||
|-----|---------|---------|
|
||
| `BOR_APP_NAME` | `Brain of Reese` | the display name everywhere (phase 39): every page `<title>`, the header brand, the chat status labels ("… is thinking"), the empty-state greeting, and the aria/placeholder text. Served to the frontend by `GET /api/config` and applied by `assets/brand.js`; a name starting `Brain of ` keeps the bold split (`Brain of <strong>rest</strong>`), any other name renders in normal weight. Unset ⇒ byte-identical to the default |
|
||
| `BOR_DATABASE_URL` | local compose URL | SQLAlchemy URL (psycopg) |
|
||
| `BOR_LLM_BASE_URL` | `https://aipi.reeseapps.com/v1` | OpenAI-compatible endpoint |
|
||
| `BOR_LLM_API_KEY` | — (falls back to `$AIPI_KEY`) | aipi API key |
|
||
| `BOR_LLM_CHAT_MODEL` | `turbo` | chat model |
|
||
| `BOR_LLM_EMBED_MODEL` | `embed` | embedding model |
|
||
| `BOR_LLM_SUMMARY_MODEL` | `lite` | one-shot (non-streaming) completions: document summaries at import (phase 30) and the KB overview (phase 31) |
|
||
| `BOR_EMBEDDING_DIM` | `768` | vector dimension (fixed at table creation) |
|
||
| `BOR_TOP_N_DOCS` | `2` | full documents fed to the LLM |
|
||
| `BOR_RELEVANCE_THRESHOLD` | `0.62` | answer when best cosine ≥ this **or** an FTS hit; below + no FTS ⇒ honest deflection |
|
||
| `BOR_HYBRID_VECTOR_CANDIDATES` | `100` | cosine list width for the RRF fusion |
|
||
| `BOR_HYBRID_LEXICAL_CANDIDATES` | `30` | FTS list width for the RRF fusion |
|
||
| `BOR_RRF_K` | `60` | RRF damping constant (`1/(k + rank)`) |
|
||
| `BOR_AGENT_MAX_ROUNDS` | `10` | hard cap on agent tool rounds per grounded turn — every call the model emits consumes a round; at the cap the loop forces one final no-tools answer (0 = no tools, the kill switch) |
|
||
| `BOR_IMPORT_EXTENSIONS` | `md,markdown,txt,yaml,yml,json,py,container,network,volume,image,pod,kube,swap,os,endpoint,j2` | csv of importable formats (may only narrow the A9 set) |
|
||
| `BOR_GIT_SOURCES` | — (empty) | csv of git repo URLs — **fallback while the admin Git sources page's list (Postgres `git_sources`) is empty**; the page is the primary management surface (see *Git-based sources*). **Git-only**: local directory sources have no env var — they are registered on the admin page (see *Local directory sources*) |
|
||
| `BOR_SOURCES_DIR` | `~/bor-sources` | where the git source repos are cloned/pulled (one subdirectory per repo) |
|
||
| `BOR_STEERING_MAX_CHARS` | `8000` | char budget for the `<tuning>` (steering notes) prompt section |
|
||
| `BOR_SUMMARY_MAX_CHARS` | `12000` | cap on document content sent to the `lite` summary model at import (see *Document summaries*) |
|
||
| `BOR_KB_OVERVIEW_MAX_CHARS` | `4000` | char budget for the `<knowledge_base>` (KB overview) prompt section |
|
||
| `BOR_OVERVIEW_INPUT_MAX_CHARS` | `40000` | cap on the document list sent to the `lite` model when generating the KB overview |
|
||
| `BOR_SUGGESTIONS` | built-in list | JSON list of onboarding chips |
|
||
| `BOR_ADMIN_PASSWORD` | *(required)* | the single admin's password (plaintext, `.env`); app refuses to start when empty |
|
||
| `BOR_SESSION_SECRET` | *(required)* | signing key for the `bor_session` cookie; `python -c 'import secrets;print(secrets.token_hex(32))'` |
|
||
| `BOR_SESSION_MAX_AGE` | `43200` | session-cookie lifetime in seconds (12 h, sliding) |
|
||
| `DEBUGPY` | `0` | `1` ⇒ attach-on-demand debugpy on `DEBUGPY_PORT` (default 5678) |
|
||
| `BOR_LOG_LEVEL` | `INFO` | app log level |
|
||
|
||
## Troubleshooting
|
||
|
||
- **`401` from aipi** — set `BOR_LLM_API_KEY` (or `$AIPI_KEY`).
|
||
- **`litellm.UnsupportedParamsError … encoding_format` from aipi** — the
|
||
aipi proxy (litellm `openai_like`) rejects the `encoding_format` parameter
|
||
that the `openai` SDK injects into every embeddings request. The app
|
||
already works around this by POSTing a minimal `{model, input}` payload
|
||
through the openai client's own httpx transport (`app/rag/llm.py` →
|
||
`LLMClient._embed_batch`). If you see this, you are likely calling the
|
||
endpoint with a different client — drop the parameter (or set
|
||
`litellm.drop_params = True` on the proxy).
|
||
- **Embedding dimension mismatch** — aipi changed models; run
|
||
`uv run python -m scripts.llm_probe`, update `BOR_EMBEDDING_DIM`, then
|
||
drop + recreate the chunks table (new migration or manual `TRUNCATE
|
||
chunks, documents`).
|
||
- **Honest deflection (the amber “I haven't done anything like that”
|
||
bubble)** — every question passes the honesty gate: deflection happens
|
||
only when the best cosine similarity is below
|
||
`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **and** no chunk matched the
|
||
question lexically (`fts_hits = 0`). A weak cosine with a lexical hit
|
||
(name-your-tool questions) still gets a grounded answer. When it does
|
||
deflect, the LLM prompt carries weak-hit *titles only* (no document
|
||
content), the reply opens with “I haven't done anything like that”, the
|
||
bubble renders amber with “Maybe try” chips derived from the closest
|
||
indexed titles, the SSE `done` event carries `deflected: true` +
|
||
`suggestions[]`, and the `query_log` row records `deflected=true` + the
|
||
weak `top_score` + `fts_hits`. This is a feature, not a bug — the KB
|
||
simply has no notes that close; the chips always point at topics Brain
|
||
really covers.
|
||
- **Answers deflect too often / too rarely** — tune
|
||
`BOR_RELEVANCE_THRESHOLD` (lower = answers more, higher = more honest
|
||
deflection): `0.0` ⇒ the gate leans entirely on FTS hits; `1.0` ⇒
|
||
everything deflects unless a chunk matches lexically. The `embed` model's
|
||
cosines cluster in a ~0.6–0.85 band on the live KB, so the default is
|
||
`0.62`; after changing it, check the real scores:
|
||
`psql … -c 'SELECT question, top_score, fts_hits, deflected FROM query_log ORDER BY created_at DESC LIMIT 20'`
|
||
- **KB offline banner in the chat** — Postgres isn't running:
|
||
`podman compose up -d db`.
|
||
- **Stuck "Thinking…"** — the LLM is slow or down; a 120s client timeout
|
||
turns it into an error banner automatically.
|