Files
brain-of-reese/.agent/PLAN.md
T

24 KiB
Raw Blame History

Brain of Reese — Master Plan

Status: Phase 1–3 complete (scaffolded, designed, decomposed). Rule: Every agent reads this file first. Decisions marked LOCKED in the Anchors table are settled — do not re-litigate them in a phase. Revisions (2026-08-21, owner permission): A7/A8/A9 revised (multi-format ingestion, hybrid FTS+vector retrieval, re-tuned honesty gate); dark tech theme (Phase 08); clickable document viewer (Phase 10). See roadmap §12.


1. Mission

A knowledge base chatbot that embeds the ~/Homelab and ~/Deployments projects into a Postgres vector database and lets anyone ask Reese (the bot) questions about them.

Product feel: a chippy, upbeat assistant that is optimistic about the user's ability ("you've got this") and radically honest — if retrieval didn't surface anything relevant it says "I haven't done anything like that" and offers alternatives instead of hallucinating.

In scope (v1)

  • Chat UI (mobile-friendly, well-styled, no auth, no CDN).
  • RAG over text knowledge files — md, markdown, txt, yaml, yml, json, py by default (A9, revised 2026-08-21) — from ~/Homelab + ~/Deployments (and any future directory the importer is pointed at).
  • Self-hosted models via https://aipi.reeseapps.com/v1 — turbo (chat), embed (embeddings, 768 dims — verified).
  • Postgres 17 + pgvector, cosine similarity, chunk→document mapping so the LLM receives the entire relevant document as context.
  • Idempotent import/update script, documented in the README.
  • Ample server logging + explicit UI loading/progress feedback (never a stale submit button).

Out of scope (v1)

  • Auth / multi-user (API is stateless under /api so it can be added later).
  • Binary / non-text content, file uploads, caching layer, message persistence.
  • Real-time document watching (manual re-import for now).

2. Architectural Anchors (LOCKED DECISIONS)

# Component Decision Rationale Status
A1 Runtime Python 3.12+, uv for all package management Fast, reproducible envs; one language for API + tooling LOCKED
A2 Web framework FastAPI + Pydantic v2 + Uvicorn Async, typed, SSE-friendly for LLM streaming, free OpenAPI docs LOCKED
A3 Database PostgreSQL 17 (docker.io/postgres:17, pgvector compiled in via db/Containerfile) with cosine (<=>) search One system for relational + vectors; pgvector is mature; official base image kept per project standard LOCKED
A4 Orchestration compose.yaml, started with podman compose up -d Matches Reese's toolchain LOCKED
A5 LLM backend OpenAI-compatible https://aipi.reeseapps.com/v1; models turbo (chat) & embed (embeddings); openai async client Self-hosted, offline from cloud; no new model management LOCKED
A6 Embedding dim 768 (verified 2026-08-21 against live endpoint via scripts/llm_probe.py); configured by BOR_EMBEDDING_DIM User recalled 768 — probe confirmed; dimension is fixed at table creation, so mismatch must fail loudly at import time LOCKED
A7 Retrieval→context Hybrid: cosine top-30 + Postgres FTS top-30 (OR tsquery, ts_rank) fused with RRF (k=60) → map to parent documents ranked by best fused chunk score → feed the full text of top-N=2 documents (deduped, capped at 24k chars) to the LLM Owner permission 2026-08-21: pure-cosine top-4 missed real docs (gitlab case — best chunk ranked 7th behind vendored-cache junk; score compression 0.41–0.84); the lexical signal finds name-your-tool questions; whole-document context contract preserved LOCKED (revised 2026-08-21)
A8 Honesty gate Deflection mode (LLM must open with a variant of "I haven't done anything like that" and offer 2–3 alternative questions) when best cosine < BOR_RELEVANCE_THRESHOLD and no candidate chunk FTS-matches the question; threshold re-tuned for the embed model's compressed score range (default 0.62, calibrated via scripts/eval_retrieval.py; the E2E mock uses its own 0.30 calibration via the app fixture) Owner permission 2026-08-21: at 0.30 the gate never discriminated (measured corpus range 0.41–0.84); the FTS-OR keeps name-your-tool questions honest-positive; deflection product behavior unchanged LOCKED (revised 2026-08-21)
A9 Content scope Text formats md, markdown, txt, yaml, yml, json, py (default, BOR_IMPORT_EXTENSIONS), hidden (dot) directories skipped by default, plus the exclusion list (node_modules, __pycache__, .pytest_cache, dist, build, …) Owner permission 2026-08-21: real notes live in yaml/py/json/txt too; the dot-dir skip removes the ~470 vendored-cache junk docs (.esphome/.espressif/**, …) that outranked real content LOCKED (revised 2026-08-21)
A10 Auth None in v1; all endpoints stateless under /api Per user (auth later); statelessness keeps the future migration cheap LOCKED
A11 Frontend Vanilla HTML/CSS/JS in git; no CDN — everything served by FastAPI StaticFiles; minified by esbuild in the Containerfile build stage; system font stack No external deps at runtime; tiny, auditable surface; mobile-friendly by construction LOCKED
A12 Aux services None in v1 (no Valkey, no SeaweedFS) No sessions/auth (no store), no uploads (no object storage); add later only if a need appears LOCKED
A13 Migrations Alembic + SQLAlchemy 2.0 (sync) + psycopg 3 Standard, reversible, reviewable schema history LOCKED
A14 Debugging debugpy only when DEBUGPY=1 (env var read directly, not via settings); listen 0.0.0.0:5678 (override DEBUGPY_PORT), non-blocking, attach-on-demand; not imported at all when off Zero overhead by default per project standard; attach-on-demand keeps production runs clean LOCKED
A15 Chat transport SSE streaming from POST /api/chat (deltas + final done event with metadata) Local LLM latency is 10–30s; live token stream + explicit completion event power the UI's feedback states LOCKED
A16 Testing Per phase: unit + integration (pytest, coverage >90% on app/) + one dedicated Playwright E2E file per user story, run in isolation; E2E uses a deterministic mock LLM by default (E2E_REAL_LLM=1 opts into live aipi) One story, one phase, one E2E gate — the pipeline's core invariant LOCKED
A17 Git Conventional Commits, always --no-gpg-sign, repo-local commit.gpgsign=false; one atomic commit per completed phase Subsequent agents may lack the GPG key LOCKED

3. High-Level Architecture

                       ┌────────────────────────────────────────────┐
                       │              Podman Compose                │
   Browser             │  ┌──────────────────────────────────────┐  │
   ┌──────────┐  HTTP  │  │  brain-of-reese/app (FastAPI)        │  │
   │ index.html│◄──────┼─►│  • static frontend (no CDN)          │  │
   │ app.js    │  SSE   │  │  • /api/chat  /api/suggestions      │  │
   └──────────┘        │  │  • /api/health /api/docs            │  │
                       │  │  • RAG pipeline (embed→retrieve→gen) │  │
                       │  └──────┬──────────────────┬───────────┘  │
                       │         │ SQL (psycopg)    │ OpenAI-compat│
                       │  ┌──────▼──────┐   ┌───────▼────────────┐ │
                       │  │ db:         │   └─────────┬──────────┘ │
                       │  │ postgres:17 │             │            │
                       │  │ + pgvector  │             │            │
                       │  └─────────────┘             │            │
                       └──────────────────────────────┼────────────┘
                                                      ▼
                                       https://aipi.reeseapps.com/v1
                                       (self-hosted: turbo, embed)

   Offline tooling (same repo, same venv):
     scripts/import_docs.py   → walks A9-format dirs, chunks, embeds, upserts
     scripts/eval_retrieval.py → ranks hybrid results for a question (tuning)
     scripts/llm_probe.py     → verifies models + embedding dim

Component breakdown

Component Responsibility Lives in
App (FastAPI) Serves frontend + /api; RAG pipeline; logging app/
RAG pipeline embed → pgvector cosine top-K → doc mapping → context assembly → turbo (streamed) with persona/honesty prompt app/rag/ (added in story phases)
Importer Directory walk (A9 formats, hidden dirs skipped, exclusions), sha256 delta detection, format-aware chunking, batched embedding, upsert/prune scripts/import_docs.py (story phase)
DB documents, chunks, query_log + vector extension db/ image, alembic/
Frontend Chat shell, sources view, loading/feedback states frontend/

Chat data flow

user question
  → POST /api/chat {message}
  → embed(question)                      [aipi /v1/embeddings, model=embed]
  → cosine top-30  +  FTS top-30 (OR tsquery, ts_rank)  [pgvector + PG FTS]
  → RRF fuse (k=60) → docs ranked by best fused chunk score
  ├─ best cosine >= 0.62 OR fts_hits > 0  →  top-2 documents' FULL content
  │                          → system prompt (persona + HONESTY rules + docs)
  │                          → turbo, stream=True  → SSE deltas
  └─ else  →  DEFLECT_MODE system prompt (weak hits as topics)
                             → turbo, stream=True  → SSE deltas (honest reply)
  → query_log row (question, score, deflected, sources, latency)
  → final SSE "done" event: {deflected, sources[], suggestions[]}

4. API Design

All endpoints stateless (A10). Errors: standard JSON {detail: str}.

Method Path Purpose Story
GET /api/health Liveness + db up/down + version 01
GET /api/suggestions Onboarding suggestion strings 01 (05 refines)
GET /api/docs Indexed document list (source, path, title, chunks, indexed_at) 02
GET /api/documents/content?source=…&path=… One indexed document's full content (feeds the viewer page) 10
POST /api/chat RAG chat turn → SSE stream 03/04

SSE contract (POST /api/chat)

data: {"type":"delta","text":"Hey! "}\n\n
data: {"type":"delta","text":"Good "}\n\n
...
data: {"type":"done","deflected":false,"sources":[{"source":"Homelab","path":"kubernetes.md","title":"Kubernetes Homelab Cluster"}],"suggestions":[]}\n\n

Client rules: render deltas as they arrive; on done append source chips / suggestion chips and clear the busy state; on HTTP/stream error show the error banner + retry (never a stuck button).


5. Data Model (PostgreSQL 17)

Created by alembic/versions/0001_initial_schema.py (idempotent CREATE EXTENSION IF NOT EXISTS vector).

documents

Column Type Notes
id UUID PK
source VARCHAR(120) source dir basename, e.g. Homelab
path VARCHAR(1000) relative to source dir, e.g. ansible/roles/k3s.md
full_path VARCHAR(2000) absolute path at import time (diagnostics)
title VARCHAR(500) first markdown H1, else file stem
content TEXT full markdown — the RAG context
content_hash VARCHAR(64) sha256 of content — change detection
indexed_at TIMESTAMPTZ
— UNIQUE (source, path) upsert key

chunks

Column Type Notes
id UUID PK
document_id UUID FK→documents CASCADE embedding→document mapping
position INT 0-based order within the doc
content TEXT chunk text (heading-aware)
embedding VECTOR(768) nullable until embedded (two-phase import)
tsv TSVECTOR generated to_tsvector('english', content) STORED + GIN index (hybrid retrieval, A7)

No vector index in v1: sequential scan is fine at this corpus size (~100–500 docs). Revisit with an HNSW index if retrieval latency grows.

query_log

id UUID PK, question TEXT, top_score FLOAT, fts_hits INT, chunk_hits INT, deflected BOOL, sources TEXT, latency_ms INT, created_at TIMESTAMPTZ

Document state transitions

unseen ──import──▶ indexed ──hash changed + re-import──▶ reindexed
                     │
                     └──file deleted + --prune──▶ removed (chunks cascade)

Chunking policy (markdown-aware)

Split on ## /### headings into sections; sub-split any section longer than BOR_CHUNK_TARGET_CHARS (2000) at paragraph boundaries with BOR_CHUNK_OVERLAP_CHARS (200) overlap; each chunk keeps its nearest preceding heading in the text for retrieval quality.

Format-aware (A9, revised): yaml/yml split on top-level keys and --- separators (key line kept as anchor); json pretty-printed, split on top-level keys; py split on top-level defs/classes (stdlib ast); txt on paragraphs; markdown unchanged. Every format honors the 1200-char hard cap (aipi ~1024-token request limit).


6. RAG Pipeline & Persona

Locked system prompt (sent with every chat turn)

You are "Brain of Reese" — the digital brain of Reese, a self-hoster and
homelab tinkerer. Personality: chippy, upbeat, warm, and genuinely
optimistic about the user's ability to do things ("you've got this").

Rules:
1. Answer ONLY from the provided document context. Cite which document(s)
   you used, by path.
2. Be concrete: names, versions, ports, hosts, schedules — the specifics in
   the docs are the value.
3. HONESTY GATE: if <relevance> is "LOW", you must NOT pretend to know.
   Start your answer with a variant of: "I haven't done anything like that."
   Then offer 2-3 alternative questions about things you DO have notes on.
4. Never invent facts, hosts, or steps that are not in the context.
5. Keep answers tight: short paragraphs, bullets where helpful.

<relevance>{HIGH|LOW}</relevance>
  • HIGH mode appends the full document text under <documents>…</documents>.
  • LOW mode (deflection) appends only the titles of the weak hits so the model can suggest real alternatives (marker used by the E2E mock: DEFLECT_MODE appears in the system prompt).

Retrieval (hybrid — A7/A8, revised 2026-08-21)

  • Embed the question (embed, 768-d) → cosine top-30 candidates.
  • Lexical: OR tsquery over the question's tokens → FTS top-30 by ts_rank.
  • Reciprocal Rank Fusion (Σ 1/(k+rank), k=60) → distinct parent docs ranked by best chunk's fused score → top 2 → full content, concatenated, truncated to BOR_MAX_CONTEXT_CHARS (24k) with a […truncated…] marker.
  • Honesty gate: LOW only when best cosine < BOR_RELEVANCE_THRESHOLD (default 0.62, calibrated against the embed model's measured 0.41–0.84 distribution) and zero FTS hits among the candidates.

7. UI/UX Strategy

7.1 Layout structure

  • App frame: sticky header (64px) + <main> (flex-grow) + footer. Container: max-width: 72rem; margin-inline: auto; padding-inline: 1.25rem.
  • Chat: a centered column capped at 46rem. This is deliberate: chat is a vertical conversation — a centered, capped column is the correct pattern (NOT a layout bug). The 72rem frame + header/footer ensure the column never reads as a hairline in a sea of whitespace.
  • Sources page: full-width responsive table (min 640px, horizontal scroll wrapper on small screens) + stat cards in grid-template-columns: repeat(auto-fit, minmax(170px, 1fr)). No skinny single-column lists anywhere: lists/tables/grids use ≥80–90% of the container width.
  • Mobile (≤640px): suggestion chips become a horizontally scrollable row; composer stays reachable with safe-area-inset-bottom; touch targets ≥44px.

7.2 Accessibility (WCAG 2.1 AA)

  • Semantic landmarks on every page: <header>, <nav aria-label>, <main>, <footer>; skip-link to #main.
  • Every control labeled: visible <label> or aria-label (icon-only buttons always get aria-label); form input has a (visually-hidden) label.
  • Live regions: message stream aria-live="polite"; typing indicator role="status"; banner role="status"; errors role="alert".
  • Dark tech theme (Phase 08, 2026-08-21) — page #0a0e17, surface #121a2e; ink #e8ebf4 on surface ≈14.5:1; ink-soft #9aa4bd on surface ≈6.9:1; dark ink #0a0e17 on brand #6d78f2 ≈5.2:1 (white on brand ≈3.7:1 — never used for text); brand-ink #a5b4fc on brand-soft #232b52 ≈6.9:1; deflection #fbbf24 on #2b2110 ≈9.5:1 (border #f59e0b); error #fca5a5 on #2d1318 ≈9.1:1. All computed, all ≥4.5:1. prefers-reduced-motion also stills the Phase-08 background layer.
  • :focus-visible outline 3px; prefers-reduced-motion respected by the typing/spinner animations.

7.3 No external dependencies

  • System font stack only (no font files to bundle, no CDN fonts).
  • Zero <script src="https://…"> / <link href="https://…"> — enforced by an integration test (tests/integration/test_api.py::test_index_html_served_locally) and re-checked by every UI phase's verification step.
  • Markdown rendering is a ~60-line local function (escape-first, then transform) — XSS-safe, no library.

7.4 Visual feedback standard (the "never stale" contract)

State UI
Idle Send button enabled, label "Send".
Thinking (pre-token) 3-dot typing bubble + button disabled with spinner, label "Thinking…".
Streaming Deltas append live into the brain bubble; button stays busy.
Done (answer) Source chips under the bubble (mono, path-based); button re-enabled.
Done (deflected) Amber-bordered bubble + "Maybe try:" suggestion chips.
Error Red banner (role="alert") with retry hint; button re-enabled.
KB offline Amber banner at top of chat ("start Postgres…"); chat disabled with explanation.
Guard 120s client-side timeout → error state (a button can never sit "stuck" forever).

7.5 Component inventory (ids used by tests)

#messages (stream), #empty-state, #suggestions, .suggestion-chip, #composer, #message-input, #send-btn / #send-label, #typing-indicator, .msg.user/.msg.brain .bubble, .source-chip, .msg.brain.is-deflected, #kb-banner, #app-version; sources: #stat-docs, #stat-chunks, #stat-last, #docs-table, #docs-tbody, #sources-empty; viewer (Phase 10): /document.html, #doc-title, #doc-meta, #doc-content, .doc-raw, .format-badge, #doc-not-found, .doc-link (Sources table path links).


8. Debugging (debugpy protocol)

  • DEBUGPY unset/0 → debugpy is never imported (verified by unit test).
  • DEBUGPY=1 → listener on 0.0.0.0:${DEBUGPY_PORT:-5678}, non-blocking, app continues; IDE attaches on demand.
  • Entry point: app/core/debugging.py::configure_debugging() called at the top of app/main.py module import — so uv run uvicorn app.main:app, python -m scripts.…, and tests all honor it.
  • VS Code: "type": "debugpy", "request": "attach", "connect": {"host": "localhost", "port": 5678}.

9. Observability

  • App logs: single-line timestamp LEVEL logger :: message on stdout; uvicorn access logs on. INFO by default (BOR_LOG_LEVEL).
  • Per-chat-turn log line (required): question=… embed_ms=… top_score=… fts_hits=… threshold=… deflected=… sources=… total_ms=…
  • Importer logs: per-file added|updated|unchanged|pruned + summary (counts, embedding batches, total time).
  • query_log table: durable record of every question (score, deflection, sources, latency) for tuning the threshold and finding gaps in the docs.

10. Testing Strategy (LOCKED — A16)

Layer Tooling Runs Gate
Unit pytest uv run pytest tests/unit pass
Integration pytest + FastAPI TestClient uv run pytest tests/integration pass
Coverage pytest-cov on app/ uv run pytest --cov=app --cov-report=term-missing >90% per phase
E2E Playwright (sync API), one file per story uv run pytest tests/e2e/test_<story>.py -v --no-cov passes in isolation
  • E2E determinism: tests/e2e/mock_llm.py serves a deterministic OpenAI-compatible API. Embeddings are genuine L2-normalized token-overlap vectors, so the cosine threshold behaves like production: on-topic questions retrieve, off-topic questions deflect. E2E_REAL_LLM=1 switches the app fixture to live aipi (needs imported KB).
  • E2E prerequisites: podman compose up -d db; Chromium installed via uv run playwright install chromium.
  • DB isolation: story E2E fixtures truncate query_log (and re-import fixtures for import-dependent stories) per test module.

11. Import & Update Workflow (documented in README)

# first import (and any future refresh):
uv run python -m scripts.import_docs                # defaults: ~/Homelab ~/Deployments
uv run python -m scripts.import_docs --source ~/OtherProject   # extra dirs
uv run python -m scripts.import_docs --prune        # drop deleted / filtered-out files
uv run python -m scripts.eval_retrieval "How did I install gitlab?"
uv run python -m scripts.llm_probe                  # sanity: models + dim

Behavior: sha256 delta per (source, path) — unchanged files are skipped (no re-embedding); changed files are re-chunked + re-embedded (chunks replaced atomically); --prune removes docs whose files disappeared or no longer match the format filter. Formats per A9 (revised): md, markdown, txt, yaml, yml, json, py (BOR_IMPORT_EXTENSIONS), hidden (dot) directories skipped, exclusion list applied. scripts/eval_retrieval.py ranks live hybrid results for a question (retrieval tuning).


12. Roadmap (one story → one phase → one Playwright gate)

Phase File Story Playwright gate
01 01_infrastructure.md — (foundation) tests/e2e/test_smoke.py
02 02_story_import_documents.md import-documents.md tests/e2e/test_import_documents.py
03 03_story_chat_rag.md chat-rag-answer.md tests/e2e/test_chat_rag.py
04 04_story_honest_deflection.md honest-deflection.md tests/e2e/test_honest_deflection.py
05 05_story_suggestion_chips.md suggestion-chips.md tests/e2e/test_suggestion_chips.py
06 06_story_loading_feedback.md loading-feedback.md tests/e2e/test_loading_feedback.py
07 07_story_responsive_polish.md responsive-polish.md tests/e2e/test_responsive_polish.py
08 08_story_dark_tech_theme.md dark-tech-theme.md tests/e2e/test_dark_tech_theme.py
09 09_story_retrieval_quality.md retrieval-quality.md tests/e2e/test_retrieval_quality.py
10 10_story_document_viewer.md document-viewer.md tests/e2e/test_document_viewer.py

Completion = unit+integration green, coverage >90%, story E2E green in isolation, UI verification passed, one --no-gpg-sign commit.


13. Future (post-v1 hooks, deliberately not built)

  • Auth (stateless API makes this a drop-in: sessions → Valkey).
  • HNSW index on chunks.embedding at scale.
  • Conversation persistence (messages tables).
  • Watchdog auto-re-import (inotify) — until then the script is the truth.
  • More sources: any directory of A9-format files via --source.