From 022da8e2bccdabe29ed9d92cd2cb7964d79f7687 Mon Sep 17 00:00:00 2001 From: ducoterra Date: Fri, 21 Aug 2026 13:42:21 -0400 Subject: [PATCH] =?UTF-8?q?feat:=20scaffold=20Brain=20of=20Reese=20?= =?UTF-8?q?=E2=80=94=20FastAPI=20RAG=20chat=20over=20Postgres=2017=20+=20p?= =?UTF-8?q?gvector?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Foundation (phase 01, verified): - FastAPI app: /api/health, /api/suggestions, /api/chat (placeholder), static frontend served locally (no CDN) - Postgres 17 + pgvector via db/Containerfile + compose.yaml (podman compose up -d db), Alembic initial migration (documents, chunks with vector(768), query_log) - LLM client targeting https://aipi.reeseapps.com/v1 (turbo/embed); scripts/llm_probe.py verified models + 768-dim embeddings live - Conditional debugpy: imported only when DEBUGPY=1 (attach on demand, :5678); logging config for clean single-line logs - Frontend shell: mobile-first chat + Sources pages, tokens, a11y baselines - Tests: 24 unit+integration (99% coverage on app/), ruff + pyright clean, Playwright smoke E2E (3 tests) against a deterministic mock LLM - Planning: .agent/PLAN.md (architecture + LOCKED decisions), AGENTS.md, 6 user stories, 7 phase files (one story / one phase / one Playwright suite each) --- .agent/PLAN.md | 368 +++++ .agent/phases/complete/.gitkeep | 0 .agent/phases/todo/01_infrastructure.md | 59 + .../phases/todo/02_story_import_documents.md | 76 + .agent/phases/todo/03_story_chat_rag.md | 74 + .../phases/todo/04_story_honest_deflection.md | 66 + .../phases/todo/05_story_suggestion_chips.md | 59 + .../phases/todo/06_story_loading_feedback.md | 64 + .../phases/todo/07_story_responsive_polish.md | 57 + .agent/user_stories/chat-rag-answer.md | 61 + .agent/user_stories/honest-deflection.md | 57 + .agent/user_stories/import-documents.md | 62 + .agent/user_stories/loading-feedback.md | 61 + .agent/user_stories/responsive-polish.md | 65 + .agent/user_stories/suggestion-chips.md | 58 + .env.example | 31 + .gitignore | 29 + AGENTS.md | 52 + Containerfile | 49 + README.md | 208 +++ alembic.ini | 40 + alembic/env.py | 50 + alembic/script.py.mako | 25 + alembic/versions/0001_initial_schema.py | 86 ++ app/__init__.py | 3 + app/api/__init__.py | 1 + app/api/chat.py | 29 + app/api/health.py | 21 + app/api/suggestions.py | 14 + app/config.py | 64 + app/core/__init__.py | 1 + app/core/debugging.py | 64 + app/core/logging.py | 26 + app/db.py | 41 + app/main.py | 47 + app/models.py | 81 ++ app/schemas.py | 45 + compose.yaml | 49 + db/Containerfile | 19 + frontend/assets/app.js | 189 +++ frontend/assets/sources.js | 69 + frontend/assets/styles.css | 460 ++++++ frontend/index.html | 76 + frontend/sources.html | 87 ++ pyproject.toml | 49 + scripts/__init__.py | 1 + scripts/entrypoint.sh | 8 + scripts/llm_probe.py | 63 + tests/conftest.py | 12 + tests/e2e/__init__.py | 1 + tests/e2e/conftest.py | 126 ++ tests/e2e/mock_llm.py | 174 +++ tests/e2e/test_smoke.py | 45 + .../fixtures/docs/deployments/new-service.md | 13 + tests/fixtures/docs/homelab/backups.md | 14 + tests/fixtures/docs/homelab/kubernetes.md | 16 + tests/integration/test_api.py | 49 + tests/unit/test_config.py | 42 + tests/unit/test_db.py | 56 + tests/unit/test_debugging.py | 60 + tests/unit/test_main.py | 21 + tests/unit/test_models.py | 46 + uv.lock | 1286 +++++++++++++++++ 63 files changed, 5225 insertions(+) create mode 100644 .agent/PLAN.md create mode 100644 .agent/phases/complete/.gitkeep create mode 100644 .agent/phases/todo/01_infrastructure.md create mode 100644 .agent/phases/todo/02_story_import_documents.md create mode 100644 .agent/phases/todo/03_story_chat_rag.md create mode 100644 .agent/phases/todo/04_story_honest_deflection.md create mode 100644 .agent/phases/todo/05_story_suggestion_chips.md create mode 100644 .agent/phases/todo/06_story_loading_feedback.md create mode 100644 .agent/phases/todo/07_story_responsive_polish.md create mode 100644 .agent/user_stories/chat-rag-answer.md create mode 100644 .agent/user_stories/honest-deflection.md create mode 100644 .agent/user_stories/import-documents.md create mode 100644 .agent/user_stories/loading-feedback.md create mode 100644 .agent/user_stories/responsive-polish.md create mode 100644 .agent/user_stories/suggestion-chips.md create mode 100644 .env.example create mode 100644 .gitignore create mode 100644 AGENTS.md create mode 100644 Containerfile create mode 100644 README.md create mode 100644 alembic.ini create mode 100644 alembic/env.py create mode 100644 alembic/script.py.mako create mode 100644 alembic/versions/0001_initial_schema.py create mode 100644 app/__init__.py create mode 100644 app/api/__init__.py create mode 100644 app/api/chat.py create mode 100644 app/api/health.py create mode 100644 app/api/suggestions.py create mode 100644 app/config.py create mode 100644 app/core/__init__.py create mode 100644 app/core/debugging.py create mode 100644 app/core/logging.py create mode 100644 app/db.py create mode 100644 app/main.py create mode 100644 app/models.py create mode 100644 app/schemas.py create mode 100644 compose.yaml create mode 100644 db/Containerfile create mode 100644 frontend/assets/app.js create mode 100644 frontend/assets/sources.js create mode 100644 frontend/assets/styles.css create mode 100644 frontend/index.html create mode 100644 frontend/sources.html create mode 100644 pyproject.toml create mode 100644 scripts/__init__.py create mode 100644 scripts/entrypoint.sh create mode 100644 scripts/llm_probe.py create mode 100644 tests/conftest.py create mode 100644 tests/e2e/__init__.py create mode 100644 tests/e2e/conftest.py create mode 100644 tests/e2e/mock_llm.py create mode 100644 tests/e2e/test_smoke.py create mode 100644 tests/fixtures/docs/deployments/new-service.md create mode 100644 tests/fixtures/docs/homelab/backups.md create mode 100644 tests/fixtures/docs/homelab/kubernetes.md create mode 100644 tests/integration/test_api.py create mode 100644 tests/unit/test_config.py create mode 100644 tests/unit/test_db.py create mode 100644 tests/unit/test_debugging.py create mode 100644 tests/unit/test_main.py create mode 100644 tests/unit/test_models.py create mode 100644 uv.lock diff --git a/.agent/PLAN.md b/.agent/PLAN.md new file mode 100644 index 0000000..1101d14 --- /dev/null +++ b/.agent/PLAN.md @@ -0,0 +1,368 @@ +# Brain of Reese — Master Plan + +> **Status:** Phase 1–3 complete (scaffolded, designed, decomposed). +> **Rule:** Every agent reads this file first. Decisions marked `LOCKED` in the +> Anchors table are settled — do not re-litigate them in a phase. + +--- + +## 1. Mission + +A **knowledge base chatbot** that embeds the `~/Homelab` and `~/Deployments` +projects into a Postgres vector database and lets anyone ask *Reese* (the +bot) questions about them. + +**Product feel:** a chippy, upbeat assistant that is optimistic about the +user's ability ("you've got this") and **radically honest** — if retrieval +didn't surface anything relevant it says *"I haven't done anything like +that"* and offers alternatives instead of hallucinating. + +### In scope (v1) +- Chat UI (mobile-friendly, well-styled, no auth, no CDN). +- RAG over `*.md` files **only** from `~/Homelab` + `~/Deployments` + (and any future directory the importer is pointed at). +- Self-hosted models via `https://aipi.reeseapps.com/v1` — `turbo` (chat), + `embed` (embeddings, **768 dims — verified**). +- Postgres 17 + pgvector, cosine similarity, chunk→document mapping so the + LLM receives the **entire relevant document** as context. +- Idempotent import/update script, documented in the README. +- Ample server logging + explicit UI loading/progress feedback (never a + stale submit button). + +### Out of scope (v1) +- Auth / multi-user (API is stateless under `/api` so it can be added later). +- Non-markdown content, file uploads, caching layer, message persistence. +- Real-time document watching (manual re-import for now). + +--- + +## 2. Architectural Anchors (LOCKED DECISIONS) + +| # | Component | Decision | Rationale | Status | +|---|-----------|----------|-----------|--------| +| A1 | Runtime | Python 3.12+, `uv` for all package management | Fast, reproducible envs; one language for API + tooling | LOCKED | +| A2 | Web framework | FastAPI + Pydantic v2 + Uvicorn | Async, typed, SSE-friendly for LLM streaming, free OpenAPI docs | LOCKED | +| A3 | Database | **PostgreSQL 17** (`docker.io/postgres:17`, pgvector compiled in via `db/Containerfile`) with **cosine** (`<=>`) search | One system for relational + vectors; pgvector is mature; official base image kept per project standard | LOCKED | +| A4 | Orchestration | `compose.yaml`, started with **`podman compose up -d`** | Matches Reese's toolchain | LOCKED | +| A5 | LLM backend | OpenAI-compatible `https://aipi.reeseapps.com/v1`; models **`turbo`** (chat) & **`embed`** (embeddings); `openai` async client | Self-hosted, offline from cloud; no new model management | LOCKED | +| A6 | Embedding dim | **768** (verified 2026-08-21 against live endpoint via `scripts/llm_probe.py`); configured by `BOR_EMBEDDING_DIM` | User recalled 768 — probe confirmed; dimension is fixed at table creation, so mismatch must fail loudly at import time | LOCKED | +| A7 | Retrieval→context | Cosine **top-K=4 chunks** → map to parent documents → feed the **full text of top-N=2 documents** (deduped, capped at 24k chars) to the LLM | User requirement: whole-document context; mapping via `chunks.document_id → documents.path` | LOCKED | +| A8 | Honesty gate | If best cosine similarity < `BOR_RELEVANCE_THRESHOLD` (0.30) → **deflection mode**: LLM must open with a variant of *"I haven't done anything like that"* and offer 2–3 alternative questions | Required product behavior; threshold is tunable without code change | LOCKED | +| A9 | Content scope | **`*.md` only**, with an exclusion list for non-content dirs (`.venv`, `node_modules`, `.git`, `__pycache__`, `.pytest_cache`, `dist`, `build`) | Simplicity per user; prevents indexing dependency license files (~1.6k junk files in `~/Homelab/.venv`) | LOCKED | +| A10 | Auth | **None in v1**; all endpoints stateless under `/api` | Per user (auth later); statelessness keeps the future migration cheap | LOCKED | +| A11 | Frontend | Vanilla HTML/CSS/JS in git; **no CDN** — everything served by FastAPI `StaticFiles`; minified by esbuild in the `Containerfile` build stage; system font stack | No external deps at runtime; tiny, auditable surface; mobile-friendly by construction | LOCKED | +| A12 | Aux services | **None in v1** (no Valkey, no SeaweedFS) | No sessions/auth (no store), no uploads (no object storage); add later only if a need appears | LOCKED | +| A13 | Migrations | Alembic + SQLAlchemy 2.0 (sync) + psycopg 3 | Standard, reversible, reviewable schema history | LOCKED | +| A14 | Debugging | `debugpy` **only when `DEBUGPY=1`** (env var read directly, not via settings); listen `0.0.0.0:5678` (override `DEBUGPY_PORT`), non-blocking, attach-on-demand; **not imported at all when off** | Zero overhead by default per project standard; attach-on-demand keeps production runs clean | LOCKED | +| A15 | Chat transport | **SSE streaming** from `POST /api/chat` (deltas + final `done` event with metadata) | Local LLM latency is 10–30s; live token stream + explicit completion event power the UI's feedback states | LOCKED | +| A16 | Testing | Per phase: unit + integration (pytest, **coverage >90%** on `app/`) + **one dedicated Playwright E2E file per user story**, run in isolation; E2E uses a deterministic mock LLM by default (`E2E_REAL_LLM=1` opts into live aipi) | One story, one phase, one E2E gate — the pipeline's core invariant | LOCKED | +| A17 | Git | Conventional Commits, **always `--no-gpg-sign`**, repo-local `commit.gpgsign=false`; one atomic commit per completed phase | Subsequent agents may lack the GPG key | LOCKED | + +--- + +## 3. High-Level Architecture + +``` + ┌────────────────────────────────────────────┐ + │ Podman Compose │ + Browser │ ┌──────────────────────────────────────┐ │ + ┌──────────┐ HTTP │ │ brain-of-reese/app (FastAPI) │ │ + │ index.html│◄──────┼─►│ • static frontend (no CDN) │ │ + │ app.js │ SSE │ │ • /api/chat /api/suggestions │ │ + └──────────┘ │ │ • /api/health /api/docs │ │ + │ │ • RAG pipeline (embed→retrieve→gen) │ │ + │ └──────┬──────────────────┬───────────┘ │ + │ │ SQL (psycopg) │ OpenAI-compat│ + │ ┌──────▼──────┐ ┌───────▼────────────┐ │ + │ │ db: │ └─────────┬──────────┘ │ + │ │ postgres:17 │ │ │ + │ │ + pgvector │ │ │ + │ └─────────────┘ │ │ + └──────────────────────────────┼────────────┘ + ▼ + https://aipi.reeseapps.com/v1 + (self-hosted: turbo, embed) + + Offline tooling (same repo, same venv): + scripts/import_docs.py → walks *.md dirs, chunks, embeds, upserts + scripts/llm_probe.py → verifies models + embedding dim +``` + +### Component breakdown +| Component | Responsibility | Lives in | +|-----------|----------------|----------| +| **App (FastAPI)** | Serves frontend + `/api`; RAG pipeline; logging | `app/` | +| **RAG pipeline** | `embed` → pgvector cosine top-K → doc mapping → context assembly → `turbo` (streamed) with persona/honesty prompt | `app/rag/` (added in story phases) | +| **Importer** | Directory walk (exclusions), sha256 delta detection, markdown chunking, batched embedding, upsert/prune | `scripts/import_docs.py` (story phase) | +| **DB** | `documents`, `chunks`, `query_log` + `vector` extension | `db/` image, `alembic/` | +| **Frontend** | Chat shell, sources view, loading/feedback states | `frontend/` | + +### Chat data flow +``` +user question + → POST /api/chat {message} + → embed(question) [aipi /v1/embeddings, model=embed] + → SELECT chunks ORDER BY embedding <=> $1 LIMIT 4 [pgvector cosine] + → best_score = max(1 - distance) + ├─ best_score >= 0.30 → top-2 documents' FULL content + │ → system prompt (persona + HONESTY rules + docs) + │ → turbo, stream=True → SSE deltas + └─ best_score < 0.30 → DEFLECT_MODE system prompt (weak hits as topics) + → turbo, stream=True → SSE deltas (honest reply) + → query_log row (question, score, deflected, sources, latency) + → final SSE "done" event: {deflected, sources[], suggestions[]} +``` + +--- + +## 4. API Design + +All endpoints stateless (A10). Errors: standard JSON `{detail: str}`. + +| Method | Path | Purpose | Story | +|--------|------|---------|-------| +| GET | `/api/health` | Liveness + db up/down + version | 01 | +| GET | `/api/suggestions` | Onboarding suggestion strings | 01 (05 refines) | +| GET | `/api/docs` | Indexed document list (source, path, title, chunks, indexed_at) | 02 | +| POST | `/api/chat` | RAG chat turn → **SSE stream** | 03/04 | + +### SSE contract (`POST /api/chat`) +``` +data: {"type":"delta","text":"Hey! "}\n\n +data: {"type":"delta","text":"Good "}\n\n +... +data: {"type":"done","deflected":false,"sources":[{"source":"Homelab","path":"kubernetes.md","title":"Kubernetes Homelab Cluster"}],"suggestions":[]}\n\n +``` +Client rules: render deltas as they arrive; on `done` append source chips / +suggestion chips and clear the busy state; on HTTP/stream error show the +error banner + retry (never a stuck button). + +--- + +## 5. Data Model (PostgreSQL 17) + +Created by `alembic/versions/0001_initial_schema.py` (idempotent +`CREATE EXTENSION IF NOT EXISTS vector`). + +### `documents` +| Column | Type | Notes | +|--------|------|-------| +| id | `UUID` PK | | +| source | `VARCHAR(120)` | source dir basename, e.g. `Homelab` | +| path | `VARCHAR(1000)` | relative to source dir, e.g. `ansible/roles/k3s.md` | +| full_path | `VARCHAR(2000)` | absolute path at import time (diagnostics) | +| title | `VARCHAR(500)` | first markdown H1, else file stem | +| content | `TEXT` | **full markdown — the RAG context** | +| content_hash | `VARCHAR(64)` | sha256 of content — change detection | +| indexed_at | `TIMESTAMPTZ` | | +| — | `UNIQUE (source, path)` | upsert key | + +### `chunks` +| Column | Type | Notes | +|--------|------|-------| +| id | `UUID` PK | | +| document_id | `UUID` FK→documents CASCADE | **embedding→document mapping** | +| position | `INT` | 0-based order within the doc | +| content | `TEXT` | chunk text (heading-aware) | +| embedding | `VECTOR(768)` | nullable until embedded (two-phase import) | + +> No vector index in v1: sequential scan is fine at this corpus size +> (~100–500 docs). Revisit with an HNSW index if retrieval latency grows. + +### `query_log` +`id UUID PK, question TEXT, top_score FLOAT, chunk_hits INT, deflected BOOL, sources TEXT, latency_ms INT, created_at TIMESTAMPTZ` + +### Document state transitions +``` +unseen ──import──▶ indexed ──hash changed + re-import──▶ reindexed + │ + └──file deleted + --prune──▶ removed (chunks cascade) +``` + +### Chunking policy (markdown-aware) +Split on `## `/`### ` headings into sections; sub-split any section longer +than `BOR_CHUNK_TARGET_CHARS` (2000) at paragraph boundaries with +`BOR_CHUNK_OVERLAP_CHARS` (200) overlap; each chunk keeps its nearest +preceding heading in the text for retrieval quality. + +--- + +## 6. RAG Pipeline & Persona + +### Locked system prompt (sent with every chat turn) +``` +You are "Brain of Reese" — the digital brain of Reese, a self-hoster and +homelab tinkerer. Personality: chippy, upbeat, warm, and genuinely +optimistic about the user's ability to do things ("you've got this"). + +Rules: +1. Answer ONLY from the provided document context. Cite which document(s) + you used, by path. +2. Be concrete: names, versions, ports, hosts, schedules — the specifics in + the docs are the value. +3. HONESTY GATE: if is "LOW", you must NOT pretend to know. + Start your answer with a variant of: "I haven't done anything like that." + Then offer 2-3 alternative questions about things you DO have notes on. +4. Never invent facts, hosts, or steps that are not in the context. +5. Keep answers tight: short paragraphs, bullets where helpful. + +{HIGH|LOW} +``` +- `HIGH` mode appends the full document text under `…`. +- `LOW` mode (deflection) appends only the **titles** of the weak hits so the + model can suggest real alternatives (marker used by the E2E mock: + `DEFLECT_MODE` appears in the system prompt). + +### Retrieval +- Embed the question (`embed`, 768-d) → `ORDER BY embedding <=> $1 LIMIT 4`. +- `score = 1 − cosine_distance`. Gate on `max(score) >= 0.30`. +- Distinct parent docs ranked by best chunk score → top 2 → full content, + concatenated, truncated to `BOR_MAX_CONTEXT_CHARS` (24k) with a + `[…truncated…]` marker. + +--- + +## 7. UI/UX Strategy + +### 7.1 Layout structure +- **App frame:** sticky header (64px) + `
` (flex-grow) + footer. + Container: `max-width: 72rem; margin-inline: auto; padding-inline: 1.25rem`. +- **Chat:** a *centered column capped at 46rem*. This is deliberate: chat is + a vertical conversation — a centered, capped column is the correct pattern + (NOT a layout bug). The 72rem frame + header/footer ensure the column + never reads as a hairline in a sea of whitespace. +- **Sources page:** full-width responsive **table** (min 640px, horizontal + scroll wrapper on small screens) + stat cards in + `grid-template-columns: repeat(auto-fit, minmax(170px, 1fr))`. + No skinny single-column lists anywhere: lists/tables/grids use ≥80–90% of + the container width. +- **Mobile (≤640px):** suggestion chips become a horizontally scrollable row; + composer stays reachable with `safe-area-inset-bottom`; touch targets ≥44px. + +### 7.2 Accessibility (WCAG 2.1 AA) +- Semantic landmarks on every page: `
`, `