Files
brain-of-reese/.agent/phases/todo/02_story_import_documents.md
T
ducoterra 022da8e2bc feat: scaffold Brain of Reese — FastAPI RAG chat over Postgres 17 + pgvector
Foundation (phase 01, verified):
- FastAPI app: /api/health, /api/suggestions, /api/chat (placeholder),
  static frontend served locally (no CDN)
- Postgres 17 + pgvector via db/Containerfile + compose.yaml
  (podman compose up -d db), Alembic initial migration (documents,
  chunks with vector(768), query_log)
- LLM client targeting https://aipi.reeseapps.com/v1 (turbo/embed);
  scripts/llm_probe.py verified models + 768-dim embeddings live
- Conditional debugpy: imported only when DEBUGPY=1 (attach on demand,
  :5678); logging config for clean single-line logs
- Frontend shell: mobile-first chat + Sources pages, tokens, a11y baselines
- Tests: 24 unit+integration (99% coverage on app/), ruff + pyright clean,
  Playwright smoke E2E (3 tests) against a deterministic mock LLM
- Planning: .agent/PLAN.md (architecture + LOCKED decisions), AGENTS.md,
  6 user stories, 7 phase files (one story / one phase / one Playwright
  suite each)
2026-08-21 13:42:21 -04:00

3.7 KiB

Phase 02 — Story: Import Documents

Story: .agent/user_stories/import-documents.md Context: .agent/PLAN.md §5 (data model), §9 (logging), §11 (import workflow)

Goal

The importer (scripts/import_docs.py) + GET /api/docs + the Sources page rendering the indexed documents — the knowledge base becomes refreshable.

Implementation steps

  1. app/rag/__init__.py, app/rag/chunker.py — markdown-aware chunker (PLAN §5 policy: heading splits, 2000-char target, 200 overlap, keep nearest heading). Pure functions, fully unit-testable.
  2. app/rag/llm.py — LLMClient (openai async) with embed(texts) -> list[list[float]] (batched, BOR_EMBED_BATCH_SIZE) and a embed_one; dimension check vs settings.embedding_dim with a loud, actionable error. (Chat streaming is added in Phase 03 on this client.)
  3. app/rag/importer.py — the core: directory walk (exclusion list, PLAN A9; *.md only), sha256 delta vs documents.content_hash, upsert-or-skip, two-phase chunk replace (insert doc → replace chunks → embed → commit), --prune support, per-file + summary logging.
  4. scripts/import_docs.py — CLI wrapper (argparse): repeatable --source (default ~/Homelab ~/Deployments, expanduser), --prune, --limit.
  5. app/api/docs.py — GET /api/docs → {"documents": [DocSummary]} (include chunks count via func.count); mount in app/main.py before the static mount.
  6. frontend/assets/sources.js + sources.html polish — wire the real endpoint (already scaffolded to expect this shape); keep the empty state.
  7. Update README.md §Knowledge Base Import with the final commands + exclusion list + "update your docs → re-run the script" workflow.

UI Verification

Compare /sources.html against the story's "UI Visualization & Structure": stat cards auto-fit minmax(170px,1fr); full-width table (≥85% container); mono path column with title ellipsis; empty state with the exact command; <caption class="visually-hidden">, scope="col", scroll wrapper role="region" tabindex="0". No CDN refs. Take a 1280px and 375px screenshot pass before finishing.

Testing & Quality

  • Unit: chunker (heading splits, overlap, short-doc single chunk, code fences kept intact), exclusion walk (temp tree with .venv junk), delta logic (unchanged/changed/pruned via tmp Postgres or in-memory fakes — real DB preferred since compose runs locally).
  • Integration: GET /api/docs empty shape + populated shape; importer end-to-end against tests/fixtures/docs/ into a test schema.
  • Coverage: uv run pytest --cov=app --cov-report=term-missing — >90% on app/ (importer + chunker + client are the bulk; test them hard).

Playwright Execution Phase

Run ONLY this story's suite (DB must be up: podman compose up -d db):

uv run pytest tests/e2e/test_import_documents.py -v --no-cov

The test file implements the story's Playwright Mapping Rule (seed via the import function against tests/fixtures/docs/ with the mock LLM; assert Sources page rows, layout width, and the empty state).

Success criteria

  • uv run python -m scripts.import_docs (fixtures) imports all 3 docs, re-run reports unchanged
  • GET /api/docs + Sources page show the docs (real run: ~/Homelab + ~/Deployments counts logged)
  • unit + integration green, coverage >90%
  • UI verification passed (screenshots attached to the phase record)
  • story E2E green in isolation
  • README import section updated
  • committed

Commit

git add -A && git commit --no-gpg-sign -m "feat(kb): markdown importer with sha256 deltas, chunking, batched embeddings, and Sources page"