Files
brain-of-reese/.agent/phases/complete/03_story_chat_rag.md
T

3.5 KiB
Raw Blame History

Phase 03 — Story: Chat RAG Answer (happy path)

Story: .agent/user_stories/chat-rag-answer.md Context: .agent/PLAN.md §3 (data flow), §4 (SSE contract), §6 (persona), §9 (logging)

Goal

The core product loop: question → embed → cosine top-4 → full top-2 documents → turbo (streamed) → chippy grounded answer with source chips.

Implementation steps

  1. app/rag/retriever.py — retrieve(db, question_embedding) -> list[RetrievedChunk] (score = 1 − distance, ORDER BY embedding <=> $1 LIMIT BOR_TOP_K_CHUNKS) + select_documents(chunks, n) -> list[Document] (distinct by document_id, ranked by best chunk score, cap content at BOR_MAX_CONTEXT_CHARS with […truncated…]).
  2. app/rag/prompts.py — locked persona + HONESTY GATE prompt builder (PLAN §6 verbatim, <relevance>HIGH|LOW</relevance>, <documents> block; LOW mode includes the DEFLECT_MODE marker + weak-hit titles).
  3. app/rag/llm.py — add chat_stream(messages) -> AsyncIterator[str] (openai async, stream=True, model=turbo, temperature 0.4, max_tokens ~700).
  4. app/api/chat.py — POST /api/chat (ChatRequest) → StreamingResponse (SSE): emit delta events from the stream, then the done event (deflected, sources, suggestions); insert query_log row (deflected= false this phase); per-turn log line (PLAN §9); structured error events ({"type":"error","detail":…}) on LLM/DB failure.
  5. frontend/assets/app.js — replace the placeholder handler: fetch + ReadableStream SSE parser; render deltas live into a brain bubble (reuse the typing-indicator → streaming handoff); on done, append .source-chips under the bubble; on error, show the banner (full state machine is Phase 06 — keep it simple-correct here).
  6. Tune settings.suggestions if the real Homelab import revealed better defaults (optional here; Phase 05 owns the chips).

UI Verification

Against the story's "UI Visualization & Structure": bubbles right/left (brand vs surface, ≥4.5:1 text), avatar 🧠, source chips mono/brand-soft with source/path and ellipsis, safe markdown (paste an answer containing <script>alert(1)</script> from the mock to prove it's escaped). Chat column 46rem centered. 1280px + 375px screenshot pass.

Testing & Quality

  • Unit: retriever ordering/dedup/cap (fake rows), prompt builder (HIGH contains documents + HIGH, LOW contains DEFLECT_MODE + titles only, persona rules present verbatim), SSE event serialization.
  • Integration: /api/chat against the mock LLM with a seeded temp schema — assert SSE delta sequence, done payload (sources non-empty, deflected false), query_log row, error event when LLM unreachable.
  • Coverage: uv run pytest --cov=app --cov-report=term-missing — >90%.

Playwright Execution Phase

Run ONLY this story's suite:

uv run pytest tests/e2e/test_chat_rag.py -v --no-cov

Implements the story mapping: streamed grounded answer + kubernetes.md source chip + button recovery; DB query_log assertion; raw SSE shape check via httpx.

Success criteria

  • end-to-end: question → streamed chippy answer citing kubernetes.md
  • query_log row per turn; per-turn log line in stdout
  • LLM-down path shows error banner, no stuck button
  • unit + integration green, coverage >90%
  • UI verification passed
  • story E2E green in isolation
  • committed

Commit

git add -A && git commit --no-gpg-sign -m "feat(rag): stream grounded chat answers via pgvector cosine retrieval with source citations"