Files
brain-of-reese/.agent/user_stories/whole-document-context.md
T

68 lines
3.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Story: Whole-Document Context — a matched document is never truncated
**Phase:** `24_whole_document_context` · **E2E:** `tests/e2e/test_whole_document_context.py`
## Bug report (TODO.md L3–L4, verbatim)
> *"Documents are truncated for some reason? This should never happen"*
>
> *"When the LLM matches a chunk it should get the entire document placed in its context so it can see the whole thing before answering the question"*
## Narrative
As **a user asking about an indexed note**, I want **the entire document
in the LLM's context** whenever retrieval matches one of its chunks — so
the answer is grounded in the whole note, not a silently cut-down version
of it.
- **Given** a question whose chunk matches a document
- **When** the turn assembles context
- **Then** the LLM receives the **entire** parent document — never a
`[…truncated…]`-cut version — so the answer is grounded in the whole
note.
## Owner-confirmed (2026-08-24, roadmap D1–D5)
1. **D1 — no cap at all (revises LOCKED A7):** `select_documents` returns
the full top-N document texts, always. The `max_context_chars` setting
and `BOR_MAX_CONTEXT_CHARS` env var are removed. If a future KB ever
makes the prompt too large for the model, the existing `LLMError` → SSE
`error` path surfaces it loudly — no silent partial context. The
emergency-valve variant (raised cap + warning log) was **explicitly
rejected**.
2. **D2 — `top_n_docs = 2` unchanged** (the TODO is about truncation, not
about how many documents).
3. **D3 — no viewer/import changes** — both already serve full content
(verified diagnosis above).
4. **D4 — E2E evidence via a deterministic mock tail-echo** (repo pattern,
cf. the phase-15 tuning-note echo); the big documents are seeded
directly in the DB inside the E2E test — `tests/fixtures/docs/` must
not grow, because other suites pin `summary.added == 8`.
5. **D5 — no `query_log` schema change** (no new columns, no migration).
## Acceptance criteria
1. No budget parameter in `select_documents` — it never truncates
(`TRUNCATION_MARKER` remains for the steering section only).
2. `BOR_MAX_CONTEXT_CHARS` gone from settings/env/README
(`app/config.py`, `.env.example`, `README.md`).
3. The story E2E's three tests green in isolation
(`uv run pytest tests/e2e/test_whole_document_context.py -v --no-cov`).
4. Steering-note truncation (phase 15, `BOR_STEERING_MAX_CHARS` + shared
marker) unchanged.
5. Unit + integration green, `app/` coverage >90%, one
`--no-gpg-sign` commit.
## Playwright Mapping Rule
**Test Scenario → `tests/e2e/test_whole_document_context.py`** (mock LLM
tail-echo, oversized documents seeded directly in the DB):
1. `test_whole_document_over_old_cap_reaches_llm` — a 30 000-char document
(past the old 24 000 cap): its tail sentinel (last line) is echoed in
the rendered answer, `[…truncated…]` is absent, the source chip renders,
`query_log` row `deflected == False`.
2. `test_second_document_of_over_cap_pair_reaches_llm` — two ~16 000-char
documents (32 000 combined — the exact case the old budget cut): the
second, lower-ranked document's tail sentinel is echoed (doc 1's is not
— it pins the rank order), both source chips render.
3. `test_small_document_path_unchanged` — regression: standard fixtures via
the real importer → grounded answer with the `kubernetes.md` chip, no
`[…truncated…]` (the under-cap path is byte-identical to before).