phase: 118_summary_seed_context
**Phase 118 final verification pass — complete.** All criteria verified; 4 pre-existing defects found and fixed.
- **Verified:** summary-seed wiring (`select_suggested` top-5 no-floor → summary blocks, no full text in HIGH prompt), all-doc markdown summaries + NULL backfill (`summary_backfilled`, no `sources_meta` bump), `read` adds full text with `read_docs`-only dedupe, `done.sources` = suggested+read / durable record = suggested+related+read + `suggested=N` log line (seen live in E2E), byte-locked PERSONA/LOW/TOOLS_SECTION, battery gate PASS recorded in `TOOL_CALLING_TESTING.md` §10 (turbo 2026-09-16: 1/2/4 GREEN, cond-3 reported 9/10 per A7, contract 21/21, caps 0).
- **Defects fixed (all pre-existing, none phase-118):** ① `ChatMessage` schema missing the phase-113 `related` key → `extra="forbid"` 422'd every done-time auto-save of grounded turns with a related tier, leaving `message_count=1` (root cause of `test_share_chat` 3F; browser-level instrumentation proved the PUT 422) — added the field + unit/integration pins; ② `test_theme_semantic_completion` pins stale vs phase-117 debox (border/chip removed) — re-targeted to assert border/chip *absence*; ③ `test_header_consistency` `<26`px pin red on 26.125px native date-input line — bound relaxed to `<34` (wrap-detection intent kept); ④ `test_navbar_refresh` bor.chat.v1 key set updated for `related`.
- **Test/lint/coverage:** `uv run pytest --cov=app --cov-report=term-missing` → **2506 passed, app/ 99%** (>90%); `uv run ruff check . && uv run pyright` → clean, 0 errors.
- **E2E:** new story suite in isolation → **2 passed**; full 103-suite matrix sweep (each isolated) → **all 103 green** after the fixes; `test_share_chat` 4 passed, `test_theme_semantic_completion` 8 passed, `test_header_consistency` 3 passed, `test_navbar_refresh` 7 passed.
- **Deviations:** none from LOCKED decisions. Note: orphaned diagnostic uvicorn processes briefly made E2E sessions exercise stale code — killed and re-verified; a sweep-regenerated tracked screenshot was restored. No commits made (harness commits).
- **Completion criteria:** all 7 ✅ (commit/phase-move is the harness's step).
- **Next pending phase:** none — `todo/` holds only this phase's overview pending the harness move.
This commit is contained in:
@@ -0,0 +1,9 @@
|
||||
**Phase 118 final verification pass — complete.** All criteria verified; 4 pre-existing defects found and fixed.
|
||||
|
||||
- **Verified:** summary-seed wiring (`select_suggested` top-5 no-floor → summary blocks, no full text in HIGH prompt), all-doc markdown summaries + NULL backfill (`summary_backfilled`, no `sources_meta` bump), `read` adds full text with `read_docs`-only dedupe, `done.sources` = suggested+read / durable record = suggested+related+read + `suggested=N` log line (seen live in E2E), byte-locked PERSONA/LOW/TOOLS_SECTION, battery gate PASS recorded in `TOOL_CALLING_TESTING.md` §10 (turbo 2026-09-16: 1/2/4 GREEN, cond-3 reported 9/10 per A7, contract 21/21, caps 0).
|
||||
- **Defects fixed (all pre-existing, none phase-118):** ① `ChatMessage` schema missing the phase-113 `related` key → `extra="forbid"` 422'd every done-time auto-save of grounded turns with a related tier, leaving `message_count=1` (root cause of `test_share_chat` 3F; browser-level instrumentation proved the PUT 422) — added the field + unit/integration pins; ② `test_theme_semantic_completion` pins stale vs phase-117 debox (border/chip removed) — re-targeted to assert border/chip *absence*; ③ `test_header_consistency` `<26`px pin red on 26.125px native date-input line — bound relaxed to `<34` (wrap-detection intent kept); ④ `test_navbar_refresh` bor.chat.v1 key set updated for `related`.
|
||||
- **Test/lint/coverage:** `uv run pytest --cov=app --cov-report=term-missing` → **2506 passed, app/ 99%** (>90%); `uv run ruff check . && uv run pyright` → clean, 0 errors.
|
||||
- **E2E:** new story suite in isolation → **2 passed**; full 103-suite matrix sweep (each isolated) → **all 103 green** after the fixes; `test_share_chat` 4 passed, `test_theme_semantic_completion` 8 passed, `test_header_consistency` 3 passed, `test_navbar_refresh` 7 passed.
|
||||
- **Deviations:** none from LOCKED decisions. Note: orphaned diagnostic uvicorn processes briefly made E2E sessions exercise stale code — killed and re-verified; a sweep-regenerated tracked screenshot was restored. No commits made (harness commits).
|
||||
- **Completion criteria:** all 7 ✅ (commit/phase-move is the harness's step).
|
||||
- **Next pending phase:** none — `todo/` holds only this phase's overview pending the harness move.
|
||||
+103
@@ -0,0 +1,103 @@
|
||||
........................................................................ [ 2%]
|
||||
........................................................................ [ 5%]
|
||||
........................................................................ [ 8%]
|
||||
........................................................................ [ 11%]
|
||||
........................................................................ [ 14%]
|
||||
........................................................................ [ 17%]
|
||||
........................................................................ [ 20%]
|
||||
........................................................................ [ 22%]
|
||||
........................................................................ [ 25%]
|
||||
........................................................................ [ 28%]
|
||||
........................................................................ [ 31%]
|
||||
........................................................................ [ 34%]
|
||||
........................................................................ [ 37%]
|
||||
........................................................................ [ 40%]
|
||||
........................................................................ [ 43%]
|
||||
........................................................................ [ 45%]
|
||||
........................................................................ [ 48%]
|
||||
........................................................................ [ 51%]
|
||||
........................................................................ [ 54%]
|
||||
........................................................................ [ 57%]
|
||||
........................................................................ [ 60%]
|
||||
........................................................................ [ 63%]
|
||||
........................................................................ [ 66%]
|
||||
........................................................................ [ 68%]
|
||||
........................................................................ [ 71%]
|
||||
........................................................................ [ 74%]
|
||||
........................................................................ [ 77%]
|
||||
........................................................................ [ 80%]
|
||||
........................................................................ [ 83%]
|
||||
........................................................................ [ 86%]
|
||||
........................................................................ [ 89%]
|
||||
........................................................................ [ 91%]
|
||||
........................................................................ [ 94%]
|
||||
........................................................................ [ 97%]
|
||||
.......................................................... [100%]
|
||||
=============================== warnings summary ===============================
|
||||
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
|
||||
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||
from starlette.testclient import TestClient as TestClient # noqa
|
||||
|
||||
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||
================================ tests coverage ================================
|
||||
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
|
||||
|
||||
Name Stmts Miss Cover
|
||||
--------------------------------------------------
|
||||
app/__init__.py 1 0 100%
|
||||
app/api/__init__.py 0 0 100%
|
||||
app/api/auth.py 52 0 100%
|
||||
app/api/chat.py 225 1 99%
|
||||
app/api/chats.py 110 0 100%
|
||||
app/api/config.py 13 0 100%
|
||||
app/api/doc_drafts.py 99 0 100%
|
||||
app/api/docs.py 156 1 99%
|
||||
app/api/git_sources.py 232 0 100%
|
||||
app/api/health.py 10 0 100%
|
||||
app/api/steering.py 42 0 100%
|
||||
app/api/suggestions.py 33 0 100%
|
||||
app/api/sync.py 139 0 100%
|
||||
app/api/tokens.py 40 0 100%
|
||||
app/api/ui_settings.py 55 0 100%
|
||||
app/config.py 224 0 100%
|
||||
app/core/__init__.py 0 0 100%
|
||||
app/core/auth.py 45 0 100%
|
||||
app/core/caching.py 124 0 100%
|
||||
app/core/debugging.py 29 2 93%
|
||||
app/core/docs_push.py 39 0 100%
|
||||
app/core/errors.py 5 0 100%
|
||||
app/core/logging.py 13 0 100%
|
||||
app/core/rate_limit.py 44 0 100%
|
||||
app/core/security_headers.py 20 0 100%
|
||||
app/core/theming.py 38 0 100%
|
||||
app/core/tokens.py 44 0 100%
|
||||
app/db.py 22 0 100%
|
||||
app/main.py 66 0 100%
|
||||
app/models.py 128 0 100%
|
||||
app/rag/__init__.py 0 0 100%
|
||||
app/rag/agent.py 317 1 99%
|
||||
app/rag/archive_upload.py 134 0 100%
|
||||
app/rag/chunker.py 206 4 98%
|
||||
app/rag/doc_dates.py 18 0 100%
|
||||
app/rag/folder_summaries.py 123 0 100%
|
||||
app/rag/git_sources.py 14 0 100%
|
||||
app/rag/importer.py 219 3 99%
|
||||
app/rag/llm.py 244 1 99%
|
||||
app/rag/overview.py 71 0 100%
|
||||
app/rag/prompts.py 100 0 100%
|
||||
app/rag/retriever.py 217 3 99%
|
||||
app/rag/scaffolding.py 55 0 100%
|
||||
app/rag/source_removal.py 41 0 100%
|
||||
app/rag/sources_meta.py 16 0 100%
|
||||
app/rag/suggestions.py 27 0 100%
|
||||
app/rag/summarizer.py 24 0 100%
|
||||
app/schemas.py 330 0 100%
|
||||
--------------------------------------------------
|
||||
TOTAL 4204 16 99%
|
||||
coverage gate: app/ 99% (>90%) OK
|
||||
All checks passed!
|
||||
0 errors, 0 warnings, 0 informations
|
||||
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||
|
||||
validation OK
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
All gates green. Task 01 is complete.
|
||||
|
||||
**Implemented**
|
||||
- `app/rag/importer.py`: markdown exclusion removed — every new/changed A9 doc gets a `lite` summary (locked A2); UNCHANGED branch backfills NULL-summary docs via `_store_summary(backfill=True)` before the `created_at_manual` early-return/date refresh (strict `is None` — owner-set summaries untouched); `ImportSummary.summary_backfilled` counter + `summary_backfilled=%d` in the PLAN §9 log line after `summary_errors`; fail-soft rollback + `summary_errors` unchanged; no added/updated/pruned count (no `sources_meta` bump)
|
||||
- Docstrings updated for the retired "non-markdown only" wording: `models.py`, `schemas.py`, `rag/llm.py`, `rag/overview.py`
|
||||
- Unit tests: markdown test inverted (`test_markdown_file_gets_stored_summary`) + 5 new backfill tests (success, no-resummarize, empty-string never backfilled, manual-date, fail-soft with date still refreshing) + log-line pin
|
||||
- Re-pinned 7 integration suites whose old-contract pins my change invalidated (`test_importer_e2e`, `test_import_extensions_env`, `test_import_extensionless`, `test_importer_ignore`, `test_importer_include_hidden`, `test_import_docs_overview`, `test_sync_folder_summaries`)
|
||||
|
||||
**Gates**
|
||||
- `uv run pytest tests/unit/test_importer.py -v --no-cov` → 35 passed
|
||||
- `uv run pytest --cov=app --cov-report=term-missing` → 2471 passed, **99%** on `app/` (>90%)
|
||||
- `uv run ruff check . && uv run pyright` → clean (0 errors)
|
||||
|
||||
**Decisions**
|
||||
- Sync-API result dict and the `scripts/import_docs.py` CLI line left unchanged (task mandates only the `log()` line + dataclass counter; the sync result is a pinned API surface)
|
||||
- E2E pins of the retired behavior (`test_document_summaries.py` etc.) deliberately left for task 06 per the phase design
|
||||
|
||||
**Next pending task:** `02_suggested_tier_selection.md`
|
||||
+103
@@ -0,0 +1,103 @@
|
||||
........................................................................ [ 2%]
|
||||
........................................................................ [ 5%]
|
||||
........................................................................ [ 8%]
|
||||
........................................................................ [ 11%]
|
||||
........................................................................ [ 14%]
|
||||
........................................................................ [ 17%]
|
||||
........................................................................ [ 20%]
|
||||
........................................................................ [ 23%]
|
||||
........................................................................ [ 26%]
|
||||
........................................................................ [ 29%]
|
||||
........................................................................ [ 32%]
|
||||
........................................................................ [ 34%]
|
||||
........................................................................ [ 37%]
|
||||
........................................................................ [ 40%]
|
||||
........................................................................ [ 43%]
|
||||
........................................................................ [ 46%]
|
||||
........................................................................ [ 49%]
|
||||
........................................................................ [ 52%]
|
||||
........................................................................ [ 55%]
|
||||
........................................................................ [ 58%]
|
||||
........................................................................ [ 61%]
|
||||
........................................................................ [ 64%]
|
||||
........................................................................ [ 67%]
|
||||
........................................................................ [ 69%]
|
||||
........................................................................ [ 72%]
|
||||
........................................................................ [ 75%]
|
||||
........................................................................ [ 78%]
|
||||
........................................................................ [ 81%]
|
||||
........................................................................ [ 84%]
|
||||
........................................................................ [ 87%]
|
||||
........................................................................ [ 90%]
|
||||
........................................................................ [ 93%]
|
||||
........................................................................ [ 96%]
|
||||
........................................................................ [ 99%]
|
||||
....................... [100%]
|
||||
=============================== warnings summary ===============================
|
||||
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
|
||||
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||
from starlette.testclient import TestClient as TestClient # noqa
|
||||
|
||||
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||
================================ tests coverage ================================
|
||||
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
|
||||
|
||||
Name Stmts Miss Cover
|
||||
--------------------------------------------------
|
||||
app/__init__.py 1 0 100%
|
||||
app/api/__init__.py 0 0 100%
|
||||
app/api/auth.py 52 0 100%
|
||||
app/api/chat.py 224 1 99%
|
||||
app/api/chats.py 110 0 100%
|
||||
app/api/config.py 13 0 100%
|
||||
app/api/doc_drafts.py 99 0 100%
|
||||
app/api/docs.py 156 1 99%
|
||||
app/api/git_sources.py 232 0 100%
|
||||
app/api/health.py 10 0 100%
|
||||
app/api/steering.py 42 0 100%
|
||||
app/api/suggestions.py 33 0 100%
|
||||
app/api/sync.py 139 0 100%
|
||||
app/api/tokens.py 40 0 100%
|
||||
app/api/ui_settings.py 55 0 100%
|
||||
app/config.py 210 0 100%
|
||||
app/core/__init__.py 0 0 100%
|
||||
app/core/auth.py 45 0 100%
|
||||
app/core/caching.py 124 0 100%
|
||||
app/core/debugging.py 29 2 93%
|
||||
app/core/docs_push.py 39 0 100%
|
||||
app/core/errors.py 5 0 100%
|
||||
app/core/logging.py 13 0 100%
|
||||
app/core/rate_limit.py 44 0 100%
|
||||
app/core/security_headers.py 20 0 100%
|
||||
app/core/theming.py 38 0 100%
|
||||
app/core/tokens.py 44 0 100%
|
||||
app/db.py 22 0 100%
|
||||
app/main.py 66 0 100%
|
||||
app/models.py 128 0 100%
|
||||
app/rag/__init__.py 0 0 100%
|
||||
app/rag/agent.py 317 1 99%
|
||||
app/rag/archive_upload.py 134 0 100%
|
||||
app/rag/chunker.py 206 4 98%
|
||||
app/rag/doc_dates.py 18 0 100%
|
||||
app/rag/folder_summaries.py 123 0 100%
|
||||
app/rag/git_sources.py 14 0 100%
|
||||
app/rag/importer.py 219 3 99%
|
||||
app/rag/llm.py 244 1 99%
|
||||
app/rag/overview.py 71 0 100%
|
||||
app/rag/prompts.py 88 0 100%
|
||||
app/rag/retriever.py 192 3 98%
|
||||
app/rag/scaffolding.py 55 0 100%
|
||||
app/rag/source_removal.py 41 0 100%
|
||||
app/rag/sources_meta.py 16 0 100%
|
||||
app/rag/suggestions.py 27 0 100%
|
||||
app/rag/summarizer.py 24 0 100%
|
||||
app/schemas.py 329 0 100%
|
||||
--------------------------------------------------
|
||||
TOTAL 4151 16 99%
|
||||
coverage gate: app/ 99% (>90%) OK
|
||||
All checks passed!
|
||||
0 errors, 0 warnings, 0 informations
|
||||
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||
|
||||
validation OK
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
Task 02 complete — all gates green.
|
||||
|
||||
**Implemented**
|
||||
- `app/rag/retriever.py`: pure `select_suggested(chunks, n=None)` — top-N distinct parent docs in the same stable score-descending walk as `select_documents_tiered` (rank fixed by first-seen chunk, dedupe by `document.id`), **no cosine floor** (A3), `n` defaults to `get_settings().suggested_docs`; rows carry full byte-identical content (A6).
|
||||
- `app/config.py`: `suggested_docs: int = 5` (`BOR_SUGGESTED_DOCS`, validator `>= 1`, `agent_max_rounds` pattern) + `.env.example` entry in the RAG-tuning block.
|
||||
- `tests/unit/test_config.py`: default 5 + env override; 0/−2 rejected, 1 accepted.
|
||||
- `tests/unit/test_retriever.py`: first-seen-chunk rank order, dedupe, cap (explicit `n` + live settings default via monkeypatched `get_settings`), no-floor contrast pin vs `select_documents_tiered`'s floored cited tier on the same input, deterministic tie-break inherited from fused order, empty + full-content pins.
|
||||
|
||||
**Results**
|
||||
- `uv run pytest tests/unit/test_retriever.py tests/unit/test_config.py -v` → 104 passed
|
||||
- `uv run pytest --cov=app --cov-report=term-missing` → 2480 passed, **99%** total (`app/config.py` 100%; `retriever.py` missing lines are pre-existing SQL paths)
|
||||
- `uv run ruff check .` → All checks passed; `uv run pyright` → 0 errors
|
||||
|
||||
**Decisions** — early-exit cap in the walk (safe: no floor/related tier to compute); settings-default test pins the live value, not a frozen constant. No defects found in prior work; no git operations (left for the harness).
|
||||
|
||||
**Next pending task:** `.agents/phases/todo/118_summary_seed_context/03_summary_document_blocks.md`
|
||||
+103
@@ -0,0 +1,103 @@
|
||||
........................................................................ [ 2%]
|
||||
........................................................................ [ 5%]
|
||||
........................................................................ [ 8%]
|
||||
........................................................................ [ 11%]
|
||||
........................................................................ [ 14%]
|
||||
........................................................................ [ 17%]
|
||||
........................................................................ [ 20%]
|
||||
........................................................................ [ 23%]
|
||||
........................................................................ [ 26%]
|
||||
........................................................................ [ 29%]
|
||||
........................................................................ [ 31%]
|
||||
........................................................................ [ 34%]
|
||||
........................................................................ [ 37%]
|
||||
........................................................................ [ 40%]
|
||||
........................................................................ [ 43%]
|
||||
........................................................................ [ 46%]
|
||||
........................................................................ [ 49%]
|
||||
........................................................................ [ 52%]
|
||||
........................................................................ [ 55%]
|
||||
........................................................................ [ 58%]
|
||||
........................................................................ [ 60%]
|
||||
........................................................................ [ 63%]
|
||||
........................................................................ [ 66%]
|
||||
........................................................................ [ 69%]
|
||||
........................................................................ [ 72%]
|
||||
........................................................................ [ 75%]
|
||||
........................................................................ [ 78%]
|
||||
........................................................................ [ 81%]
|
||||
........................................................................ [ 84%]
|
||||
........................................................................ [ 87%]
|
||||
........................................................................ [ 90%]
|
||||
........................................................................ [ 92%]
|
||||
........................................................................ [ 95%]
|
||||
........................................................................ [ 98%]
|
||||
................................ [100%]
|
||||
=============================== warnings summary ===============================
|
||||
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
|
||||
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||
from starlette.testclient import TestClient as TestClient # noqa
|
||||
|
||||
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||
================================ tests coverage ================================
|
||||
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
|
||||
|
||||
Name Stmts Miss Cover
|
||||
--------------------------------------------------
|
||||
app/__init__.py 1 0 100%
|
||||
app/api/__init__.py 0 0 100%
|
||||
app/api/auth.py 52 0 100%
|
||||
app/api/chat.py 224 1 99%
|
||||
app/api/chats.py 110 0 100%
|
||||
app/api/config.py 13 0 100%
|
||||
app/api/doc_drafts.py 99 0 100%
|
||||
app/api/docs.py 156 1 99%
|
||||
app/api/git_sources.py 232 0 100%
|
||||
app/api/health.py 10 0 100%
|
||||
app/api/steering.py 42 0 100%
|
||||
app/api/suggestions.py 33 0 100%
|
||||
app/api/sync.py 139 0 100%
|
||||
app/api/tokens.py 40 0 100%
|
||||
app/api/ui_settings.py 55 0 100%
|
||||
app/config.py 217 0 100%
|
||||
app/core/__init__.py 0 0 100%
|
||||
app/core/auth.py 45 0 100%
|
||||
app/core/caching.py 124 0 100%
|
||||
app/core/debugging.py 29 2 93%
|
||||
app/core/docs_push.py 39 0 100%
|
||||
app/core/errors.py 5 0 100%
|
||||
app/core/logging.py 13 0 100%
|
||||
app/core/rate_limit.py 44 0 100%
|
||||
app/core/security_headers.py 20 0 100%
|
||||
app/core/theming.py 38 0 100%
|
||||
app/core/tokens.py 44 0 100%
|
||||
app/db.py 22 0 100%
|
||||
app/main.py 66 0 100%
|
||||
app/models.py 128 0 100%
|
||||
app/rag/__init__.py 0 0 100%
|
||||
app/rag/agent.py 317 1 99%
|
||||
app/rag/archive_upload.py 134 0 100%
|
||||
app/rag/chunker.py 206 4 98%
|
||||
app/rag/doc_dates.py 18 0 100%
|
||||
app/rag/folder_summaries.py 123 0 100%
|
||||
app/rag/git_sources.py 14 0 100%
|
||||
app/rag/importer.py 219 3 99%
|
||||
app/rag/llm.py 244 1 99%
|
||||
app/rag/overview.py 71 0 100%
|
||||
app/rag/prompts.py 88 0 100%
|
||||
app/rag/retriever.py 205 3 99%
|
||||
app/rag/scaffolding.py 55 0 100%
|
||||
app/rag/source_removal.py 41 0 100%
|
||||
app/rag/sources_meta.py 16 0 100%
|
||||
app/rag/suggestions.py 27 0 100%
|
||||
app/rag/summarizer.py 24 0 100%
|
||||
app/schemas.py 329 0 100%
|
||||
--------------------------------------------------
|
||||
TOTAL 4171 16 99%
|
||||
coverage gate: app/ 99% (>90%) OK
|
||||
All checks passed!
|
||||
0 errors, 0 warnings, 0 informations
|
||||
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||
|
||||
validation OK
|
||||
+19
@@ -0,0 +1,19 @@
|
||||
All gates verified. Task complete.
|
||||
|
||||
**Task 03 — "start here" suggestion blocks in the HIGH prompt**
|
||||
|
||||
Implemented:
|
||||
- `app/rag/prompts.py`: `build_high_prompt` now seeds each `<document>` block with the doc's **summary** (markup `source/path/title/date` byte-stable); new pinned `SUGGEST_INTRO` start-here line before the first block; `_document_body()` A5 fallback — NULL/whitespace summary → first `suggestion_preview_chars` chars + `TRUNCATION_MARKER` (marker only when content was actually cut; settings read on fallback path only)
|
||||
- `app/config.py` + `.env.example`: `suggestion_preview_chars` (default 400, env `BOR_SUGGESTION_PREVIEW_CHARS`, validator `> 0`)
|
||||
- `app/rag/retriever.py`: module docstring + `TRUNCATION_MARKER` comment + `select_documents`/`select_documents_tiered` docstrings re-pointed to the A6 re-revised contract (functions untouched)
|
||||
- Tests: summary-seed sentinels (5-doc pin, zero full-content), preview-fallback matrix (None/whitespace/short/exact-cap/cap-override/no-settings-read), `date`-attribute pin on summary bodies, `SUGGEST_INTRO` sha+prefix lock, LOW-prompt build sha pin (byte-identical); updated the one retired full-text pin in `tests/integration/test_chat_api.py` to pin the summary contract
|
||||
|
||||
Gates:
|
||||
- `uv run pytest tests/unit/test_prompts.py tests/unit/test_prompts_dates.py tests/unit/test_prompt_lock.py -v --no-cov` → 59 passed
|
||||
- `uv run pytest` (unit+integration) → exit 0, 0 failures
|
||||
- `uv run pytest --cov=app --cov-report=term-missing` → TOTAL 99% (>90%)
|
||||
- `uv run ruff check .` → clean; `uv run pyright` → 0 errors
|
||||
|
||||
Decisions: empty `<documents>` section keeps the fallback line with **no** intro (intro rides only on present blocks); LOW/`DEFLECT_MODE`/`PERSONA`/`TOOLS_SECTION` untouched.
|
||||
|
||||
Next pending task: `.agents/phases/todo/118_summary_seed_context/04_read_suggested_docs.md`
|
||||
+103
@@ -0,0 +1,103 @@
|
||||
........................................................................ [ 2%]
|
||||
........................................................................ [ 5%]
|
||||
........................................................................ [ 8%]
|
||||
........................................................................ [ 11%]
|
||||
........................................................................ [ 14%]
|
||||
........................................................................ [ 17%]
|
||||
........................................................................ [ 20%]
|
||||
........................................................................ [ 23%]
|
||||
........................................................................ [ 25%]
|
||||
........................................................................ [ 28%]
|
||||
........................................................................ [ 31%]
|
||||
........................................................................ [ 34%]
|
||||
........................................................................ [ 37%]
|
||||
........................................................................ [ 40%]
|
||||
........................................................................ [ 43%]
|
||||
........................................................................ [ 46%]
|
||||
........................................................................ [ 49%]
|
||||
........................................................................ [ 51%]
|
||||
........................................................................ [ 54%]
|
||||
........................................................................ [ 57%]
|
||||
........................................................................ [ 60%]
|
||||
........................................................................ [ 63%]
|
||||
........................................................................ [ 66%]
|
||||
........................................................................ [ 69%]
|
||||
........................................................................ [ 72%]
|
||||
........................................................................ [ 75%]
|
||||
........................................................................ [ 77%]
|
||||
........................................................................ [ 80%]
|
||||
........................................................................ [ 83%]
|
||||
........................................................................ [ 86%]
|
||||
........................................................................ [ 89%]
|
||||
........................................................................ [ 92%]
|
||||
........................................................................ [ 95%]
|
||||
........................................................................ [ 98%]
|
||||
.............................................. [100%]
|
||||
=============================== warnings summary ===============================
|
||||
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
|
||||
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||
from starlette.testclient import TestClient as TestClient # noqa
|
||||
|
||||
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||
================================ tests coverage ================================
|
||||
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
|
||||
|
||||
Name Stmts Miss Cover
|
||||
--------------------------------------------------
|
||||
app/__init__.py 1 0 100%
|
||||
app/api/__init__.py 0 0 100%
|
||||
app/api/auth.py 52 0 100%
|
||||
app/api/chat.py 224 1 99%
|
||||
app/api/chats.py 110 0 100%
|
||||
app/api/config.py 13 0 100%
|
||||
app/api/doc_drafts.py 99 0 100%
|
||||
app/api/docs.py 156 1 99%
|
||||
app/api/git_sources.py 232 0 100%
|
||||
app/api/health.py 10 0 100%
|
||||
app/api/steering.py 42 0 100%
|
||||
app/api/suggestions.py 33 0 100%
|
||||
app/api/sync.py 139 0 100%
|
||||
app/api/tokens.py 40 0 100%
|
||||
app/api/ui_settings.py 55 0 100%
|
||||
app/config.py 224 0 100%
|
||||
app/core/__init__.py 0 0 100%
|
||||
app/core/auth.py 45 0 100%
|
||||
app/core/caching.py 124 0 100%
|
||||
app/core/debugging.py 29 2 93%
|
||||
app/core/docs_push.py 39 0 100%
|
||||
app/core/errors.py 5 0 100%
|
||||
app/core/logging.py 13 0 100%
|
||||
app/core/rate_limit.py 44 0 100%
|
||||
app/core/security_headers.py 20 0 100%
|
||||
app/core/theming.py 38 0 100%
|
||||
app/core/tokens.py 44 0 100%
|
||||
app/db.py 22 0 100%
|
||||
app/main.py 66 0 100%
|
||||
app/models.py 128 0 100%
|
||||
app/rag/__init__.py 0 0 100%
|
||||
app/rag/agent.py 317 1 99%
|
||||
app/rag/archive_upload.py 134 0 100%
|
||||
app/rag/chunker.py 206 4 98%
|
||||
app/rag/doc_dates.py 18 0 100%
|
||||
app/rag/folder_summaries.py 123 0 100%
|
||||
app/rag/git_sources.py 14 0 100%
|
||||
app/rag/importer.py 219 3 99%
|
||||
app/rag/llm.py 244 1 99%
|
||||
app/rag/overview.py 71 0 100%
|
||||
app/rag/prompts.py 100 0 100%
|
||||
app/rag/retriever.py 205 3 99%
|
||||
app/rag/scaffolding.py 55 0 100%
|
||||
app/rag/source_removal.py 41 0 100%
|
||||
app/rag/sources_meta.py 16 0 100%
|
||||
app/rag/suggestions.py 27 0 100%
|
||||
app/rag/summarizer.py 24 0 100%
|
||||
app/schemas.py 329 0 100%
|
||||
--------------------------------------------------
|
||||
TOTAL 4190 16 99%
|
||||
coverage gate: app/ 99% (>90%) OK
|
||||
All checks passed!
|
||||
0 errors, 0 warnings, 0 informations
|
||||
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||
|
||||
validation OK
|
||||
+17
@@ -0,0 +1,17 @@
|
||||
All green. Final report:
|
||||
|
||||
**Task 04 complete — `read` on a suggested doc adds its full text; copy rewritten; lock re-cut**
|
||||
|
||||
- `app/rag/agent.py`: `_execute_tool` read dedupe set → `holder.read_docs` only (seeds drop out; first read of a suggested doc succeeds through the unchanged path — cap/marker/notice/date-line intact); `ALREADY_IN_CONTEXT` constant byte-identical; `AGENT_TOOLS` `read` description + `path` param rewritten (identity, bare-path, truncation-notice, one-call sentences byte-preserved); module docstring A7-clarification → A6 re-revised contract; point-3 + `run_agent` `seed_docs` docstring paragraphs rewritten (parameters kept, callers unchanged)
|
||||
- `app/rag/prompts.py`: `TOOLS_SECTION` `read` clause rewritten for summary-seed mode; `ls`/`grep` clauses + discipline rules byte-identical; comment block updated
|
||||
- `tests/unit/test_prompt_lock.py`: `TOOLS_SECTION_SHA256`/`LEN` (2465) re-cut, comment cites phase-118 (A6, 2026-09-15); PREFIX/SUFFIX verified to survive; `PERSONA`/LOW-body anchors untouched and passing
|
||||
- `tests/unit/test_agent.py`: suggested-read success + re-read refusal (no DB lookup, counters untouched) + over-cap suggested doc truncation re-target + description pins; `tests/unit/test_prompts.py`: new phase-118 read-clause pin (new copy, byte-pinned ls/grep/discipline, old copy gone, HIGH-vs-LOW)
|
||||
|
||||
**Results**
|
||||
- `uv run pytest tests/unit/test_agent.py tests/unit/test_prompt_lock.py tests/unit/test_prompts.py -v` → 161 passed
|
||||
- `uv run pytest --no-cov` → 2496 passed; `uv run pytest --cov=app` → 2496 passed, app coverage 99% (>90%)
|
||||
- `uv run ruff check .` → clean; `uv run pyright` → 0 errors, 0 warnings
|
||||
|
||||
**Decisions:** E2E mock + old pins (`mock_llm.py`, `test_ls_tree_drilldown.py`, `test_read_truncation_cap.py`) intentionally untouched — task 06's explicit scope (not in the default unit+integration suite).
|
||||
|
||||
Next pending task: `.agents/phases/todo/118_summary_seed_context/05_chat_turn_wiring.md`
|
||||
+103
@@ -0,0 +1,103 @@
|
||||
........................................................................ [ 2%]
|
||||
........................................................................ [ 5%]
|
||||
........................................................................ [ 8%]
|
||||
........................................................................ [ 11%]
|
||||
........................................................................ [ 14%]
|
||||
........................................................................ [ 17%]
|
||||
........................................................................ [ 20%]
|
||||
........................................................................ [ 23%]
|
||||
........................................................................ [ 25%]
|
||||
........................................................................ [ 28%]
|
||||
........................................................................ [ 31%]
|
||||
........................................................................ [ 34%]
|
||||
........................................................................ [ 37%]
|
||||
........................................................................ [ 40%]
|
||||
........................................................................ [ 43%]
|
||||
........................................................................ [ 46%]
|
||||
........................................................................ [ 49%]
|
||||
........................................................................ [ 51%]
|
||||
........................................................................ [ 54%]
|
||||
........................................................................ [ 57%]
|
||||
........................................................................ [ 60%]
|
||||
........................................................................ [ 63%]
|
||||
........................................................................ [ 66%]
|
||||
........................................................................ [ 69%]
|
||||
........................................................................ [ 72%]
|
||||
........................................................................ [ 75%]
|
||||
........................................................................ [ 77%]
|
||||
........................................................................ [ 80%]
|
||||
........................................................................ [ 83%]
|
||||
........................................................................ [ 86%]
|
||||
........................................................................ [ 89%]
|
||||
........................................................................ [ 92%]
|
||||
........................................................................ [ 95%]
|
||||
........................................................................ [ 98%]
|
||||
................................................ [100%]
|
||||
=============================== warnings summary ===============================
|
||||
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
|
||||
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||
from starlette.testclient import TestClient as TestClient # noqa
|
||||
|
||||
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||
================================ tests coverage ================================
|
||||
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
|
||||
|
||||
Name Stmts Miss Cover
|
||||
--------------------------------------------------
|
||||
app/__init__.py 1 0 100%
|
||||
app/api/__init__.py 0 0 100%
|
||||
app/api/auth.py 52 0 100%
|
||||
app/api/chat.py 224 1 99%
|
||||
app/api/chats.py 110 0 100%
|
||||
app/api/config.py 13 0 100%
|
||||
app/api/doc_drafts.py 99 0 100%
|
||||
app/api/docs.py 156 1 99%
|
||||
app/api/git_sources.py 232 0 100%
|
||||
app/api/health.py 10 0 100%
|
||||
app/api/steering.py 42 0 100%
|
||||
app/api/suggestions.py 33 0 100%
|
||||
app/api/sync.py 139 0 100%
|
||||
app/api/tokens.py 40 0 100%
|
||||
app/api/ui_settings.py 55 0 100%
|
||||
app/config.py 224 0 100%
|
||||
app/core/__init__.py 0 0 100%
|
||||
app/core/auth.py 45 0 100%
|
||||
app/core/caching.py 124 0 100%
|
||||
app/core/debugging.py 29 2 93%
|
||||
app/core/docs_push.py 39 0 100%
|
||||
app/core/errors.py 5 0 100%
|
||||
app/core/logging.py 13 0 100%
|
||||
app/core/rate_limit.py 44 0 100%
|
||||
app/core/security_headers.py 20 0 100%
|
||||
app/core/theming.py 38 0 100%
|
||||
app/core/tokens.py 44 0 100%
|
||||
app/db.py 22 0 100%
|
||||
app/main.py 66 0 100%
|
||||
app/models.py 128 0 100%
|
||||
app/rag/__init__.py 0 0 100%
|
||||
app/rag/agent.py 317 1 99%
|
||||
app/rag/archive_upload.py 134 0 100%
|
||||
app/rag/chunker.py 206 4 98%
|
||||
app/rag/doc_dates.py 18 0 100%
|
||||
app/rag/folder_summaries.py 123 0 100%
|
||||
app/rag/git_sources.py 14 0 100%
|
||||
app/rag/importer.py 219 3 99%
|
||||
app/rag/llm.py 244 1 99%
|
||||
app/rag/overview.py 71 0 100%
|
||||
app/rag/prompts.py 100 0 100%
|
||||
app/rag/retriever.py 205 3 99%
|
||||
app/rag/scaffolding.py 55 0 100%
|
||||
app/rag/source_removal.py 41 0 100%
|
||||
app/rag/sources_meta.py 16 0 100%
|
||||
app/rag/suggestions.py 27 0 100%
|
||||
app/rag/summarizer.py 24 0 100%
|
||||
app/schemas.py 329 0 100%
|
||||
--------------------------------------------------
|
||||
TOTAL 4190 16 99%
|
||||
coverage gate: app/ 99% (>90%) OK
|
||||
All checks passed!
|
||||
0 errors, 0 warnings, 0 informations
|
||||
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||
|
||||
validation OK
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
Task 05 complete. All validation gates are green.
|
||||
|
||||
**Implemented (task 05 — `plan_turn` wiring):**
|
||||
- `app/rag/retriever.py`: new pure `select_related(chunks, excluded_ids, cap)` — same stable score-descending walk, skips excluded ids, no floor, capped; plus "phase 118 retired the seeding role (A6)" docstring notes on dormant `select_documents`/`select_documents_tiered`
|
||||
- `app/api/chat.py`: `plan_turn` now computes `select_suggested` (top-5, no floor, honors `suggested_docs`) + `select_related` (rank 6+, honors `related_max_docs`) once for BOTH branches; HIGH prompt seeds the suggested summaries; `TurnPlan.docs` → `suggested_docs` (same position); `summary_hits` redefined to the suggested set; `run_agent(seed_docs=plan.suggested_docs)`; `cited_docs` = suggested+read, `record_docs` = suggested+related+read; log line gains `suggested=%d` after `summary_hits=%d` (phase-114 tail stays last)
|
||||
- `app/config.py`: retirement notes on `top_n_docs`/`source_usefulness_floor` (kept, no `app/` consumer); script ripples: `plan.suggested_docs` in `scripts/agent_realmodel_check.py` + `scripts/load_test_kb.py`
|
||||
- Tests: 8 `select_related` unit tests; `test_chat_gate.py` re-pinned (7-doc fixture: 5 summaries seeded, zero full content, related = rank 6–7; deflected prompt sha256-pinned byte-identical + `build_deflect_prompt` equality; `summary_hits` redefinition; suggested/related on both branches; `suggested=N` log-slot + full field-order pin; retired-settings behavioral pin); `test_source_chip_quality.py` shapes re-pinned to no-floor tiering; `test_chat_api.py`: `done.sources` = 5 suggested (+read deduped last), `done.related` = rank 6–7 minus cited, `query_log.sources` = full 7-doc retrieval, deflection = `[]`
|
||||
|
||||
**Results:**
|
||||
- `uv run pytest tests/unit/test_chat_gate.py tests/integration/test_chat_api.py -v` → 94 passed
|
||||
- `uv run pytest --cov=app` → 2506 passed, app/ coverage 99% (>90% gate)
|
||||
- `uv run ruff check .` → clean; `uv run pyright` → 0 errors
|
||||
|
||||
**Decisions:** `select_suggested(chunks, n=settings.suggested_docs)` passes the turn's settings (unit tests control the cap without env); deflected-frame "byte-identical" pinned as prompt-sha + frame-shape (related-row content is rank 6+ by design, per task item 7); one deflection wire test's scripted read re-targeted from `backups.md` (now rank 6, durably recorded) to `tables.md` (rank 11) to keep the "never read" pin distinct.
|
||||
|
||||
**Next pending task:** `118_summary_seed_context/06_e2e_summary_context.md`
|
||||
+103
@@ -0,0 +1,103 @@
|
||||
........................................................................ [ 2%]
|
||||
........................................................................ [ 5%]
|
||||
........................................................................ [ 8%]
|
||||
........................................................................ [ 11%]
|
||||
........................................................................ [ 14%]
|
||||
........................................................................ [ 17%]
|
||||
........................................................................ [ 20%]
|
||||
........................................................................ [ 22%]
|
||||
........................................................................ [ 25%]
|
||||
........................................................................ [ 28%]
|
||||
........................................................................ [ 31%]
|
||||
........................................................................ [ 34%]
|
||||
........................................................................ [ 37%]
|
||||
........................................................................ [ 40%]
|
||||
........................................................................ [ 43%]
|
||||
........................................................................ [ 45%]
|
||||
........................................................................ [ 48%]
|
||||
........................................................................ [ 51%]
|
||||
........................................................................ [ 54%]
|
||||
........................................................................ [ 57%]
|
||||
........................................................................ [ 60%]
|
||||
........................................................................ [ 63%]
|
||||
........................................................................ [ 66%]
|
||||
........................................................................ [ 68%]
|
||||
........................................................................ [ 71%]
|
||||
........................................................................ [ 74%]
|
||||
........................................................................ [ 77%]
|
||||
........................................................................ [ 80%]
|
||||
........................................................................ [ 83%]
|
||||
........................................................................ [ 86%]
|
||||
........................................................................ [ 89%]
|
||||
........................................................................ [ 91%]
|
||||
........................................................................ [ 94%]
|
||||
........................................................................ [ 97%]
|
||||
.......................................................... [100%]
|
||||
=============================== warnings summary ===============================
|
||||
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
|
||||
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||
from starlette.testclient import TestClient as TestClient # noqa
|
||||
|
||||
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||
================================ tests coverage ================================
|
||||
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
|
||||
|
||||
Name Stmts Miss Cover
|
||||
--------------------------------------------------
|
||||
app/__init__.py 1 0 100%
|
||||
app/api/__init__.py 0 0 100%
|
||||
app/api/auth.py 52 0 100%
|
||||
app/api/chat.py 225 1 99%
|
||||
app/api/chats.py 110 0 100%
|
||||
app/api/config.py 13 0 100%
|
||||
app/api/doc_drafts.py 99 0 100%
|
||||
app/api/docs.py 156 1 99%
|
||||
app/api/git_sources.py 232 0 100%
|
||||
app/api/health.py 10 0 100%
|
||||
app/api/steering.py 42 0 100%
|
||||
app/api/suggestions.py 33 0 100%
|
||||
app/api/sync.py 139 0 100%
|
||||
app/api/tokens.py 40 0 100%
|
||||
app/api/ui_settings.py 55 0 100%
|
||||
app/config.py 224 0 100%
|
||||
app/core/__init__.py 0 0 100%
|
||||
app/core/auth.py 45 0 100%
|
||||
app/core/caching.py 124 0 100%
|
||||
app/core/debugging.py 29 2 93%
|
||||
app/core/docs_push.py 39 0 100%
|
||||
app/core/errors.py 5 0 100%
|
||||
app/core/logging.py 13 0 100%
|
||||
app/core/rate_limit.py 44 0 100%
|
||||
app/core/security_headers.py 20 0 100%
|
||||
app/core/theming.py 38 0 100%
|
||||
app/core/tokens.py 44 0 100%
|
||||
app/db.py 22 0 100%
|
||||
app/main.py 66 0 100%
|
||||
app/models.py 128 0 100%
|
||||
app/rag/__init__.py 0 0 100%
|
||||
app/rag/agent.py 317 1 99%
|
||||
app/rag/archive_upload.py 134 0 100%
|
||||
app/rag/chunker.py 206 4 98%
|
||||
app/rag/doc_dates.py 18 0 100%
|
||||
app/rag/folder_summaries.py 123 0 100%
|
||||
app/rag/git_sources.py 14 0 100%
|
||||
app/rag/importer.py 219 3 99%
|
||||
app/rag/llm.py 244 1 99%
|
||||
app/rag/overview.py 71 0 100%
|
||||
app/rag/prompts.py 100 0 100%
|
||||
app/rag/retriever.py 217 3 99%
|
||||
app/rag/scaffolding.py 55 0 100%
|
||||
app/rag/source_removal.py 41 0 100%
|
||||
app/rag/sources_meta.py 16 0 100%
|
||||
app/rag/suggestions.py 27 0 100%
|
||||
app/rag/summarizer.py 24 0 100%
|
||||
app/schemas.py 329 0 100%
|
||||
--------------------------------------------------
|
||||
TOTAL 4203 16 99%
|
||||
coverage gate: app/ 99% (>90%) OK
|
||||
All checks passed!
|
||||
0 errors, 0 warnings, 0 informations
|
||||
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||
|
||||
validation OK
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
All gates verified and the repo is clean. Final report:
|
||||
|
||||
**Task 06 (phase 118) complete — summary-seed E2E contract.**
|
||||
|
||||
- NEW `tests/e2e/test_summary_seed_context.py` (mock-only, own app :8138): 7-doc markdown KB (token-controlled strength gradient) proves cases a–e — tail-echo pins the LAST suggested doc's mock-digest summary tail in the prompt (a/d, markdown included), NO tail sentinels (b, full text never seeded), scripted `SUMMARY_SEED_READ_TRIGGER` flow reads rank-1 suggested doc → verbatim read-echo lands the sentinel (c), chips = suggested+read deduped, related row = rank 6+ (e). Setup pins tier order via the app's real `retrieve`/`select_suggested`/`select_related` (tail: a,d,b,c,e+f,g; read: a,b,c,d,e+f,g).
|
||||
- Deleted `test_whole_document_context.py`; re-targeted drill (read SUCCEEDS — seed-read refusal gone; mock drill quote now skips the D5 `date:` line → byte-identical content quote, fixing a pre-existing drill break), read-cap (removed the not-a-seed guard + `embed_text` import), source-chip-quality (chips = suggested top-5 no-floor; related = rank 6+), document-dates (4 chips, no related row, exact `row.sources`), document-summaries (summaries 1→2; md HAS the digest; sentinel pin inverted to absent), summary-in-viewer (md gets the panel, API summary non-null), mock docstrings (ALREADY_IN_CONTEXT → already-read case).
|
||||
- Defect found + fixed: `test_document_dates` focus-trap pin broke under A2 (markdown now has a summary → `doc-summary-edit` joins the trap): Tab chain re-targeted (revert → summary-edit → wrap to `doc-modal-open`).
|
||||
- Audit sweep: no test pins the retired full-text-seed/seed-read-refusal contract (comment-only updates in agent_document_tools / theme_semantic_completion / share_chat; unit pins were task 04's).
|
||||
- Gates: `uv run pytest tests/e2e/test_summary_seed_context.py -v --no-cov` → 2 passed; re-targeted E2E in isolation → drill 3, read-cap 3, chip-quality 2, dates 6, doc-summaries 2, viewer 4 (all passed); tool-flow regressions (agent_document_tools, grep_regex_teaching, harness_aligned, +7) all passed; `uv run pytest --cov=app --cov-report=term-missing` → 2506 passed, app TOTAL **99%** (>90%); `uv run ruff check .` clean; `uv run pyright` → 0 errors.
|
||||
- Noted, not fixed (pre-existing, out of scope): the COMBINED `pytest tests/e2e` batch is poisoned by `test_extensionless_import.py`'s module-level `os.environ["BOR_IMPORT_EXTENSIONS"]` pin (phase 102) leaking through collection into every other suite's in-process `Settings` (65F+110E "5 vs 13 docs" cascade) — the project gate is per-story isolation (AGENTS.md rule 9), which is green; a proper fix needs that suite to own a module app (separate task).
|
||||
- Next pending task: none — task 06 is the last task of phase 118 (01–05 already in `complete/`).
|
||||
+103
@@ -0,0 +1,103 @@
|
||||
........................................................................ [ 2%]
|
||||
........................................................................ [ 5%]
|
||||
........................................................................ [ 8%]
|
||||
........................................................................ [ 11%]
|
||||
........................................................................ [ 14%]
|
||||
........................................................................ [ 17%]
|
||||
........................................................................ [ 20%]
|
||||
........................................................................ [ 22%]
|
||||
........................................................................ [ 25%]
|
||||
........................................................................ [ 28%]
|
||||
........................................................................ [ 31%]
|
||||
........................................................................ [ 34%]
|
||||
........................................................................ [ 37%]
|
||||
........................................................................ [ 40%]
|
||||
........................................................................ [ 43%]
|
||||
........................................................................ [ 45%]
|
||||
........................................................................ [ 48%]
|
||||
........................................................................ [ 51%]
|
||||
........................................................................ [ 54%]
|
||||
........................................................................ [ 57%]
|
||||
........................................................................ [ 60%]
|
||||
........................................................................ [ 63%]
|
||||
........................................................................ [ 66%]
|
||||
........................................................................ [ 68%]
|
||||
........................................................................ [ 71%]
|
||||
........................................................................ [ 74%]
|
||||
........................................................................ [ 77%]
|
||||
........................................................................ [ 80%]
|
||||
........................................................................ [ 83%]
|
||||
........................................................................ [ 86%]
|
||||
........................................................................ [ 89%]
|
||||
........................................................................ [ 91%]
|
||||
........................................................................ [ 94%]
|
||||
........................................................................ [ 97%]
|
||||
.......................................................... [100%]
|
||||
=============================== warnings summary ===============================
|
||||
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
|
||||
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||
from starlette.testclient import TestClient as TestClient # noqa
|
||||
|
||||
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||
================================ tests coverage ================================
|
||||
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
|
||||
|
||||
Name Stmts Miss Cover
|
||||
--------------------------------------------------
|
||||
app/__init__.py 1 0 100%
|
||||
app/api/__init__.py 0 0 100%
|
||||
app/api/auth.py 52 0 100%
|
||||
app/api/chat.py 225 1 99%
|
||||
app/api/chats.py 110 0 100%
|
||||
app/api/config.py 13 0 100%
|
||||
app/api/doc_drafts.py 99 0 100%
|
||||
app/api/docs.py 156 1 99%
|
||||
app/api/git_sources.py 232 0 100%
|
||||
app/api/health.py 10 0 100%
|
||||
app/api/steering.py 42 0 100%
|
||||
app/api/suggestions.py 33 0 100%
|
||||
app/api/sync.py 139 0 100%
|
||||
app/api/tokens.py 40 0 100%
|
||||
app/api/ui_settings.py 55 0 100%
|
||||
app/config.py 224 0 100%
|
||||
app/core/__init__.py 0 0 100%
|
||||
app/core/auth.py 45 0 100%
|
||||
app/core/caching.py 124 0 100%
|
||||
app/core/debugging.py 29 2 93%
|
||||
app/core/docs_push.py 39 0 100%
|
||||
app/core/errors.py 5 0 100%
|
||||
app/core/logging.py 13 0 100%
|
||||
app/core/rate_limit.py 44 0 100%
|
||||
app/core/security_headers.py 20 0 100%
|
||||
app/core/theming.py 38 0 100%
|
||||
app/core/tokens.py 44 0 100%
|
||||
app/db.py 22 0 100%
|
||||
app/main.py 66 0 100%
|
||||
app/models.py 128 0 100%
|
||||
app/rag/__init__.py 0 0 100%
|
||||
app/rag/agent.py 317 1 99%
|
||||
app/rag/archive_upload.py 134 0 100%
|
||||
app/rag/chunker.py 206 4 98%
|
||||
app/rag/doc_dates.py 18 0 100%
|
||||
app/rag/folder_summaries.py 123 0 100%
|
||||
app/rag/git_sources.py 14 0 100%
|
||||
app/rag/importer.py 219 3 99%
|
||||
app/rag/llm.py 244 1 99%
|
||||
app/rag/overview.py 71 0 100%
|
||||
app/rag/prompts.py 100 0 100%
|
||||
app/rag/retriever.py 217 3 99%
|
||||
app/rag/scaffolding.py 55 0 100%
|
||||
app/rag/source_removal.py 41 0 100%
|
||||
app/rag/sources_meta.py 16 0 100%
|
||||
app/rag/suggestions.py 27 0 100%
|
||||
app/rag/summarizer.py 24 0 100%
|
||||
app/schemas.py 329 0 100%
|
||||
--------------------------------------------------
|
||||
TOTAL 4203 16 99%
|
||||
coverage gate: app/ 99% (>90%) OK
|
||||
All checks passed!
|
||||
0 errors, 0 warnings, 0 informations
|
||||
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||
|
||||
validation OK
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
All completion criteria are met. Final report:
|
||||
|
||||
**Phase 118, task 07 — fixture-battery gate: complete (PASS)**
|
||||
|
||||
- Rebuilt `tests/fixtures/test_kb.dump.sql` via `scripts.load_test_kb` — all 8 markdown fixture docs now carry a `lite` summary + embedded `is_summary` chunk (8 docs, 17 chunks: 9 content + 8 summary); `summary_errors=0`, 0 failed folder-summary batches, retrieval report all 10 GROUNDED, round-trip checksum verified.
|
||||
- Fixed a pre-existing defect blocking the rebuild: `load_test_kb._TABLES` omitted `documents.created_at`/`created_at_manual`, `git_sources.ignore_paths`/`include_hidden`, `folder_summaries.manually_edited` → restore re-defaulted `created_at` and the build's own round-trip check failed on `documents`. Full column sets now dumped; header reports content + summary chunk counts.
|
||||
- Verdict run (`uv run python -m scripts.agent_realmodel_check --restore --mode fixture`, model `turbo`): `gate: turbo PASS turns=10 answered=10 caps=0 tool-turns=9 calls 21/21 executed (100%) contract 21/21 (100%) 2026-09-16 (wall 118.1s)` — conditions 1/2/4 GREEN, condition 3 reported (9/10, also ≥6); no copy iteration needed.
|
||||
- Appended dated §10 to `TOOL_CALLING_TESTING.md` (telemetry, contract line, wall, dump rebuild note, locked-A7 semantics note, final `read` copy); prior sections untouched.
|
||||
- Results: `uv run pytest --cov=app` → 2506 passed, 99% on `app/` (no suite pins the old 9-chunk dump — it is script-owned); `uv run ruff check . && uv run pyright` → clean; `uv run pytest tests/e2e/test_summary_seed_context.py -v --no-cov` → 2 passed (sanity).
|
||||
- Notable: gate script's `evaluate()` left unchanged per the task file — A7 interpretation (condition 3 reported, not gated) is applied to the recorded verdict; the script exited 0 here anyway. No commits made (harness commits).
|
||||
- Next pending task: none — task 07 was the last task of phase 118; the phase is complete.
|
||||
+103
@@ -0,0 +1,103 @@
|
||||
........................................................................ [ 2%]
|
||||
........................................................................ [ 5%]
|
||||
........................................................................ [ 8%]
|
||||
........................................................................ [ 11%]
|
||||
........................................................................ [ 14%]
|
||||
........................................................................ [ 17%]
|
||||
........................................................................ [ 20%]
|
||||
........................................................................ [ 22%]
|
||||
........................................................................ [ 25%]
|
||||
........................................................................ [ 28%]
|
||||
........................................................................ [ 31%]
|
||||
........................................................................ [ 34%]
|
||||
........................................................................ [ 37%]
|
||||
........................................................................ [ 40%]
|
||||
........................................................................ [ 43%]
|
||||
........................................................................ [ 45%]
|
||||
........................................................................ [ 48%]
|
||||
........................................................................ [ 51%]
|
||||
........................................................................ [ 54%]
|
||||
........................................................................ [ 57%]
|
||||
........................................................................ [ 60%]
|
||||
........................................................................ [ 63%]
|
||||
........................................................................ [ 66%]
|
||||
........................................................................ [ 68%]
|
||||
........................................................................ [ 71%]
|
||||
........................................................................ [ 74%]
|
||||
........................................................................ [ 77%]
|
||||
........................................................................ [ 80%]
|
||||
........................................................................ [ 83%]
|
||||
........................................................................ [ 86%]
|
||||
........................................................................ [ 89%]
|
||||
........................................................................ [ 91%]
|
||||
........................................................................ [ 94%]
|
||||
........................................................................ [ 97%]
|
||||
.......................................................... [100%]
|
||||
=============================== warnings summary ===============================
|
||||
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
|
||||
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||
from starlette.testclient import TestClient as TestClient # noqa
|
||||
|
||||
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||
================================ tests coverage ================================
|
||||
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
|
||||
|
||||
Name Stmts Miss Cover
|
||||
--------------------------------------------------
|
||||
app/__init__.py 1 0 100%
|
||||
app/api/__init__.py 0 0 100%
|
||||
app/api/auth.py 52 0 100%
|
||||
app/api/chat.py 225 1 99%
|
||||
app/api/chats.py 110 0 100%
|
||||
app/api/config.py 13 0 100%
|
||||
app/api/doc_drafts.py 99 0 100%
|
||||
app/api/docs.py 156 1 99%
|
||||
app/api/git_sources.py 232 0 100%
|
||||
app/api/health.py 10 0 100%
|
||||
app/api/steering.py 42 0 100%
|
||||
app/api/suggestions.py 33 0 100%
|
||||
app/api/sync.py 139 0 100%
|
||||
app/api/tokens.py 40 0 100%
|
||||
app/api/ui_settings.py 55 0 100%
|
||||
app/config.py 224 0 100%
|
||||
app/core/__init__.py 0 0 100%
|
||||
app/core/auth.py 45 0 100%
|
||||
app/core/caching.py 124 0 100%
|
||||
app/core/debugging.py 29 2 93%
|
||||
app/core/docs_push.py 39 0 100%
|
||||
app/core/errors.py 5 0 100%
|
||||
app/core/logging.py 13 0 100%
|
||||
app/core/rate_limit.py 44 0 100%
|
||||
app/core/security_headers.py 20 0 100%
|
||||
app/core/theming.py 38 0 100%
|
||||
app/core/tokens.py 44 0 100%
|
||||
app/db.py 22 0 100%
|
||||
app/main.py 66 0 100%
|
||||
app/models.py 128 0 100%
|
||||
app/rag/__init__.py 0 0 100%
|
||||
app/rag/agent.py 317 1 99%
|
||||
app/rag/archive_upload.py 134 0 100%
|
||||
app/rag/chunker.py 206 4 98%
|
||||
app/rag/doc_dates.py 18 0 100%
|
||||
app/rag/folder_summaries.py 123 0 100%
|
||||
app/rag/git_sources.py 14 0 100%
|
||||
app/rag/importer.py 219 3 99%
|
||||
app/rag/llm.py 244 1 99%
|
||||
app/rag/overview.py 71 0 100%
|
||||
app/rag/prompts.py 100 0 100%
|
||||
app/rag/retriever.py 217 3 99%
|
||||
app/rag/scaffolding.py 55 0 100%
|
||||
app/rag/source_removal.py 41 0 100%
|
||||
app/rag/sources_meta.py 16 0 100%
|
||||
app/rag/suggestions.py 27 0 100%
|
||||
app/rag/summarizer.py 24 0 100%
|
||||
app/schemas.py 329 0 100%
|
||||
--------------------------------------------------
|
||||
TOTAL 4203 16 99%
|
||||
coverage gate: app/ 99% (>90%) OK
|
||||
All checks passed!
|
||||
0 errors, 0 warnings, 0 informations
|
||||
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||
|
||||
validation OK
|
||||
@@ -38,6 +38,8 @@ BOR_RELEVANCE_THRESHOLD=0.62 # answer when best cosine >= this OR an FTS hi
|
||||
# BOR_LEXICAL_SUPPORT_FLOOR=0.35 # cosine floor for FTS hits to flip HIGH (A8 revised 2026-09-14); 0 <= floor <= relevance_threshold
|
||||
# BOR_SOURCE_USEFULNESS_FLOOR=0.35 # cosine bar for a citation slot (phase 113): a doc earns done.sources only when its best-chunk cosine clears it; 0 <= floor <= relevance_threshold; 0 = no bar
|
||||
# BOR_RELATED_MAX_DOCS=2 # cap on the related-doc tier in the done frame (phase 113: the UI's de-emphasized "nearby docs" row, never a citation chip); 0 = no related docs
|
||||
# BOR_SUGGESTED_DOCS=5 # cap on the "start here" suggestion tier (phase 118): top-N doc summaries seeded into the grounded prompt; NO cosine floor (opt-in starting points, not citations); >= 1
|
||||
# BOR_SUGGESTION_PREVIEW_CHARS=400 # NULL-summary fallback preview in a suggestion block (phase 118): first N content chars + the truncation marker; > 0
|
||||
BOR_MAX_OUTPUT_TOKENS=32768 # max answer length in tokens (answers must not be cut off)
|
||||
BOR_STEERING_MAX_CHARS=8000 # char budget for the <tuning> (steering notes) prompt section
|
||||
BOR_SUMMARY_MAX_CHARS=12000 # cap on document content sent to the lite summary model (phase 30)
|
||||
|
||||
@@ -589,3 +589,136 @@ listing answers to reach the file level by default, the lever is the
|
||||
`ls` description / `TOOLS_SECTION` copy (teach that a folder's file
|
||||
lines appear only after drilling into it) — flagged here rather than
|
||||
acted on unilaterally.
|
||||
|
||||
---
|
||||
|
||||
## 10. Phase 118 — summary-seed gate, 2026-09-15 (run 2026-09-16)
|
||||
|
||||
Phase 118 re-worked the grounded-turn seed: the `<documents>` section
|
||||
no longer carries the **full text** of the top-2 cited documents — it
|
||||
carries the **summaries** of the top 5 ranked documents (the
|
||||
"start here if these summaries seem right to you" tier), and a
|
||||
document's full text enters the context only through the capped `read`
|
||||
tool. The fixture battery had to be re-run against the real configured
|
||||
chat model under the re-cut gate semantics (locked A7, owner-confirmed
|
||||
2026-09-15) before the phase could complete — the phase-70/72/94
|
||||
tool-copy gate precedent.
|
||||
|
||||
**Gate semantics changed (locked A7).** The phase gate is now
|
||||
condition **1** (all turns answer) AND condition **2** (zero round-cap
|
||||
hits) AND condition **4** (fixture-mode contract accuracy ≥ 0.90).
|
||||
Condition **3** (≥6/10 turns emit ≥1 tool call) is **REPORTED, not
|
||||
gated** — a summary-only answer is now the *intended* fast path, not
|
||||
tool abandonment. The gate script still gates condition 3 on a full
|
||||
(non-`--turns`) run and exits 1 if it is the only miss; the phase
|
||||
verdict reads the four printed conditions under A7, so a run that
|
||||
misses **only** condition 3 with 1/2/4 green is a **PASS** (recorded
|
||||
as such, below). The full telemetry is recorded regardless.
|
||||
|
||||
**Fixture dump rebuild (a legitimate re-run trigger).** The import
|
||||
pipeline's summary policy changed (phase 118, task 01 — every document
|
||||
gets a stored summary + one embedded `is_summary` chunk, markdown
|
||||
included), so the dump was rebuilt with
|
||||
`uv run python -m scripts.load_test_kb` (the real pipeline). All eight
|
||||
hand-written fixture documents are markdown and each now carries a
|
||||
`doc.summary` + one `is_summary` chunk: **8 documents, 17 chunks**
|
||||
(9 content + 8 summary). The build verified
|
||||
`summary_errors=0`, zero failed folder-summary batches, and an
|
||||
all-GROUNDED retrieval report (the honesty gate is untouched by this
|
||||
phase — grounded/deflected is unchanged). The dump is tracked and
|
||||
committed with the phase.
|
||||
|
||||
The rebuild also surfaced and fixed a latent dump bug: the
|
||||
`_TABLES` column lists in `scripts/load_test_kb.py` predated three
|
||||
migrations, so the snapshot omitted `documents.created_at` /
|
||||
`documents.created_at_manual` (phase 106),
|
||||
`git_sources.ignore_paths` / `git_sources.include_hidden`, and
|
||||
`folder_summaries.manually_edited` (phase 97). Restoring such a dump
|
||||
silently re-defaulted `created_at` to `now()`, so the build's own
|
||||
round-trip checksum verification FAILED on `documents`. The column
|
||||
lists now cover every column of each table, and the round-trip
|
||||
verification passes.
|
||||
|
||||
**The verdict run.** `uv run python -m scripts.agent_realmodel_check
|
||||
--restore --mode fixture` — the configured chat model (`turbo` per
|
||||
`.env`), the refreshed fixture KB, the real grounded path, the full
|
||||
10-question battery (the verdict is always the full battery —
|
||||
`--turns N` is the copy-iteration micro-loop only):
|
||||
|
||||
```
|
||||
$ uv run python -m scripts.agent_realmodel_check --restore --mode fixture
|
||||
restore: ok in 0.04s (8 docs, 2 sources)
|
||||
turn 01 | emitted=1 executed=1 cap=no defl=no | 9.49s | List the files in this directory.
|
||||
turn 02 | emitted=4 executed=4 cap=no defl=no | 15.22s | List the documents you have in the …
|
||||
turn 03 | emitted=8 executed=8 cap=no defl=no | 24.16s | List every document you have indexed.
|
||||
turn 04 | emitted=1 executed=1 cap=no defl=no | 12.55s | Open the document …
|
||||
turn 05 | emitted=1 executed=1 cap=no defl=no | 10.45s | Read …
|
||||
turn 06 | emitted=1 executed=1 cap=no defl=no | 7.99s | Open the document …
|
||||
turn 07 | emitted=1 executed=1 cap=no defl=no | 6.63s | Find the exact string "rbm-8842" in …
|
||||
turn 08 | emitted=0 executed=0 cap=no defl=no | 9.92s | Which document has the title "Lab …
|
||||
turn 09 | emitted=1 executed=1 cap=no defl=no | 10.57s | What do you know about the qwen 3.8 …
|
||||
turn 10 | emitted=3 executed=3 cap=no defl=no | 11.00s | List the files in the deployments …
|
||||
gate: turbo PASS turns=10 answered=10 caps=0 tool-turns=9 calls 21/21 executed (100%) contract 21/21 (100%) 2026-09-16 (wall 118.1s)
|
||||
```
|
||||
|
||||
**The four conditions (locked A7) and metrics.**
|
||||
|
||||
| condition | gated under A7 | result |
|
||||
|---|---|---|
|
||||
| 1. all turns answer | yes | 10/10 answered — **GREEN** |
|
||||
| 2. zero round-cap hits | yes | caps=0 — **GREEN** |
|
||||
| 3. ≥6/10 turns emit ≥1 tool call | **no (reported)** | 9/10 tool-turns |
|
||||
| 4. contract accuracy ≥ 0.90 | yes | 21/21 (100%) — **GREEN** |
|
||||
| executed / emitted (reported) | no | 21/21 (100%) |
|
||||
|
||||
Contract line: **contract 21/21 (100%)**. Wall time: **118.1 s**.
|
||||
Model: **turbo** (the configured chat model).
|
||||
|
||||
**Per-turn reading.** The three designed read turns (04/05/06) each
|
||||
emitted exactly one contract-correct `read` in the combined
|
||||
`source/path` form — and each **executed** (the seeded document is now
|
||||
a *summary*, so `read` adds its full text rather than being
|
||||
dedupe-refused: 100 % executed, not the ~58–73 % the full-text seed
|
||||
produced). The discipline turns behave exactly as the new contract
|
||||
wants: turn 08 (title lookup, target seeded as a summary) answered
|
||||
straight from the `<documents>` context with **zero** tool calls, and
|
||||
turn 09 (the qwen launch arguments) answered from the seeded summary
|
||||
context. The listing turns drill the tree as under phase 94 (turn 03:
|
||||
top level + both sources + the leaf folders in 8 calls). Zero cap
|
||||
hits, zero contract violations, 10/10 answered.
|
||||
|
||||
**Baseline comparison.** Against the phase-94 `turbo` baseline
|
||||
(98.7 s wall, 100 % contract, 100 % executed on 19 calls): contract
|
||||
and executed are unchanged at 100 %, caps remain 0, and the wall time
|
||||
is 118.1 s (+19.7 %) — inside the ~20 % band the phase-94 gate used
|
||||
as its slowdown tripwire, and within the normal endpoint-load
|
||||
variance of the 97–135 s `turbo` range recorded in §3. The gate is
|
||||
met on the conditions that gate (1/2/4 all GREEN) and condition 3 is
|
||||
reported (9/10).
|
||||
|
||||
**`read` copy — final shape (no copy iteration was needed).**
|
||||
Condition 4 did not degrade and condition 2 was never hit, so no
|
||||
`--turns` micro-loop iteration was required. The final copy is the
|
||||
task-04 rewrite, byte-pinned in `tests/unit/test_prompt_lock.py` /
|
||||
`test_agent.py`: the `read` description in `AGENT_TOOLS`
|
||||
(`app/rag/agent.py`) teaches that `<documents>` shows the *summaries*
|
||||
of the top-ranked documents (their full texts are NOT in the prompt
|
||||
yet) and that `read` adds one of them — or any other document — by
|
||||
its combined `source/path`, and must not re-read an already-read
|
||||
document (that refusal's copy stays accurate: "the full text is
|
||||
already in your prompt"); the `read` clause of `TOOLS_SECTION`
|
||||
(`app/rag/prompts.py`) mirrors it ("The `<documents>` section holds
|
||||
SUMMARIES — the full text of a suggested document is not in your
|
||||
prompt until you `read` it"). `classify_call` needed **no**
|
||||
re-baselining: a `read` of a suggested document is a well-formed call
|
||||
targeting a resolvable document — contract-correct under both the old
|
||||
and the new app semantics (its docstring's `ALREADY_IN_CONTEXT` note
|
||||
already excludes dedupe re-reads from the metric).
|
||||
|
||||
**Conclusion.** The phase-118 summary-seed gate is **PASS**: conditions
|
||||
1, 2, and 4 GREEN (condition 3 reported at 9/10, not gated), contract
|
||||
accuracy 100 %, zero cap hits, the full battery all-answered, wall
|
||||
within the ~20 % band. The summary-seed behavior is proven against the
|
||||
real configured chat model — full text enters the context only through
|
||||
the capped `read` tool, and a summary-only answer is the intended fast
|
||||
path.
|
||||
|
||||
+104
-57
@@ -37,13 +37,30 @@ Steering (phase 15): the owner's stored tuning notes are loaded per turn
|
||||
section — both the HIGH and the LOW prompt carry it. The per-turn log
|
||||
line records ``tuning=N`` (the number of injected notes).
|
||||
|
||||
Summaries (phase 30): a lite-model summary chunk's parent *is* the
|
||||
source document, so a summary hit resolves to the full source document
|
||||
through the unchanged chunk→document mapping (A7 revised) — context
|
||||
assembly is untouched. ``TurnPlan.summary_hits`` counts the hit chunks
|
||||
with ``is_summary`` whose parent document landed in the selected
|
||||
top-N context, and the per-turn log line records ``summary_hits=N``
|
||||
after ``fts_hits`` (PLAN §9 line extension).
|
||||
Summaries (phase 30; phase 118 redefinition): a lite-model summary
|
||||
chunk's parent *is* the source document, so a summary hit resolves to
|
||||
the full source document through the unchanged chunk→document mapping
|
||||
(A7 revised). ``TurnPlan.summary_hits`` counts the hit chunks with
|
||||
``is_summary`` whose parent document is in the SUGGESTED set (phase 118:
|
||||
redefined from the phase-113 cited set — the suggestion tier is the
|
||||
seeded context now), and the per-turn log line records ``summary_hits=N``
|
||||
after ``fts_hits`` (PLAN §9 line extension; phase 118 adds
|
||||
``suggested=N`` after ``summary_hits=N``).
|
||||
|
||||
Summary seed context (phase 118, TODO L3 — the owner directive that
|
||||
re-revises A7, LOCKED A6): a grounded turn's ``<documents>`` section
|
||||
seeds the top-5 suggested documents' SUMMARIES (never their full
|
||||
texts) — the "start here if these summaries seem right to you" starting
|
||||
point — and the LLM extends its context by ``read``-ing what it needs
|
||||
(the capped ``read`` tool is the ONLY full-text path). Both tiers are
|
||||
computed once per turn, for BOTH branches: ``suggested_docs`` (no
|
||||
floor, A3) and ``related_docs`` (rank 6+, the done frame's row).
|
||||
``done.sources`` is the citation surface — suggested + agent-read,
|
||||
deduped (A4: a grounded turn always shows chips); ``query_log.sources``
|
||||
and the per-turn log line keep recording the full retrieval
|
||||
(suggested + related + read, LOCKED A3). The deflected branch's prompt
|
||||
is byte-identical (weak-hit titles only — A8 untouched); its TurnPlan
|
||||
still carries both tiers for the durable record.
|
||||
|
||||
KB overview (phase 31): the lite-generated outline of the knowledge
|
||||
base (single ``kb_overview`` row) is read per turn (one indexed PK
|
||||
@@ -195,7 +212,13 @@ from app.rag.llm import (
|
||||
)
|
||||
from app.rag.overview import load_kb_overview
|
||||
from app.rag.prompts import build_deflect_prompt, build_high_prompt, history_to_messages
|
||||
from app.rag.retriever import RetrievedChunk, retrieve, select_documents_tiered, weak_hit_titles
|
||||
from app.rag.retriever import (
|
||||
RetrievedChunk,
|
||||
retrieve,
|
||||
select_related,
|
||||
select_suggested,
|
||||
weak_hit_titles,
|
||||
)
|
||||
from app.rag.scaffolding import ScaffoldingFilter # phase 71: the streaming filter
|
||||
from app.rag.suggestions import derive_suggestions
|
||||
from app.schemas import (
|
||||
@@ -258,12 +281,19 @@ class TurnPlan:
|
||||
fts_hits: int # lexical (OR-tsquery) candidates matched
|
||||
deflected: bool
|
||||
system_prompt: str
|
||||
docs: list[Document] # cited sources (phase 113: the bar-clearing tier)
|
||||
related_docs: list[Document] = field(default_factory=list) # phase 113
|
||||
# The summary-seeded suggestion tier (phase 118, A6: top-N distinct
|
||||
# documents, NO floor, A3 — the HIGH prompt carries their SUMMARIES;
|
||||
# the done frame's citation surface + run_agent's seed_docs, A4).
|
||||
suggested_docs: list[Document]
|
||||
# Rank 6+ after the suggested set (phase 118: the next ranked docs
|
||||
# that are not already suggested, at most related_max_docs — the
|
||||
# done frame's de-emphasized "nearby docs" row, never a citation).
|
||||
related_docs: list[Document] = field(default_factory=list)
|
||||
suggestions: list[str] = field(default_factory=list) # "Maybe try" chips (deflected turns only)
|
||||
tuning_count: int = 0 # steering notes injected into the system prompt
|
||||
#: Hit chunks with ``is_summary`` whose parent document made it into
|
||||
#: *docs* (phase 30; per-turn log line ``summary_hits=N``).
|
||||
#: Hit chunks with ``is_summary`` whose parent document is in the
|
||||
#: SUGGESTED set (phase 30; redefined from the phase-113 cited set
|
||||
#: in phase 118; per-turn log line ``summary_hits=N``).
|
||||
summary_hits: int = 0
|
||||
#: Length of the stored KB overview injected as the
|
||||
#: ``<knowledge_base>`` section (phase 31; per-turn log line
|
||||
@@ -281,7 +311,8 @@ def plan_turn(
|
||||
|
||||
* **HIGH (grounded)** when ``best_cosine >= threshold`` **or**
|
||||
(``fts_hits > 0`` **and** ``best_cosine >= lexical_support_floor``):
|
||||
HIGH prompt with the full top-N documents, no suggestions. A cosine
|
||||
HIGH prompt seeded with the suggested documents' SUMMARIES (phase
|
||||
118, A6 — never their full texts), no suggestions. A cosine
|
||||
exactly at the threshold is an answer — the gate is strict
|
||||
(``< threshold``). An FTS hit alone, without vector corroboration
|
||||
(cosine < lexical_support_floor), stays LOW (A8 revised 2026-09-14).
|
||||
@@ -305,42 +336,52 @@ def plan_turn(
|
||||
prompts are byte-identical to the pre-phase text and ``kb_chars``
|
||||
is 0.
|
||||
|
||||
``summary_hits`` (phase 30) counts the hit chunks with
|
||||
``is_summary`` whose parent document is among the selected
|
||||
top-N documents — both the HIGH and the LOW branch record it.
|
||||
``summary_hits`` (phase 30; phase 118 redefinition) counts the hit
|
||||
chunks with ``is_summary`` whose parent document is in the
|
||||
SUGGESTED set — both the HIGH and the LOW branch record it.
|
||||
|
||||
Phase 113 (the usefulness bar, LOCKED A2): retrieval documents are
|
||||
tiered before either branch — ``docs`` (cited) are the distinct
|
||||
parent documents whose best hit-chunk cosine clears
|
||||
``settings.source_usefulness_floor`` (at most ``top_n_docs`` — the
|
||||
ceiling, never a quota); ``related_docs`` are the next ranked
|
||||
documents (at most ``related_max_docs``) that did not earn a cited
|
||||
slot. On a deflected turn the weak hits fall to ``related_docs``
|
||||
(the cited tier is usually empty — nothing below the bar earned a
|
||||
citation slot); the LOW prompt itself is unchanged (weak-hit titles
|
||||
only).
|
||||
Phase 118 (the summary seeding, LOCKED A3/A6): retrieval documents
|
||||
are tiered once, before either branch — ``suggested_docs`` are the
|
||||
top-N distinct parent documents in fused rank order with NO cosine
|
||||
floor (``settings.suggested_docs``, default 5 — the "start here"
|
||||
seeding: the HIGH prompt carries their SUMMARIES, never their full
|
||||
texts; the floor never filters, so a lexical-only hit is suggested
|
||||
when it ranks) and ``related_docs`` are the next ranked documents
|
||||
that are not already suggested (rank 6+ for the contiguous top-5
|
||||
suggestion set, at most ``related_max_docs`` — the done frame's
|
||||
de-emphasized "nearby docs" row). On a deflected turn the weak hits
|
||||
are suggested too (no floor); the LOW prompt itself is unchanged
|
||||
(weak-hit titles only), and both tiers still ride the TurnPlan for
|
||||
the durable record (LOCKED A3: query_log records retrieval, not
|
||||
citations).
|
||||
|
||||
``settings.top_n_docs`` and ``settings.source_usefulness_floor`` are
|
||||
NOT consulted here (phase 118 retired their seeding role, A6) —
|
||||
they remain settings for env back-compat only.
|
||||
"""
|
||||
steering = list(notes or [])
|
||||
kb_text = (kb_overview or "").strip()
|
||||
kb_chars = len(kb_text)
|
||||
best_cosine = max((c.cosine for c in chunks), default=0.0)
|
||||
fts_hits = sum(1 for c in chunks if c.fts_hit)
|
||||
docs, related_docs = select_documents_tiered(
|
||||
chunks,
|
||||
n=settings.top_n_docs,
|
||||
floor=settings.source_usefulness_floor,
|
||||
related_cap=settings.related_max_docs,
|
||||
# Phase 118 (A3/A6): BOTH tiers, computed once, for BOTH branches —
|
||||
# the suggested tier (top-N, NO floor) seeds the HIGH prompt and the
|
||||
# agent; the related tier (rank 6+ after the suggested set) feeds
|
||||
# the done frame's row and the durable record.
|
||||
suggested = select_suggested(chunks, n=settings.suggested_docs)
|
||||
related_docs = select_related(
|
||||
chunks, {d.id for d in suggested}, settings.related_max_docs
|
||||
)
|
||||
selected_ids = {d.id for d in docs}
|
||||
summary_hits = sum(1 for c in chunks if c.is_summary and c.document.id in selected_ids)
|
||||
suggested_ids = {d.id for d in suggested}
|
||||
summary_hits = sum(1 for c in chunks if c.is_summary and c.document.id in suggested_ids)
|
||||
lexical_supported = fts_hits > 0 and best_cosine >= settings.lexical_support_floor
|
||||
if best_cosine >= settings.relevance_threshold or lexical_supported:
|
||||
return TurnPlan(
|
||||
best_cosine,
|
||||
fts_hits,
|
||||
False,
|
||||
build_high_prompt(docs, notes=steering, kb_overview=kb_text),
|
||||
docs,
|
||||
build_high_prompt(suggested, notes=steering, kb_overview=kb_text),
|
||||
suggested,
|
||||
related_docs,
|
||||
[],
|
||||
len(steering),
|
||||
@@ -353,7 +394,7 @@ def plan_turn(
|
||||
fts_hits,
|
||||
True,
|
||||
build_deflect_prompt(titles, notes=steering, kb_overview=kb_text),
|
||||
docs,
|
||||
suggested,
|
||||
related_docs,
|
||||
derive_suggestions(titles, settings.suggestions),
|
||||
len(steering),
|
||||
@@ -606,7 +647,7 @@ async def chat(
|
||||
db_factory, # SEC-14-04: session factory, not a long-lived session
|
||||
system_prompt=plan.system_prompt,
|
||||
user_message=request.message,
|
||||
seed_docs=plan.docs,
|
||||
seed_docs=plan.suggested_docs, # phase 118 (A4): the suggestion tier
|
||||
settings=settings,
|
||||
holder=holder,
|
||||
history=hist, # phase 74: the same trimmed prior turns
|
||||
@@ -804,30 +845,33 @@ async def chat(
|
||||
# 4. Durable record + required per-turn log line (PLAN §9).
|
||||
# Phase 37: the agent's read documents join the
|
||||
# retrieval's — deduped by (source, path), order
|
||||
# preserved (the read doc stays last). Phase 113:
|
||||
# the retrieval now arrives in two tiers — the cited
|
||||
# docs (``plan.docs``) and the related docs
|
||||
# (``plan.related_docs``, the scored-but-below-the-bar
|
||||
# documents). The DURABLE record keeps the full
|
||||
# preserved (a read doc not already suggested stays
|
||||
# last). Phase 118: the retrieval arrives in two
|
||||
# tiers — the suggested docs
|
||||
# (``plan.suggested_docs``, the summary-seeded
|
||||
# "start here" tier, no floor, A3) and the related
|
||||
# docs (``plan.related_docs``, rank 6+ after the
|
||||
# suggested set). The DURABLE record keeps the full
|
||||
# retrieval (LOCKED A3: query_log records retrieval,
|
||||
# not citations — even on deflected turns, where the
|
||||
# weak hits live in the related tier). Phase 112:
|
||||
# weak hits are suggested). Phase 112/118:
|
||||
# done.sources is the CITATION surface — it carries the
|
||||
# cited docs + the agent-read docs on grounded turns
|
||||
# and [] on deflected ones (a deflected answer cites
|
||||
# nothing; the weak hits stay in the durable record).
|
||||
# suggested docs + the agent-read docs (deduped,
|
||||
# LOCKED A4) on grounded turns and [] on deflected
|
||||
# ones (a deflected answer cites nothing; the weak
|
||||
# hits stay in the durable record).
|
||||
# A cancelled turn (the generator closed by the
|
||||
# consumer) never reaches this step — no query_log row.
|
||||
cited_docs: list[Document] = []
|
||||
cited_seen: set[tuple[str, str]] = set()
|
||||
for doc in [*plan.docs, *holder.read_docs]:
|
||||
for doc in [*plan.suggested_docs, *holder.read_docs]:
|
||||
key = (doc.source, doc.path)
|
||||
if key not in cited_seen:
|
||||
cited_seen.add(key)
|
||||
cited_docs.append(doc)
|
||||
record_docs: list[Document] = []
|
||||
seen: set[tuple[str, str]] = set()
|
||||
for doc in [*plan.docs, *plan.related_docs, *holder.read_docs]:
|
||||
for doc in [*plan.suggested_docs, *plan.related_docs, *holder.read_docs]:
|
||||
key = (doc.source, doc.path)
|
||||
if key not in seen:
|
||||
seen.add(key)
|
||||
@@ -852,7 +896,8 @@ async def chat(
|
||||
logger.exception("chat: failed to write query_log question=%r", request.message)
|
||||
|
||||
logger.info(
|
||||
"question=%r embed_ms=%d top_score=%.3f fts_hits=%d summary_hits=%d tuning=%d "
|
||||
"question=%r embed_ms=%d top_score=%.3f fts_hits=%d summary_hits=%d "
|
||||
"suggested=%d tuning=%d "
|
||||
"kb_chars=%d history_msgs=%d threshold=%.2f deflected=%s sources=%r "
|
||||
"thinking_chars=%d tool_calls=%d total_ms=%d retries=%d scaffold_stripped=%d",
|
||||
request.message,
|
||||
@@ -860,6 +905,7 @@ async def chat(
|
||||
plan.top_score,
|
||||
plan.fts_hits,
|
||||
plan.summary_hits,
|
||||
len(plan.suggested_docs), # phase 118: the seeded suggestion tier size
|
||||
plan.tuning_count,
|
||||
plan.kb_chars,
|
||||
len(hist),
|
||||
@@ -879,14 +925,15 @@ async def chat(
|
||||
# citations). The retrieval stays durably recorded above
|
||||
# (query_log.sources + the log line — observability
|
||||
# unchanged).
|
||||
# Phase 113 (A2/A4): done.related carries the related
|
||||
# tier — the scored documents that did not clear the
|
||||
# usefulness bar (deduped against the cited list, the
|
||||
# same (source, path) pattern as cited_docs: an agent-
|
||||
# read related doc is a citation, never a "nearby doc"
|
||||
# — and capped by related_max_docs in the tiering). The
|
||||
# UI renders it as the de-emphasized related-docs row,
|
||||
# never a citation chip; old clients ignore the field.
|
||||
# Phase 118 (A3/A4): done.related carries the related
|
||||
# tier — the ranked documents beyond the suggested set
|
||||
# (rank 6+ after the contiguous top-N suggestion, capped
|
||||
# by related_max_docs in the tiering), deduped against
|
||||
# the cited list, the same (source, path) pattern as
|
||||
# cited_docs: an agent-read related doc is a citation,
|
||||
# never a "nearby doc". The UI renders it as the
|
||||
# de-emphasized related-docs row, never a citation chip;
|
||||
# old clients ignore the field.
|
||||
cited_refs: list[SourceRef] = []
|
||||
if not plan.deflected:
|
||||
cited_refs = [
|
||||
|
||||
@@ -111,6 +111,9 @@ class Settings(BaseSettings):
|
||||
|
||||
# --- RAG tuning ---
|
||||
embedding_dim: int = 768 # verified against aipi /v1 (embed model)
|
||||
#: Phase 118 retired the full-text seeding role (A6); the suggested
|
||||
#: tier (``select_suggested``) seeds the prompt now — kept for env
|
||||
#: back-compat (no ``app/`` consumer left).
|
||||
top_n_docs: int = 2
|
||||
# Honesty gate (A8, re-tuned 2026-08-21): the ``embed`` model's cosine
|
||||
# scores compress into 0.41–0.84 on the real corpus, so the old 0.30
|
||||
@@ -126,6 +129,9 @@ class Settings(BaseSettings):
|
||||
# ``relevance_threshold`` (a floor above the threshold is a typo that
|
||||
# would make every FTS hit require a HIGH cosine anyway).
|
||||
lexical_support_floor: float = 0.35
|
||||
#: Phase 118 retired the full-text seeding role (A6); the suggested
|
||||
#: tier (``select_suggested``) seeds the prompt now (no floor, A3) —
|
||||
#: kept for env back-compat (no ``app/`` consumer left).
|
||||
#: Usefulness bar for the citation slot (phase 113, LOCKED A2): a
|
||||
#: retrieved document earns ``done.sources`` (the UI's citation chip)
|
||||
#: only when the **cosine** of its best hit chunk clears this floor —
|
||||
@@ -149,6 +155,26 @@ class Settings(BaseSettings):
|
||||
#: docs at all (the kill switch); a negative value fails startup
|
||||
#: loudly (the ``agent_max_rounds`` pattern).
|
||||
related_max_docs: int = 2
|
||||
#: Cap on the "start here" suggestion tier (phase 118, LOCKED A3 —
|
||||
#: the owner directive, TODO L3): a grounded turn seeds the top-N
|
||||
#: related documents into the prompt as SUMMARY blocks (opt-in
|
||||
#: starting points, never citations) and the LLM extends its context
|
||||
#: by reading only what it needs. NO cosine floor applies — unlike
|
||||
#: the cited tier's ``source_usefulness_floor``, a lexical-only hit
|
||||
#: (cosine 0.0) is a valid starting point when it ranks. Default 5;
|
||||
#: tunable via ``BOR_SUGGESTED_DOCS``. A value below 1 is a typo —
|
||||
#: the validator fails startup loudly (the ``agent_max_rounds``
|
||||
#: pattern).
|
||||
suggested_docs: int = 5
|
||||
#: Preview cap for a suggestion block whose document summary is missing
|
||||
#: (phase 118, task 03, LOCKED A5): a NULL/blank ``doc.summary`` (a
|
||||
#: fail-soft import miss) falls back to the first ``suggestion_preview_chars``
|
||||
#: characters of the document content plus the shared
|
||||
#: ``[…truncated…]`` marker — deterministic, no LLM call at chat time.
|
||||
#: Tolerates content at or under the cap whole (no marker — nothing was
|
||||
#: cut). ``0``/negative is a typo (empty preview) — the validator fails
|
||||
#: startup loudly (the ``agent_max_rounds`` pattern).
|
||||
suggestion_preview_chars: int = 400
|
||||
#: Maximum output tokens a chat answer may use (owner instruction
|
||||
#: 2026-08-22: answers must run to their natural end — the old hard
|
||||
#: 700-token cap cut long answers off mid-sentence).
|
||||
@@ -387,6 +413,25 @@ class Settings(BaseSettings):
|
||||
raise ValueError("related_max_docs must be >= 0 (0 = no related docs)")
|
||||
return v
|
||||
|
||||
@field_validator("suggested_docs")
|
||||
@classmethod
|
||||
def _suggested_docs_at_least_one(cls, v: int) -> int:
|
||||
"""The suggestion tier always seeds at least one summary block —
|
||||
``0`` (no starting points) and negatives are typos (the
|
||||
``agent_max_rounds`` pattern, phase 118)."""
|
||||
if v < 1:
|
||||
raise ValueError("suggested_docs must be >= 1")
|
||||
return v
|
||||
|
||||
@field_validator("suggestion_preview_chars")
|
||||
@classmethod
|
||||
def _suggestion_preview_chars_positive(cls, v: int) -> int:
|
||||
"""``0``/negative would preview an empty/absent prefix — fail loud at
|
||||
startup (the ``agent_max_rounds`` pattern, phase 118)."""
|
||||
if v <= 0:
|
||||
raise ValueError("suggestion_preview_chars must be > 0 (chars)")
|
||||
return v
|
||||
|
||||
@field_validator("import_extensions")
|
||||
@classmethod
|
||||
def _import_extensions_known(cls, v: str) -> str:
|
||||
|
||||
+5
-4
@@ -133,10 +133,11 @@ class Document(Base):
|
||||
Boolean, default=False, server_default=text("false"), nullable=False
|
||||
)
|
||||
#: Lite-model summary, phase 30. Natural-language summary of the
|
||||
#: document (non-markdown A9 docs only, generated at import time by the
|
||||
#: aipi ``lite`` model). NULL for markdown docs, pre-phase-30 rows, and
|
||||
#: the fail-soft path where summary generation failed but the document
|
||||
#: was still indexed.
|
||||
#: document (every A9 doc — markdown included since phase 118, A2 —
|
||||
#: generated at import time by the aipi ``lite`` model). NULL for
|
||||
#: pre-phase-30 rows, the fail-soft path where summary generation
|
||||
#: failed, and until the phase-118 backfill stores one on the next
|
||||
#: sync.
|
||||
summary: Mapped[str | None] = mapped_column(Text, default=None)
|
||||
|
||||
chunks: Mapped[list[Chunk]] = relationship(
|
||||
|
||||
+44
-32
@@ -97,12 +97,13 @@ task 04):
|
||||
context; ``grep`` — which searches the whole document — is the
|
||||
follow-up), with the truncation recorded on the holder so the loop
|
||||
yields a :class:`app.rag.llm.ToolResultPiece` (task 02 → the SSE
|
||||
``tool_result`` frame + UI marker). **A7 scope clarification:** the
|
||||
never-truncated contract is for the retrieval ``<documents>`` path
|
||||
(the top-2 seed documents stay whole — "this should never happen");
|
||||
the ``read`` TOOL path is the only capped read, per the owner's
|
||||
explicit request — the two paths are distinct (retrieval seeds vs.
|
||||
agent-requested additions). And ``grep`` greps the
|
||||
``tool_result`` frame + UI marker). **A6 re-revised contract
|
||||
(phase 118, owner directive 2026-09-15):** the retrieval
|
||||
``<documents>`` path seeds SUMMARIES only — a suggested document's
|
||||
full text never enters the prompt on the retrieval path; full text
|
||||
enters the context ONLY through this ``read`` TOOL path, which is
|
||||
the only capped read (the phase-95 cap unchanged). And ``grep``
|
||||
greps the
|
||||
indexed documents (or the one document a combined ``source/path``
|
||||
names) for a case-insensitive fixed substring and returns up to 20
|
||||
``source/path:line: text`` match lines (owner-locked A5, phase 68),
|
||||
@@ -138,7 +139,9 @@ task 04):
|
||||
:data:`NOT_A_FOLDER`, the drill-down teaching with the argument
|
||||
echoed and the deepest existing ancestor's direct subfolders
|
||||
listed, so the model self-corrects in the next round; a document
|
||||
already in context (seed or previously read) →
|
||||
already READ into full-text context (phase 118: the suggested
|
||||
seeds are summary blocks in the prompt, not full text — only an
|
||||
already-read document is refused) →
|
||||
:data:`ALREADY_IN_CONTEXT` (phase 72, task 05 gate iteration:
|
||||
the line names the correct action — answer from the text already
|
||||
in the prompt, do not call read again — so a fired refusal ends
|
||||
@@ -301,19 +304,21 @@ AGENT_TOOLS: list[dict[str, Any]] = [
|
||||
"function": {
|
||||
"name": "read",
|
||||
"description": (
|
||||
"Do not call this tool for a document already shown in "
|
||||
"the <documents> section, even when the user asks you to "
|
||||
"open or read it — its full text is already in your "
|
||||
"prompt; answer directly from it. Use it only to add a "
|
||||
"document NOT already in <documents> to your context, "
|
||||
"by its combined `source/path` string. Very large "
|
||||
"documents are truncated: you receive the first part "
|
||||
"plus a TRUNCATED notice naming how many more characters "
|
||||
"exist — the notice is authoritative, the document did "
|
||||
"NOT end where it stopped. Follow it and use `grep` "
|
||||
"(pattern) to locate the rest — it searches the whole "
|
||||
"document. Call one tool at a time — wait for this "
|
||||
"result before your next call."
|
||||
"The <documents> section shows the SUMMARIES of the "
|
||||
"top-ranked documents — their full texts are NOT in "
|
||||
"your prompt yet. Use this tool to add one of them (or "
|
||||
"any other document) to your context, by its combined "
|
||||
"`source/path` string, exactly as shown in the `ls` "
|
||||
"output or the <documents> blocks. Do not re-read a "
|
||||
"document you have already read — its full text is "
|
||||
"already in your prompt. Very large documents are "
|
||||
"truncated: you receive the first part plus a TRUNCATED "
|
||||
"notice naming how many more characters exist — the "
|
||||
"notice is authoritative, the document did NOT end "
|
||||
"where it stopped. Follow it and use `grep` (pattern) "
|
||||
"to locate the rest — it searches the whole document. "
|
||||
"Call one tool at a time — wait for this result before "
|
||||
"your next call."
|
||||
),
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
@@ -326,10 +331,9 @@ AGENT_TOOLS: list[dict[str, Any]] = [
|
||||
"shown in the `ls` output (e.g. "
|
||||
"'homelab/active/container_caddy/caddy.md'). "
|
||||
"A bare document path (without the source "
|
||||
"name) will not resolve. Only pass a document "
|
||||
"NOT already shown in the <documents> "
|
||||
"section — it is already in your context; do "
|
||||
"not re-read it."
|
||||
"name) will not resolve. Do not re-read a "
|
||||
"document you have already read — its full "
|
||||
"text is already in your prompt."
|
||||
),
|
||||
}
|
||||
},
|
||||
@@ -1152,10 +1156,15 @@ def _execute_tool(
|
||||
arg = raw_path.strip() if isinstance(raw_path, str) else ""
|
||||
if not arg:
|
||||
return MISSING_READ_ARGS
|
||||
known = {(doc.source, doc.path) for doc in (*seed_docs, *holder.read_docs)}
|
||||
# Phase 118 (A6): the dedupe set is ``holder.read_docs`` ONLY —
|
||||
# the ``seed_docs`` are SUMMARY blocks in the prompt, not full
|
||||
# text, so a FIRST read of a suggested document adds its full
|
||||
# text through the path below; only a document ALREADY READ is
|
||||
# refused.
|
||||
known = {(doc.source, doc.path) for doc in holder.read_docs}
|
||||
# The dedupe check needs no DB: the split pair of a combined
|
||||
# identity that is in context is in `known` as-is (the resolve
|
||||
# below would find the same document).
|
||||
# identity that is already in full-text context is in `known`
|
||||
# as-is (the resolve below would find the same document).
|
||||
if "/" in arg:
|
||||
src, _, p = arg.partition("/")
|
||||
if (src, p) in known:
|
||||
@@ -1326,11 +1335,14 @@ async def run_agent(
|
||||
the capture mechanism for new registry entries; *holder* accumulates
|
||||
the turn's ``scaffold_stripped`` total for the API layer's log line.
|
||||
|
||||
``seed_docs`` are the documents the retrieval already put in context
|
||||
(they shape the *system_prompt* the caller built); re-reading one of
|
||||
them is rejected with :data:`ALREADY_IN_CONTEXT` (the phase-72
|
||||
teaching line — answer from the text already in the prompt) — the
|
||||
rejection counts in nothing, but it still consumes a round.
|
||||
``seed_docs`` are the suggested documents whose SUMMARY blocks the
|
||||
caller put in the *system_prompt* (phase 118: the retrieval seeds
|
||||
summaries, never full texts); reading one of them ADDS its full
|
||||
text to the context through the ordinary ``read`` path — only a
|
||||
document ALREADY READ is rejected with :data:`ALREADY_IN_CONTEXT`
|
||||
(the phase-72 teaching line — answer from the text already in the
|
||||
prompt); the rejection counts in nothing, but it still consumes a
|
||||
round.
|
||||
|
||||
DB sessions (SEC-14-04): *db_factory* is a callable that returns a
|
||||
new :class:`sqlalchemy.orm.Session` (e.g. ``lambda: SessionLocal()``).
|
||||
|
||||
+60
-10
@@ -10,7 +10,8 @@ the two-phase upsert:
|
||||
2. embed the new chunks in batches and attach the vectors
|
||||
3. commit — one transaction per file, so a failed embedding leaves the
|
||||
database untouched and the file is simply retried on the next run
|
||||
4. non-markdown files only (phase 30): generate a ``lite``-model summary
|
||||
4. every file (phase 30; phase 118, A2: markdown included — the
|
||||
non-markdown-only scope is retired): generate a ``lite``-model summary
|
||||
and, best-effort, store it on ``documents.summary`` plus one extra
|
||||
embedded chunk (``is_summary``, position −1). The document row and its
|
||||
content chunks are already committed at this point, so a summary
|
||||
@@ -42,6 +43,17 @@ guard) and counted in ``summary.dates_updated`` — unless the row
|
||||
carries the owner's manual correction (``created_at_manual``, D1), which
|
||||
the sync never touches.
|
||||
|
||||
NULL-summary backfill (phase 118, A2): an UNCHANGED file (same
|
||||
``content_hash``) whose ``documents.summary`` is still NULL — a
|
||||
pre-phase-30 row, or an earlier fail-soft miss — gets the same
|
||||
best-effort summary pass on every sync until it sticks. A success counts
|
||||
``summary_backfilled`` (never ``summaries``) and touches nothing else: no
|
||||
content re-embed, no added/updated/pruned count — so no
|
||||
``sources_meta`` bump, no KB-overview/folder-summary regeneration. The
|
||||
backfill runs BEFORE the ``created_at_manual`` early-return (the manual
|
||||
flag protects the DATE only, D1) and the strict ``is None`` check leaves
|
||||
owner-set summaries (even empty strings, phase 57) alone.
|
||||
|
||||
``import_sources`` accepts an optional per-file ``progress`` callback
|
||||
(phase 64, task 01) reporting the file being processed right now.
|
||||
"""
|
||||
@@ -99,12 +111,18 @@ class ImportSummary:
|
||||
errors: int = 0
|
||||
chunks: int = 0
|
||||
embed_batches: int = 0
|
||||
#: Non-markdown files whose lite summary was generated + indexed
|
||||
#: (phase 30). One ``is_summary`` chunk per success.
|
||||
#: Files whose lite summary was generated + indexed (phase 30; phase
|
||||
#: 118, A2: every A9 format, markdown included). One ``is_summary``
|
||||
#: chunk per success.
|
||||
summaries: int = 0
|
||||
#: Non-markdown files whose summary generation failed (best-effort —
|
||||
#: the document is still indexed, without a summary).
|
||||
#: Files whose summary generation failed (best-effort — the document
|
||||
#: is still indexed, without a summary).
|
||||
summary_errors: int = 0
|
||||
#: Unchanged docs whose NULL summary was backfilled (phase 118, A2) —
|
||||
#: one ``is_summary`` chunk per success; the content is untouched, so
|
||||
#: a backfill NEVER counts added/updated/pruned (no
|
||||
#: ``sources_meta`` bump, no overview/folder-summary regeneration).
|
||||
summary_backfilled: int = 0
|
||||
#: Files whose ``created_at`` was refreshed on the UNCHANGED path —
|
||||
#: content untouched, date re-sourced (phase 106, D4: the date may
|
||||
#: go OLDER; a date-only refresh NEVER counts added/updated/pruned,
|
||||
@@ -125,7 +143,7 @@ class ImportSummary:
|
||||
logger.info(
|
||||
"import: summary files=%d added=%d updated=%d unchanged=%d pruned=%d "
|
||||
"errors=%d chunks=%d embed_batches=%d summaries=%d summary_errors=%d "
|
||||
"dates_updated=%d formats=%s",
|
||||
"summary_backfilled=%d dates_updated=%d formats=%s",
|
||||
self.files,
|
||||
self.added,
|
||||
self.updated,
|
||||
@@ -136,6 +154,7 @@ class ImportSummary:
|
||||
self.embed_batches,
|
||||
self.summaries,
|
||||
self.summary_errors,
|
||||
self.summary_backfilled,
|
||||
self.dates_updated,
|
||||
self.format_counts(),
|
||||
)
|
||||
@@ -442,6 +461,20 @@ async def _index_file(
|
||||
if doc is not None and doc.content_hash == digest:
|
||||
summary.unchanged += 1
|
||||
logger.info("import: unchanged source=%s path=%s", source, rel)
|
||||
# Phase 118 (A2): an unchanged doc whose summary is still NULL
|
||||
# (a pre-phase-30 row, or an earlier fail-soft miss) gets a
|
||||
# summary-only backfill — one ``is_summary`` chunk, no content
|
||||
# re-embed, and NEVER an added/updated/pruned count (so no
|
||||
# ``sources_meta`` bump, no overview/folder-summary
|
||||
# regeneration). Strict ``is None``: an empty-string summary is
|
||||
# owner-set (phase 57) and is never overwritten. BEFORE the
|
||||
# manual-date early-return: ``created_at_manual`` protects the
|
||||
# DATE only (phase 106, D1), not the summary.
|
||||
if doc.summary is None:
|
||||
await _store_summary(
|
||||
session, doc=doc, source=source, rel=rel, content=content,
|
||||
llm=llm, summary=summary, backfill=True,
|
||||
)
|
||||
if doc.created_at_manual:
|
||||
# D1/D4: the owner's correction survives the sync — no
|
||||
# write at all (the phase-97 ``manually_edited`` precedent).
|
||||
@@ -536,10 +569,11 @@ async def _index_file(
|
||||
summary.chunks += len(chunks_text)
|
||||
logger.info("import: %s source=%s path=%s chunks=%d", verb, source, rel, len(chunks_text))
|
||||
|
||||
# Phase 30: markdown is already natural language, so only the other A9
|
||||
# formats (txt, yaml, yml, json, py) get a ``lite``-model summary.
|
||||
if full_path.suffix.lower() in (".md", ".markdown"):
|
||||
return
|
||||
# Phase 30; phase 118 (A2, 2026-09-15): EVERY new/changed document
|
||||
# gets a ``lite``-model summary — markdown included. Phase 30's
|
||||
# "markdown is already natural language" exclusion is retired: the
|
||||
# summary is the retrieval seed context (the phase-118 suggestion
|
||||
# blocks), not a formatting convenience.
|
||||
await _store_summary(
|
||||
session, doc=doc, source=source, rel=rel, content=content, llm=llm, summary=summary
|
||||
)
|
||||
@@ -554,6 +588,7 @@ async def _store_summary(
|
||||
content: str,
|
||||
llm: Embedder,
|
||||
summary: ImportSummary,
|
||||
backfill: bool = False,
|
||||
) -> None:
|
||||
"""Best-effort ``lite`` summary for one already-committed document.
|
||||
|
||||
@@ -569,6 +604,11 @@ async def _store_summary(
|
||||
:class:`EmbeddingError` only rolls back the summary rows — the file
|
||||
stays indexed, without a summary, and the failure is counted in
|
||||
``summary_errors`` (PLAN phase 30).
|
||||
|
||||
``backfill`` (phase 118, A2): the unchanged-doc NULL-summary path —
|
||||
a success counts ``summary_backfilled`` instead of ``summaries``
|
||||
(the doc content is untouched, so the import's KB-change signal must
|
||||
not move); the rest of the mechanics are identical.
|
||||
"""
|
||||
try:
|
||||
text = await generate_summary(llm, source=source, path=rel, content=content)
|
||||
@@ -588,6 +628,16 @@ async def _store_summary(
|
||||
# ``expire_on_commit=False`` — reflects the committed state.
|
||||
doc.chunks.append(chunk)
|
||||
session.commit()
|
||||
if backfill:
|
||||
# Phase 118 (A2): the backfill counts itself apart from fresh
|
||||
# imports — the doc content is unchanged, so ``summaries``
|
||||
# (a KB-change signal) must not move.
|
||||
summary.summary_backfilled += 1
|
||||
logger.info(
|
||||
"import: summary-backfill source=%s path=%s chars=%d",
|
||||
source, rel, len(text),
|
||||
)
|
||||
else:
|
||||
summary.summaries += 1
|
||||
logger.info("import: summary source=%s path=%s chars=%d", source, rel, len(text))
|
||||
except (LLMError, EmbeddingError) as e:
|
||||
|
||||
+3
-2
@@ -1,8 +1,9 @@
|
||||
"""Async OpenAI-compatible client for the self-hosted aipi endpoint (PLAN A5).
|
||||
|
||||
Provides the embeddings surface (importer, retrieval), one-shot chat
|
||||
completions (phase 30: the ``lite`` model summarizes non-markdown
|
||||
documents at import time), and chat streaming (PLAN A15) for the RAG
|
||||
completions (phase 30: the ``lite`` model summarizes documents at import
|
||||
time — every A9 format, markdown included since phase 118), and chat
|
||||
streaming (PLAN A15) for the RAG
|
||||
pipeline. Chat streaming yields typed :class:`StreamPiece` values
|
||||
(phase 17) and — when the caller passes a ``tools`` list —
|
||||
:class:`ToolCallPiece` values (phase 37): aipi's ``turbo`` model streams
|
||||
|
||||
+3
-2
@@ -100,8 +100,9 @@ def build_overview_prompt(
|
||||
* ``user`` — one line per document,
|
||||
``source — path — title — {first line of summary}``, joined with
|
||||
newlines. The summary field is omitted when the document has no
|
||||
summary (markdown docs and the fail-soft path — no dangling
|
||||
dash). The list is capped at *max_chars* (default
|
||||
summary (pre-phase-30 rows, the fail-soft path, and pre-backfill
|
||||
NULL rows — no dangling dash). The list is capped at *max_chars*
|
||||
(default
|
||||
``BOR_OVERVIEW_INPUT_MAX_CHARS``): overflow is cut exactly at the
|
||||
cap and the shared ``[…truncated…]`` marker is appended on its own
|
||||
line, so the model never sees more than the cap and the cut is
|
||||
|
||||
+99
-41
@@ -7,7 +7,14 @@ no mandated deflection opening; the honesty gate itself is unchanged.)
|
||||
|
||||
Two modes:
|
||||
|
||||
* ``HIGH`` — grounded turn: full top-document texts under ``<documents>``.
|
||||
* ``HIGH`` — grounded turn: the top-ranked documents' SUMMARIES under
|
||||
``<documents>`` (phase 118, LOCKED A6 re-revising A7 — the "start here"
|
||||
suggestion seeding: each block is the document's stored summary, never
|
||||
the full content; the LLM extends its context by ``read``-ing a
|
||||
document's full text through the capped ``read`` tool, the ONLY full-
|
||||
text path. A NULL/blank summary — a fail-soft import miss — falls back
|
||||
to a ``suggestion_preview_chars`` content preview + the shared
|
||||
``[…truncated…]`` marker; no LLM call at chat time, LOCKED A5).
|
||||
* ``LOW`` — deflection turn: weak-hit *titles only* plus the
|
||||
``DEFLECT_MODE`` marker (the E2E mock LLM keys on that marker).
|
||||
|
||||
@@ -124,6 +131,24 @@ _KB_INTRO = (
|
||||
"(generated at import time):\n"
|
||||
)
|
||||
|
||||
#: The ``<documents>`` section's start-here intro (phase 118, task 03 —
|
||||
#: the owner directive, TODO L3): the seeded blocks are the SUMMARIES of
|
||||
#: the top-ranked documents for the question, opt-in starting points —
|
||||
#: start here if one seems right; ``read`` the document's combined
|
||||
#: ``source/path`` to add its FULL content before answering specifics
|
||||
#: (the full text is NOT in the prompt until you read it); cite the
|
||||
#: document(s) you used, by path. Rendered BEFORE the first block, only
|
||||
#: when at least one block is present. The ``<documents>`` tag and the
|
||||
#: per-block ``<document>`` markup stay byte-stable around it (the E2E
|
||||
#: mock's block parser and the ``read``-tool copy key off both).
|
||||
SUGGEST_INTRO = (
|
||||
"The blocks below are the summaries of the top-ranked documents for "
|
||||
"your question — start here if one seems right to you: call `read` "
|
||||
"with that document's combined `source/path` to add its full content "
|
||||
"before answering specifics (its full text is not in the prompt until "
|
||||
"you read it). Cite the document(s) you used, by path."
|
||||
)
|
||||
|
||||
#: The ``<tools>`` instructions section — **HIGH prompt only** (phase 37,
|
||||
#: task 03; phase 70: the copy is rewritten for the harness-aligned
|
||||
#: ``ls`` / ``read`` / ``grep`` shapes, names/args exactly as the
|
||||
@@ -158,12 +183,15 @@ _KB_INTRO = (
|
||||
#: the user's "open it / read it" and ``read`` seed-context documents,
|
||||
#: then repeated the refused call; run 2: 8/18, 44% — the repeat is
|
||||
#: gone, but a grep turn batched two calls per reply and the harness
|
||||
#: runs only the first of a batch): the do-not-read rule names the
|
||||
#: user-command scenario (a document already in the ``<documents>``
|
||||
#: section: do not call ``read``, answer from the text already in the
|
||||
#: prompt — an anchor on the concrete ``<document path="...">`` markup
|
||||
#: was tried and REVERTED: it primed the model to latch the seed
|
||||
#: documents' paths as ``ls`` scopes, regressing the incident turn);
|
||||
#: runs only the first of a batch): the ``read`` clause carries the
|
||||
#: phase-118 summary-seed contract (the ``<documents>`` section holds
|
||||
#: SUMMARIES — a suggested document's full text is not in the prompt
|
||||
#: until ``read`` adds it; do not re-read a document you have already
|
||||
#: read — answer from the text already in the prompt; if the user
|
||||
#: asks to open or read a suggested document, ``read`` it — that is
|
||||
#: the point of the section — the phase-72 do-not-read rule and the
|
||||
#: reverted ``<document path="...">`` anchor are retired with the
|
||||
#: full-text seeds);
|
||||
#: the one-call rule names the consequence (a batched second call is
|
||||
#: discarded — runs only the first); the never-repeat rule says why
|
||||
#: (the refusal already told you the correct form); the ``grep`` clause
|
||||
@@ -171,10 +199,11 @@ _KB_INTRO = (
|
||||
#: live runs showed the model scoping ``grep`` with an ``ls``-style
|
||||
#: source name — the incident shape, but on grep). The behavioral
|
||||
#: contract lives in the ``AGENT_TOOLS`` descriptions as well (the most
|
||||
#: local text at call time): ``read`` must not be called for a
|
||||
#: ``<documents>`` document at all; ``grep`` with only ``pattern``
|
||||
#: searches the whole knowledge base, and a source name is not a
|
||||
#: document.
|
||||
#: local text at call time): ``read`` adds a document's FULL text by
|
||||
#: its combined ``source/path`` identity — the ``<documents>``
|
||||
#: summaries are the starting points, not the content; ``grep`` with
|
||||
#: only ``pattern`` searches the whole knowledge base, and a source
|
||||
#: name is not a document.
|
||||
TOOLS_SECTION: str = (
|
||||
"<tools>\n"
|
||||
"You may extend your context with three tools. `ls` lists the "
|
||||
@@ -189,12 +218,16 @@ TOOLS_SECTION: str = (
|
||||
"`source: X | path: Y | title: Z`; to find one specific document "
|
||||
"without listing, use `grep`. `read` pulls in one document by its "
|
||||
"combined `source/path` string, exactly as shown in the `ls` "
|
||||
"output — including the source name — adding its full content to "
|
||||
"your context. Do not call `read` for a document already shown in "
|
||||
"the <documents> section, even when the user asks you to open or "
|
||||
"read it — its full text is already in your prompt; answer "
|
||||
"directly from it. For `read`, a bare document path (without the "
|
||||
"source name) will not resolve. Very large documents are capped: a "
|
||||
"output — including the source name — or in the <documents> "
|
||||
"summary blocks — adding its full content to your context. The "
|
||||
"<documents> section holds SUMMARIES — the full text of a "
|
||||
"suggested document is not in your prompt until you `read` it. Do "
|
||||
"not re-read a document you have already read — its full text is "
|
||||
"already in your prompt; answer directly from it. If the user "
|
||||
"asks you to open or read a suggested document, `read` it — that "
|
||||
"is the point of the section. For `read`, a bare document path "
|
||||
"(without the source name) will not resolve. Very large documents "
|
||||
"are capped: a "
|
||||
"cut read returns the first part plus a TRUNCATED notice — the "
|
||||
"document did not end where it stopped; use `grep` (pattern) to "
|
||||
"find the rest, it searches the whole document. `grep` locates an exact string "
|
||||
@@ -355,17 +388,42 @@ def build_kb_section(overview: str, max_chars: int | None = None) -> str:
|
||||
return ""
|
||||
|
||||
|
||||
def _document_body(doc: Document) -> str:
|
||||
"""The body of one ``<document>`` suggestion block (phase 118, LOCKED
|
||||
A6): the document's stored summary (stripped) — NEVER the full
|
||||
content.
|
||||
|
||||
Defensive fallback ONLY when the summary is missing (``None`` or
|
||||
whitespace — a fail-soft import miss, LOCKED A5): the first
|
||||
``suggestion_preview_chars`` characters of the content plus the
|
||||
shared :data:`TRUNCATION_MARKER` on its own line — the settings read
|
||||
happens on this fallback path ONLY (a prompt built from
|
||||
summary-bearing docs makes no settings call for the cap). Content at
|
||||
or under the cap rides whole, unmarked (nothing was cut). No LLM
|
||||
call at chat time — the preview is deterministic.
|
||||
"""
|
||||
summary = (doc.summary or "").strip()
|
||||
if summary:
|
||||
return summary
|
||||
limit = get_settings().suggestion_preview_chars
|
||||
content = doc.content
|
||||
if len(content) > limit:
|
||||
return content[:limit] + "\n" + TRUNCATION_MARKER
|
||||
return content
|
||||
|
||||
|
||||
def build_high_prompt(
|
||||
documents: Sequence[Document],
|
||||
notes: Sequence[str] | None = None,
|
||||
kb_overview: str | None = None,
|
||||
) -> str:
|
||||
"""Grounded turn: locked persona (+ steering, + KB overview) + full
|
||||
texts of the top documents + the ``<tools>`` instructions (phase 37;
|
||||
phase 70: the harness-aligned ``ls`` / ``read`` / ``grep`` shapes;
|
||||
phase 72: the copy states the document-identity contract — the
|
||||
source-name ``ls`` scope, the combined ``source/path`` identity for
|
||||
``read``/``grep`` — up front).
|
||||
"""Grounded turn: locked persona (+ steering, + KB overview) + the
|
||||
top-ranked documents' SUMMARY blocks + the ``<tools>`` instructions
|
||||
(phase 37; phase 70: the harness-aligned ``ls`` / ``read`` /
|
||||
``grep`` shapes; phase 72: the document-identity contract up front;
|
||||
phase 118, LOCKED A6 re-revising A7: the ``<documents>`` section
|
||||
seeds SUMMARIES, never full texts — full text enters the context
|
||||
only through the capped ``read`` tool).
|
||||
|
||||
Section order: ``<relevance>`` → ``<knowledge_base>`` → ``<tuning>``
|
||||
→ ``<documents>`` → ``<tools>``; empty steering/overview omit their
|
||||
@@ -373,42 +431,42 @@ source-name ``ls`` scope, the combined ``source/path`` identity for
|
||||
cap — not the prompt — decides whether the tools are actually
|
||||
offered to the model, see :mod:`app.rag.agent`).
|
||||
|
||||
When at least one block is present, the section leads with the
|
||||
start-here :data:`SUGGEST_INTRO` line (before the first block — the
|
||||
phase-15 ``_STEERING_INTRO`` / phase-31 ``_KB_INTRO`` precedent): the
|
||||
blocks are the summaries of the top-ranked documents, ``read`` adds
|
||||
the full text, and the answer cites the document(s) used by path.
|
||||
This is NOT the reverted phase-72 in-context reminder (that copy
|
||||
taught the seed documents as already-read context; this one teaches
|
||||
the summary-as-starting-point contract the A6 re-revision requires).
|
||||
|
||||
Each ``<document>`` block carries the identity attributes
|
||||
``source`` / ``path`` / ``title`` — and, since phase 106 (D5),
|
||||
``date`` (the row's ``created_at`` UTC date part, ``YYYY-MM-DD``,
|
||||
APPENDED after ``title`` — the only position; always present,
|
||||
``created_at`` is NOT NULL) — plus the document's full text.
|
||||
|
||||
Gate-iteration note (task 05, 2026-09-03/04): an in-context reminder
|
||||
LEADING this section (the document texts are already context — do
|
||||
not ``read`` one the user asked to open) was tried and REVERTED:
|
||||
it never flipped the seed-doc reads (15/15 across gate runs 1-5)
|
||||
and correlated with the incident-turn regression (the model latched
|
||||
the seed documents' paths as ``ls`` scopes — cap reached on the
|
||||
"list the files in this directory" turn) whenever the copy named
|
||||
the ``<document>`` blocks explicitly.
|
||||
``created_at`` is NOT NULL) — plus the document's summary
|
||||
(:func:`_document_body`; the NULL/whitespace-summary preview
|
||||
fallback is LOCKED A5).
|
||||
"""
|
||||
# 2026-09-04 (controlled tool-calling fast loop): the do-not-read
|
||||
# rule for seed documents lives in TOOLS_SECTION and the ``read``
|
||||
# tool descriptions (the copy levers that stuck — see the gate's
|
||||
# telemetry in TOOL_CALLING_TESTING.md). A per-block instruction
|
||||
# attribute at the ``source``/``path`` copy site was TRIED and
|
||||
# REVERTED the same day (no improvement across runs; the block stays
|
||||
# exactly the document identity + full text).
|
||||
# Phase 106 (D5): every document the model sees carries its
|
||||
# creation date — the block's ``date`` attribute (the row's
|
||||
# ``created_at`` UTC date part, appended after ``title`` — the
|
||||
# only position; always present, ``created_at`` is NOT NULL).
|
||||
# Phase 118 (A6): the block body is the summary — the full content
|
||||
# stays out of the prompt (the capped ``read`` tool is the only
|
||||
# full-text path).
|
||||
blocks = [
|
||||
f'<document source="{doc.source}" path="{doc.path}" title="{doc.title}" '
|
||||
f'date="{doc.created_at:%Y-%m-%d}">\n'
|
||||
f"{doc.content}\n"
|
||||
f"{_document_body(doc)}\n"
|
||||
"</document>"
|
||||
for doc in documents
|
||||
]
|
||||
body = "\n\n".join(blocks) if blocks else (
|
||||
"(no documents matched — do not invent specifics)"
|
||||
)
|
||||
if blocks:
|
||||
body = SUGGEST_INTRO + "\n\n" + body
|
||||
prompt = _base("HIGH")
|
||||
for part in (build_kb_section(kb_overview or ""), build_steering_section(notes or [])):
|
||||
if part:
|
||||
|
||||
+113
-13
@@ -35,12 +35,16 @@
|
||||
ranked documents (up to ``BOR_RELATED_MAX_DOCS``) become the related
|
||||
tier.
|
||||
|
||||
The product requirement (LOCKED A7, revised 2026-08-24): the LLM receives
|
||||
the **entire relevant document**, not just the chunk — chunk hits map back
|
||||
to their parents, dedupe, rank by best fused score, and the full text of
|
||||
the top-N documents is always fed through, never truncated. If a future KB
|
||||
ever makes the prompt too large for the model, the ``LLMError`` → SSE
|
||||
``error`` path surfaces it loudly — no silent partial context.
|
||||
The product requirement (A7, re-revised by the phase-118 owner directive,
|
||||
LOCKED A6, 2026-09-15): the retrieval path seeds **summaries** — the
|
||||
suggestion tier (:func:`select_suggested`, top-N distinct documents, no
|
||||
cosine floor, LOCKED A3) whose summary blocks are the grounded prompt's
|
||||
``<documents>`` starting points. The full text of a document enters the
|
||||
context ONLY through the agent's capped ``read`` tool
|
||||
(:mod:`app.rag.agent`; ``BOR_READ_MAX_CHARS`` + :data:`TRUNCATION_MARKER`),
|
||||
never through the retrieval seeding. A document's content itself is still
|
||||
carried on its rows byte-identical — the ``read`` tool serves it whole,
|
||||
un-truncated up to its cap.
|
||||
|
||||
Deterministic tie-break for equal fused scores:
|
||||
``(−fused, −cosine, document.path, chunk.position)``.
|
||||
@@ -61,9 +65,12 @@ from sqlalchemy.orm import Session
|
||||
from app.config import get_settings
|
||||
from app.models import Chunk, Document
|
||||
|
||||
#: Shared overflow marker — now used by the steering (<tuning>) section
|
||||
#: only (phase 15; imported by ``app.rag.prompts``). The document context
|
||||
#: path never truncates (A7 revised, owner permission 2026-08-24).
|
||||
#: Shared overflow marker (phase 15; imported by ``app.rag.prompts``)
|
||||
#: — used by the steering (<tuning>) section, the phase-118 NULL-summary
|
||||
#: suggestion preview fallback (A5), and the capped agent ``read``
|
||||
#: result. The seeded summary blocks and the ``read``-served document
|
||||
#: content never truncate silently (A6 re-revised): full text enters the
|
||||
#: context only through the capped ``read`` tool.
|
||||
TRUNCATION_MARKER = "[…truncated…]"
|
||||
|
||||
#: Alphanumeric tokens of a question (``to_tsquery`` input, OR-joined),
|
||||
@@ -574,6 +581,11 @@ def select_documents_tiered(
|
||||
"""Tier chunk hits into the cited and the related parent documents
|
||||
(phase 113, LOCKED A2/A4 — the usefulness bar).
|
||||
|
||||
Phase 118 retired the full-text seeding role (A6); the suggested
|
||||
tier (:func:`select_suggested`) seeds the prompt now — this helper
|
||||
stays as a dormant public helper (env back-compat for the settings
|
||||
it was calibrated by).
|
||||
|
||||
Distinct parent documents are ranked exactly like :func:`select_documents`
|
||||
(best fused score first — the same stable score-descending walk, so a
|
||||
document's rank position is fixed by its FIRST seen chunk) and each
|
||||
@@ -603,7 +615,11 @@ def select_documents_tiered(
|
||||
|
||||
The returned rows carry the full document content, byte-identical —
|
||||
a matched parent document is **never truncated** (A7 revised, owner
|
||||
permission 2026-08-24).
|
||||
permission 2026-08-24; A6 re-revised 2026-09-15: the retrieval path
|
||||
seeds SUMMARIES — the cited tier's full texts no longer ride the
|
||||
grounded prompt, full text enters the context only through the
|
||||
capped ``read`` tool; the rows themselves still carry the whole
|
||||
content).
|
||||
"""
|
||||
top_n = n if n is not None else get_settings().top_n_docs
|
||||
no_bar = floor <= 0.0
|
||||
@@ -642,12 +658,20 @@ def select_documents(
|
||||
) -> list[Document]:
|
||||
"""Map chunk hits to distinct parent documents, ranked by best fused score.
|
||||
|
||||
Phase 118 retired the full-text seeding role (A6); the suggested
|
||||
tier (:func:`select_suggested`) seeds the prompt now — this helper
|
||||
stays as a dormant public helper (env back-compat for the settings
|
||||
it was calibrated by).
|
||||
|
||||
At most *n* documents are returned (default ``BOR_TOP_N_DOCS``). The
|
||||
returned rows carry the full document content, byte-identical — a
|
||||
matched parent document is **never truncated** (A7 revised, owner
|
||||
permission 2026-08-24). There is deliberately no context budget: an
|
||||
oversized prompt must fail loudly through the ``LLMError`` → SSE
|
||||
``error`` path, never arrive as silent partial context.
|
||||
permission 2026-08-24; A6 re-revised 2026-09-15: the seeded prompt
|
||||
now carries SUMMARIES — the full text reaches the context only
|
||||
through the capped ``read`` tool, not through this selection). There
|
||||
is deliberately no context budget: an oversized prompt must fail
|
||||
loudly through the ``LLMError`` → SSE ``error`` path, never arrive
|
||||
as silent partial context.
|
||||
|
||||
Phase 113: a thin wrapper on :func:`select_documents_tiered` — the
|
||||
legacy "any score, top-N" behavior is the cited tier with a zero
|
||||
@@ -656,3 +680,79 @@ def select_documents(
|
||||
"""
|
||||
cited, _ = select_documents_tiered(chunks, n, 0.0, 0)
|
||||
return cited
|
||||
|
||||
|
||||
def select_suggested(
|
||||
chunks: Sequence[RetrievedChunk],
|
||||
n: int | None = None,
|
||||
) -> list[Document]:
|
||||
"""Top-N distinct parent documents in fused rank order — the phase-118
|
||||
"start here" suggestion tier (LOCKED A3), with NO cosine floor.
|
||||
|
||||
Distinct parent documents are walked in the SAME stable score-
|
||||
descending order as :func:`select_documents_tiered` (a document's rank
|
||||
position is fixed by its FIRST seen chunk; dedupe by ``document.id``),
|
||||
and at most *n* of them are returned (default the
|
||||
``BOR_SUGGESTED_DOCS`` setting, 5). Unlike the phase-113 cited tier,
|
||||
the usefulness bar NEVER filters here: a lexical-only hit with
|
||||
cosine 0.0 is suggested when it ranks. Suggestions are opt-in
|
||||
starting points, not citations — the seeded prompt carries the
|
||||
document's summary, and the LLM decides whether to extend its context
|
||||
by reading the document's full text.
|
||||
|
||||
The returned rows carry the full document content, byte-identical —
|
||||
the content is what the agent's ``read`` tool serves later (never
|
||||
truncated; A6 re-revises A7: full text enters the context only through
|
||||
the capped ``read`` tool).
|
||||
"""
|
||||
top_n = n if n is not None else get_settings().suggested_docs
|
||||
|
||||
order: list[Document] = []
|
||||
seen: set[uuid.UUID] = set()
|
||||
for rc in sorted(chunks, key=lambda c: c.score, reverse=True):
|
||||
if len(order) >= top_n:
|
||||
break
|
||||
doc = rc.document
|
||||
if doc.id in seen:
|
||||
continue
|
||||
seen.add(doc.id)
|
||||
order.append(doc)
|
||||
return order
|
||||
|
||||
|
||||
def select_related(
|
||||
chunks: Sequence[RetrievedChunk],
|
||||
excluded_ids: set[uuid.UUID],
|
||||
cap: int,
|
||||
) -> list[Document]:
|
||||
"""The documents ranked AFTER *excluded_ids* — the phase-118 related
|
||||
tier (rank 6+ for the contiguous top-5 suggestion set), up to *cap*
|
||||
(``BOR_RELATED_MAX_DOCS``).
|
||||
|
||||
The SAME stable score-descending walk as
|
||||
:func:`select_documents_tiered` / :func:`select_suggested` (a
|
||||
document's rank position is fixed by its FIRST seen chunk; dedupe by
|
||||
``document.id``), skipping every document whose id is in
|
||||
*excluded_ids* and admitting at most *cap* documents. There is NO
|
||||
cosine floor: the related tier is the ranked remainder (a lexical-
|
||||
only cosine 0.0 hit is included) — its job on the ``done`` frame is
|
||||
visibility (the UI's de-emphasized "nearby docs" row), not
|
||||
citation. With the turn wiring's exclusion — exactly the suggested
|
||||
tier's document ids (LOCKED A3: a contiguous top-N, no floor) —
|
||||
"excluding the suggested" is exactly "rank 6+".
|
||||
|
||||
The returned rows carry the full document content, byte-identical
|
||||
(the tier is metadata for the ``done`` frame and the durable
|
||||
record; the prompt and ``read`` contract are untouched).
|
||||
"""
|
||||
out: list[Document] = []
|
||||
seen: set[uuid.UUID] = set()
|
||||
for rc in sorted(chunks, key=lambda c: c.score, reverse=True):
|
||||
if len(out) >= cap:
|
||||
break
|
||||
doc = rc.document
|
||||
if doc.id in seen or doc.id in excluded_ids:
|
||||
continue
|
||||
seen.add(doc.id)
|
||||
out.append(doc)
|
||||
return out
|
||||
|
||||
+22
-7
@@ -388,9 +388,10 @@ class DocContent(BaseModel):
|
||||
path: str
|
||||
title: str
|
||||
format: str
|
||||
#: Lite-model summary (phase 30) — non-markdown A9 docs only; None for
|
||||
#: markdown documents, pre-phase-30 rows, and the fail-soft path where
|
||||
#: summary generation failed but the document was still indexed.
|
||||
#: Lite-model summary (phase 30) — every A9 doc, markdown included
|
||||
#: since phase 118 A2; None for pre-phase-30 rows, the fail-soft path
|
||||
#: where summary generation failed, and until the phase-118 backfill
|
||||
#: stores one on the next sync.
|
||||
summary: str | None = None
|
||||
#: The document's creation date (phase 106, D8) — ISO-8601, verbatim
|
||||
#: from the row; the viewer's top meta row renders the ``Created``
|
||||
@@ -742,14 +743,26 @@ class ChatMessage(BaseModel):
|
||||
"""One conversation record in the ``bor.chat.v1`` localStorage shape
|
||||
(phase 14) — the stored ``messages`` payload of a saved chat (phase 50).
|
||||
|
||||
``{who, text, sources?, deflected?, suggestions?, thinking?, tools?,
|
||||
stopped?}`` — raw text, never HTML, so a saved chat restores
|
||||
pixel-identical through the existing ``renderStoredMessage`` path.
|
||||
``{who, text, sources?, related?, deflected?, suggestions?,
|
||||
thinking?, tools?, stopped?}`` — raw text, never HTML, so a saved
|
||||
chat restores pixel-identical through the existing
|
||||
``renderStoredMessage`` path.
|
||||
``extra="forbid"`` rejects unknown keys (a corrupted or HTML-shaped
|
||||
payload, e.g. a stray ``<b>``-ish extra key) at the boundary with a
|
||||
422, so nothing outside this shape can poison a restored
|
||||
conversation.
|
||||
|
||||
``related`` (phase 113, the related-doc tier the UI persists with
|
||||
every grounded brain record — the restore path re-renders the
|
||||
de-emphasized row from it): the same :class:`SourceRef` list shape
|
||||
as ``sources``, the same cap. Its absence from this model was a
|
||||
phase-113 omission — the ``extra="forbid"`` boundary 422'd every
|
||||
done-time auto-save carrying the key (the A2 quiet-failure path
|
||||
swallowed it), leaving grounded turns' brain messages unsaved;
|
||||
pinned by ``tests/integration/test_chats_api.py`` (the full brain
|
||||
record carries ``related``) and the E2E ``test_share_chat``
|
||||
auto-save count.
|
||||
|
||||
Phase 83 (SEC-05) bounds the anonymous write surface (``POST/PUT
|
||||
/api/chats`` is public — the row id is the credential, phase 55 A1):
|
||||
``text`` / ``thinking`` carry :class:`HistoryTurn`'s 32 000 caps
|
||||
@@ -757,7 +770,8 @@ class ChatMessage(BaseModel):
|
||||
path, so a saved chat can never legitimately carry more), and the
|
||||
nested lists get length caps (``max_length``) sized to the realistic
|
||||
``bor.chat.v1`` record the UI produces (``sources`` ≤ 20 — top-N docs
|
||||
+ agent reads; ``suggestions`` ≤ 50 chips of ≤ 200 chars; ``tools``
|
||||
+ agent reads; ``related`` ≤ 20 — the related-doc tier, same ref
|
||||
shape; ``suggestions`` ≤ 50 chips of ≤ 200 chars; ``tools``
|
||||
≤ 50 — one entry per tool call, the round cap is 10). Only value
|
||||
bounds were added — the accepted/rejected KEYS are unchanged.
|
||||
"""
|
||||
@@ -767,6 +781,7 @@ class ChatMessage(BaseModel):
|
||||
who: Literal["user", "brain"]
|
||||
text: str = Field(min_length=1, max_length=32_000)
|
||||
sources: list[SourceRef] | None = Field(default=None, max_length=20)
|
||||
related: list[SourceRef] | None = Field(default=None, max_length=20)
|
||||
deflected: bool | None = None
|
||||
suggestions: list[_Chip] | None = Field(default=None, max_length=50)
|
||||
thinking: str | None = Field(default=None, max_length=32_000)
|
||||
|
||||
@@ -376,7 +376,7 @@ async def run_turn(
|
||||
lambda: SessionLocal(),
|
||||
system_prompt=plan.system_prompt,
|
||||
user_message=question,
|
||||
seed_docs=plan.docs,
|
||||
seed_docs=plan.suggested_docs, # phase 118: the suggestion tier
|
||||
settings=settings,
|
||||
holder=holder,
|
||||
)
|
||||
|
||||
+14
-8
@@ -35,9 +35,11 @@ guessable by any model. This script:
|
||||
comparing a per-table checksum**.
|
||||
|
||||
Re-run it only when the fixture documents, the chunker, the embedding
|
||||
model, or the folder-summary prompt (its ``lite`` output is baked into
|
||||
the dump) change — everyday iterations restore the dump in sub-second
|
||||
time (``scripts/restore_test_kb`` / the gate's ``--restore``), never
|
||||
model, the folder-summary prompt (its ``lite`` output is baked into
|
||||
the dump), or the importer's summary policy (phase 118: every document
|
||||
gets a stored summary + an embedded ``is_summary`` chunk) change —
|
||||
everyday iterations restore the dump in sub-second time
|
||||
(``scripts/restore_test_kb`` / the gate's ``--restore``), never
|
||||
re-embedding (see ``TOOL_CALLING_TESTING.md``).
|
||||
|
||||
Exit codes: **0** built + verified, **1** build/verification failure,
|
||||
@@ -105,10 +107,12 @@ FIXTURE_KB_OVERVIEW: str = (
|
||||
#: tsvector column is excluded; Postgres recomputes it).
|
||||
_TABLES: tuple[tuple[str, type, tuple[str, ...]], ...] = (
|
||||
("documents", Document, ("id", "source", "path", "full_path", "title",
|
||||
"content", "content_hash", "indexed_at", "summary")),
|
||||
"content", "content_hash", "indexed_at",
|
||||
"created_at", "created_at_manual", "summary")),
|
||||
("chunks", Chunk, ("id", "document_id", "position", "content",
|
||||
"embedding", "is_summary")),
|
||||
("git_sources", GitSource, ("id", "url", "kind", "path", "added_at")),
|
||||
("git_sources", GitSource, ("id", "url", "kind", "path", "ignore_paths",
|
||||
"include_hidden", "added_at")),
|
||||
("kb_overview", KbOverview, ("id", "content", "updated_at")),
|
||||
("sources_meta", SourcesMeta, ("id", "version", "updated_at")),
|
||||
("steering_notes", SteeringNote, ("id", "note", "created_at")),
|
||||
@@ -125,7 +129,8 @@ _TABLES: tuple[tuple[str, type, tuple[str, ...]], ...] = (
|
||||
# step 5 of :func:`_build` (the build truncates first, so the sync
|
||||
# paths' table-empty first-run trigger holds).
|
||||
("folder_summaries", FolderSummary, ("source", "folder_path",
|
||||
"summary", "updated_at")),
|
||||
"summary", "manually_edited",
|
||||
"updated_at")),
|
||||
)
|
||||
|
||||
|
||||
@@ -209,7 +214,7 @@ async def _retrieval_report(llm: LLMClient, battery: list[str]) -> int:
|
||||
vec = await llm.embed_one(question)
|
||||
chunks = retrieve(db, question, vec)
|
||||
plan = plan_turn(chunks, settings)
|
||||
seed = ", ".join(f"{d.source}/{d.path}" for d in plan.docs) or "—"
|
||||
seed = ", ".join(f"{d.source}/{d.path}" for d in plan.suggested_docs) or "—"
|
||||
if plan.deflected:
|
||||
deflected += 1
|
||||
print(
|
||||
@@ -317,7 +322,8 @@ async def _build(kb_dir: Path, dump_path: Path) -> int:
|
||||
"-- Data-only snapshot (the schema stays alembic-managed; the",
|
||||
"-- generated chunks.tsv column is recomputed on restore).",
|
||||
f"-- Sources: {', '.join(FIXTURE_SOURCES)} "
|
||||
f"({summary.added} documents, {summary.chunks} chunks).",
|
||||
f"({summary.added} documents, {summary.chunks} content "
|
||||
f"+ {summary.summaries + summary.summary_backfilled} summary chunks)",
|
||||
"-- Restore (one transaction, sub-second):",
|
||||
"-- uv run python -m scripts.restore_test_kb",
|
||||
"-- psql \"$BOR_DATABASE_URL\" --single-transaction -f "
|
||||
|
||||
+211
-46
@@ -55,14 +55,20 @@ Implements just enough of the aipi surface:
|
||||
- system prompt containing ``<knowledge_base>`` (phase 31, KB overview)
|
||||
-> the composed answer ends with `` (kb: <first bullet line>)`` —
|
||||
the same echo convention for the overview's prompt injection.
|
||||
- user message containing ``show the end of your notes`` (phase 24,
|
||||
whole-document context) -> the answer quotes the **last 160 chars of
|
||||
the ``<documents>`` block** — a tail echo, byte-stable across runs, so
|
||||
a sentinel placed at the *end* of a document appears in the rendered
|
||||
answer iff the whole document was in the prompt. (Phase 37: the HIGH
|
||||
prompt now ends with a ``<tools>`` section after ``</documents>``, so
|
||||
the echo targets the block itself; its tail still includes the
|
||||
closing tag — same sentinel semantics.)
|
||||
- user message containing ``show the end of your notes`` (phase 24;
|
||||
phase 118 re-targeted — the story's dedicated suite
|
||||
``tests/e2e/test_summary_seed_context.py``) -> the answer quotes
|
||||
the **last 160 chars of the ``<documents>`` block** — a tail echo,
|
||||
byte-stable across runs. (Phase 37: the HIGH prompt now ends with
|
||||
a ``<tools>`` section after ``</documents>``, so the echo targets
|
||||
the block itself; its tail still includes the closing tag.)
|
||||
Phase 118 (A6): the block carries the suggested documents'
|
||||
SUMMARIES (never full texts), so the echoed tail is the LAST
|
||||
suggested document's SUMMARY tail (its digest + the
|
||||
``Source: <source>/<path>`` pointer line) — a sentinel on a
|
||||
document's *last line* appears in the rendered answer iff the
|
||||
FULL content (not the summary) was in the prompt, which under the
|
||||
summary-seed contract is only through a ``read`` tool result.
|
||||
- user message containing ``use your tools`` (phase 37, agent document
|
||||
tools; phase 70: the flow emits the harness-aligned names — ``ls``
|
||||
/ ``read`` with the combined ``source/path`` identity; phase 94:
|
||||
@@ -235,16 +241,19 @@ Implements just enough of the aipi surface:
|
||||
content>`` (the phase-37 single-read shape — the grounded-
|
||||
turn citation contract);
|
||||
* the last tool result is the agent's ALREADY_IN_CONTEXT dedupe
|
||||
refusal (the read target is already a top-2 retrieval
|
||||
document — with the drill fixture that is DETERMINISTIC:
|
||||
the read question names the file's path, so the file
|
||||
self-matches the hybrid gate and its FULL text is in the
|
||||
``<documents>`` prompt): the model answers FROM THE PROMPT —
|
||||
the deterministic answer ``Already in context: Read
|
||||
<source/path>. <first 80 chars of the target document's text
|
||||
as it appears in the ``<documents>`` block>`` (same citation
|
||||
shape as the read-result branch — the document text reached
|
||||
the model either way, and the answer proves it);
|
||||
refusal (phase 118: the read target is a document ALREADY
|
||||
READ into full-text context earlier in the same turn — the
|
||||
seeds are summaries, so a first read of any document
|
||||
succeeds and only a re-read is refused): the model answers
|
||||
FROM THE PROMPT — the deterministic answer ``Already in
|
||||
context: Read <source/path>. <first 80 chars of the target
|
||||
document's text as it appears in the ``<documents>`` block>``
|
||||
(same citation shape as the read-result branch — the document
|
||||
text reached the model either way, and the answer proves it;
|
||||
the block it quotes now carries the document's SUMMARY).
|
||||
No suite exercises this branch today (the drill questions
|
||||
read once per turn) — it is kept for the still-real
|
||||
already-read refusal;
|
||||
* any other last result (a top-level or folder LISTING landed):
|
||||
the deterministic ECHO answer ``Here's the level I listed:\n
|
||||
<the listing, verbatim>`` — the mock echoes what it received
|
||||
@@ -287,6 +296,38 @@ Implements just enough of the aipi surface:
|
||||
needs the ``<tools>`` section, so deflected turns never hit it);
|
||||
verified 2026-09-10: no existing E2E question or fixture file
|
||||
contains the phrase, so every other suite is unaffected.
|
||||
- user message containing ``read the suggested document``
|
||||
(``SUMMARY_SEED_READ_TRIGGER``, phase 118 task 06 — the
|
||||
summary-seed context's dedicated story suite
|
||||
``tests/e2e/test_summary_seed_context.py``) **and** the system
|
||||
prompt carries the ``<tools>`` section -> the deterministic
|
||||
SCRIPTED SUMMARY-READ flow: the question carries its own tool call
|
||||
after the colon — ``read the suggested document: read source/path``
|
||||
— parsed by ``_SUMMARY_SEED_READ_CALL_RE`` from the RAW user
|
||||
message (the target keeps its case), then discriminated
|
||||
statelessly from the tool results (streaming only):
|
||||
* request 1 (``tools`` offered, no ``tool``-role result in the
|
||||
messages yet): the scripted call — ``read`` with the parsed
|
||||
target (synthetic id ``call_0``);
|
||||
* a ``tool``-role result is in the messages: the deterministic
|
||||
ECHO — the answer carries the LAST tool result VERBATIM
|
||||
(``Here's what the read returned:\n<result>``): under the
|
||||
phase-118 summary-seed contract a first ``read`` of ANY
|
||||
document succeeds (the seeds are summaries, not full text), so
|
||||
a read result (``"Document <source/path>:…`` — header + the
|
||||
phase-106 D5 ``date:`` line + the FULL content) lands in the
|
||||
answer with its tail intact — the story suite's lens on the
|
||||
full text the ``read`` tool delivered (a tail sentinel on the
|
||||
document's last line appears in the answer iff the full content
|
||||
reached the model through the read, not the seed); a refusal
|
||||
(the premise broke) lands just as visibly, so the suite fails
|
||||
loudly on it. The mock is the only E2E lens on the LLM's
|
||||
context, so the echo is the assertion surface.
|
||||
Checked BEFORE the plain ``TOOLS_TRIGGER`` flow (disjoint trigger
|
||||
phrases — the phase-71/72/94 ordering convention; the trigger
|
||||
needs the ``<tools>`` section, so deflected turns never hit it);
|
||||
no existing E2E question or fixture file contains the phrase, so
|
||||
every other suite is unaffected.
|
||||
- user message containing ``what are the correct llama.cpp
|
||||
arguments`` (``GREP_TEACH_TRIGGER``, the 2026-09-05 incident —
|
||||
the harness prior is that grep takes a REGEX; this app's grep is a
|
||||
@@ -802,10 +843,12 @@ _DRILL_CALL_RE = re.compile(
|
||||
_NOT_A_FOLDER_MARKER = "is not a folder"
|
||||
|
||||
#: The stable substring of the harness-owned dedupe refusal the
|
||||
#: drill-down flow's ``ctx_answer`` branch keys on (a read target that
|
||||
#: is already a top-2 retrieval document — the mock then answers from
|
||||
#: the document's text in the ``<documents>`` prompt block, exactly
|
||||
#: what the refusal instructs). Keyed on a substring (not the whole
|
||||
#: drill-down flow's ``ctx_answer`` branch keys on (phase 118: a read
|
||||
#: target ALREADY READ into full-text context in the same turn — the
|
||||
#: seeds are summaries, so a first read of any document succeeds and
|
||||
#: only a re-read is refused; the mock then answers from the document's
|
||||
#: summary in the ``<documents>`` prompt block, exactly what the
|
||||
#: refusal instructs). Keyed on a substring (not the whole
|
||||
#: constant) so a re-wrap of the constant cannot silently re-route the
|
||||
#: mock; the module-level assert below fails loudly if the substring
|
||||
#: ever leaves the constant (the mock must never drift from
|
||||
@@ -819,11 +862,13 @@ assert _ALREADY_IN_CONTEXT_MARKER in ALREADY_IN_CONTEXT, (
|
||||
#: section (``app.rag.prompts.build_high_prompt``): the block is the
|
||||
#: document identity (``source``/``path``/``title`` attributes — plus,
|
||||
#: since phase 106 D5, the ``date`` attribute, the row's ``created_at``
|
||||
#: UTC date part, APPENDED after ``title``) plus the document's FULL
|
||||
#: text (never truncated on the retrieval path, owner-locked A7)
|
||||
#: between the tags. The ``date`` group is OPTIONAL so the mock
|
||||
#: tolerates the pre- and post-phase block shapes (house rule: the
|
||||
#: marker/regex lands with the prompt change).
|
||||
#: UTC date part, APPENDED after ``title``) plus the document's
|
||||
#: SUMMARY (phase 118, A6 — the summary-seed contract: the seeded
|
||||
#: blocks are summaries, never full text; full text enters the context
|
||||
#: only through the capped ``read`` tool) between the tags. The
|
||||
#: ``date`` group is OPTIONAL so the mock tolerates the pre- and
|
||||
#: post-phase block shapes (house rule: the marker/regex lands with
|
||||
#: the prompt change).
|
||||
_DOCUMENT_BLOCK_RE = re.compile(
|
||||
r'<document source="(?P<source>[^"]+)" path="(?P<path>[^"]+)" '
|
||||
r'title="[^"]*"(\sdate="[^"]*")?>\n(?P<content>.*?)\n</document>',
|
||||
@@ -835,10 +880,14 @@ def _document_block(system: str, source: str, path: str) -> str | None:
|
||||
"""The stored text of one ``<document>`` block (or ``None``).
|
||||
|
||||
The drill-down flow's ``ctx_answer`` branch: when the agent's read
|
||||
of a top-2 retrieval document gets the ALREADY_IN_CONTEXT dedupe,
|
||||
the document's full text is in the ``<documents>`` prompt — the
|
||||
mock (the model) extracts it by the block's identity attributes
|
||||
and quotes it, answering from the prompt as the refusal instructs.
|
||||
of a document ALREADY READ in the same turn gets the
|
||||
ALREADY_IN_CONTEXT dedupe (phase 118 — the seeds are summaries, so
|
||||
a first read of any document succeeds and only a re-read is
|
||||
refused), the mock (the model) extracts the block's text by the
|
||||
identity attributes and quotes it, answering from the prompt as
|
||||
the refusal instructs. The block carries the document's SUMMARY
|
||||
(phase 118 A6) — no suite exercises the branch today; it is kept
|
||||
for the still-real already-read refusal.
|
||||
"""
|
||||
for block in _DOCUMENT_BLOCK_RE.finditer(system):
|
||||
if block.group("source") == source and block.group("path") == path:
|
||||
@@ -1377,16 +1426,22 @@ def _drill_flow(body: dict[str, Any]) -> tuple[str, ...] | None:
|
||||
last = results[-1]
|
||||
if last.startswith(_READ_RESULT_PREFIX):
|
||||
head, _, content = last.partition("\n")
|
||||
# The read result is ``"Document <source/path>:\n<content>"`` —
|
||||
# the head carries the server's appended ``:`` (removed here;
|
||||
# a document path never legitimately ends with one).
|
||||
# The read result is ``"Document <source/path>:\ndate: …\n<content>"``
|
||||
# — the head carries the server's appended ``:`` (removed here;
|
||||
# a document path never legitimately ends with one); the
|
||||
# phase-106 D5 ``date:`` line sits between the header and the
|
||||
# document text — skipped so the quote stays pure document
|
||||
# content (the suite's byte-identical pin).
|
||||
combined = head[len(_READ_RESULT_PREFIX):].strip().removesuffix(":")
|
||||
if content.startswith("date: "):
|
||||
content = content.partition("\n")[2]
|
||||
return ("read_answer", combined, content[:80])
|
||||
if _ALREADY_IN_CONTEXT_MARKER in last and verb == "read" and "/" in target:
|
||||
# The dedupe fired: the read target is already a top-2
|
||||
# retrieval document, so its FULL text is in the
|
||||
# ``<documents>`` prompt — answer from the prompt (the
|
||||
# refusal's instruction), quoting the block's text.
|
||||
# Phase 118: the dedupe fires only for a document ALREADY READ
|
||||
# in the same turn (the seeds are summaries — a first read of
|
||||
# a suggested document succeeds). Answer from the prompt (the
|
||||
# refusal's instruction), quoting the block's text (now the
|
||||
# document's summary).
|
||||
src, _, p = target.partition("/")
|
||||
content = _document_block(_system(body), src, p)
|
||||
if content is not None:
|
||||
@@ -1469,6 +1524,81 @@ def _read_cap_flow(body: dict[str, Any]) -> tuple[str, ...] | None:
|
||||
return ("echo", results[-1])
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Phase 118 (task 06, the summary-seed context's dedicated story suite):
|
||||
# the deterministic SCRIPTED summary read — see the module docstring
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
#: A user message containing this substring (case-insensitive) —
|
||||
#: combined with the ``<tools>`` section in the system prompt — drives
|
||||
#: the scripted SUMMARY-READ flow (the summary-seed context's story
|
||||
#: suite, ``tests/e2e/test_summary_seed_context.py``): the question
|
||||
#: carries its own tool call after the colon — ``read the suggested
|
||||
#: document: read source/path`` — the mock emits the scripted ``read``
|
||||
#: (phase 118: the seeds are summaries, so a first read of a suggested
|
||||
#: document SUCCEEDS — the full text arrives through the read), then
|
||||
#: ECHOES the ENTIRE tool result into its answer (the house
|
||||
#: scripted-turn lens on the LLM's context — the full content's tail
|
||||
#: reaches the rendered answer iff the read delivered it). Checked
|
||||
#: BEFORE the plain ``TOOLS_TRIGGER`` flow (disjoint trigger phrases —
|
||||
#: the phase-71/72/94 ordering convention); no existing E2E question
|
||||
#: or fixture file contains the phrase, so every other suite is
|
||||
#: unaffected.
|
||||
SUMMARY_SEED_READ_TRIGGER = "read the suggested document"
|
||||
|
||||
#: The scripted call in the summary-read question (case-insensitive —
|
||||
#: the suite's questions capitalize the trigger's first letter): the
|
||||
#: verb (``read``) plus the target — a combined ``source/path``, parsed
|
||||
#: from the RAW user message so the target keeps its case. The target
|
||||
#: is a ``[a-z0-9_./-]`` run (case-insensitively), so the suite's
|
||||
#: `` — `` flavor separator (em dash) can never bleed into it (the
|
||||
#: read-cap/drill convention, ``_READ_CAP_CALL_RE`` / ``_DRILL_CALL_RE``).
|
||||
_SUMMARY_SEED_READ_CALL_RE = re.compile(
|
||||
r"read the suggested document:\s*read\s+(?P<arg>[a-z0-9_./-]+)",
|
||||
re.I,
|
||||
)
|
||||
|
||||
|
||||
def _summary_seed_read_flow(body: dict[str, Any]) -> tuple[str, ...] | None:
|
||||
"""Classify a phase-118 scripted summary-read request (see the
|
||||
module docstring). The question carries the scripted call (``read
|
||||
the suggested document: read source/path``); the step is then
|
||||
discriminated statelessly from the tool results, like the other
|
||||
marker flows:
|
||||
|
||||
* ``("call", target, "call_0")`` — ``tools`` are offered and no
|
||||
``tool``-role result is in the messages yet: the scripted
|
||||
``read`` on the parsed target (synthetic id ``call_0``).
|
||||
* ``("echo", result)`` — a ``tool``-role result is in the messages:
|
||||
the deterministic ECHO — the answer carries the LAST tool result
|
||||
VERBATIM (``Here's what the read returned:\n<result>``): a read
|
||||
result (``"Document <source/path>:…`` — header + date line +
|
||||
FULL content) lands in the answer with its tail intact (the
|
||||
full text the ``read`` tool delivered — the phase-118 contract:
|
||||
a first read of a suggested document succeeds, the seed was a
|
||||
summary); a refusal (the premise broke) lands just as visibly,
|
||||
so the suite fails loudly on it.
|
||||
* ``None`` — not the flow: the trigger is absent, the ``<tools>``
|
||||
section is missing (deflected turns never carry it), the scripted
|
||||
call is unparseable, or ``tools`` are not offered and no tool
|
||||
results are in the messages yet (e.g. ``agent_max_rounds=0``).
|
||||
"""
|
||||
user = _user(body)
|
||||
if SUMMARY_SEED_READ_TRIGGER not in user.lower():
|
||||
return None
|
||||
if "<tools>" not in _system(body):
|
||||
return None
|
||||
match = _SUMMARY_SEED_READ_CALL_RE.search(user)
|
||||
if match is None:
|
||||
return None
|
||||
results = _tool_results(body)
|
||||
if not results:
|
||||
if not body.get("tools"):
|
||||
return None
|
||||
return ("call", match.group("arg"), "call_0")
|
||||
return ("echo", results[-1])
|
||||
|
||||
|
||||
def long_answer() -> str:
|
||||
"""~900-word deterministic walkthrough (phase 11): numbered steps plus
|
||||
a unique final line that must survive the stream untruncated."""
|
||||
@@ -1924,14 +2054,19 @@ def compose_answer(body: dict[str, Any]) -> str:
|
||||
"You've got this!"
|
||||
)
|
||||
elif END_OF_NOTES_TRIGGER in user.lower():
|
||||
# Whole-document-context story (phase 24): echo the tail of the
|
||||
# document context. Byte-stable across runs — a sentinel on the
|
||||
# document's last line appears in the answer iff the whole
|
||||
# document was in the prompt. (The tail includes the closing
|
||||
# </documents> — harmless for the E2E sentinel assertions.)
|
||||
# Phase 37: the HIGH prompt now ends with the <tools> section
|
||||
# after </documents>, so the echo targets the <documents> block
|
||||
# itself — the sentinel semantics are unchanged.
|
||||
# Phase 24 (whole-document context) — phase 118 re-targeted
|
||||
# (the summary-seed context's story suite): echo the tail of
|
||||
# the <documents> block. Byte-stable across runs — under the
|
||||
# phase-118 summary-seed contract the block carries the
|
||||
# suggested docs' SUMMARIES, so the echoed tail is the LAST
|
||||
# suggested doc's summary tail (digest + Source: pointer line),
|
||||
# and a sentinel on a document's last line appears in the
|
||||
# answer only if the FULL content reached the model (through a
|
||||
# read tool result — never the seed). (The tail includes the
|
||||
# closing </documents> — harmless for the E2E sentinel
|
||||
# assertions.) Phase 37: the HIGH prompt ends with the <tools>
|
||||
# section after </documents>, so the echo targets the
|
||||
# <documents> block itself.
|
||||
block = _DOCUMENTS_BLOCK_RE.search(_system(body))
|
||||
tail_source = block.group(0) if block else _context(body)
|
||||
answer = (
|
||||
@@ -2461,6 +2596,36 @@ def chat_completions(body: dict[str, Any]) -> Any:
|
||||
media_type="text/event-stream",
|
||||
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
|
||||
)
|
||||
# Phase 118 (task 06): the deterministic SCRIPTED summary read
|
||||
# (the question carries its own call — ``read the suggested
|
||||
# document: read source/path``): the scripted ``read`` (phase
|
||||
# 118: a first read of a suggested document succeeds — the
|
||||
# seeds are summaries), then the answer that ECHOES the whole
|
||||
# tool result (the full text's tail reaches the rendered answer
|
||||
# iff the read delivered it — the suite's lens on the LLM's
|
||||
# context). Checked BEFORE the plain TOOLS_TRIGGER flow
|
||||
# (disjoint trigger phrases — the phase-71/72/94 ordering
|
||||
# convention; the trigger needs the ``<tools>`` section, so
|
||||
# deflected turns never hit it).
|
||||
seed_read = _summary_seed_read_flow(body)
|
||||
if seed_read is not None:
|
||||
if seed_read[0] == "call":
|
||||
stream = _tool_call_stream(
|
||||
"read", {"path": seed_read[1]}, seed_read[2]
|
||||
)
|
||||
else: # "echo" — the last tool result verbatim (the lens)
|
||||
stream = _sse_stream(
|
||||
_apply_max_tokens(
|
||||
f"Here's what the read returned:\n{seed_read[1]}",
|
||||
body.get("max_tokens"),
|
||||
),
|
||||
0.0,
|
||||
)
|
||||
return StreamingResponse(
|
||||
stream,
|
||||
media_type="text/event-stream",
|
||||
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
|
||||
)
|
||||
# Phase 109 (task 03, never-frozen-turn story suite): the
|
||||
# deterministic REPORTED-REPRO turn (delta → tool →
|
||||
# thinking-after-delta — TODO.md L3): the scripted first answer,
|
||||
|
||||
@@ -44,8 +44,13 @@ reseelink.json" but does not include it):
|
||||
(a ``documents`` row: it is in the agent's catalog, readable, and a
|
||||
source-chip target) but seeded WITHOUT chunks. In a real
|
||||
hundreds-of-document KB the file would simply fail to rank into the
|
||||
top-2 context; with a two-document corpus every chunk would rank, so
|
||||
"not in context" is expressed as "no retrieval candidates". Its
|
||||
suggested (summary-seeded) context; with a two-document corpus every
|
||||
chunk would rank, so "not in context" is expressed as "no retrieval
|
||||
candidates". (Phase 118, A6: a read of a suggested doc now succeeds
|
||||
anyway — the ALREADY_IN_CONTEXT refusal fires only for a document
|
||||
already READ in the turn — so the catalog-only design stands on the
|
||||
mock's first-catalog-line parse, not on the retired seed-read
|
||||
refusal.) Its
|
||||
``(source, path)`` also sorts FIRST in the catalog
|
||||
(``Deployments`` < ``Homelab``) — which is exactly the line the mock
|
||||
parses out of the listing and reads.
|
||||
|
||||
@@ -46,9 +46,14 @@ Test → contract mapping (six tests, one per contract bullet):
|
||||
the ``Indexed`` badge (D8 — the date at the top of the clicked
|
||||
document), with the source/format/indexed/chunks badges intact.
|
||||
4. ``test_old_correct_beats_new_similar`` — THE OWNER SCENARIO end to
|
||||
end: the real retriever + the DEFAULT recency boost (0.0007 / 365 d)
|
||||
over the mock's token-overlap embeddings ranks the OLDER correct
|
||||
document as the first cited source over the newer similar one.
|
||||
end (phase 118, A4): the real retriever + the DEFAULT recency boost
|
||||
(0.0007 / 365 d) over the mock's token-overlap embeddings ranks the
|
||||
OLDER correct document first in the suggested tier — the chip row IS
|
||||
the suggested tier (top-5, NO floor; this four-doc KB therefore
|
||||
chips ALL four docs in rank order: retention, draft, forward,
|
||||
old-doc) and the related row is absent (no rank-6+ doc exists).
|
||||
The newer similar doc (the boost's intended beneficiary) stays
|
||||
second — the boost never lets it outrank the one that answers.
|
||||
5. ``test_date_edit_and_sync_preserves`` — the admin-only editor in
|
||||
the real UI: set → Save → the badge re-renders from the RESPONSE
|
||||
(never the optimistic input) → the API round-trips; a re-import
|
||||
@@ -69,9 +74,12 @@ deterministic for fixed text — measured with
|
||||
``app.rag.retriever._vector_candidates`` / ``_lexical_candidates`` /
|
||||
``retrieve()`` against a real Postgres + the mock:
|
||||
|
||||
* ``retention.md`` (correct, 2020) — rank 1 in BOTH lists (cosine
|
||||
0.6222; the lexical tsquery ``how|did|i|configure|backup|retention|
|
||||
policy`` after stopword removal matches it most densely).
|
||||
* ``retention.md`` (correct, 2020) — rank 1 in BOTH lists (content
|
||||
chunk cosine 0.6222; the lexical tsquery ``how|did|i|configure|
|
||||
backup|retention|policy`` after stopword removal matches it most
|
||||
densely). Phase 118 (A2): its EMBEDDED summary is a retrieval
|
||||
candidate too (best-chunk cosine 0.7133 — the seeded summary ranks,
|
||||
which is the point of seeding it).
|
||||
* ``retention-draft.md`` (similar, now) — rank 4 in the vector list
|
||||
(cosine 0.1443) and rank 2 in the lexical list. Its wording was
|
||||
tuned for exactly this: it shares ONLY the three "backup retention
|
||||
@@ -85,14 +93,18 @@ deterministic for fixed text — measured with
|
||||
(0.1875) rank 2–3 in the vector list (unrelated content) and match
|
||||
the tsquery not at all.
|
||||
|
||||
Fused (RRF k=60) + the default boost (0.0007 · exp(−age/365d)):
|
||||
retention.md 1/61+1/61 = 0.0327878 (+ ≈ 0, 6.7 half-lives old) vs
|
||||
retention-draft.md 1/64+1/62 = 0.0317540 (+ the FULL zero-age 0.0007
|
||||
= 0.0324540) → the older correct doc wins by 0.000334 WITH the boost
|
||||
on (it would win by 0.001034 with the boost off — the scenario holds
|
||||
both ways; the boost never lets the newer similar doc outrank the one
|
||||
that answers the question). ``select_documents`` (top-2) cites
|
||||
retention.md first, the draft second. The boost defaults are owned by
|
||||
Fused (RRF k=60) + the default boost (0.0007 · exp(−age/365d)),
|
||||
probe-verified against the current candidate set (the embedded summary
|
||||
chunks join the walk — phase 118 A2): retention.md 0.032523 vs
|
||||
retention-draft.md 0.031498 (the full zero-age boost included) → the
|
||||
older correct doc wins by ≈ 0.001025 WITH the boost on (the margin
|
||||
only grew once the summary chunks ranked — the scenario holds both
|
||||
ways; the boost never lets the newer similar doc outrank the one that
|
||||
answers the question). The phase-118 suggested tier (top-5, NO floor)
|
||||
carries all four docs in that rank order — retention.md first, the
|
||||
draft second, then forward.md (0.016325) and old-doc.md (0.015874)
|
||||
(no floor filters the unrelated docs); the related tier is EMPTY
|
||||
(no rank-6+ doc in a four-doc KB). The boost defaults are owned by
|
||||
task 07 — untouched here.
|
||||
|
||||
DB isolation: every test TRUNCATEs the KB tables (the
|
||||
@@ -529,11 +541,14 @@ def test_old_correct_beats_new_similar(
|
||||
) -> None:
|
||||
"""``How did I configure the backup retention policy?`` → grounded,
|
||||
and the FIRST cited source is the OLDER correct doc (2020) — the
|
||||
NEWER similar one (now, "under review") is cited second. The real
|
||||
hybrid retriever + the DEFAULT recency boost (0.0007 / 365 d) over
|
||||
the mock's token-overlap embeddings (the module docstring records
|
||||
the measured fused scores: 0.0327878 vs 0.0324540 — margin
|
||||
0.000334 WITH the full zero-age boost on the newer doc)."""
|
||||
NEWER similar one (now, "under review") is cited second. Phase 118
|
||||
(A4): the chip row IS the suggested tier (top-5, NO floor) — this
|
||||
four-doc KB chips ALL four docs in rank order — and the related row
|
||||
is absent (no rank-6+ doc). The real hybrid retriever + the DEFAULT
|
||||
recency boost (0.0007 / 365 d) over the mock's token-overlap
|
||||
embeddings (the module docstring records the measured fused scores:
|
||||
0.032523 vs 0.031498 — margin ≈ 0.001025 WITH the full zero-age
|
||||
boost on the newer doc)."""
|
||||
_reset_db(mock_llm, dates_tree)
|
||||
page.set_default_timeout(30_000)
|
||||
login(page, app_url, next="/") # phase 79: chat is require_user-gated
|
||||
@@ -548,22 +563,34 @@ def test_old_correct_beats_new_similar(
|
||||
# top_score 0.6222 ≥ the e2e threshold 0.30).
|
||||
expect(page.locator(".msg.brain.is-deflected")).to_have_count(0)
|
||||
|
||||
# The OLDER correct doc is the FIRST source chip; the NEWER similar
|
||||
# one (the boost's intended beneficiary) is cited — but second.
|
||||
# Phase 118 (A4): the chip row IS the suggested tier — top-5, NO
|
||||
# floor, and this KB has exactly four docs, so ALL four are
|
||||
# suggested (chipped) in fused rank order: the OLDER correct doc
|
||||
# first, the NEWER similar one (the boost's intended beneficiary)
|
||||
# second, then the two unrelated docs (no floor filters them — the
|
||||
# LLM decides what the summaries earn).
|
||||
chips = page.locator(".msg.brain .source-chip")
|
||||
expect(chips).to_have_count(2, timeout=30_000)
|
||||
expect(chips).to_have_count(4, timeout=30_000)
|
||||
assert chips.nth(0).inner_text() == f"{dates_tree.name}/{RETENTION_MD}"
|
||||
assert chips.nth(1).inner_text() == f"{dates_tree.name}/{DRAFT_MD}"
|
||||
assert chips.nth(2).inner_text() == f"{dates_tree.name}/{FORWARD_MD}"
|
||||
assert chips.nth(3).inner_text() == f"{dates_tree.name}/{OLDDOC_MD}"
|
||||
# No rank-6+ doc exists in this four-doc KB → the related row is
|
||||
# absent (the de-emphasized row renders only when it has entries).
|
||||
expect(page.locator(".msg.brain .related-docs")).to_have_count(0)
|
||||
|
||||
# Durable record: one row, grounded, both docs cited in rank order.
|
||||
# Durable record: one row, grounded, the FULL retrieval (suggested
|
||||
# tier + related + read, deduped — here: all four docs) in rank
|
||||
# order (LOCKED A3).
|
||||
with SessionLocal() as db:
|
||||
row = db.scalars(select(QueryLog)).one()
|
||||
assert row.question == QUESTION
|
||||
assert row.deflected is False
|
||||
assert row.top_score >= 0.30 # the e2e mock-calibrated threshold
|
||||
assert (row.fts_hits or 0) >= 1
|
||||
assert f"{dates_tree.name}/{RETENTION_MD}" in row.sources
|
||||
assert row.sources.index(RETENTION_MD) < row.sources.index(DRAFT_MD)
|
||||
assert row.sources == ", ".join(
|
||||
f"{dates_tree.name}/{p}" for p in (RETENTION_MD, DRAFT_MD, FORWARD_MD, OLDDOC_MD)
|
||||
), row.sources
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -728,12 +755,20 @@ def test_anonymous_gate_and_editor_a11y(
|
||||
" === document.querySelector('#doc-modal .doc-date-input')"
|
||||
)
|
||||
# Keyboard traversal runs through the editor's controls (the modal
|
||||
# focus trap): Tab from Save → Cancel, and Tab from Revert wraps to
|
||||
# the panel's first control (the trap's edge behavior). (The CDP
|
||||
# key dispatch of headless Chromium does NOT perform the native
|
||||
# focus move off <input type="date"> itself — a harness artifact,
|
||||
# not a product defect: the same Tab works from every text input
|
||||
# and button in the editor, pinned here through the buttons.)
|
||||
# focus trap — the visible focusable order is "Full page" → Close →
|
||||
# the content's controls, in DOM order): Tab from Save → Cancel,
|
||||
# Tab from Revert → the summary editor's button (phase 118, A2: the
|
||||
# markdown doc is summarized TOO, so the modal's summary-edit
|
||||
# button is visible and joins the trap after the date editor — it
|
||||
# did not exist in the focusable set when this pin was written,
|
||||
# which is why the wrap below is now reached from IT), and Tab from
|
||||
# the summary editor wraps to the panel's FIRST control, the
|
||||
# "Full page" link (the trap's edge behavior the pin originally
|
||||
# carried). (The CDP key dispatch of headless Chromium does NOT
|
||||
# perform the native focus move off <input type="date"> itself — a
|
||||
# harness artifact, not a product defect: the same Tab works from
|
||||
# every text input and button in the editor, pinned here through
|
||||
# the buttons.)
|
||||
page.evaluate("() => document.querySelector('#doc-modal .doc-date-save').focus()")
|
||||
page.keyboard.press("Tab")
|
||||
assert page.evaluate(
|
||||
@@ -742,6 +777,12 @@ def test_anonymous_gate_and_editor_a11y(
|
||||
)
|
||||
page.evaluate("() => document.querySelector('#doc-modal .doc-date-revert').focus()")
|
||||
page.keyboard.press("Tab")
|
||||
assert page.evaluate(
|
||||
"() => !!document.activeElement"
|
||||
" && document.activeElement.classList.contains('doc-summary-edit')"
|
||||
)
|
||||
page.evaluate("() => document.querySelector('#doc-modal .doc-summary-edit').focus()")
|
||||
page.keyboard.press("Tab")
|
||||
assert page.evaluate(
|
||||
"() => !!document.activeElement"
|
||||
" && document.activeElement.classList.contains('doc-modal-open')"
|
||||
|
||||
@@ -1,4 +1,5 @@
|
||||
"""Phase 30 E2E (Playwright): a summary hit delivers the full source doc.
|
||||
"""Phase 30 E2E (Playwright) — phase 118 re-targeted (A2/A6): a summary
|
||||
hit seeds the SUMMARY into the prompt — never the full source doc.
|
||||
|
||||
Story: ``.agents/user_stories/document-summaries.md``
|
||||
Run in isolation (DB must be up: ``podman compose up -d db``):
|
||||
@@ -7,27 +8,33 @@ Run in isolation (DB must be up: ``podman compose up -d db``):
|
||||
|
||||
The fixture KB is a story-dedicated directory
|
||||
(``tests/fixtures/summary_kb/`` — the shared ``tests/fixtures/docs/``
|
||||
stays at its 9 pinned files) with two documents:
|
||||
stays at its pinned files) with two documents:
|
||||
|
||||
* ``quadlet/qwen-llamacpp.yaml`` — a non-markdown A9 doc. At import the
|
||||
mock ``lite`` model (``SUMMARY_MODE`` marker, ``tests/e2e/mock_llm.py``)
|
||||
reduces it to a deterministic 24-token digest, stored on
|
||||
``documents.summary`` and indexed as one ``is_summary`` chunk. The raw
|
||||
yaml body is deliberately token-diluted, so the document's best fused
|
||||
chunk is its summary chunk. The sentinel ``RESE-SUMMARY-SENTINEL-7f3a``
|
||||
sits on the document's LAST line — outside the 24-token digest,
|
||||
unreachable from the summary.
|
||||
* ``notes/qwen-llamacpp-notes.md`` — a markdown control doc (never
|
||||
summarized) that ranks first, which puts the yaml document LAST inside
|
||||
``<documents>``.
|
||||
* ``quadlet/qwen-llamacpp.yaml`` — at import the mock ``lite`` model
|
||||
(``SUMMARY_MODE`` marker, ``tests/e2e/mock_llm.py``) reduces it to a
|
||||
deterministic 24-token digest, stored on ``documents.summary`` and
|
||||
indexed as one ``is_summary`` chunk. The raw yaml body is deliberately
|
||||
token-diluted, so the document's top fused chunks are the summary and
|
||||
the one lexical-hitting raw chunk. The sentinel
|
||||
``RESE-SUMMARY-SENTINEL-7f3a`` sits on the document's LAST line —
|
||||
outside the 24-token digest, unreachable from the summary.
|
||||
* ``notes/qwen-llamacpp-notes.md`` — a markdown doc (phase 118 A2: it is
|
||||
summarized TOO — the phase-30 non-markdown-only scope is retired) that
|
||||
ranks first, which puts the yaml document LAST inside
|
||||
``<documents>`` (a two-doc KB → both docs are suggested, the related
|
||||
tier is empty).
|
||||
|
||||
The mock LLM's tail-echo trigger (``END_OF_NOTES_TRIGGER``) makes the
|
||||
answer quote the last 160 chars of the document context — the tail of
|
||||
the LAST selected document. The sentinel therefore appears in the
|
||||
rendered answer **iff the entire yaml source document (not the summary
|
||||
digest) reached the LLM prompt** — the summary→parent-document resolution
|
||||
through the unchanged chunk→document mapping (A7 revised: never
|
||||
truncated), which is what this story is about.
|
||||
answer quote the last 160 chars of the seeded ``<documents>`` block —
|
||||
under the phase-118 summary-seed contract (A6) that block carries the
|
||||
suggested docs' SUMMARIES, never their full texts, so the echoed tail is
|
||||
the LAST suggested doc's summary (the yaml doc's byte-stable digest tail
|
||||
+ ``Source:`` pointer line). The sentinel therefore appears in the
|
||||
rendered answer **only if the entire yaml source document (not the
|
||||
summary) reached the LLM prompt** — under the locked contract it must be
|
||||
ABSENT (full text enters the context only through the capped ``read``
|
||||
tool), which is the inverse of the retired phase-30/24 full-text pin
|
||||
and what this story is about now.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -47,7 +54,7 @@ from app.models import Document, QueryLog
|
||||
from app.rag.chunker import chunk_document
|
||||
from app.rag.importer import ImportSummary, import_sources
|
||||
from app.rag.llm import LLMClient
|
||||
from app.rag.retriever import RetrievedChunk, retrieve
|
||||
from app.rag.retriever import RetrievedChunk, retrieve, select_suggested
|
||||
from e2e.auth_helpers import login
|
||||
from tests.e2e.mock_llm import TOKEN_RE, embed_text
|
||||
|
||||
@@ -156,17 +163,20 @@ def _chunks_by_path(chunks: Sequence[RetrievedChunk], path: str) -> list[Retriev
|
||||
# --- Story tests -------------------------------------------------------------
|
||||
|
||||
|
||||
def test_summary_hit_retrieves_full_source_document(
|
||||
def test_summary_hit_seeds_the_summary_not_the_full_text(
|
||||
page: Page, app_url: str, mock_llm: int, db_ready: None
|
||||
) -> None:
|
||||
"""A question whose best yaml match is the summary chunk yields an
|
||||
answer grounded in the FULL yaml source document: its tail sentinel —
|
||||
which the summary digest cannot contain — is echoed back, and the
|
||||
source chip cites the yaml path (deflected: false)."""
|
||||
"""A question whose best yaml match is its summary chunk yields a
|
||||
grounded answer seeded from the SUMMARY — the mock's tail echo
|
||||
quotes the last suggested doc's summary tail (digest + pointer), and
|
||||
the sentinel (the yaml doc's last line, outside the digest) is
|
||||
ABSENT: the full source doc never reached the prompt (A6). Both
|
||||
fixture docs are suggested (two-doc KB, no floor) and both chips
|
||||
render (deflected: false)."""
|
||||
_reset_db()
|
||||
summary = _run_in_thread(_import_fixtures(mock_llm))
|
||||
assert summary.added == 2 # yaml + md control
|
||||
assert summary.summaries == 1 and summary.summary_errors == 0
|
||||
assert summary.summaries == 2 and summary.summary_errors == 0 # A2: md too
|
||||
assert summary.errors == 0
|
||||
|
||||
# Import state: exactly one embedded ``is_summary`` chunk (position
|
||||
@@ -182,30 +192,48 @@ def test_summary_hit_retrieves_full_source_document(
|
||||
assert schunks[0].embedding is not None
|
||||
assert yaml_doc.summary == _expected_summary(yaml_content, SOURCE, YAML_PATH)
|
||||
|
||||
# Retrieval state: the summary chunk is the yaml document's best fused
|
||||
# chunk — the document enters the context through its summary, not
|
||||
# through the diluted raw yaml chunks.
|
||||
# Retrieval state (phase 118, A2/A6): the EMBEDDED summary chunk is
|
||||
# a retrieval candidate (the seeded text ranks on its own), and the
|
||||
# suggested tier is the two docs in rank order — md first, yaml
|
||||
# LAST (so the yaml's summary is the tail of the <documents> block,
|
||||
# the mock echo's target). The document enters the prompt through
|
||||
# its SUMMARY block, not through the diluted raw yaml chunks.
|
||||
with SessionLocal() as db:
|
||||
chunks = retrieve(db, QUESTION, embed_text(QUESTION))
|
||||
yaml_chunks = _chunks_by_path(chunks, YAML_PATH)
|
||||
best_yaml = max(yaml_chunks, key=lambda c: c.score)
|
||||
assert best_yaml.is_summary
|
||||
assert any(c.is_summary for c in yaml_chunks) # the embedded summary ranks
|
||||
assert len(yaml_chunks) >= 2 # summary + at least one raw candidate
|
||||
assert [d.path for d in select_suggested(chunks)] == [MD_PATH, YAML_PATH]
|
||||
|
||||
bubble = _ask(page, app_url, QUESTION)
|
||||
|
||||
# The tail sentinel exists only on the document's last line and
|
||||
# cannot be in the summary digest — its presence proves the entire
|
||||
# source document was in the LLM prompt (summary→parent resolution).
|
||||
expect(bubble).to_contain_text(SENTINEL, timeout=30_000)
|
||||
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER)
|
||||
# Phase 118 (A6): the seed is the SUMMARY, not the full text — the
|
||||
# mock's tail echo quotes the last 160 chars of the <documents>
|
||||
# block, which end in the LAST suggested doc's summary: the yaml
|
||||
# doc's byte-stable digest tail + pointer line. The sentinel
|
||||
# (document's last line, outside the digest) is therefore ABSENT —
|
||||
# the full source doc never reached the prompt (full text enters
|
||||
# only through the capped read tool; the inverse of the retired
|
||||
# phase-30/24 full-text pin). The bubble renders the answer as
|
||||
# markdown, which collapses the summary's newline — so pin each
|
||||
# LINE separately (the digest line's tail sits inside the echoed
|
||||
# 160 chars; the pointer line is single-line too).
|
||||
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER, timeout=30_000)
|
||||
yaml_summary = _expected_summary(yaml_content, SOURCE, YAML_PATH)
|
||||
yaml_digest_line = yaml_summary.split("\n", 1)[0]
|
||||
expect(bubble).to_contain_text(f"Source: {SOURCE}/{YAML_PATH}")
|
||||
expect(bubble).to_contain_text(yaml_digest_line[-80:])
|
||||
expect(bubble).not_to_contain_text(SENTINEL)
|
||||
|
||||
# Grounded: the yaml source chip renders (the md doc ranks first, so
|
||||
# both fixtures are cited).
|
||||
# Grounded: both fixture docs are suggested (two-doc KB — no floor)
|
||||
# and both chips render, in rank order (md first, yaml last).
|
||||
chip = page.locator(".msg.brain .source-chip", has_text=YAML_PATH)
|
||||
expect(chip).to_have_count(1)
|
||||
expect(chip.first).to_contain_text(f"{SOURCE}/{YAML_PATH}")
|
||||
expect(page.locator(".msg.brain .source-chip", has_text=MD_PATH)).to_have_count(1)
|
||||
expect(page.locator(".msg.brain .source-chip")).to_have_count(2)
|
||||
# No rank-6+ doc in a two-doc KB → the related row is absent.
|
||||
expect(page.locator(".msg.brain .related-docs")).to_have_count(0)
|
||||
|
||||
# Button recovers (never stale) and the turn was grounded, not
|
||||
# deflected.
|
||||
@@ -214,20 +242,26 @@ def test_summary_hit_retrieves_full_source_document(
|
||||
row = _last_query_log()
|
||||
assert row.question == QUESTION
|
||||
assert row.deflected is False
|
||||
assert f"{SOURCE}/{YAML_PATH}" in row.sources
|
||||
assert f"{SOURCE}/{MD_PATH}" in row.sources
|
||||
# The durable record: suggested + related + read (deduped) — here
|
||||
# exactly the two suggested docs, in rank order (LOCKED A3).
|
||||
assert row.sources == (
|
||||
f"{SOURCE}/{MD_PATH}, {SOURCE}/{YAML_PATH}"
|
||||
), row.sources
|
||||
|
||||
|
||||
def test_markdown_control_doc_gets_no_summary_chunk(
|
||||
def test_markdown_control_doc_gets_a_summary_chunk(
|
||||
mock_llm: int, db_ready: None
|
||||
) -> None:
|
||||
"""Control: in the same KB the markdown doc gets no summary at all —
|
||||
its chunk count is exactly the raw chunks; the yaml doc has exactly
|
||||
one ``is_summary`` row and its raw chunk count is untouched by the
|
||||
summary."""
|
||||
"""Phase 118 (A2): in the same KB the markdown doc gets a summary TOO
|
||||
— the phase-30 non-markdown-only scope is retired (markdown docs
|
||||
backfill + summarize like every other doc): the same byte-stable
|
||||
digest + pointer line, exactly one ``is_summary`` chunk, and the raw
|
||||
chunk count untouched. The yaml doc keeps its exactly-one summary
|
||||
row with its raw chunks untouched."""
|
||||
_reset_db()
|
||||
summary = _run_in_thread(_import_fixtures(mock_llm))
|
||||
assert summary.added == 2
|
||||
assert summary.summaries == 2 # A2: the markdown doc is summarized too
|
||||
|
||||
md_content = (FIXTURES / MD_PATH).read_text(encoding="utf-8")
|
||||
yaml_content = (FIXTURES / YAML_PATH).read_text(encoding="utf-8")
|
||||
@@ -235,11 +269,20 @@ def test_markdown_control_doc_gets_no_summary_chunk(
|
||||
md_doc = _doc(db, MD_PATH)
|
||||
yaml_doc = _doc(db, YAML_PATH)
|
||||
md_chunks = [c for c in md_doc.chunks if not c.is_summary]
|
||||
md_summary = [c for c in md_doc.chunks if c.is_summary]
|
||||
yaml_raw = [c for c in yaml_doc.chunks if not c.is_summary]
|
||||
yaml_summary = [c for c in yaml_doc.chunks if c.is_summary]
|
||||
|
||||
# Markdown: never summarized (phase 30 scope — A9 non-markdown only).
|
||||
assert md_doc.summary is None
|
||||
# Markdown: NOW summarized (phase 118 A2) — the same byte-stable
|
||||
# digest + deterministic pointer line, exactly one is_summary chunk
|
||||
# (position −1, embedded), raw chunks untouched.
|
||||
expected_md_summary = _expected_summary(md_content, SOURCE, MD_PATH)
|
||||
assert md_doc.summary == expected_md_summary
|
||||
assert len(md_summary) == 1
|
||||
assert md_summary[0].position == -1
|
||||
assert md_summary[0].embedding is not None
|
||||
assert md_summary[0].content == expected_md_summary
|
||||
assert expected_md_summary.endswith(f"\nSource: {SOURCE}/{MD_PATH}")
|
||||
assert len(md_chunks) == len(
|
||||
chunk_document(md_content, MD_PATH, CHUNK_TARGET, CHUNK_OVERLAP)
|
||||
)
|
||||
|
||||
@@ -201,7 +201,13 @@ def test_viewer_header_content_still_fits(
|
||||
assert _box_height(page, "#doc-title") < 30 # one line at either size
|
||||
|
||||
# Meta row: badges + path present and visible, single line.
|
||||
assert _box_height(page, ".doc-meta") < 26
|
||||
# The bound catches a WRAP (a second line would be ≥ ~44px),
|
||||
# not a pixel-exact line height: the row is content-sized by
|
||||
# its tallest child — the native type=date input (phase 106
|
||||
# D8), whose metrics are Chromium/font-dependent and render
|
||||
# the single line at ~26px on this host (the pre-phase 26px
|
||||
# pin red-lined on the 26.125px measurement, 2026-09-16).
|
||||
assert _box_height(page, ".doc-meta") < 34
|
||||
expect(page.locator(".doc-source-badge", has_text="docs")).to_be_visible()
|
||||
expect(page.locator(".format-badge", has_text="md")).to_be_visible()
|
||||
expect(page.locator(".doc-path", has_text="homelab/kubernetes.md")).to_be_visible()
|
||||
|
||||
@@ -57,14 +57,18 @@ Test → story mapping (Playwright Mapping Rule):
|
||||
``source: X | path: Y | title: Z`` file lines), and the grounded
|
||||
``read(alpha/two/two-a.md)`` (the ``📄 Reading …`` line + the answer
|
||||
citing the document, the phase-37 assertion pattern). The read
|
||||
target is DELIBERATELY a top-2 retrieval document for its question
|
||||
(the question names the file's path, so the file self-matches the
|
||||
hybrid gate deterministically): the agent's phase-72 dedupe returns
|
||||
``ALREADY_IN_CONTEXT`` (a refusal — counts in nothing), and the
|
||||
mock answers FROM THE ``<documents>`` PROMPT with the same citation
|
||||
shape (``Already in context: Read <sp>. <first 80 chars>``) — the
|
||||
document text reached the model either way, and the prefix pins
|
||||
that the dedupe notice itself reached it.
|
||||
target is DELIBERATELY a suggested (summary-seeded) document for
|
||||
its question (the question names the file's path, so the file
|
||||
self-matches the hybrid gate and takes rank 1 deterministically):
|
||||
phase 118 (A6) — the seed is a SUMMARY, not full text, so the read
|
||||
SUCCEEDS (the phase-72 ALREADY_IN_CONTEXT dedupe fires only for a
|
||||
document already READ in the turn — the retired top-2 seed-read
|
||||
refusal is gone) and the full text arrives through the ``read``
|
||||
tool; the mock answers from the READ RESULT with the phase-37
|
||||
citation shape (``Read <sp>. <first 80 chars of the document
|
||||
content>`` — the mock skips the phase-106 D5 ``date:`` line, so the
|
||||
quote is pure content, byte-identical to the retired
|
||||
answer-from-prompt quote).
|
||||
2. ``test_wide_folder_holds_the_fifty_line_cap`` — ``ls(alpha/wide)``
|
||||
on the 51-file folder: the mock's echo carries exactly 50 file lines
|
||||
+ the ``…and 1 more documents in this folder — use grep (pattern)…``
|
||||
@@ -226,12 +230,16 @@ READ_QUESTION = f"Drill down the tree: read {READ_SP} — read the file"
|
||||
WIDE_QUESTION = f"Drill down the tree: ls {ALPHA}/wide — how many files does this folder hold?"
|
||||
NOPE_QUESTION = f"Drill down the tree: ls {ALPHA}/nope — is there such a folder?"
|
||||
|
||||
#: The read target is a top-2 retrieval document for its question (the
|
||||
#: question names the path — the file self-matches the hybrid gate
|
||||
#: deterministically), so the read gets the phase-72 ALREADY_IN_CONTEXT
|
||||
#: dedupe and the mock answers from the ``<documents>`` prompt with the
|
||||
#: same citation shape, prefixed (the suite pins the dedupe path).
|
||||
READ_ANSWER_PREFIX = f"Already in context: Read {READ_SP}."
|
||||
#: The read target is a suggested (summary-seeded) document for its
|
||||
#: question (the question names the path — the file self-matches the
|
||||
#: hybrid gate and takes rank 1 deterministically). Phase 118 (A6):
|
||||
#: the seed is a summary, not full text — a first read of a suggested
|
||||
#: document SUCCEEDS, the full text arrives through the read tool, and
|
||||
#: the mock answers from the READ RESULT with the phase-37 citation
|
||||
#: shape (the mock skips the phase-106 D5 ``date:`` line, so the quote
|
||||
#: is pure document content — byte-identical to the retired
|
||||
#: answer-from-prompt quote, which the suite pins).
|
||||
READ_ANSWER_PREFIX = f"Read {READ_SP}."
|
||||
READ_ANSWER_QUOTE = TWO_A_CONTENT[:80]
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
@@ -658,14 +666,15 @@ def test_drill_down_sources_folders_files_and_read(
|
||||
{"type": "tool", "name": "read", "argument": READ_SP}
|
||||
], _tool_frames(frames)
|
||||
bubble = _last_brain(page).locator(".bubble")
|
||||
# The dedupe notice reached the model (the prefix pins it — the read
|
||||
# was refused as ALREADY_IN_CONTEXT because the target is a top-2
|
||||
# retrieval document)…
|
||||
# Phase 118 (A6): the read SUCCEEDED — the seeds are summaries,
|
||||
# not full text, so the phase-72 ALREADY_IN_CONTEXT refusal (only
|
||||
# for a document already READ in the turn) did not fire —
|
||||
expect(bubble).to_contain_text(READ_ANSWER_PREFIX)
|
||||
# …and the answer still cites the document: the mock quotes the
|
||||
# FIRST 80 chars of the target's text from the ``<documents>``
|
||||
# prompt (the refusal's instruction — answer from that text; the
|
||||
# quote is newline-free, pinned above).
|
||||
expect(bubble).not_to_contain_text("Already in context")
|
||||
# — and the answer cites the document from the READ RESULT: the mock
|
||||
# quotes the FIRST 80 chars of the document's own content (the full
|
||||
# text arrived through the read tool; the quote is newline-free,
|
||||
# pinned above).
|
||||
expect(bubble).to_contain_text(READ_ANSWER_QUOTE)
|
||||
done = next(f for f in frames if f.get("type") == "done")
|
||||
assert done["deflected"] is False, done
|
||||
|
||||
@@ -623,8 +623,11 @@ def test_api_created_chats_carry_the_bor_chat_v1_shape(
|
||||
assert body["message_count"] == 2
|
||||
# The bor.chat.v1 record shape: who/text present; only the schema's
|
||||
# optional keys may accompany them (explicit nulls are preserved —
|
||||
# the plain model_dump round-trips byte-identical).
|
||||
allowed = {"who", "text", "sources", "deflected", "suggestions",
|
||||
# the plain model_dump round-trips byte-identical). ``related`` is
|
||||
# the phase-113 related-doc tier the UI persists with grounded
|
||||
# brain records (accepted by the ChatMessage schema — its earlier
|
||||
# absence 422'd the done-time auto-save, phase-118 verification fix).
|
||||
allowed = {"who", "text", "sources", "related", "deflected", "suggestions",
|
||||
"thinking", "tools", "stopped"}
|
||||
assert all({"who", "text"} <= set(m) <= allowed for m in body["messages"])
|
||||
# The History row renders it (the Open link's text IS the title).
|
||||
|
||||
@@ -36,16 +36,14 @@ the same host) registered as a local-directory source (the
|
||||
``test_local_directory_sources.py`` registration + real-Sync pattern —
|
||||
registration through the authenticated API, the real in-process
|
||||
``POST /api/sync`` pipeline; no git anywhere), with THREE documents
|
||||
whose bodies are token-controlled so the phase-72 ALREADY_IN_CONTEXT
|
||||
dedupe (the read target is refused when it is a top-2 retrieval seed)
|
||||
NEVER fires — each scripted read target must actually execute, not be
|
||||
refused:
|
||||
whose bodies are token-controlled so every scripted turn stays
|
||||
GROUNDED (the hybrid gate HIGH — the mock's read-cap flow needs the
|
||||
``<tools>`` section):
|
||||
|
||||
* ``anchor-2024.md`` — the retrieval ANCHOR: a digit-bearing name.
|
||||
Both questions name "anchor 2024", whose joined normalized token
|
||||
(``anchor2024``) name-hits this document — the name-hit list LEADS
|
||||
the lexical side, so the anchor is the #1 seed of BOTH turns
|
||||
(grounding: the name hit's ``fts_hit`` keeps the gate HIGH, the
|
||||
(``anchor2024``) name-hits this document — the name hit's
|
||||
``fts_hit`` keeps the gate HIGH for every turn (grounding: the
|
||||
``<tools>`` section is present, and the anchor is never a read
|
||||
target);
|
||||
* ``zz-capped.md`` — the SUBJECT: ~3 100 chars of varied rotation
|
||||
@@ -53,18 +51,19 @@ refused:
|
||||
``chars_total`` the assertions use is ``len(CAP_DOC)`` of this very
|
||||
string, and ``synced_kb`` pins the stored content byte-identical to
|
||||
it). Its body avoids EVERY token of both questions, so on its own
|
||||
(turn 1) question it has zero lexical hits and only the common-word
|
||||
cosine — it is the #3 fused candidate, NOT a seed;
|
||||
question it has zero lexical hits and only the common-word cosine —
|
||||
it ranks below the anchor (irrelevant to the cap's contract: phase
|
||||
118, A6 — the seeds are SUMMARIES, so a first ``read`` of ANY
|
||||
document, suggested or not, succeeds, and the cap applies to
|
||||
everything ``read`` returns — the retired top-2 seed-read refusal
|
||||
that used to constrain this design is gone);
|
||||
* ``aa-short.md`` — the CONTROL: ~200 chars, under the cap.
|
||||
|
||||
The turn-specific flavor words pick the #2 seed: turn 1's question
|
||||
carries ``note, long form`` (planted in the SHORT doc's body) and turn
|
||||
2's carries ``quick pass`` (planted in the CAPPED doc's body), so each
|
||||
turn's NON-target document wins the second seed slot on its own
|
||||
question and the target stays #3/#6. ``synced_kb`` pins this design
|
||||
with the app's real hybrid retrieval (``_assert_target_not_seed`` — a
|
||||
fixture-text regression that makes a target a seed fails at setup with
|
||||
a clear message, not at the wire assertions).
|
||||
The turn-specific flavor words (``note, long form`` planted in the
|
||||
SHORT doc's body, ``quick pass`` in the CAPPED doc's) are inert
|
||||
leftovers of the retired not-a-seed design — they no longer select a
|
||||
seed slot (phase 118: a read of a suggested document succeeds, so the
|
||||
target's seed status never changes the wire contract).
|
||||
|
||||
Test → story mapping (Playwright Mapping Rule; the story is the owner
|
||||
TODO item — one Playwright file per story, A16):
|
||||
@@ -116,7 +115,6 @@ from e2e.conftest import (
|
||||
USE_REAL_LLM,
|
||||
_wait_http,
|
||||
)
|
||||
from e2e.mock_llm import embed_text
|
||||
|
||||
REPO = Path(__file__).resolve().parents[2]
|
||||
|
||||
@@ -163,9 +161,10 @@ def _cap_doc() -> str:
|
||||
every token of BOTH questions (``read``, ``capped``, ``document``,
|
||||
``capkb``, ``zz``, ``md``, ``anchor``, ``2024``, ``note``,
|
||||
``long``, ``form``, ``aa``, ``short``, ``quick``, ``pass``,
|
||||
``wire``, ``save``, ``control``, ``check``) — the target must have
|
||||
zero lexical hits on its own turn so the dedupe cannot refuse the
|
||||
read."""
|
||||
``wire``, ``save``, ``control``, ``check``) — the target has zero
|
||||
lexical hits on its own turn, so it ranks below the anchor (phase
|
||||
118: its seed status is irrelevant — a first read of any document
|
||||
succeeds; the anchor keeps the gate HIGH)."""
|
||||
topics = [
|
||||
"vault", "mirror", "raid", "pool", "drive",
|
||||
"chain", "slot", "cycle", "guard", "probe",
|
||||
@@ -226,7 +225,8 @@ NOTICE = READ_TRUNCATION_NOTICE.format(shown=READ_CAP, total=CAP_DOC_TOTAL)
|
||||
|
||||
# The scripted turns (the mock's ``READ_CAP_TRIGGER`` questions — each
|
||||
# carries its own tool call after the colon; the turn-specific flavor
|
||||
# words are the #2-seed selectors, see the module docstring).
|
||||
# words are inert leftovers of the retired #2-seed design — see the
|
||||
# module docstring).
|
||||
Q_WIRE = (
|
||||
f"Read the capped document: read {CAPPED_SP} — "
|
||||
"anchor 2024 note, long form, wire check"
|
||||
@@ -359,26 +359,6 @@ def _truncate_all() -> None:
|
||||
db.commit()
|
||||
|
||||
|
||||
def _assert_target_not_seed(question: str, target_rel: str) -> None:
|
||||
"""Pin the retrieval-anchor design (see the module docstring) with
|
||||
the app's REAL hybrid retrieval over the mock's embeddings
|
||||
(deterministic): the scripted read target must NOT be a top-2 seed
|
||||
for its own question — the phase-72 ALREADY_IN_CONTEXT dedupe would
|
||||
refuse the read and the cap would never fire (the turn would echo
|
||||
the refusal instead). A fixture-text regression that breaks this
|
||||
fails here, at setup, with a clear message."""
|
||||
from app.rag.retriever import retrieve, select_documents
|
||||
|
||||
with SessionLocal() as db:
|
||||
seeds = select_documents(retrieve(db, question, embed_text(question)))
|
||||
paths = [f"{d.source}/{d.path}" for d in seeds]
|
||||
assert f"{SOURCE}/{target_rel}" not in paths, (
|
||||
f"the read target {SOURCE}/{target_rel} is a top-2 seed for its own "
|
||||
f"question — the ALREADY_IN_CONTEXT dedupe would refuse the read and "
|
||||
f"the cap would never fire (seeds: {paths})"
|
||||
)
|
||||
|
||||
|
||||
def _wait_sync_done_http(client: httpx.Client, timeout_s: float = 180.0) -> dict[str, Any]:
|
||||
"""Poll the (cookie-authenticated) status endpoint until the run
|
||||
reaches a terminal state (the test_ls_tree_drilldown pattern, over
|
||||
@@ -403,8 +383,10 @@ def synced_kb(app_server: str, cap_dirs: Path) -> None:
|
||||
runs the REAL in-process sync (``POST /api/sync`` — walk → chunk →
|
||||
embed → overview → version bump), pins the stored content
|
||||
byte-identical to the fixture strings (the ``chars_total``
|
||||
assumption), and pins the retrieval-anchor design for all three
|
||||
scripted questions (the dedupe never fires)."""
|
||||
assumption). Phase 118: the retrieval-anchor design only pins the
|
||||
gate's grounding (the anchor keeps every turn HIGH) — a read of a
|
||||
suggested document succeeds, so the targets' seed status no longer
|
||||
constrains the suite."""
|
||||
_truncate_all()
|
||||
with httpx.Client(base_url=app_server, timeout=30.0) as client:
|
||||
r = client.post("/api/login", json={"password": ADMIN_PASSWORD})
|
||||
@@ -436,11 +418,6 @@ def synced_kb(app_server: str, cap_dirs: Path) -> None:
|
||||
)
|
||||
)
|
||||
assert stored == expected, f"stored content drifted for {rel}"
|
||||
# The retrieval-anchor design (the module docstring): every scripted
|
||||
# read target stays OUT of its own question's top-2 seeds.
|
||||
_assert_target_not_seed(Q_WIRE, CAPPED_REL)
|
||||
_assert_target_not_seed(Q_SAVE, CAPPED_REL)
|
||||
_assert_target_not_seed(Q_CONTROL, SHORT_REL)
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
@@ -681,9 +658,10 @@ def test_truncated_read_frame_order_live_marker_and_llm_notice(
|
||||
assert i_tool < i_result < i_delta, (i_tool, i_result, i_delta)
|
||||
done = next(f for f in frames if f.get("type") == "done")
|
||||
assert done["deflected"] is False, done
|
||||
# The read document is the turn's cited source even though it was
|
||||
# NOT a retrieval seed (the anchor design) — retrieval + agent-read,
|
||||
# deduped (the grounded-turn record).
|
||||
# The read document is the turn's cited source (retrieval +
|
||||
# agent-read, deduped — the grounded-turn record; phase 118: the
|
||||
# target may be a suggested (summary-seeded) document — its seed
|
||||
# status never changes the cap's wire contract).
|
||||
assert any(
|
||||
s["path"] == CAPPED_REL and s["source"] == SOURCE for s in done["sources"]
|
||||
), done["sources"]
|
||||
|
||||
@@ -346,9 +346,10 @@ def test_anonymous_shared_view(
|
||||
expect(think.first).not_to_have_attribute("open")
|
||||
|
||||
# Source chips are PLAIN TEXT: the on-topic turn carries its
|
||||
# source chips (top-2 docs — the hybrid retrieval), but every
|
||||
# one as a <span>: zero <a.source-chip> anywhere (a guest
|
||||
# cannot open documents; the documents API is admin-only).
|
||||
# source chips (the suggested docs — the hybrid retrieval's
|
||||
# summary-seed tier, phase 118), but every one as a <span>:
|
||||
# zero <a.source-chip> anywhere (a guest cannot open documents;
|
||||
# the documents API is admin-only).
|
||||
assert (
|
||||
anon.locator(".msg.brain .source-chip").count() >= 1
|
||||
), "the grounded turn must carry its source chips"
|
||||
|
||||
@@ -1,35 +1,44 @@
|
||||
"""Phase 113 E2E (Playwright): the source-chip quality contract (TODO L5 +
|
||||
L2c) — the usefulness bar + the de-emphasized related-docs row, as VISIBLE
|
||||
chip counts.
|
||||
"""Phase 113 E2E (Playwright) — phase 118 re-targeted: the citation-surface
|
||||
contract (LOCKED A4) as VISIBLE chip counts — the summary-seed contract
|
||||
replaced the phase-112/113 usefulness bar: the chip row is the
|
||||
SUGGESTED tier (top-5 distinct docs, NO floor) + agent reads (none on
|
||||
these turns), the related row is rank 6+ (capped at
|
||||
``related_max_docs`` = 2), and a deflected turn still cites nothing
|
||||
(done.sources = [] — its weak hits are suggested for the durable record
|
||||
but are never citation chips).
|
||||
|
||||
Run in isolation (DB must be up: ``podman compose up -d db``):
|
||||
|
||||
uv run pytest tests/e2e/test_source_chip_quality.py -v --no-cov
|
||||
|
||||
Acceptance (TODO L144–146): "for a single-document question, the turn
|
||||
shows one citation chip"; a weak 2nd doc renders only in the
|
||||
de-emphasized related row (``.related-doc`` links, NEVER ``.source-chip``);
|
||||
a deflected turn shows zero citation chips (its weak hits, if any, live in
|
||||
the related row).
|
||||
Acceptance (TODO L144–146, phase-118 shape): a grounded turn's chip row
|
||||
is the suggested tier — the 5 docs the model was seeded with (its
|
||||
"start here" set — the old "one chip per bar-clearing doc" is retired
|
||||
with the full-text seeds; the LLM decides what the summaries earn);
|
||||
a deflected turn shows zero citation chips (its weak hits, if any, live
|
||||
in the de-emphasized related row — ``.related-doc`` links, NEVER
|
||||
``.source-chip``; L2c: never render weak hits as answer citations).
|
||||
|
||||
The fixture KB's tier shapes are deterministic under the E2E mock's
|
||||
bag-of-words embeddings + the mock-calibrated bar (conftest:
|
||||
``BOR_SOURCE_USEFULNESS_FLOOR=0.15``, half the 0.30 threshold — like the
|
||||
lexical floor):
|
||||
bag-of-words embeddings + the fused rank walk (probe-verified, pinned
|
||||
here against the wire):
|
||||
|
||||
* **single-source question** — "What SSH aliases do I have?":
|
||||
``ssh_aliases.txt`` is the ONLY doc whose best-chunk cosine clears the
|
||||
bar (0.352 ≥ 0.15; grounded at 0.352 ≥ 0.30) → the done frame carries
|
||||
exactly ONE cited ref; the two below-bar docs (gitlab-compose.yaml
|
||||
0.025, uptime_probe.py 0.113) ride the related tier. This IS the
|
||||
strong+weak two-tier shape on the wire (the four OBSERVED live shapes
|
||||
are unit-pinned at plan level in
|
||||
``tests/unit/test_source_chip_quality.py`` — the fixture KB reproduces
|
||||
the same shape live, so no docstring caveat is needed).
|
||||
* **grounded question** — "What SSH aliases do I have?": best cosine
|
||||
0.352 ≥ 0.30 → grounded; the suggested tier (top-5, NO floor) is
|
||||
ssh_aliases.txt > gitlab-compose.yaml > uptime_probe.py > kubernetes.md
|
||||
> backups.md (fused rank — the lexical-only cosine-0.0 docs rank when
|
||||
they rank: no floor filters them, LOCKED A3) and the done frame
|
||||
carries EXACTLY those five cited refs; the rank-6+ remainder
|
||||
(compose.container, static-dns.json) rides the related row, capped at
|
||||
two. (The four OBSERVED live shapes are unit-pinned at plan level in
|
||||
``tests/unit/test_source_chip_quality.py`` — the fixture KB
|
||||
reproduces the tiered shape live, so no docstring caveat is needed.)
|
||||
* **deflected question** — "How do I bake sourdough bread?": best cosine
|
||||
0.109 < 0.30 and zero FTS hits → honest deflection; nothing clears the
|
||||
bar → zero cited refs; the weak hits (new-service.md 0.109,
|
||||
ssh_aliases.txt 0.050) ride the related row.
|
||||
0.109 < 0.30 and zero FTS hits → honest deflection → zero cited refs
|
||||
(done.sources = []); the weak hits ARE suggested (no floor) for the
|
||||
durable record — new-service.md > ssh_aliases.txt > lan.network >
|
||||
uptime_probe.py > compose.container — and the rank-6+ remainder
|
||||
(backups.md, kubernetes.md) rides the related row.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -97,12 +106,14 @@ def _ask(page: Page, message: str) -> None:
|
||||
def test_single_source_question_shows_exactly_one_citation_chip(
|
||||
page: Page, app_url: str, mock_llm: int, db_ready: None
|
||||
) -> None:
|
||||
"""The acceptance pin: a single-document question → the done bubble
|
||||
carries EXACTLY ONE ``.source-chip`` (the bar-clearing doc) and the
|
||||
below-bar docs render only in the de-emphasized ``.related-docs`` row
|
||||
(``.related-doc`` links — never ``.source-chip``), labeled
|
||||
"Nearby docs, in case:". The durable record keeps the FULL retrieval
|
||||
(LOCKED A3)."""
|
||||
"""The acceptance pin (phase 118, A4): a grounded question → the done
|
||||
bubble carries EXACTLY the suggested tier as ``.source-chip``s (the
|
||||
five docs the model was seeded with — top-5, NO floor: the old
|
||||
usefulness-bar "one chip" is retired with the full-text seeds) and
|
||||
the rank-6+ remainder renders only in the de-emphasized
|
||||
``.related-docs`` row (``.related-doc`` links — never
|
||||
``.source-chip``), labeled "Nearby docs, in case:". The durable
|
||||
record keeps the FULL retrieval (LOCKED A3)."""
|
||||
_reset_db(mock_llm)
|
||||
page.set_default_timeout(30_000)
|
||||
login(page, app_url, next="/") # phase 79: chat is require_user-gated
|
||||
@@ -116,52 +127,64 @@ def test_single_source_question_shows_exactly_one_citation_chip(
|
||||
# Grounded: no deflected bubble at all.
|
||||
expect(page.locator(".msg.brain.is-deflected")).to_have_count(0)
|
||||
|
||||
# EXACTLY ONE citation chip — the bar-clearing doc (the 2nd chip of
|
||||
# the pre-phase turn demoted to the row; the acceptance criterion).
|
||||
# The citation surface IS the suggested tier (LOCKED A4 — top-5,
|
||||
# NO floor): five chips, in fused rank order — the model was
|
||||
# seeded with exactly these five docs' summaries.
|
||||
chip = page.locator(".msg.brain .source-chip")
|
||||
expect(chip).to_have_count(1, timeout=30_000)
|
||||
expect(chip.first).to_contain_text("ssh_aliases.txt")
|
||||
expect(chip).to_have_count(5, timeout=30_000)
|
||||
expect(chip.nth(0)).to_contain_text("ssh_aliases.txt")
|
||||
expect(chip.nth(1)).to_contain_text("gitlab-compose.yaml")
|
||||
expect(chip.nth(2)).to_contain_text("uptime_probe.py")
|
||||
expect(chip.nth(3)).to_contain_text("kubernetes.md")
|
||||
expect(chip.nth(4)).to_contain_text("backups.md")
|
||||
|
||||
# The below-bar docs ride the related row: a labeled, de-emphasized
|
||||
# list — one .related-doc link per doc (rank order, capped at
|
||||
# related_max_docs = 2), never a .source-chip.
|
||||
# The rank-6+ remainder rides the related row: a labeled,
|
||||
# de-emphasized list — one .related-doc link per doc (rank order,
|
||||
# capped at related_max_docs = 2), never a .source-chip.
|
||||
row = page.locator(".msg.brain .related-docs")
|
||||
expect(row).to_have_count(1)
|
||||
expect(row.first).to_have_attribute("aria-label", "Nearby docs, in case")
|
||||
expect(row.first.locator(".related-docs-label")).to_have_text("Nearby docs, in case:")
|
||||
links = page.locator(".msg.brain .related-docs .related-doc")
|
||||
expect(links).to_have_count(2, timeout=30_000)
|
||||
expect(links.nth(0)).to_contain_text("gitlab-compose.yaml")
|
||||
expect(links.nth(1)).to_contain_text("uptime_probe.py")
|
||||
expect(links.nth(0)).to_contain_text("compose.container")
|
||||
expect(links.nth(1)).to_contain_text("static-dns.json")
|
||||
expect(page.locator(".msg.brain .related-docs .source-chip")).to_have_count(0)
|
||||
# The related links keep the chip's /document.html href + identity.
|
||||
expect(links.first).to_have_attribute(
|
||||
"title", "docs/homelab/container_gitlab/gitlab-compose.yaml"
|
||||
"title", "docs/homelab/quadlet/compose.container"
|
||||
)
|
||||
|
||||
# Durable record: not deflected; the FULL retrieval (cited + related)
|
||||
# is logged — query_log records retrieval, not citations (LOCKED A3).
|
||||
# Durable record: not deflected; the FULL retrieval (suggested +
|
||||
# related + read, deduped) is logged — query_log records retrieval,
|
||||
# not citations (LOCKED A3).
|
||||
with SessionLocal() as db:
|
||||
row_log = db.scalars(select(QueryLog)).one()
|
||||
assert row_log.question == SINGLE_SOURCE_QUESTION
|
||||
assert row_log.deflected is False
|
||||
for path in (
|
||||
"homelab/ssh/ssh_aliases.txt",
|
||||
"homelab/container_gitlab/gitlab-compose.yaml",
|
||||
"homelab/scripts/uptime_probe.py",
|
||||
):
|
||||
assert path in row_log.sources
|
||||
# Suggested tier + related remainder, in the logged order.
|
||||
assert row_log.sources == (
|
||||
"docs/homelab/ssh/ssh_aliases.txt, "
|
||||
"docs/homelab/container_gitlab/gitlab-compose.yaml, "
|
||||
"docs/homelab/scripts/uptime_probe.py, "
|
||||
"docs/homelab/kubernetes.md, "
|
||||
"docs/homelab/backups.md, "
|
||||
"docs/homelab/quadlet/compose.container, "
|
||||
"docs/homelab/networking/static-dns.json"
|
||||
), row_log.sources
|
||||
|
||||
|
||||
def test_deflected_question_shows_zero_citation_chips(
|
||||
page: Page, app_url: str, mock_llm: int, db_ready: None
|
||||
) -> None:
|
||||
"""The acceptance pin: a deflected turn (known-out-of-KB) → ZERO
|
||||
``.source-chip`` elements under the bubble; the weak hits live in the
|
||||
related row only — its links are ``.related-doc``, never
|
||||
``.source-chip`` (L2c: "at minimum: never render them as answer
|
||||
citations"). The "Maybe try" chips and the durable record are
|
||||
unchanged."""
|
||||
"""The acceptance pin (phase 118, A4): a deflected turn
|
||||
(known-out-of-KB) → ZERO ``.source-chip`` elements under the bubble
|
||||
(done.sources = [] — a deflected answer cites nothing; the weak hits
|
||||
are suggested for the durable record but are scored docs, not
|
||||
citations); the rank-6+ weak hits live in the related row only —
|
||||
its links are ``.related-doc``, never ``.source-chip`` (L2c: "at
|
||||
minimum: never render them as answer citations"). The "Maybe try"
|
||||
chips are unchanged."""
|
||||
_reset_db(mock_llm)
|
||||
page.set_default_timeout(30_000)
|
||||
login(page, app_url, next="/") # phase 79: chat is require_user-gated
|
||||
@@ -181,21 +204,35 @@ def test_deflected_question_shows_zero_citation_chips(
|
||||
row = page.locator(".msg.brain .related-docs")
|
||||
expect(row).to_have_count(1, timeout=30_000)
|
||||
links = page.locator(".msg.brain .related-docs .related-doc")
|
||||
expect(links).to_have_count(2) # the weak hits, rank order, capped at 2
|
||||
expect(links.nth(0)).to_contain_text("new-service.md")
|
||||
expect(links.nth(1)).to_contain_text("ssh_aliases.txt")
|
||||
# The rank-6+ remainder of the weak hits, rank order, capped at 2 —
|
||||
# the top five weak hits are suggested (no floor) for the durable
|
||||
# record, but a deflected turn cites nothing, so they never render
|
||||
# as chips.
|
||||
expect(links).to_have_count(2)
|
||||
expect(links.nth(0)).to_contain_text("backups.md")
|
||||
expect(links.nth(1)).to_contain_text("kubernetes.md")
|
||||
|
||||
# ZERO citation chips under the bubble — the weak hits are scored
|
||||
# docs, not citations (the phase-112/113 contract on the wire).
|
||||
expect(page.locator(".msg.brain .source-chip")).to_have_count(0)
|
||||
expect(page.locator(".msg.brain .related-docs .source-chip")).to_have_count(0)
|
||||
|
||||
# Durable record: deflected, weak top score — the retrieval stays
|
||||
# logged for threshold tuning (LOCKED A3: observability unchanged).
|
||||
# Durable record: deflected, weak top score — the FULL retrieval
|
||||
# (the weak hits ARE suggested — no floor — plus the rank-6+ related
|
||||
# remainder) stays logged for threshold tuning (LOCKED A3:
|
||||
# observability unchanged).
|
||||
with SessionLocal() as db:
|
||||
row_log = db.scalars(select(QueryLog)).one()
|
||||
assert row_log.question == OFF_TOPIC
|
||||
assert row_log.deflected is True
|
||||
assert 0.0 < row_log.top_score < get_settings().relevance_threshold
|
||||
assert row_log.fts_hits == 0
|
||||
assert row_log.sources # the weak-hit paths, for threshold tuning
|
||||
assert row_log.sources == (
|
||||
"docs/deployments/new-service.md, "
|
||||
"docs/homelab/ssh/ssh_aliases.txt, "
|
||||
"docs/homelab/quadlet/lan.network, "
|
||||
"docs/homelab/scripts/uptime_probe.py, "
|
||||
"docs/homelab/quadlet/compose.container, "
|
||||
"docs/homelab/backups.md, "
|
||||
"docs/homelab/kubernetes.md"
|
||||
), row_log.sources
|
||||
|
||||
@@ -16,9 +16,10 @@ reused here as closely as possible):
|
||||
``documents.summary``. The sentinel ``RESE-SUMMARY-SENTINEL-7f3a`` sits
|
||||
on the document's LAST line — **outside** the 24-token digest — so it
|
||||
is a marker for "the original, not the summary".
|
||||
* ``notes/qwen-llamacpp-notes.md`` — a markdown control doc, never
|
||||
summarized (phase 30 scope): the viewer must render it exactly as
|
||||
before, with no Summary panel.
|
||||
* ``notes/qwen-llamacpp-notes.md`` — a markdown doc that phase 118 (A2)
|
||||
summarizes TOO (the phase-30 non-markdown-only scope is retired): the
|
||||
viewer must show its Summary panel exactly like the yaml's, above the
|
||||
rendered markdown.
|
||||
|
||||
This phase only surfaces the stored field: the content endpoint returns
|
||||
``summary`` (task 01) and the shared ``renderDocument`` core draws the
|
||||
@@ -27,8 +28,8 @@ labeled ``.doc-summary`` panel above the content on BOTH surfaces (task
|
||||
exactly that contract. Phase 79 supersedes the phase-16 soft rule: the
|
||||
content endpoint is ``require_user``-gated, so the viewer surfaces and
|
||||
the API shape pin run under a signed-in session (the shape itself —
|
||||
``summary`` for the yaml, ``null`` for the markdown control — is
|
||||
unchanged).
|
||||
``summary`` for BOTH fixture docs, phase 118 A2 — is what the pins
|
||||
carry).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -98,13 +99,14 @@ def _run_in_thread(coro: Any) -> Any:
|
||||
|
||||
def _reset_db_and_import(mock_llm: int) -> ImportSummary:
|
||||
"""Truncate the KB (and query log + steering) and re-import the
|
||||
summary_kb fixtures — yaml summarized, md control not."""
|
||||
summary_kb fixtures — both docs summarized (phase 118, A2: the
|
||||
markdown doc too)."""
|
||||
with SessionLocal() as db:
|
||||
db.execute(text("TRUNCATE chunks, documents, query_log, steering_notes"))
|
||||
db.commit()
|
||||
summary = _run_in_thread(_import_fixtures(mock_llm))
|
||||
assert summary.added == 2 # yaml + md control
|
||||
assert summary.summaries == 1 and summary.summary_errors == 0
|
||||
assert summary.added == 2 # yaml + md
|
||||
assert summary.summaries == 2 and summary.summary_errors == 0 # A2: md too
|
||||
assert summary.errors == 0
|
||||
return summary
|
||||
|
||||
@@ -279,19 +281,30 @@ def test_modal_shows_panel_and_full_page_agrees(
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_markdown_doc_has_no_summary_panel(
|
||||
def test_markdown_doc_has_a_summary_panel(
|
||||
page: Page, app_url: str, mock_llm: int, db_ready: None
|
||||
) -> None:
|
||||
"""Control doc (markdown — never summarized): NO .doc-summary
|
||||
element on either surface, and the content renders exactly as
|
||||
before (first child of the content container is the doc body)."""
|
||||
"""Phase 118 (A2): the markdown doc is summarized TOO — the SAME
|
||||
labeled ``.doc-summary`` panel (the deterministic mock digest +
|
||||
pointer line) renders above the rendered markdown on BOTH surfaces
|
||||
(the phase-30 "markdown never summarized" control is retired with
|
||||
the non-markdown-only scope)."""
|
||||
_reset_db_and_import(mock_llm)
|
||||
page.set_default_timeout(30_000)
|
||||
login(page, app_url, next="/") # phase 79: the viewer content is gated
|
||||
|
||||
# Full page: no panel, markdown column untouched.
|
||||
digest_line, pointer_line = _summary_lines(SOURCE, MD_PATH)
|
||||
|
||||
# Full page: the labeled panel above the markdown column (the same
|
||||
# contract as the yaml — one shared renderer).
|
||||
page.goto(f"{app_url}/document.html?source={SOURCE}&path={MD_URL_PATH}")
|
||||
expect(page.locator(".doc-summary")).to_have_count(0)
|
||||
panel = page.locator(".doc-summary")
|
||||
expect(panel).to_have_count(1)
|
||||
expect(panel).to_be_visible()
|
||||
expect(panel).to_have_attribute("aria-label", "Summary")
|
||||
expect(panel.locator(".doc-summary-title")).to_have_text("Summary")
|
||||
expect(panel).to_contain_text(digest_line)
|
||||
expect(panel).to_contain_text(pointer_line)
|
||||
expect(page.locator("#doc-content .doc-md")).to_have_count(1)
|
||||
expect(page.locator("#doc-content")).to_contain_text(
|
||||
"came out of a week of"
|
||||
@@ -300,10 +313,10 @@ def test_markdown_doc_has_no_summary_panel(
|
||||
"() => [...document.querySelector('#doc-content').children]"
|
||||
".map((el) => el.className)"
|
||||
)
|
||||
assert order == ["doc-md"], f"markdown doc gained children: {order}"
|
||||
assert order == ["doc-summary", "doc-md"], f"panel not first: {order}"
|
||||
|
||||
# Modal: same story — no panel, .doc-md is the sole content child.
|
||||
# (the session is already signed in — the form login above)
|
||||
# Modal: same story — the panel above .doc-md (the session is
|
||||
# already signed in — the form login above).
|
||||
page.goto(f"{app_url}/sources.html")
|
||||
# Phase 97: the re-mount lands on the tree's top level — drill to
|
||||
# the notes level where the row lives (the drill is the only
|
||||
@@ -314,7 +327,11 @@ def test_markdown_doc_has_no_summary_panel(
|
||||
expect(row).to_have_count(1)
|
||||
row.locator("td:nth-child(2) a.doc-link").click()
|
||||
expect(page.locator(".doc-modal")).to_be_visible()
|
||||
expect(page.locator("#doc-modal .doc-summary")).to_have_count(0)
|
||||
modal_panel = page.locator("#doc-modal .doc-summary")
|
||||
expect(modal_panel).to_have_count(1)
|
||||
expect(modal_panel).to_be_visible()
|
||||
expect(modal_panel).to_contain_text(digest_line)
|
||||
expect(modal_panel).to_contain_text(pointer_line)
|
||||
expect(page.locator("#doc-modal .doc-md")).to_have_count(1)
|
||||
expect(page.locator("#doc-modal-content")).to_contain_text(
|
||||
"came out of a week of"
|
||||
@@ -323,7 +340,7 @@ def test_markdown_doc_has_no_summary_panel(
|
||||
"() => [...document.querySelector('#doc-modal-content').children]"
|
||||
".map((el) => el.className)"
|
||||
)
|
||||
assert order == ["doc-md"], f"markdown modal gained children: {order}"
|
||||
assert order == ["doc-summary", "doc-md"], f"markdown modal order: {order}"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -336,10 +353,9 @@ def test_content_api_summary_shape(
|
||||
page: Page, app_url: str, mock_llm: int, db_ready: None
|
||||
) -> None:
|
||||
"""``GET /api/documents/content`` carries ``summary`` — the string
|
||||
for the summarized yaml, ``null`` for the markdown control. Phase 79
|
||||
superseded the phase-16 soft rule: the endpoint is require_user, so
|
||||
the pin runs under the form-login session (the shape itself is
|
||||
unchanged)."""
|
||||
for BOTH fixture docs (phase 118, A2: the markdown doc too). Phase
|
||||
79 superseded the phase-16 soft rule: the endpoint is require_user,
|
||||
so the pin runs under the form-login session."""
|
||||
_reset_db_and_import(mock_llm)
|
||||
login(page, app_url, next="/")
|
||||
|
||||
@@ -364,5 +380,10 @@ def test_content_api_summary_shape(
|
||||
assert resp_md.status == 200
|
||||
md_body = resp_md.json()
|
||||
assert md_body["format"] == "md"
|
||||
assert md_body["summary"] is None, "markdown docs never carry a summary"
|
||||
# Phase 118 (A2): the markdown doc carries its summary too — the
|
||||
# same byte-stable digest + pointer line.
|
||||
md_digest_line, _ = _summary_lines(SOURCE, MD_PATH)
|
||||
assert md_body["summary"] == (
|
||||
f"{md_digest_line}\nSource: {SOURCE}/{MD_PATH}"
|
||||
)
|
||||
assert "came out of a week of" in md_body["content"]
|
||||
|
||||
@@ -0,0 +1,671 @@
|
||||
"""Phase 118 task 06 E2E (Playwright, mock-only): the summary-seed
|
||||
context — grounded turns seed SUMMARIES, full text arrives only via
|
||||
``read`` (the owner directive, TODO.md L3; the retired phase-24
|
||||
full-text seeding is gone — locked A1/A6).
|
||||
|
||||
Run in isolation (DB must be up: ``podman compose up -d db``):
|
||||
|
||||
uv run pytest tests/e2e/test_summary_seed_context.py -v --no-cov
|
||||
|
||||
MOCK-ONLY suite: ``E2E_REAL_LLM=1`` is not supported — the story's gate
|
||||
is the deterministic contract, not the real model's behavior: the mock
|
||||
LLM's tail-echo trigger (``END_OF_NOTES_TRIGGER``) quotes the last 160
|
||||
chars of the seeded ``<documents>`` block, and the scripted
|
||||
summary-read flow (``SUMMARY_SEED_READ_TRIGGER``) echoes the whole read
|
||||
tool result — both are the house scripted-turn lenses on the LLM's
|
||||
context (the mock is the only E2E lens), so what reached the prompt is
|
||||
assertable byte-exactly.
|
||||
|
||||
KB fixture — one host temp dir (``tmp_path_factory``; the app runs on
|
||||
the same host) registered as a local-directory source (the
|
||||
``test_ls_tree_drilldown.py`` / ``test_read_truncation_cap.py``
|
||||
registration + real-Sync pattern — registration through the
|
||||
authenticated API, the real in-process ``POST /api/sync`` pipeline; no
|
||||
git anywhere), with SEVEN markdown documents (every file markdown —
|
||||
the locked A2 end-to-end proof: the mock ``SUMMARY_MODE`` digest must
|
||||
land on every doc, markdown included) whose bodies are
|
||||
token-controlled so the hybrid gate picks the intended suggested set
|
||||
deterministically (``synced_kb`` pins it with the app's REAL
|
||||
retrieval — a fixture-text regression that breaks the design fails at
|
||||
setup with a clear message, not at the wire assertions):
|
||||
|
||||
* each doc = a neutral 24-token prelude (no question tokens — the
|
||||
mock's first-24-token digest therefore shares NOTHING with either
|
||||
question, so a doc's embedded summary chunk carries no strength
|
||||
tokens and never outranks that doc's own content chunks) +
|
||||
``seed vault rotation notes`` repeated *i* times (the strength
|
||||
gradient) + six unique filler tokens + a unique tail SENTINEL on
|
||||
the document's LAST line (outside the digest — the full-text
|
||||
marker);
|
||||
* the measured rank order (probe, stable across re-imports) is
|
||||
``doc-a > doc-d > doc-b > doc-c > doc-e > doc-f > doc-g`` for the
|
||||
tail question (×1 strength gradient — md5-collision-reordered) and a
|
||||
clean ``doc-a > doc-b > doc-c > doc-d > doc-e > doc-f > doc-g`` for
|
||||
the read question (×3 gradient — the collision noise cannot cross
|
||||
the wider gaps); the suggested tier (LOCKED A3: top-5, NO floor) is
|
||||
the first five of each, the related tier (rank 6+, ``related_max_docs``
|
||||
= 2) is ``doc-f``/``doc-g`` in both.
|
||||
|
||||
Test → contract mapping (the task's cases a–e; the story is the owner
|
||||
TODO item — one Playwright file per story, A16):
|
||||
1. ``test_summaries_seed_the_prompt_not_the_full_text`` — cases (a) +
|
||||
(b) + (d): the tail-echo question's answer quotes the last 160
|
||||
chars of the seeded ``<documents>`` block, which end in the LAST
|
||||
suggested document's SUMMARY — the mock's byte-stable
|
||||
``SUMMARY_MODE`` digest tail + pointer line (case a: summaries
|
||||
reached the prompt; case d: the markdown doc's block carries the
|
||||
digest, not a content preview — the phase-30 digest shape, and the
|
||||
``synced_kb`` pin proves every doc's stored summary IS the digest).
|
||||
The inverse of the retired phase-24 pin: NO document's tail
|
||||
sentinel is in the echoed context — the full content of no
|
||||
suggested doc ever reached the model (case b). The grounded turn
|
||||
cites the 5 suggested docs (case e's chip surface, no read yet) and
|
||||
renders the de-emphasized related row (rank 6+); the durable
|
||||
record carries suggested + related.
|
||||
2. ``test_read_suggested_doc_adds_full_text_and_cites`` — cases (c) +
|
||||
(e): the scripted flow ``read``s the rank-1 suggested doc — the
|
||||
read SUCCEEDS (phase 118: the seeds are summaries, not full text —
|
||||
the retired top-2 seed-read refusal is gone) and the mock's
|
||||
verbatim echo of the read result lands the tail sentinel in the
|
||||
answer (case c: the full text now arrives through the ``read``
|
||||
tool, not the seed); the citation chips = suggested + read,
|
||||
deduped — the read doc is among the suggested, so the chip row is
|
||||
exactly the 5 suggested docs, and the related row renders rank 6+
|
||||
(case e); the durable record carries suggested + related + read.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
from collections.abc import Iterator
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
from playwright.sync_api import Locator, Page, expect
|
||||
from sqlalchemy import select, text
|
||||
|
||||
from app.config import Settings as _Settings
|
||||
from app.db import SessionLocal
|
||||
from app.models import Document, QueryLog
|
||||
from app.rag.retriever import TRUNCATION_MARKER, retrieve, select_related, select_suggested
|
||||
from e2e.auth_helpers import login
|
||||
from e2e.conftest import (
|
||||
ADMIN_PASSWORD,
|
||||
SESSION_SECRET,
|
||||
USE_REAL_LLM,
|
||||
_wait_http,
|
||||
)
|
||||
from e2e.mock_llm import TOKEN_RE, embed_text
|
||||
|
||||
REPO = Path(__file__).resolve().parents[2]
|
||||
|
||||
# Phase 79 (task 04, full inventory): the conftest session app owns its
|
||||
# port in a combined run — this module app binds its own port instead
|
||||
# (a same-port second uvicorn dies on bind and would drive the wrong
|
||||
# server). Env-overridable.
|
||||
APP_PORT = int(os.environ.get("E2E_APP_PORT_SEEDCTX", "8138"))
|
||||
APP_URL = f"http://127.0.0.1:{APP_PORT}"
|
||||
|
||||
SOURCE = "seedkb" # the local directory's basename = the source name
|
||||
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# Fixture documents (deterministic, token-controlled — see the module
|
||||
# docstring for the design and the measured rank order)
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
#: 22 neutral tokens — together with the 2-token title they ARE the
|
||||
#: mock's first-24-token digest, so no summary chunk shares a token
|
||||
#: with either question (the summary chunks rank below every content
|
||||
#: chunk; the digest is the byte-stable assertion surface).
|
||||
PRELUDE = (
|
||||
"fixture preamble block one two three four five six seven eight nine ten "
|
||||
"eleven twelve thirteen fourteen fifteen sixteen seventeen eighteen nineteen"
|
||||
)
|
||||
assert len(TOKEN_RE.findall(PRELUDE.lower())) == 22
|
||||
|
||||
#: The strength gradient (each token a question token — ``seed``,
|
||||
#: ``vault``, ``rotation``, ``notes``) — repeated *i* times in doc *i*.
|
||||
PHRASE = "seed vault rotation notes"
|
||||
|
||||
#: (path, title, phrase repeats, filler token, tail sentinel) — the
|
||||
#: MEASURED rank order (probe, stable across re-imports) is
|
||||
#: doc-a > doc-d > doc-b > doc-c > doc-e > doc-f > doc-g: the suggested
|
||||
#: tier is the first five, the related tier the last two.
|
||||
DOCS: list[tuple[str, str, int, str, str]] = [
|
||||
("doc-a.md", "Zeta Alfa", 7, "fillera", "SEEDA-TAIL-7f3a"),
|
||||
("doc-b.md", "Zeta Bravo", 6, "fillerb", "SEEDB-TAIL-8c4d"),
|
||||
("doc-c.md", "Zeta Charlie", 5, "fillerc", "SEEDC-TAIL-9d5e"),
|
||||
("doc-d.md", "Zeta Delta", 4, "fillerd", "SEEDD-TAIL-0e6f"),
|
||||
("doc-e.md", "Zeta Echo", 3, "fillere", "SEEDE-TAIL-1f7a"),
|
||||
("doc-f.md", "Zeta Foxtrot", 2, "fillerf", "SEEDF-TAIL-2a8b"),
|
||||
("doc-g.md", "Zeta Golf", 1, "fillerg", "SEEDG-TAIL-3b9c"),
|
||||
]
|
||||
#: The suggested tier per scripted question (the MEASURED rank order —
|
||||
#: probe-verified, stable across re-imports): the tail question's ×1
|
||||
#: strength gradient is md5-collision-reordered (d before b), the read
|
||||
#: question's ×3 gradient is clean. Related is the rank-6+ remainder in
|
||||
#: both (the gradient never puts a strength doc below the filler docs).
|
||||
SUGGESTED_TAIL = ["doc-a.md", "doc-d.md", "doc-b.md", "doc-c.md", "doc-e.md"]
|
||||
SUGGESTED_READ = ["doc-a.md", "doc-b.md", "doc-c.md", "doc-d.md", "doc-e.md"]
|
||||
RELATED_PATHS = ["doc-f.md", "doc-g.md"]
|
||||
LAST_SUGGESTED = "doc-e.md" # the tail-echo target (case a/d)
|
||||
READ_TARGET = "doc-a.md" # rank 1 — a suggested doc (case c)
|
||||
SENTINELS = [sentinel for _p, _t, _i, _f, sentinel in DOCS]
|
||||
READ_SENTINEL = next(s for p, _t, _i, _f, s in DOCS if p == READ_TARGET)
|
||||
|
||||
|
||||
def _doc_content(title: str, i: int, filler: str, sentinel: str) -> str:
|
||||
return (
|
||||
f"# {title}\n"
|
||||
f"\n{PRELUDE}\n"
|
||||
f"\n{' '.join([PHRASE] * i)}\n"
|
||||
f"\n{' '.join([filler] * 6)}\n"
|
||||
f"\n{sentinel}\n"
|
||||
)
|
||||
|
||||
|
||||
def _expected_summary(content: str, path: str) -> str:
|
||||
"""The mock lite model's byte-stable digest + the code pointer line
|
||||
(mirrors ``mock_llm.compose_answer``'s ``SUMMARY_MODE`` branch —
|
||||
first 24 tokens of the document content — plus the summarizer's
|
||||
deterministic ``Source:`` line; no model output is ever trusted)."""
|
||||
digest = " ".join(TOKEN_RE.findall(content.lower())[:24])
|
||||
return f"This document covers {digest}.\nSource: {SOURCE}/{path}"
|
||||
|
||||
|
||||
#: The turn's questions (the mock's trigger phrases — see the module
|
||||
#: docstring). The tail question carries ``END_OF_NOTES_TRIGGER``; its
|
||||
#: strength tokens ground the turn (best cosine ≈ 0.55 ≥ the E2E 0.30
|
||||
#: threshold). The read question carries ``SUMMARY_SEED_READ_TRIGGER``
|
||||
#: with its own scripted call after the colon (the house scripted-turn
|
||||
#: convention — ``_READ_CAP_CALL_RE`` / ``_DRILL_CALL_RE`` shape),
|
||||
#: ``seed vault rotation notes`` × 3 for the same grounding.
|
||||
TAIL_QUESTION = "Show the end of your notes about the seed vault rotation, please."
|
||||
READ_QUESTION = (
|
||||
f"Read the suggested document: read {SOURCE}/{READ_TARGET} — "
|
||||
f"{PHRASE} {PHRASE} {PHRASE}"
|
||||
)
|
||||
assert "show the end of your notes" in TAIL_QUESTION.lower()
|
||||
assert "read the suggested document" in READ_QUESTION.lower()
|
||||
assert re.fullmatch(r"[a-z0-9_./-]+", f"{SOURCE}/{READ_TARGET}")
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# Fixtures
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def seed_dirs(tmp_path_factory: pytest.TempPathFactory) -> Path:
|
||||
"""The story's local-directory source: one host temp dir (the app
|
||||
server runs on the same host, so the path is visible to it) holding
|
||||
the seven token-controlled markdown documents. The directory's
|
||||
basename is the source name (``kind=local``, phase 38)."""
|
||||
root = tmp_path_factory.mktemp("bor_seed_ctx") / SOURCE
|
||||
root.mkdir()
|
||||
for path, title, i, filler, sentinel in DOCS:
|
||||
(root / path).write_text(
|
||||
_doc_content(title, i, filler, sentinel), encoding="utf-8"
|
||||
)
|
||||
assert (root / READ_TARGET).is_file()
|
||||
return root
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def app_server(mock_llm: int, seed_dirs: Path) -> Iterator[str]:
|
||||
"""The real app under test — per-module app (the conftest pattern,
|
||||
cf. ``test_read_truncation_cap.py``): NO ``BOR_GIT_SOURCES`` (the
|
||||
env fallback is git-only — the source here is a DB-registered local
|
||||
directory), the mock LLM, the mock-calibrated threshold, and the
|
||||
leak-guarded code defaults (the suggested/related tier settings
|
||||
ride their code defaults — 5 / 2 — exactly like the production
|
||||
``.env``-free defaults). The session app is never started in this
|
||||
isolated run, so no port clash."""
|
||||
env = dict(os.environ)
|
||||
env.pop("DEBUGPY", None)
|
||||
env["BOR_ENVIRONMENT"] = "e2e"
|
||||
env["BOR_STATIC_DIR"] = str(REPO / "frontend")
|
||||
env["BOR_LLM_BASE_URL"] = (
|
||||
"https://aipi.reeseapps.com/v1"
|
||||
if USE_REAL_LLM
|
||||
else f"http://127.0.0.1:{mock_llm}/v1"
|
||||
)
|
||||
# Mock-calibrated gate (conftest pattern): the strength gradient
|
||||
# keeps every scripted turn grounded (best cosine ≈ 0.55 / ≈ 0.80).
|
||||
env["BOR_RELEVANCE_THRESHOLD"] = "0.30"
|
||||
env["BOR_LEXICAL_SUPPORT_FLOOR"] = "0.15"
|
||||
# Phase 67: instant retry waits + the code-default budget (the
|
||||
# conftest leak-guard pattern).
|
||||
env["BOR_LLM_RETRY_DELAY"] = "0"
|
||||
env["BOR_LLM_RETRIES"] = str(_Settings.model_fields["llm_retries"].default)
|
||||
env.setdefault(
|
||||
"BOR_DATABASE_URL",
|
||||
"postgresql+psycopg://reese:reese@localhost:5432/brain_of_reese",
|
||||
)
|
||||
# Phase 16: admin auth must be set or create_app() refuses to boot.
|
||||
env["BOR_ADMIN_PASSWORD"] = ADMIN_PASSWORD
|
||||
env["BOR_SESSION_SECRET"] = SESSION_SECRET
|
||||
# The repo's .env file carries the owner's BOR_GIT_SOURCES (the app
|
||||
# reads it from cwd) — override it with an EMPTY value (the env var
|
||||
# beats the .env file): the registry must hold EXACTLY the local
|
||||
# directory this suite registers (a leftover env git list would
|
||||
# pollute the KB the scripted turns run against).
|
||||
env["BOR_GIT_SOURCES"] = ""
|
||||
# Leak guards (conftest pattern): an operator's local (gitignored)
|
||||
# .env cannot leak corpus-specific settings into the app under test.
|
||||
env["BOR_DOCS_REPO"] = ""
|
||||
env["BOR_SUGGESTIONS"] = json.dumps(
|
||||
_Settings.model_fields["suggestions"].default
|
||||
)
|
||||
env["BOR_INPUT_PLACEHOLDER"] = _Settings.model_fields["input_placeholder"].default
|
||||
env["BOR_FOOTER_TEXT"] = _Settings.model_fields["footer_text"].default
|
||||
proc = subprocess.Popen(
|
||||
[sys.executable, "-m", "uvicorn", "app.main:app",
|
||||
"--host", "127.0.0.1", "--port", str(APP_PORT), "--log-level", "warning"],
|
||||
cwd=REPO,
|
||||
env=env,
|
||||
)
|
||||
try:
|
||||
_wait_http(f"{APP_URL}/api/health")
|
||||
yield APP_URL
|
||||
finally:
|
||||
proc.terminate()
|
||||
try:
|
||||
proc.wait(timeout=10)
|
||||
except subprocess.TimeoutExpired:
|
||||
proc.kill()
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def app_url(app_server: str) -> str:
|
||||
return app_server
|
||||
|
||||
|
||||
def _truncate_all() -> None:
|
||||
"""Fresh registry + KB (the E2E isolation pattern): the E2E suites
|
||||
share one Postgres, so a leftover git_sources row or document would
|
||||
pollute the retrieval the scripted turns run against (the strength
|
||||
gradient's margins are pinned against EXACTLY these seven
|
||||
documents)."""
|
||||
with SessionLocal() as db:
|
||||
db.execute(
|
||||
text(
|
||||
"TRUNCATE chunks, documents, query_log, steering_notes, "
|
||||
"kb_overview, git_sources, folder_summaries"
|
||||
)
|
||||
)
|
||||
db.commit()
|
||||
|
||||
|
||||
def _wait_sync_done_http(client: httpx.Client, timeout_s: float = 180.0) -> dict[str, Any]:
|
||||
"""Poll the (cookie-authenticated) status endpoint until the run
|
||||
reaches a terminal state (the test_ls_tree_drilldown pattern, over
|
||||
plain httpx — this fixture has no browser page yet)."""
|
||||
deadline = time.monotonic() + timeout_s
|
||||
body: dict[str, Any] = {}
|
||||
while time.monotonic() < deadline:
|
||||
r = client.get("/api/sync/status")
|
||||
assert r.status_code == 200, r.text
|
||||
body = r.json()
|
||||
if body["state"] in ("success", "failed"):
|
||||
return body
|
||||
time.sleep(0.5)
|
||||
raise AssertionError(f"sync did not reach a terminal state: {body}")
|
||||
|
||||
|
||||
def _assert_tiers(question: str, suggested_paths: list[str]) -> None:
|
||||
"""Pin the strength-gradient design with the app's REAL hybrid
|
||||
retrieval over the mock's embeddings (deterministic): the suggested
|
||||
tier is exactly the five strength docs in the MEASURED rank order
|
||||
for THIS question (LOCKED A3 — top-5, NO floor) and the related
|
||||
tier is the rank-6+ remainder (``related_max_docs`` = 2). A
|
||||
fixture-text regression that breaks the design fails here, at
|
||||
setup, with a clear message."""
|
||||
with SessionLocal() as db:
|
||||
chunks = retrieve(db, question, embed_text(question))
|
||||
suggested = [f"{d.source}/{d.path}" for d in select_suggested(chunks)]
|
||||
related = [
|
||||
f"{d.source}/{d.path}"
|
||||
for d in select_related(
|
||||
chunks,
|
||||
{d.id for d in select_suggested(chunks)},
|
||||
_Settings.model_fields["related_max_docs"].default,
|
||||
)
|
||||
]
|
||||
expected = [f"{SOURCE}/{p}" for p in suggested_paths]
|
||||
assert suggested == expected, (
|
||||
f"suggested tier drifted for {question!r}: {suggested} "
|
||||
f"(expected {expected})"
|
||||
)
|
||||
assert related == [f"{SOURCE}/{p}" for p in RELATED_PATHS], (
|
||||
f"related tier drifted for {question!r}: {related}"
|
||||
)
|
||||
|
||||
|
||||
@pytest.fixture(scope="module")
|
||||
def synced_kb(app_server: str, seed_dirs: Path) -> None:
|
||||
"""The story's precondition: the one-source KB synced under the
|
||||
deterministic mock. Registers the temp directory through the
|
||||
authenticated API (the ``test_local_directory_sources.py``
|
||||
pattern), runs the REAL in-process sync (``POST /api/sync`` — walk
|
||||
→ chunk → embed → summaries → overview → folder summaries →
|
||||
version bump), pins the stored content byte-identical to the
|
||||
fixture strings, pins LOCKED A2 end-to-end (every doc — markdown
|
||||
included — stores the mock's byte-stable digest + exactly one
|
||||
embedded ``is_summary`` chunk), and pins the tier design for both
|
||||
scripted questions."""
|
||||
_truncate_all()
|
||||
with httpx.Client(base_url=app_server, timeout=30.0) as client:
|
||||
r = client.post("/api/login", json={"password": ADMIN_PASSWORD})
|
||||
assert r.status_code == 204, r.text
|
||||
r = client.post(
|
||||
"/api/git-sources", json={"kind": "local", "path": str(seed_dirs)}
|
||||
)
|
||||
assert r.status_code == 201, r.text
|
||||
r = client.post("/api/sync")
|
||||
assert r.status_code == 202, r.text
|
||||
body = _wait_sync_done_http(client)
|
||||
assert body["state"] == "success", body
|
||||
detail = body["detail"]
|
||||
assert detail["added"] == len(DOCS), detail
|
||||
assert detail["pruned"] == 0, detail
|
||||
# The import stored the fixture strings BYTE-IDENTICALLY and, for
|
||||
# EVERY doc (markdown included — locked A2), the mock's byte-stable
|
||||
# digest: the deterministic assertion surface of the whole suite.
|
||||
with SessionLocal() as db:
|
||||
for path, title, i, filler, sentinel in DOCS:
|
||||
stored = db.scalar(
|
||||
select(Document).where(
|
||||
Document.source == SOURCE, Document.path == path
|
||||
)
|
||||
)
|
||||
assert stored is not None, f"fixture doc {path} was not imported"
|
||||
expected = _doc_content(title, i, filler, sentinel)
|
||||
assert stored.content == expected, f"stored content drifted for {path}"
|
||||
assert stored.summary == _expected_summary(expected, path), (
|
||||
f"summary is not the mock digest for {path}: {stored.summary!r}"
|
||||
)
|
||||
schunks = [c for c in stored.chunks if c.is_summary]
|
||||
assert len(schunks) == 1 and schunks[0].position == -1, (
|
||||
f"expected exactly one is_summary chunk for {path}"
|
||||
)
|
||||
assert schunks[0].embedding is not None, (
|
||||
f"the is_summary chunk of {path} is not embedded"
|
||||
)
|
||||
_assert_tiers(TAIL_QUESTION, SUGGESTED_TAIL)
|
||||
_assert_tiers(READ_QUESTION, SUGGESTED_READ)
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _clean(db_ready: None) -> Iterator[None]:
|
||||
"""Per-test query_log isolation (the KB itself is module-scoped —
|
||||
the scripted turns never change it, so the registry and the KB
|
||||
persist across the tests of this module)."""
|
||||
with SessionLocal() as db:
|
||||
db.execute(text("TRUNCATE query_log"))
|
||||
db.commit()
|
||||
yield
|
||||
with SessionLocal() as db:
|
||||
db.execute(text("TRUNCATE query_log"))
|
||||
db.commit()
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# Page helpers (the test_ls_tree_drilldown / test_read_truncation_cap
|
||||
# house patterns)
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
#: Captures the raw SSE ``data:`` payloads of the /api/chat stream
|
||||
#: (a response clone read in the background) — wire-level assertions
|
||||
#: for the ``tool`` / ``done`` frames, independent of the UI rendering.
|
||||
SSE_HOOK = """
|
||||
() => {
|
||||
if (window.__sseInstalled) return;
|
||||
window.__sseInstalled = true;
|
||||
window.__sseFrames = [];
|
||||
const origFetch = window.fetch;
|
||||
window.fetch = async function (...args) {
|
||||
const res = await origFetch.apply(this, args);
|
||||
try {
|
||||
const url = typeof args[0] === 'string' ? args[0] : args[0].url;
|
||||
if (url.includes('/api/chat')) {
|
||||
res.clone().text().then((bodyText) => {
|
||||
for (const block of bodyText.split('\\n\\n')) {
|
||||
const line = block.trim();
|
||||
if (line.startsWith('data: ')) {
|
||||
window.__sseFrames.push(line.slice(6));
|
||||
}
|
||||
}
|
||||
});
|
||||
}
|
||||
} catch (e) { /* non-clonable responses: ignored */ }
|
||||
return res;
|
||||
};
|
||||
}
|
||||
"""
|
||||
|
||||
|
||||
def _install_page_hooks(page: Page) -> None:
|
||||
page.evaluate(SSE_HOOK)
|
||||
|
||||
|
||||
def _frames(page: Page) -> list[dict]:
|
||||
"""The SSE frames captured since the last submit (``_submit``
|
||||
clears the buffer), once the hook's background read settles."""
|
||||
deadline = time.monotonic() + 30.0
|
||||
while True:
|
||||
raw = page.evaluate("() => window.__sseFrames || []")
|
||||
parsed = [json.loads(line) for line in raw if line]
|
||||
if any(f.get("type") == "done" for f in parsed):
|
||||
return parsed
|
||||
if time.monotonic() > deadline:
|
||||
raise AssertionError(
|
||||
f"SSE hook captured no `done` frame (frames so far: "
|
||||
f"{len(parsed)}) — hook install failed?"
|
||||
)
|
||||
time.sleep(0.05)
|
||||
|
||||
|
||||
def _tool_frames(frames: list[dict]) -> list[dict]:
|
||||
return [f for f in frames if f.get("type") == "tool"]
|
||||
|
||||
|
||||
def _submit(page: Page, question: str) -> None:
|
||||
page.evaluate("window.__sseFrames = []")
|
||||
page.fill("#message-input", question)
|
||||
page.click("#send-btn")
|
||||
# The user bubble lands synchronously with the submit handler.
|
||||
expect(page.locator(".msg.user .bubble").last).to_contain_text(question)
|
||||
|
||||
|
||||
def _wait_settled(page: Page) -> None:
|
||||
"""The turn is complete: answer text in the bubble, button recovered
|
||||
(the phase-48 settle wait, the test_agent_document_tools helper)."""
|
||||
expect(page.locator(".msg.brain .bubble").last).not_to_have_text("", timeout=30_000)
|
||||
expect(page.locator("#send-btn")).to_be_enabled(timeout=30_000)
|
||||
expect(page.locator("#send-label")).to_have_text("Send", timeout=30_000)
|
||||
|
||||
|
||||
def _last_brain(page: Page) -> Locator:
|
||||
return page.locator(".msg.brain").last
|
||||
|
||||
|
||||
def _last_query_log() -> QueryLog:
|
||||
with SessionLocal() as db:
|
||||
rows = db.scalars(select(QueryLog)).all()
|
||||
assert len(rows) == 1, f"expected exactly one query_log row, got {len(rows)}"
|
||||
return rows[0]
|
||||
|
||||
|
||||
def _assert_suggested_chips_and_related_row(page: Page, suggested_paths: list[str]) -> None:
|
||||
"""The citation surface of a grounded turn (LOCKED A4): the chip row
|
||||
is the suggested set (+ agent reads, deduped — asserted per test)
|
||||
and the de-emphasized ``related-docs`` row carries the rank-6+
|
||||
remainder (phase-113 UI, never a citation chip)."""
|
||||
chips = page.locator(".msg.brain .source-chip")
|
||||
expect(chips).to_have_count(len(suggested_paths))
|
||||
for path in suggested_paths:
|
||||
expect(
|
||||
chips.filter(has_text=path), message=f"chip for {path}"
|
||||
).to_have_count(1)
|
||||
row = page.locator(".msg.brain .related-docs")
|
||||
expect(row).to_have_count(1)
|
||||
expect(row.first).to_have_attribute("aria-label", "Nearby docs, in case")
|
||||
links = page.locator(".msg.brain .related-docs .related-doc")
|
||||
expect(links).to_have_count(len(RELATED_PATHS))
|
||||
for i, path in enumerate(RELATED_PATHS):
|
||||
expect(links.nth(i)).to_contain_text(f"{SOURCE}/{path}")
|
||||
# The related links are never citation chips.
|
||||
expect(page.locator(".msg.brain .related-docs .source-chip")).to_have_count(0)
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# 1. Cases (a) + (b) + (d): summaries seed the prompt — no full text
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_summaries_seed_the_prompt_not_the_full_text(
|
||||
page: Page, app_url: str, synced_kb: None, db_ready: None
|
||||
) -> None:
|
||||
"""The tail-echo question quotes the last 160 chars of the seeded
|
||||
``<documents>`` block: they end in the LAST suggested doc's SUMMARY
|
||||
(the mock's byte-stable digest tail + pointer — cases a + d: the
|
||||
markdown doc's block carries the digest, not a content preview),
|
||||
and NO document's tail sentinel is in the echoed context (case b:
|
||||
the full text of no suggested doc reached the model — the inverse
|
||||
of the retired phase-24 pin)."""
|
||||
page.set_default_timeout(30_000)
|
||||
login(page, app_url, next="/")
|
||||
_install_page_hooks(page)
|
||||
|
||||
_submit(page, TAIL_QUESTION)
|
||||
_wait_settled(page)
|
||||
|
||||
# No tools on this turn — the answer is the mock's direct tail echo
|
||||
# of the seeded context (the summary-seed lens).
|
||||
frames = _frames(page)
|
||||
assert _tool_frames(frames) == [], _tool_frames(frames)
|
||||
done = next(f for f in frames if f.get("type") == "done")
|
||||
assert done["deflected"] is False, done
|
||||
# The grounded turn cites the 5 suggested docs (LOCKED A4 — no read
|
||||
# yet: the chip row IS the suggested set, in rank order).
|
||||
assert [
|
||||
(s["source"], s["path"]) for s in done["sources"]
|
||||
] == [(SOURCE, p) for p in SUGGESTED_TAIL], done["sources"]
|
||||
|
||||
bubble = _last_brain(page).locator(".bubble")
|
||||
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER, timeout=30_000)
|
||||
|
||||
# Case (a) + (d): the echoed tail ends in the LAST suggested doc's
|
||||
# summary — the markdown doc-e's byte-stable digest tail + pointer
|
||||
# line (the digest is the mock's SUMMARY_MODE shape, pinned in
|
||||
# ``synced_kb`` — a content preview would carry the prelude/phrase
|
||||
# text instead, and the pointer line only ever exists on a stored
|
||||
# summary). The bubble renders the answer as markdown, which
|
||||
# collapses the summary's newline — so pin each LINE separately
|
||||
# (the digest line's tail sits inside the echoed 160 chars; the
|
||||
# pointer line is single-line too).
|
||||
last_doc = next(d for d in DOCS if d[0] == LAST_SUGGESTED)
|
||||
last_summary = _expected_summary(_doc_content(*last_doc[1:]), LAST_SUGGESTED)
|
||||
digest_line = last_summary.split("\n", 1)[0]
|
||||
expect(bubble).to_contain_text("Source: seedkb/doc-e.md")
|
||||
expect(bubble).to_contain_text(digest_line[-100:])
|
||||
# The digest is NOT the raw content: the echoed tail cannot carry
|
||||
# any document's tail sentinel — the inverse of the retired
|
||||
# phase-24 full-text pin (case b).
|
||||
for sentinel in SENTINELS:
|
||||
expect(bubble).not_to_contain_text(sentinel)
|
||||
expect(bubble).not_to_contain_text(TRUNCATION_MARKER)
|
||||
|
||||
# Case (e)'s chip surface (no read yet): chips = the 5 suggested
|
||||
# docs; the related row renders rank 6+ (the de-emphasized row).
|
||||
_assert_suggested_chips_and_related_row(page, SUGGESTED_TAIL)
|
||||
|
||||
# Durable record: grounded; suggested + related (LOCKED A3 — the log
|
||||
# records retrieval, not citations).
|
||||
row = _last_query_log()
|
||||
assert row.question == TAIL_QUESTION
|
||||
assert row.deflected is False
|
||||
assert row.sources == ", ".join(f"{SOURCE}/{p}" for p in SUGGESTED_TAIL + RELATED_PATHS)
|
||||
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# 2. Cases (c) + (e): ``read`` adds the full text; chips = suggested +
|
||||
# read (deduped); related row = rank 6+
|
||||
# --------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_read_suggested_doc_adds_full_text_and_cites(
|
||||
page: Page, app_url: str, synced_kb: None, db_ready: None
|
||||
) -> None:
|
||||
"""The scripted flow ``read``s the rank-1 suggested doc: the read
|
||||
SUCCEEDS (phase 118 — the seeds are summaries, not full text; the
|
||||
retired top-2 seed-read refusal is gone) and the mock's verbatim
|
||||
echo of the read result lands the document's tail sentinel in the
|
||||
answer (case c: the full text reached the model through the
|
||||
``read`` tool, not the seed). The citation chips = suggested +
|
||||
read, deduped — the read doc is among the suggested, so the chip
|
||||
row is exactly the 5 suggested docs; the related row renders rank
|
||||
6+ (case e)."""
|
||||
page.set_default_timeout(30_000)
|
||||
login(page, app_url, next="/")
|
||||
_install_page_hooks(page)
|
||||
|
||||
_submit(page, READ_QUESTION)
|
||||
_wait_settled(page)
|
||||
|
||||
# Exactly ONE executed tool call: the scripted read of the
|
||||
# rank-1 suggested doc.
|
||||
frames = _frames(page)
|
||||
assert _tool_frames(frames) == [
|
||||
{"type": "tool", "name": "read", "argument": f"{SOURCE}/{READ_TARGET}"}
|
||||
], _tool_frames(frames)
|
||||
done = next(f for f in frames if f.get("type") == "done")
|
||||
assert done["deflected"] is False, done
|
||||
# LOCKED A4: done.sources = suggested + read, deduped — the read doc
|
||||
# is already among the suggested, so the citation surface is
|
||||
# exactly the 5 suggested docs (in rank order).
|
||||
assert [
|
||||
(s["source"], s["path"]) for s in done["sources"]
|
||||
] == [(SOURCE, p) for p in SUGGESTED_READ], done["sources"]
|
||||
# done.related = the rank-6+ remainder (deduped against the cited).
|
||||
assert [
|
||||
(s["source"], s["path"]) for s in done["related"]
|
||||
] == [(SOURCE, p) for p in RELATED_PATHS], done["related"]
|
||||
|
||||
# The live Reading line carries the combined identity.
|
||||
line = _last_brain(page).locator(".tool-call")
|
||||
expect(line).to_have_count(1)
|
||||
expect(line).to_contain_text(f"Reading {SOURCE}/{READ_TARGET}")
|
||||
|
||||
# Case (c): the answer is the mock's verbatim echo of the READ
|
||||
# RESULT — its header and the document's tail sentinel (the last
|
||||
# line, outside the digest): the full text reached the model
|
||||
# through the read, not the seed.
|
||||
bubble = _last_brain(page).locator(".bubble")
|
||||
expect(bubble).to_contain_text(f"Document {SOURCE}/{READ_TARGET}:", timeout=30_000)
|
||||
expect(bubble).to_contain_text(READ_SENTINEL)
|
||||
expect(bubble).not_to_contain_text(TRUNCATION_MARKER)
|
||||
# The seed itself carried no full text: none of the OTHER docs'
|
||||
# sentinels are in the answer either.
|
||||
for path, _t, _i, _f, sentinel in DOCS:
|
||||
if path != READ_TARGET:
|
||||
expect(bubble).not_to_contain_text(sentinel)
|
||||
|
||||
# Case (e): the UI chip row = suggested + read (deduped — the read
|
||||
# doc is among the suggested, so chips = the 5 suggested docs), and
|
||||
# the related row renders rank 6+.
|
||||
_assert_suggested_chips_and_related_row(page, SUGGESTED_READ)
|
||||
|
||||
# Durable record: grounded; suggested + related + read (deduped,
|
||||
# LOCKED A3).
|
||||
row = _last_query_log()
|
||||
assert row.question == READ_QUESTION
|
||||
assert row.deflected is False
|
||||
assert row.sources == ", ".join(f"{SOURCE}/{p}" for p in SUGGESTED_READ + RELATED_PATHS)
|
||||
@@ -71,8 +71,11 @@ computed-style assertion):
|
||||
gray ok-ink on gray ok-bg, text intact.
|
||||
7. ``test_tool_call_lines_gray`` — screenshot 1: a mock-LLM turn that
|
||||
executes the tools renders the "Listing documents" / "Reading
|
||||
<source/path>" lines — gray accent-ink text, the gray accent-line
|
||||
left border, and the gray brand-soft path chip, all text intact.
|
||||
<source/path>" lines — gray accent-ink text, all text intact.
|
||||
(Phase 117 deboxed the line and dechipped the path on the owner's
|
||||
visual-glitch report — the accent rides the TEXT: the pins now
|
||||
assert the old accent-line border and the brand-soft chip
|
||||
background are GONE, not gray.)
|
||||
8. ``test_reset_removes_tag_byte_identical`` — the no-op contract: the
|
||||
gray tag is present in the live document pre-reset; Reset to
|
||||
defaults removes it from the LIVE document, a fresh load serves NO
|
||||
@@ -790,10 +793,12 @@ def test_local_badge_gray_labeled(
|
||||
#: document (a chunk carrying the mock's own bag-of-words embedding) that
|
||||
#: grounds the turn, and one CATALOG-ONLY document (indexed, no chunks) the
|
||||
#: mock's single-read flow reads. The catalog-only document sorts FIRST
|
||||
#: ("Checklist" < "ThemeNotes") — the mock reads the first catalog line, and
|
||||
#: it must NOT be the in-context retrieval document: the agent's read tool
|
||||
#: refuses documents already in the prompt (ALREADY_IN_CONTEXT), and a
|
||||
#: refused single-read flow would re-loop ls/read to the round cap.
|
||||
#: ("Checklist" < "ThemeNotes") — the mock reads the first catalog line.
|
||||
#: (Phase 118, A6: the read target's seed status no longer matters — a
|
||||
#: first read of ANY document succeeds; the ALREADY_IN_CONTEXT refusal
|
||||
#: fires only for a document already READ in the same turn. The
|
||||
#: catalog-only design stands because the flow reads the first catalog
|
||||
#: line, which must be a real, readable document.)
|
||||
READ_SOURCE = "Checklist"
|
||||
READ_PATH = "read-me.md"
|
||||
READ_SP = f"{READ_SOURCE}/{READ_PATH}"
|
||||
@@ -929,11 +934,15 @@ def test_tool_call_lines_gray(page: Page, app_url: str, db_ready: None) -> None:
|
||||
expect(page.locator("#send-label")).to_have_text("Send", timeout=60_000)
|
||||
|
||||
# Screenshot 1: the yellow "Listing documents" / "Reading" lines —
|
||||
# now gray accent-ink text with the gray accent-line left border,
|
||||
# the lines' text intact. Phase 94: the drill-down ls adds a THIRD
|
||||
# line between them — the drill ls scoped to the first source of
|
||||
# the top level (registry order: Checklist — the read target's
|
||||
# source).
|
||||
# now gray accent-ink TEXT, the lines' text intact. Phase 94: the
|
||||
# drill-down ls adds a THIRD line between them — the drill ls scoped
|
||||
# to the first source of the top level (registry order: Checklist —
|
||||
# the read target's source). Phase 117 (owner visual-glitch report)
|
||||
# deboxed the line and dechipped the path: the accent rides the
|
||||
# text color — there is NO left border and NO chip background to
|
||||
# be gray, so the pins assert their ABSENCE (a colored border or
|
||||
# chip background returning would fail both the width/alpha
|
||||
# assertion and the theme completeness contract).
|
||||
lines = page.locator(".msg.brain .tool-call")
|
||||
expect(lines).to_have_count(3)
|
||||
expect(lines.nth(0)).to_contain_text("Listing documents")
|
||||
@@ -942,14 +951,16 @@ def test_tool_call_lines_gray(page: Page, app_url: str, db_ready: None) -> None:
|
||||
expect(lines.nth(2)).to_contain_text("Reading")
|
||||
expect(lines.nth(2)).to_contain_text(READ_SP)
|
||||
_assert_gray(lines.nth(0), "color", GRAY["accent_ink"], label="ls line text")
|
||||
_assert_gray(
|
||||
lines.nth(0), "borderLeftColor", GRAY["accent_line"], label="ls line border"
|
||||
)
|
||||
assert (
|
||||
lines.nth(0).evaluate("el => getComputedStyle(el).borderLeftWidth") == "0px"
|
||||
), "phase 117 debox: the tool line must carry no left border"
|
||||
_assert_gray(lines.nth(2), "color", GRAY["accent_ink"], label="read line text")
|
||||
# The path chip on the Reading line: gray brand-soft background +
|
||||
# gray ink (the screenshot's code chip — still gray under the ramp).
|
||||
# The path `code` on the Reading line: phase 117 dechipped it —
|
||||
# transparent background, gray ink (still gray under the ramp).
|
||||
code = lines.nth(2).locator("code")
|
||||
_assert_gray(code, "backgroundColor", GRAY["brand_soft"], label="read chip bg")
|
||||
assert (
|
||||
code.evaluate("el => getComputedStyle(el).backgroundColor") == "rgba(0, 0, 0, 0)"
|
||||
), "phase 117 dechip: the path must have no chip background"
|
||||
_assert_gray(code, "color", GRAY["ink"], label="read chip text")
|
||||
|
||||
|
||||
|
||||
@@ -1,316 +0,0 @@
|
||||
"""Phase 24 E2E (Playwright): a matched document reaches the LLM whole.
|
||||
|
||||
Story: ``.agents/user_stories/whole-document-context.md``
|
||||
Run in isolation (DB must be up: ``podman compose up -d db``):
|
||||
|
||||
uv run pytest tests/e2e/test_whole_document_context.py -v --no-cov
|
||||
|
||||
The mock LLM's tail-echo trigger (``END_OF_NOTES_TRIGGER``, see
|
||||
``tests/e2e/mock_llm.py``) makes the model quote the last 160 chars of
|
||||
the document context. A sentinel placed on the *last line* of a document
|
||||
therefore appears in the rendered answer **iff the entire document was in
|
||||
the prompt** — which is what makes the no-truncation contract (A7 revised,
|
||||
owner permission 2026-08-24: matched parent documents are never cut)
|
||||
provable end-to-end.
|
||||
|
||||
The oversized documents are seeded directly via SQLAlchemy (a
|
||||
``documents`` row + 2–3 ``chunks`` rows whose embeddings are the mock's
|
||||
own deterministic bag-of-words vectors, so the question's live mock
|
||||
embedding genuinely overlaps — no fixture files added:
|
||||
``tests/fixtures/docs/`` stays at its 13 files (phase 47), other suites
|
||||
pin ``summary.added == 13``).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import hashlib
|
||||
from collections.abc import Callable, Sequence
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from threading import Thread
|
||||
from typing import Any
|
||||
|
||||
from playwright.sync_api import Page, expect
|
||||
from sqlalchemy import select, text
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
from app.config import Settings
|
||||
from app.db import SessionLocal
|
||||
from app.models import Chunk, Document, QueryLog
|
||||
from app.rag.importer import ImportSummary, import_sources
|
||||
from app.rag.llm import LLMClient
|
||||
from app.rag.retriever import TRUNCATION_MARKER
|
||||
from e2e.auth_helpers import login
|
||||
from tests.e2e.mock_llm import embed_text
|
||||
|
||||
REPO = Path(__file__).resolve().parents[2]
|
||||
FIXTURES = REPO / "tests" / "fixtures" / "docs"
|
||||
|
||||
#: The pre-phase-24 ``BOR_MAX_CONTEXT_CHARS`` default — the budget this
|
||||
#: suite proves is gone from the document path.
|
||||
OLD_CONTEXT_CAP = 24_000
|
||||
|
||||
QUESTION = "Show the end of your notes about the gitlab install playbook, please."
|
||||
SMALL_QUESTION = "How is my Kubernetes cluster set up?"
|
||||
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
|
||||
|
||||
|
||||
# --- Content builders (deterministic, token-controlled) -------------------
|
||||
|
||||
|
||||
def _repeated(line: str, min_chars: int) -> str:
|
||||
""""line" (newline-terminated) repeated until at least min_chars chars."""
|
||||
unit = line + "\n"
|
||||
return unit * max(1, -(-min_chars // len(unit)))
|
||||
|
||||
|
||||
def _doc_slices(content: str, n: int) -> list[str]:
|
||||
"""Even slices of *content* (the last slice keeps the final line)."""
|
||||
step = len(content) // n
|
||||
return [content[i * step : (i + 1) * step] for i in range(n - 1)] + [
|
||||
content[(n - 1) * step :]
|
||||
]
|
||||
|
||||
|
||||
def _gitlab_30k_doc(sentinel: str) -> str:
|
||||
"""A ~30 000-char document (past the old 24k cap): a body of repeated
|
||||
"gitlab install playbook" lines — the same tokens the question carries,
|
||||
so hybrid retrieval genuinely hits — whose LAST line is a unique
|
||||
sentinel only a tail echo can surface."""
|
||||
body = _repeated(
|
||||
"gitlab install playbook: run the gitlab install playbook on the homelab host.",
|
||||
OLD_CONTEXT_CAP + 6_000,
|
||||
)
|
||||
return body + sentinel + "\n"
|
||||
|
||||
|
||||
def _pair_doc(strong_line: str, filler_line: str, sentinel: str) -> tuple[str, list[str]]:
|
||||
"""A ~16 000-char document: a ~2 000-char first chunk carrying the
|
||||
question's key tokens, a ~14 000-char low-overlap remainder, and a
|
||||
unique sentinel as the last line. Returns (content, chunk_texts)."""
|
||||
chunk0 = _repeated(strong_line, 2_000)
|
||||
chunk1 = _repeated(filler_line, 14_000)
|
||||
return chunk0 + chunk1 + sentinel + "\n", [chunk0, chunk1]
|
||||
|
||||
|
||||
# --- DB seeding (TRUNCATE-then-seed, cf. test_chat_rag.py) -----------------
|
||||
|
||||
|
||||
def _seed_doc(
|
||||
db: Session,
|
||||
source: str,
|
||||
path: str,
|
||||
title: str,
|
||||
content: str,
|
||||
chunk_texts: Sequence[str],
|
||||
) -> None:
|
||||
"""One ``documents`` row + one ``chunks`` row per chunk text.
|
||||
|
||||
Each chunk's embedding is the mock's own ``embed_text`` vector, so the
|
||||
app's live mock embedding of the question genuinely overlaps.
|
||||
"""
|
||||
doc = Document(
|
||||
source=source,
|
||||
path=path,
|
||||
full_path=f"/tmp/{path}",
|
||||
title=title,
|
||||
content=content,
|
||||
content_hash=hashlib.sha256(content.encode("utf-8")).hexdigest(),
|
||||
indexed_at=datetime.now(UTC),
|
||||
)
|
||||
db.add(doc)
|
||||
db.flush()
|
||||
db.add_all(
|
||||
Chunk(document_id=doc.id, position=i, content=chunk, embedding=embed_text(chunk))
|
||||
for i, chunk in enumerate(chunk_texts)
|
||||
)
|
||||
|
||||
|
||||
def _reset_db(seed: Callable[[Session], None] | None = None) -> None:
|
||||
"""Truncate the KB (and query log), then optionally run *seed*."""
|
||||
with SessionLocal() as db:
|
||||
db.execute(text("TRUNCATE chunks, documents, query_log"))
|
||||
db.commit()
|
||||
if seed is not None:
|
||||
seed(db)
|
||||
db.commit()
|
||||
|
||||
|
||||
# --- Importer + thread helpers (test_chat_rag.py pattern) ------------------
|
||||
|
||||
|
||||
async def _import_fixtures(mock_port: int) -> ImportSummary:
|
||||
kwargs: dict[str, Any] = {
|
||||
"_env_file": None,
|
||||
"llm_base_url": f"http://127.0.0.1:{mock_port}/v1",
|
||||
}
|
||||
settings = Settings(**kwargs) # pyright: ignore[reportCallIssue]
|
||||
return await import_sources([FIXTURES], LLMClient(settings))
|
||||
|
||||
|
||||
def _run_in_thread(coro: Any) -> Any:
|
||||
"""Run a coroutine on a worker thread.
|
||||
|
||||
Playwright's sync API keeps an asyncio loop running on the test thread,
|
||||
so ``asyncio.run`` cannot be called directly from a test body.
|
||||
"""
|
||||
box: dict[str, Any] = {}
|
||||
|
||||
def runner() -> None:
|
||||
try:
|
||||
box["value"] = asyncio.run(coro)
|
||||
except BaseException as e: # noqa: BLE001 — re-raised on the test thread
|
||||
box["error"] = e
|
||||
|
||||
t = Thread(target=runner)
|
||||
t.start()
|
||||
t.join()
|
||||
if "error" in box:
|
||||
raise box["error"]
|
||||
return box["value"]
|
||||
|
||||
|
||||
def _ask(page: Page, app_url: str, question: str) -> Any:
|
||||
"""Submit *question* and wait for the streamed brain bubble."""
|
||||
page.set_default_timeout(30_000)
|
||||
login(page, app_url, next="/")
|
||||
page.fill("#message-input", question)
|
||||
page.click("#send-btn")
|
||||
bubble = page.locator(".msg.brain .bubble")
|
||||
bubble.first.wait_for(state="visible", timeout=30_000)
|
||||
return bubble.first
|
||||
|
||||
|
||||
def _last_query_log() -> QueryLog:
|
||||
with SessionLocal() as db:
|
||||
rows = db.scalars(select(QueryLog)).all()
|
||||
assert len(rows) == 1, f"expected exactly one query_log row, got {len(rows)}"
|
||||
return rows[0]
|
||||
|
||||
|
||||
# --- Story tests -------------------------------------------------------------
|
||||
|
||||
|
||||
def test_whole_document_over_old_cap_reaches_llm(
|
||||
page: Page, app_url: str, mock_llm: int, db_ready: None
|
||||
) -> None:
|
||||
"""A 30k document (past the old 24k cap) reaches the LLM whole: its
|
||||
tail sentinel — the last 160 chars of the context — is echoed back."""
|
||||
sentinel = "WHOLE-DOC-TAIL-GITLAB-30K"
|
||||
content = _gitlab_30k_doc(sentinel)
|
||||
assert len(content) > OLD_CONTEXT_CAP # this is the point of the test
|
||||
|
||||
def seed(db: Session) -> None:
|
||||
_seed_doc(
|
||||
db,
|
||||
"Homelab",
|
||||
"gitlab-30k.md",
|
||||
"GitLab Install Playbook (30k)",
|
||||
content,
|
||||
_doc_slices(content, 3),
|
||||
)
|
||||
|
||||
_reset_db(seed)
|
||||
bubble = _ask(page, app_url, QUESTION)
|
||||
|
||||
# The tail sentinel exists only on the document's last line — its
|
||||
# presence proves the entire 30k document was in the LLM prompt.
|
||||
expect(bubble).to_contain_text(sentinel, timeout=30_000)
|
||||
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER)
|
||||
expect(bubble).not_to_contain_text(TRUNCATION_MARKER)
|
||||
|
||||
# Grounded: the document's source chip renders under the bubble.
|
||||
chip = page.locator(".msg.brain .source-chip", has_text="gitlab-30k.md")
|
||||
expect(chip).to_have_count(1)
|
||||
expect(chip.first).to_contain_text("Homelab/gitlab-30k.md")
|
||||
|
||||
# Button recovers (never stale) and the turn was grounded.
|
||||
expect(page.locator("#send-btn")).to_be_enabled()
|
||||
expect(page.locator("#send-label")).to_have_text("Send")
|
||||
row = _last_query_log()
|
||||
assert row.question == QUESTION
|
||||
assert row.deflected is False
|
||||
assert "Homelab/gitlab-30k.md" in row.sources
|
||||
|
||||
|
||||
def test_second_document_of_over_cap_pair_reaches_llm(
|
||||
page: Page, app_url: str, mock_llm: int, db_ready: None
|
||||
) -> None:
|
||||
"""Two ~16k documents (32k combined — under the old budget the
|
||||
lower-ranked one was truncated in place): the SECOND document, the
|
||||
last block inside <documents>, reaches the LLM whole — its sentinel is
|
||||
the one the tail echo surfaces."""
|
||||
sentinel_a = "WHOLE-DOC-TAIL-PAIR-A"
|
||||
sentinel_b = "WHOLE-DOC-TAIL-PAIR-B"
|
||||
# Doc A's first chunk carries the question's key tokens → it ranks
|
||||
# first in both candidate lists → it comes first in <documents>.
|
||||
content_a, chunks_a = _pair_doc(
|
||||
"gitlab install playbook: run the gitlab install playbook end to end.",
|
||||
"the server room keeps a steady temperature and the racks are labelled.",
|
||||
sentinel_a,
|
||||
)
|
||||
# Doc B's first chunk has only a weaker overlap ("playbook", "notes")
|
||||
# → it ranks second → it is the LAST block inside <documents>.
|
||||
content_b, chunks_b = _pair_doc(
|
||||
"playbook notes: the playbook notes track what changed and where.",
|
||||
"the rack elevation drawing shows cable trays and pdu positions.",
|
||||
sentinel_b,
|
||||
)
|
||||
assert len(content_a) + len(content_b) > OLD_CONTEXT_CAP # 32k > 24k
|
||||
|
||||
def seed(db: Session) -> None:
|
||||
_seed_doc(
|
||||
db, "Homelab", "gitlab-install-playbook.md",
|
||||
"GitLab Install Playbook", content_a, chunks_a,
|
||||
)
|
||||
_seed_doc(
|
||||
db, "Homelab", "playbook-notes.md",
|
||||
"Playbook Notes", content_b, chunks_b,
|
||||
)
|
||||
|
||||
_reset_db(seed)
|
||||
bubble = _ask(page, app_url, QUESTION)
|
||||
|
||||
# The tail echo quotes doc B's sentinel (the last block's tail) — doc B
|
||||
# was in the prompt whole, past the old budget. Doc A's sentinel sits
|
||||
# mid-prompt, so it must NOT be in the quoted tail: that is what pins
|
||||
# the rank order (A first, B last inside <documents>).
|
||||
expect(bubble).to_contain_text(sentinel_b, timeout=30_000)
|
||||
expect(bubble).not_to_contain_text(sentinel_a)
|
||||
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER)
|
||||
expect(bubble).not_to_contain_text(TRUNCATION_MARKER)
|
||||
|
||||
# Both documents are cited (top-2), in rank order.
|
||||
chips = page.locator(".msg.brain .source-chip")
|
||||
expect(chips).to_have_count(2)
|
||||
expect(page.locator(".msg.brain .source-chip",
|
||||
has_text="gitlab-install-playbook.md")).to_have_count(1)
|
||||
expect(page.locator(".msg.brain .source-chip",
|
||||
has_text="playbook-notes.md")).to_have_count(1)
|
||||
|
||||
row = _last_query_log()
|
||||
assert row.question == QUESTION
|
||||
assert row.deflected is False
|
||||
assert "Homelab/gitlab-install-playbook.md" in row.sources
|
||||
assert "Homelab/playbook-notes.md" in row.sources
|
||||
|
||||
|
||||
def test_small_document_path_unchanged(
|
||||
page: Page, app_url: str, mock_llm: int, db_ready: None
|
||||
) -> None:
|
||||
"""Regression: the standard (small) fixtures still take the grounded
|
||||
path, byte-identical to before — no marker, kubernetes.md cited."""
|
||||
_reset_db(None)
|
||||
summary = _run_in_thread(_import_fixtures(mock_llm))
|
||||
assert summary.added == 13 # A9 formats (phase 47 added quadlet+j2)
|
||||
|
||||
bubble = _ask(page, app_url, SMALL_QUESTION)
|
||||
expect(bubble).to_contain_text(SMALL_QUESTION, timeout=30_000)
|
||||
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER)
|
||||
expect(bubble).not_to_contain_text(TRUNCATION_MARKER)
|
||||
|
||||
chip = page.locator(".msg.brain .source-chip", has_text="kubernetes.md")
|
||||
expect(chip).to_have_count(1)
|
||||
row = _last_query_log()
|
||||
assert row.deflected is False
|
||||
assert "docs/homelab/kubernetes.md" in row.sources
|
||||
Vendored
+78
-50
File diff suppressed because one or more lines are too long
@@ -282,17 +282,28 @@ def test_chat_streams_deltas_then_done_with_sources(client, db, seeded_kb: FakeR
|
||||
assert done[0]["suggestions"] == []
|
||||
sources = done[0]["sources"]
|
||||
assert sources, "done must carry the cited sources"
|
||||
# Phase 118 (A4): the citation surface is the suggested tier (top-5,
|
||||
# no floor) + the agent's reads (none on this turn) — deduped.
|
||||
assert len(sources) == get_settings().suggested_docs
|
||||
assert sources[0]["path"] == "homelab/kubernetes.md"
|
||||
assert sources[0]["source"] == "docs"
|
||||
assert sources[0]["title"] == "Kubernetes Homelab Cluster"
|
||||
|
||||
# The LLM received the locked HIGH prompt with the FULL document text.
|
||||
# The LLM received the locked HIGH prompt — the ``<documents>`` block
|
||||
# seeds the document's stored SUMMARY (phase 118, LOCKED A6: summary
|
||||
# seeding re-revises the pre-phase full-text contract; the full text
|
||||
# reaches the context only through the capped ``read`` tool). The
|
||||
# summarizer's code-appended pointer line proves the summary block is
|
||||
# present; the doc's full body is no longer seeded.
|
||||
(system, user) = seeded_kb.seen_messages[0][0], seeded_kb.seen_messages[0][1]
|
||||
assert user["content"] == QUESTION
|
||||
assert "<relevance>HIGH</relevance>" in system["content"]
|
||||
assert "DEFLECT_MODE" not in system["content"]
|
||||
assert "<documents>" in system["content"]
|
||||
assert "Talos Linux" in system["content"] # full doc, not just the chunk
|
||||
section = system["content"].split("<documents>", 1)[1].split("</documents>", 1)[0]
|
||||
assert "Summary of" in section # the fake lite model's summary text
|
||||
assert "Source: docs/homelab/kubernetes.md" in section # code-appended pointer
|
||||
assert "Talos Linux" not in section # full doc no longer seeded (A6)
|
||||
assert "HONESTY GATE" in system["content"]
|
||||
|
||||
|
||||
@@ -382,7 +393,19 @@ def test_chat_writes_query_log_row(client, db, seeded_kb: FakeRagLLM) -> None:
|
||||
assert 1 <= row.chunk_hits <= total_chunks
|
||||
assert row.top_score > 0.0 # genuine token-overlap cosine, best hit
|
||||
assert row.top_score <= 1.0
|
||||
assert "docs/homelab/kubernetes.md" in row.sources
|
||||
# Phase 118 (LOCKED A3): the durable record is the FULL retrieval —
|
||||
# the suggested tier (ranks 1–5) + the related tier (ranks 6–7) +
|
||||
# the agent's reads (none on this turn), for this question.
|
||||
for path in (
|
||||
"docs/homelab/kubernetes.md", # rank 1
|
||||
"docs/homelab/templates/deploy.j2", # rank 2
|
||||
"docs/homelab/ssh/ssh_aliases.txt", # rank 3
|
||||
"docs/homelab/container_gitlab/gitlab.md", # rank 4
|
||||
"docs/deployments/new-service.md", # rank 5
|
||||
"docs/homelab/quadlet/cache.volume", # rank 6 (related)
|
||||
"docs/homelab/quadlet/compose.container", # rank 7 (related)
|
||||
):
|
||||
assert path in row.sources
|
||||
assert row.latency_ms >= 0
|
||||
# Why the gate answered (A8 revised): cosine over the threshold OR a
|
||||
# lexical hit. The mock-calibrated threshold (0.30, see tests/conftest.py)
|
||||
@@ -446,10 +469,12 @@ def test_off_topic_question_deflects_honestly(client, db, seeded_kb: FakeRagLLM)
|
||||
def test_done_frame_carries_related_tier_on_grounded_turn(
|
||||
client, db, seeded_kb: FakeRagLLM
|
||||
) -> None:
|
||||
"""Phase 113 (LOCKED A4): a grounded turn's done frame carries the
|
||||
related tier — the ranked docs beyond the cited ceiling, capped at
|
||||
``related_max_docs`` (2), disjoint from the cited list. The durable
|
||||
record keeps the FULL retrieval (cited + related, LOCKED A3)."""
|
||||
"""Phase 118 (LOCKED A3/A4): a grounded turn's done frame carries the
|
||||
suggested tier in ``sources`` (top-5, no floor) and the related
|
||||
tier — the ranked docs from rank 6+ after the suggested set, capped
|
||||
at ``related_max_docs`` (2) — in ``related``, disjoint from the
|
||||
citation surface. The durable record keeps the FULL retrieval
|
||||
(suggested + related + read, LOCKED A3)."""
|
||||
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb
|
||||
try:
|
||||
_, _, frames = _stream_chat(client, QUESTION)
|
||||
@@ -458,49 +483,54 @@ def test_done_frame_carries_related_tier_on_grounded_turn(
|
||||
|
||||
done = frames[-1]
|
||||
assert done["deflected"] is False
|
||||
sources = {(s["source"], s["path"]) for s in done["sources"]}
|
||||
# A4: the citation surface is the suggested tier (5, no read on this
|
||||
# turn) — ranks 1–5 for the Kubernetes question.
|
||||
sources = [(s["source"], s["path"]) for s in done["sources"]]
|
||||
assert len(sources) == get_settings().suggested_docs
|
||||
assert sources[0] == ("docs", "homelab/kubernetes.md")
|
||||
related = done["related"]
|
||||
assert related, "the 2nd-and-lower scored docs ride the related tier"
|
||||
# Rank 6–7 for the Kubernetes question (after the top-5 suggested
|
||||
# set), capped at related_max_docs.
|
||||
assert [(s["source"], s["path"]) for s in related] == [
|
||||
("docs", "homelab/quadlet/cache.volume"),
|
||||
("docs", "homelab/quadlet/compose.container"),
|
||||
]
|
||||
assert len(related) <= get_settings().related_max_docs
|
||||
# The related tier never overlaps the cited list (the dedupe is by
|
||||
# (source, path) — the same pattern as the cited docs).
|
||||
# The related tier never overlaps the citation surface (the dedupe is
|
||||
# by (source, path) — the same pattern as the cited docs).
|
||||
related_keys = {(s["source"], s["path"]) for s in related}
|
||||
assert sources.isdisjoint(related_keys)
|
||||
# Rank order: the cited top-2 are the kubernetes doc and the template;
|
||||
# the next ranked doc is the ssh aliases file.
|
||||
assert related[0]["path"] == "homelab/ssh/ssh_aliases.txt"
|
||||
assert set(sources).isdisjoint(related_keys)
|
||||
# Every ref carries the chip identity fields (the UI row reuses them).
|
||||
assert all(s["title"] for s in related)
|
||||
|
||||
# Durable record: the full retrieval (cited + related) is logged.
|
||||
# Durable record: the full retrieval (suggested + related) is logged.
|
||||
row = db.scalars(select(QueryLog)).one()
|
||||
assert "docs/homelab/ssh/ssh_aliases.txt" in row.sources
|
||||
assert "docs/homelab/quadlet/cache.volume" in row.sources
|
||||
assert "docs/homelab/kubernetes.md" in row.sources
|
||||
|
||||
|
||||
def test_deflected_done_frame_carries_weak_hits_in_related(
|
||||
client, db, seeded_kb: FakeRagLLM, monkeypatch: pytest.MonkeyPatch
|
||||
client, db, seeded_kb: FakeRagLLM
|
||||
) -> None:
|
||||
"""Phase 113: on a deflected turn nothing clears the bar — the cited
|
||||
tier is empty (done.sources stays [], the phase-112 contract) and the
|
||||
weak hits fall to the related tier (their visibility home). The
|
||||
durable record still carries the retrieval (LOCKED A3)."""
|
||||
monkeypatch.setenv("BOR_SOURCE_USEFULNESS_FLOOR", "0.20")
|
||||
get_settings.cache_clear()
|
||||
"""Phase 118: on a deflected turn done.sources stays [] (the
|
||||
phase-112 contract — a deflected answer cites nothing) and
|
||||
done.related carries rank 6+ after the suggested set (capped at
|
||||
``related_max_docs``) — the weak hits' visibility home; the weak
|
||||
hits themselves are the suggested tier (no floor, A3). The durable
|
||||
record still carries the retrieval (LOCKED A3)."""
|
||||
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb
|
||||
try:
|
||||
_, _, frames = _stream_chat(client, OFF_TOPIC)
|
||||
done = frames[-1]
|
||||
assert done["deflected"] is True
|
||||
assert done["sources"] == [] # a deflected answer cites nothing
|
||||
# The weak hits (the sourdough question's best mock cosines are
|
||||
# ~0.11/0.04 — both below the 0.20 bar) ride the related tier, in
|
||||
# rank order, capped at related_max_docs.
|
||||
# Rank 6–7 for the sourdough question (after the top-5 suggested
|
||||
# set), capped at related_max_docs.
|
||||
related = done["related"]
|
||||
assert len(related) <= get_settings().related_max_docs
|
||||
assert [s["path"] for s in related][:2] == [
|
||||
"deployments/new-service.md",
|
||||
"homelab/quadlet/lan.network",
|
||||
assert [s["path"] for s in related] == [
|
||||
"homelab/backups.md",
|
||||
"homelab/container_gitlab/gitlab-compose.yaml",
|
||||
]
|
||||
assert all(s["title"] for s in related)
|
||||
assert done["suggestions"] # the "Maybe try" chips are unchanged
|
||||
@@ -510,19 +540,16 @@ def test_deflected_done_frame_carries_weak_hits_in_related(
|
||||
assert row.deflected is True
|
||||
assert row.sources # the weak-hit paths, for threshold tuning
|
||||
finally:
|
||||
# The cache clear is LAST — an assertion that calls get_settings()
|
||||
# after the clear would re-populate the lru_cache with the
|
||||
# monkeypatched value and leak it into the next test.
|
||||
fastapi_app.dependency_overrides.clear()
|
||||
get_settings.cache_clear()
|
||||
|
||||
|
||||
def test_related_doc_read_by_agent_is_cited_not_related(
|
||||
client, db, seeded_kb: FakeRagLLM
|
||||
) -> None:
|
||||
"""Phase 113 × phase 37: an agent-read doc is a citation by definition
|
||||
— when the agent ``read``s a doc that would otherwise ride the related
|
||||
tier, it joins done.sources (deduped, last) and is EXCLUDED from
|
||||
"""Phase 118 × phase 37: an agent-read doc is a citation by definition
|
||||
(LOCKED A4) — when the agent ``read``s a rank-6+ doc (the related
|
||||
tier, "nearby docs"), it joins done.sources (deduped, last — it was
|
||||
not suggested, so the read appends it) and is EXCLUDED from
|
||||
done.related (a "nearby doc" that was actually used must not read as
|
||||
nearby)."""
|
||||
scripted = FakeRagLLM(
|
||||
@@ -531,7 +558,7 @@ def test_related_doc_read_by_agent_is_cited_not_related(
|
||||
ToolCallPiece(
|
||||
id="call_1",
|
||||
name="read",
|
||||
arguments={"path": "docs/homelab/ssh/ssh_aliases.txt"},
|
||||
arguments={"path": "docs/homelab/quadlet/cache.volume"},
|
||||
)
|
||||
]
|
||||
]
|
||||
@@ -545,12 +572,14 @@ def test_related_doc_read_by_agent_is_cited_not_related(
|
||||
done = frames[-1]
|
||||
assert done["deflected"] is False
|
||||
sources = [(s["source"], s["path"]) for s in done["sources"]]
|
||||
assert sources[-1] == ("docs", "homelab/ssh/ssh_aliases.txt") # read ⇒ cited
|
||||
# A4: suggested (5) + the read doc (last).
|
||||
assert len(sources) == get_settings().suggested_docs + 1
|
||||
assert sources[-1] == ("docs", "homelab/quadlet/cache.volume") # read ⇒ cited
|
||||
related = [(s["source"], s["path"]) for s in done["related"]]
|
||||
assert ("docs", "homelab/ssh/ssh_aliases.txt") not in related
|
||||
assert ("docs", "homelab/quadlet/cache.volume") not in related
|
||||
assert set(sources).isdisjoint(set(related))
|
||||
# The OTHER related-tier doc (gitlab, rank 4) stays in the tier.
|
||||
assert ("docs", "homelab/container_gitlab/gitlab.md") in related
|
||||
# The OTHER related-tier doc (compose.container, rank 7) stays in the tier.
|
||||
assert ("docs", "homelab/quadlet/compose.container") in related
|
||||
|
||||
|
||||
def test_keyword_question_grounded_by_lexical_hit_despite_weak_cosine(
|
||||
@@ -1341,6 +1370,11 @@ def test_deflected_turn_stays_byte_identical_without_tools(
|
||||
direct-``chat_stream`` output even for a fake scripted to call tools
|
||||
(its script is never consumed). The LLM was called once, without a
|
||||
``tools`` key."""
|
||||
# The scripted read targets a doc OUTSIDE the OFF_TOPIC retrieval
|
||||
# top-7 (tables.md ranks 11th — not suggested, not rank 6+ related),
|
||||
# so "never read" stays distinguishable from "retrieved" in the
|
||||
# durable record below (phase 118: backups.md — the pre-phase read
|
||||
# target — now rides the rank-6+ related tier, durably recorded).
|
||||
scripted = FakeRagLLM(
|
||||
tool_script=[
|
||||
[ToolCallPiece(id="call_1", name="ls", arguments={})],
|
||||
@@ -1348,7 +1382,7 @@ def test_deflected_turn_stays_byte_identical_without_tools(
|
||||
ToolCallPiece(
|
||||
id="call_2",
|
||||
name="read",
|
||||
arguments={"path": "docs/homelab/backups.md"},
|
||||
arguments={"path": "docs/homelab/tables.md"},
|
||||
)
|
||||
],
|
||||
[StreamPiece("content", "never used — the agent never runs")],
|
||||
@@ -1380,7 +1414,7 @@ def test_deflected_turn_stays_byte_identical_without_tools(
|
||||
if r.question == OFF_TOPIC
|
||||
][-1:]
|
||||
assert row.deflected is True
|
||||
assert "backups.md" not in row.sources
|
||||
assert "tables.md" not in row.sources
|
||||
|
||||
|
||||
def test_zero_max_rounds_reproduce_pre_phase_single_request(
|
||||
|
||||
@@ -63,6 +63,16 @@ FULL_BRAIN: dict[str, Any] = {
|
||||
"sources": [
|
||||
{"source": "Homelab", "path": "kubernetes.md", "title": "Kubernetes Cluster"}
|
||||
],
|
||||
# Phase 113's related-doc tier — the UI persists it with every
|
||||
# grounded brain record (the restore path re-renders the row from
|
||||
# it). It must be an ACCEPTED key: the phase-113 omission (the key
|
||||
# missing from ChatMessage) made the extra="forbid" boundary 422
|
||||
# every done-time auto-save carrying it, so grounded turns' brain
|
||||
# messages never persisted (the A2 quiet failure swallowed the
|
||||
# 422). This round-trip is the regression pin.
|
||||
"related": [
|
||||
{"source": "Homelab", "path": "traefik.md", "title": "Traefik Notes"}
|
||||
],
|
||||
"deflected": False,
|
||||
"suggestions": ["What ports does Traefik expose?"],
|
||||
"thinking": "The kubernetes doc covers the cluster layout…",
|
||||
@@ -175,6 +185,7 @@ def _expect(records: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||
"who": m["who"],
|
||||
"text": m["text"],
|
||||
"sources": m.get("sources"),
|
||||
"related": m.get("related"),
|
||||
"deflected": m.get("deflected"),
|
||||
"suggestions": m.get("suggestions"),
|
||||
"thinking": m.get("thinking"),
|
||||
@@ -411,8 +422,8 @@ def test_create_round_trips_full_brain_record(admin_client: TestClient) -> None:
|
||||
)
|
||||
assert r.status_code == 201
|
||||
# The bor.chat.v1-shaped payload round-trips losslessly: every
|
||||
# optional key (sources/deflected/suggestions/thinking/tools/
|
||||
# stopped) survives identical.
|
||||
# optional key (sources/related/deflected/suggestions/thinking/
|
||||
# tools/stopped) survives identical.
|
||||
assert r.json()["messages"][1] == FULL_BRAIN
|
||||
|
||||
|
||||
@@ -516,9 +527,9 @@ def test_get_returns_full_payload_round_trip(admin_client: TestClient) -> None:
|
||||
assert body["id"] == created["id"]
|
||||
assert body["title"] == EXPLICIT_TITLE
|
||||
assert body["message_count"] == 2
|
||||
# Byte-identical payload: the brain record with sources/thinking/
|
||||
# tools/stopped (incl. the `argument: null` tool) survives the trip
|
||||
# to Postgres and back.
|
||||
# Byte-identical payload: the brain record with sources/related/
|
||||
# thinking/tools/stopped (incl. the `argument: null` tool) survives
|
||||
# the trip to Postgres and back.
|
||||
assert body["messages"] == _expect([_user(FIRST_QUESTION), FULL_BRAIN])
|
||||
|
||||
|
||||
|
||||
@@ -30,8 +30,10 @@ ends with the folder-summary stats —
|
||||
(this fixture's 2-doc source holds exactly ONE qualifying subtree: the
|
||||
source root) or ``folder_summaries=skipped`` otherwise — so the line
|
||||
pinned here gains that token, and a KB-changing run burns exactly ONE
|
||||
extra ``lite`` call (the source-root folder summary, markdown files
|
||||
never get a document summary).
|
||||
extra ``lite`` call beyond the phase-118 document summaries (the
|
||||
source-root folder summary; markdown files get document summaries too
|
||||
since phase 118, A2 — so this 2-doc markdown source burns TWO doc-summary
|
||||
calls on a fresh import).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -97,7 +99,8 @@ def _run_main(
|
||||
|
||||
@pytest.fixture()
|
||||
def src(tmp_path: Path) -> Path:
|
||||
"""A source dir with two markdown docs (md → no summary chat calls)."""
|
||||
"""A source dir with two markdown docs (phase 118, A2: both get
|
||||
document summaries — two extra ``chat`` calls over pre-118)."""
|
||||
root = tmp_path / "MyDocs"
|
||||
root.mkdir()
|
||||
(root / "alpha.md").write_text("# Alpha\n\nFirst document.\n", encoding="utf-8")
|
||||
@@ -164,16 +167,18 @@ def test_changed_import_writes_overview_row(
|
||||
"overview=updated sources_version=1 folder_summaries=1/0/0"
|
||||
)
|
||||
assert _version(db) == 1 # phase 53: a changed import bumps exactly once
|
||||
# Exactly two lite calls — the overview + the source-root folder
|
||||
# summary (markdown files never get a document summary, so nothing
|
||||
# else may touch ``chat``).
|
||||
assert len(llm.chat_calls) == 2
|
||||
by_role = {m["role"]: m["content"] for m in llm.chat_calls[0]}
|
||||
# Exactly four lite calls — the two phase-118 document summaries
|
||||
# (markdown included) + the overview + the source-root folder summary
|
||||
# (nothing else may touch ``chat``).
|
||||
assert len(llm.chat_calls) == 4
|
||||
by_role = {m["role"]: m["content"] for m in llm.chat_calls[2]}
|
||||
assert "KB_OVERVIEW_MODE" in by_role["system"]
|
||||
# One line per doc: source — path — title (no summary for markdown).
|
||||
assert "MyDocs — alpha.md — Alpha" in by_role["user"]
|
||||
assert "MyDocs — beta.md — Beta" in by_role["user"]
|
||||
by_role = {m["role"]: m["content"] for m in llm.chat_calls[1]}
|
||||
# One line per doc: source — path — title — first summary line
|
||||
# (phase 118: the markdown docs are summarized too — the fake's
|
||||
# deterministic digest for each).
|
||||
assert "MyDocs — alpha.md — Alpha — Summary of #" in by_role["user"]
|
||||
assert "MyDocs — beta.md — Beta — Summary of #" in by_role["user"]
|
||||
by_role = {m["role"]: m["content"] for m in llm.chat_calls[3]}
|
||||
assert "FOLDER_SUMMARY_MODE" in by_role["system"]
|
||||
assert by_role["user"].splitlines()[0] == "Folder: MyDocs"
|
||||
# The model's outline lands in the single row.
|
||||
@@ -197,17 +202,19 @@ def test_unchanged_reimport_does_not_call_lite(
|
||||
assert out.rstrip().endswith(
|
||||
"overview=updated sources_version=1 folder_summaries=1/0/0"
|
||||
)
|
||||
assert len(llm.chat_calls) == 2 # overview + source-root folder summary
|
||||
# 2 doc summaries (phase 118) + overview + source-root folder summary.
|
||||
assert len(llm.chat_calls) == 4
|
||||
assert _row(db) is not None
|
||||
|
||||
# Same hashes → no KB change → no lite call, previous outline kept.
|
||||
# Same hashes → no KB change → no lite call, previous outline kept —
|
||||
# and nothing to backfill (both summaries are already stored).
|
||||
rc, out = _run_main(monkeypatch, llm, ["--source", str(src)], capsys)
|
||||
assert rc == 0
|
||||
assert "unchanged=2" in out
|
||||
assert out.rstrip().endswith(
|
||||
"overview=skipped sources_version=skipped folder_summaries=skipped"
|
||||
)
|
||||
assert len(llm.chat_calls) == 2 # no new lite call
|
||||
assert len(llm.chat_calls) == 4 # no new lite call
|
||||
row = _row(db)
|
||||
assert row is not None and row.content == "Summary of MyDocs"
|
||||
assert _version(db) == 1 # phase 53: an unchanged re-run never bumps
|
||||
@@ -241,7 +248,10 @@ def test_lite_failure_is_fail_soft(
|
||||
assert out.rstrip().endswith(
|
||||
"overview=failed sources_version=2 folder_summaries=0/1/0"
|
||||
)
|
||||
assert len(bad.chat_calls) == 2 # the (failed) attempts were made
|
||||
# The three (failed) attempts: the changed doc's summary (phase 118),
|
||||
# the overview, and the folder summary — the unchanged, already-
|
||||
# summarized doc burns no backfill.
|
||||
assert len(bad.chat_calls) == 3
|
||||
row = _row(db)
|
||||
assert row is not None
|
||||
assert row.content == previous_content # previous row untouched
|
||||
@@ -262,7 +272,7 @@ def test_limit_run_skips_overview(
|
||||
assert out.rstrip().endswith(
|
||||
"overview=updated sources_version=1 folder_summaries=1/0/0"
|
||||
)
|
||||
assert len(llm.chat_calls) == 2
|
||||
assert len(llm.chat_calls) == 4 # 2 doc summaries + overview + folder
|
||||
|
||||
# An incomplete walk must not rewrite the outline (mirrors the
|
||||
# --prune-with---limit guard) — and must not advance the version.
|
||||
@@ -273,7 +283,9 @@ def test_limit_run_skips_overview(
|
||||
assert out.rstrip().endswith(
|
||||
"overview=skipped sources_version=skipped folder_summaries=skipped"
|
||||
)
|
||||
assert len(llm.chat_calls) == 2 # --limit never burns a lite call
|
||||
# --limit walks only alpha.md: its changed summary is the sole new
|
||||
# lite call; the overview + folder gates skip under --limit.
|
||||
assert len(llm.chat_calls) == 5
|
||||
row = _row(db)
|
||||
assert row is not None and row.content == "Summary of MyDocs"
|
||||
assert _version(db) == 1 # phase 53: --limit debug runs never bump
|
||||
@@ -323,6 +335,7 @@ def test_prune_only_run_bumps_sources_version(
|
||||
|
||||
# Delete one file; a --prune run drops exactly it: no add/update,
|
||||
# but pruned=1 → the version still bumps while the overview skips.
|
||||
assert len(llm.chat_calls) == 4 # 2 doc summaries + overview + folder
|
||||
(src / "alpha.md").unlink()
|
||||
rc, out = _run_main(monkeypatch, llm, ["--source", str(src), "--prune"], capsys)
|
||||
assert rc == 0
|
||||
|
||||
@@ -16,8 +16,8 @@ fixture of phase 56 stays pinned by its own suite):
|
||||
* positive — ``md,dockerfile,containerfile`` walks ``Dockerfile``,
|
||||
``Containerfile`` and ``notes.md`` (the ``Makefile`` negative control
|
||||
stays out): rows + embedded chunks + the deterministic mock
|
||||
``SUMMARY_MODE`` digest for both build files (non-markdown → the
|
||||
phase-30 ``lite`` path), ``formats == {"dockerfile": 1,
|
||||
``SUMMARY_MODE`` digest for EVERY doc (phase 118, A2: markdown
|
||||
included — the phase-30 ``lite`` path), ``formats == {"dockerfile": 1,
|
||||
"containerfile": 1, "md": 1}`` with NO ``unknown`` key;
|
||||
* case — an on-disk ``DOCKERFILE`` imports under the ``dockerfile``
|
||||
token, its row keeps the on-disk case;
|
||||
@@ -146,9 +146,9 @@ def test_name_token_files_import_end_to_end(mock_llm_port: int, db: Session) ->
|
||||
# D2: extensionless files count under their matched token.
|
||||
assert summary.formats == {"dockerfile": 1, "containerfile": 1, "md": 1}
|
||||
assert "unknown" not in summary.formats
|
||||
# Phase 30: both build files are non-markdown → lite summaries;
|
||||
# the markdown control never is.
|
||||
assert (summary.summaries, summary.summary_errors) == (2, 0)
|
||||
# Phase 118 (A2): EVERY doc gets a lite summary — both build
|
||||
# files AND the markdown control.
|
||||
assert (summary.summaries, summary.summary_errors) == (3, 0)
|
||||
|
||||
for rel, sentinel, tokens in (
|
||||
(DOCKER_REL, DOCKER_SENTINEL, DOCKER_SENTINEL_TOKENS),
|
||||
@@ -179,13 +179,18 @@ def test_name_token_files_import_end_to_end(mock_llm_port: int, db: Session) ->
|
||||
assert len(schunks) == 1 and schunks[0].position == -1
|
||||
assert schunks[0].embedding is not None
|
||||
|
||||
# The markdown control doc imported too — but markdown never
|
||||
# gets a summary (phase 30).
|
||||
# The markdown control doc imported too — AND got the mock
|
||||
# ``SUMMARY_MODE`` digest (phase 118, A2: markdown summarized).
|
||||
note = db.scalar(
|
||||
select(Document).where(Document.source == SOURCE, Document.path == NOTES_REL)
|
||||
)
|
||||
assert note is not None
|
||||
assert note.summary is None
|
||||
assert note.summary is not None
|
||||
assert note.summary.startswith("This document covers extensionless fixture notes")
|
||||
assert f"Source: {SOURCE}/{NOTES_REL}" in note.summary
|
||||
note_schunks = [c for c in note.chunks if c.is_summary]
|
||||
assert len(note_schunks) == 1 and note_schunks[0].position == -1
|
||||
assert note_schunks[0].embedding is not None
|
||||
assert [c for c in note.chunks if not c.is_summary]
|
||||
|
||||
# The negative control: no token names Makefile exactly — never
|
||||
@@ -243,7 +248,8 @@ def test_without_tokens_the_extensionless_files_stay_out(
|
||||
summary.files, summary.added, summary.unchanged, summary.updated, summary.errors
|
||||
) == (1, 1, 0, 0, 0)
|
||||
assert summary.formats == {"md": 1}
|
||||
assert summary.summaries == 0
|
||||
# Phase 118 (A2): the md-only run still summarizes the note.
|
||||
assert summary.summaries == 1
|
||||
for rel in (DOCKER_REL, CONTAINER_REL, "Makefile"):
|
||||
assert (
|
||||
db.scalar(
|
||||
|
||||
@@ -117,8 +117,9 @@ def test_novel_extension_imports_end_to_end(mock_llm_port: int, db: Session) ->
|
||||
summary.files, summary.added, summary.unchanged, summary.updated, summary.errors
|
||||
) == (2, 2, 0, 0, 0)
|
||||
assert summary.formats == {"sh": 1, "md": 1}
|
||||
# The .sh file is non-markdown → exactly one lite summary (phase 30).
|
||||
assert (summary.summaries, summary.summary_errors) == (1, 0)
|
||||
# Phase 118 (A2): EVERY doc gets a lite summary — the novel
|
||||
# ``.sh`` file AND the markdown control.
|
||||
assert (summary.summaries, summary.summary_errors) == (2, 0)
|
||||
|
||||
sh = db.scalar(
|
||||
select(Document).where(
|
||||
@@ -148,15 +149,20 @@ def test_novel_extension_imports_end_to_end(mock_llm_port: int, db: Session) ->
|
||||
assert len(schunks) == 1 and schunks[0].position == -1
|
||||
assert schunks[0].embedding is not None
|
||||
|
||||
# The markdown control doc imported too — but markdown never
|
||||
# gets a summary (phase 30).
|
||||
# The markdown control doc imported too — AND got the mock
|
||||
# ``SUMMARY_MODE`` digest (phase 118, A2: markdown summarized).
|
||||
note = db.scalar(
|
||||
select(Document).where(
|
||||
Document.source == SOURCE, Document.path == MD_REL
|
||||
)
|
||||
)
|
||||
assert note is not None
|
||||
assert note.summary is None
|
||||
assert note.summary is not None
|
||||
assert note.summary.startswith("This document covers extension fixture note")
|
||||
assert f"Source: {SOURCE}/{MD_REL}" in note.summary
|
||||
note_schunks = [c for c in note.chunks if c.is_summary]
|
||||
assert len(note_schunks) == 1 and note_schunks[0].position == -1
|
||||
assert note_schunks[0].embedding is not None
|
||||
assert [c for c in note.chunks if not c.is_summary]
|
||||
finally:
|
||||
_cleanup_source(db, SOURCE)
|
||||
@@ -175,7 +181,8 @@ def test_narrowing_to_md_still_excludes_the_novel_extension(
|
||||
summary.files, summary.added, summary.unchanged, summary.updated, summary.errors
|
||||
) == (1, 1, 0, 0, 0)
|
||||
assert summary.formats == {"md": 1}
|
||||
assert summary.summaries == 0
|
||||
# Phase 118 (A2): the md-only run still summarizes the note.
|
||||
assert summary.summaries == 1
|
||||
assert db.scalar(
|
||||
select(Document).where(
|
||||
Document.source == SOURCE, Document.path == SH_REL
|
||||
|
||||
@@ -78,8 +78,9 @@ def test_import_fixtures_end_to_end(admin_client, db) -> None:
|
||||
k8s = next(d for d in docs if d.path == "homelab/kubernetes.md")
|
||||
assert "Talos Linux" in k8s.content and k8s.content_hash
|
||||
|
||||
# Phase 30: the four non-markdown fixtures each gained one embedded
|
||||
# ``is_summary`` chunk, so the DB holds content + summary chunks.
|
||||
# Phase 30; phase 118 (A2): EVERY fixture gained one embedded
|
||||
# ``is_summary`` chunk (markdown included), so the DB holds content +
|
||||
# summary chunks.
|
||||
n_chunks = db.scalar(select(func.count()).select_from(Chunk))
|
||||
assert n_chunks == summary.chunks + summary.summaries
|
||||
for c in db.scalars(select(Chunk)).all():
|
||||
@@ -87,21 +88,15 @@ def test_import_fixtures_end_to_end(admin_client, db) -> None:
|
||||
|
||||
assert summary.summary_errors == 0
|
||||
for d in docs:
|
||||
non_md = Path(d.path).suffix.lower() not in (".md", ".markdown")
|
||||
# Phase 118 (A2): EVERY document — markdown included — carries a
|
||||
# stored summary + exactly one embedded summary chunk (−1).
|
||||
schunks = [c for c in d.chunks if c.is_summary]
|
||||
if non_md:
|
||||
# Lite summary stored + exactly one embedded summary chunk (−1).
|
||||
assert d.summary is not None, f"{d.path} should have a summary"
|
||||
assert len(schunks) == 1
|
||||
assert schunks[0].position == -1
|
||||
assert schunks[0].content == d.summary
|
||||
assert schunks[0].embedding is not None
|
||||
else:
|
||||
# Markdown docs never get a summary (phase 30 scope).
|
||||
assert d.summary is None and not schunks
|
||||
assert summary.summaries == sum(
|
||||
1 for d in docs if Path(d.path).suffix.lower() not in (".md", ".markdown")
|
||||
)
|
||||
assert summary.summaries == len(docs) # one per doc, no exceptions
|
||||
|
||||
# The Sources page consumes exactly this shape.
|
||||
r = admin_client.get("/api/docs") # phase 16: the catalog is admin-only
|
||||
|
||||
@@ -60,8 +60,9 @@ def test_ignored_files_never_indexed(db: Session, tmp_path: Path) -> None:
|
||||
# invisible to the pipeline.
|
||||
assert summary.files == 2
|
||||
assert summary.errors == 0
|
||||
# The kept non-markdown file IS summarized; the ignored .txt is not.
|
||||
assert summary.summaries == 1
|
||||
# Both kept files ARE summarized (phase 118, A2: markdown too);
|
||||
# the ignored files are never walked, hence never summarized.
|
||||
assert summary.summaries == 2
|
||||
assert summary.summary_errors == 0
|
||||
|
||||
docs = db.scalars(select(Document)).all()
|
||||
@@ -78,12 +79,17 @@ def test_ignored_files_never_indexed(db: Session, tmp_path: Path) -> None:
|
||||
assert not any(
|
||||
"SECRET-CONTENT" in t or "IGNORED-TEXT-CONTENT" in t for t in texts
|
||||
)
|
||||
# Exactly one summary call (top.txt) — the ignored files never reached
|
||||
# the lite model.
|
||||
assert len(llm.chat_calls) == 1
|
||||
user = next(m["content"] for m in llm.chat_calls[0] if m["role"] == "user")
|
||||
assert "TOP-TEXT-CONTENT" in user
|
||||
assert "SECRET-CONTENT" not in user and "IGNORED-TEXT-CONTENT" not in user
|
||||
# Exactly one summary call per kept file (keep.md + top.txt) — the
|
||||
# ignored files never reached the lite model.
|
||||
assert len(llm.chat_calls) == 2
|
||||
users = [
|
||||
next(m["content"] for m in msgs if m["role"] == "user")
|
||||
for msgs in llm.chat_calls
|
||||
]
|
||||
assert any("TOP-TEXT-CONTENT" in u for u in users)
|
||||
assert any("kept body" in u for u in users)
|
||||
for u in users:
|
||||
assert "SECRET-CONTENT" not in u and "IGNORED-TEXT-CONTENT" not in u
|
||||
top = next(d for d in docs if d.path == "top.txt")
|
||||
assert top.summary is not None
|
||||
_reset(db)
|
||||
|
||||
@@ -62,7 +62,8 @@ def test_hidden_paths_not_indexed_by_default(db: Session, tmp_path: Path) -> Non
|
||||
assert summary.files == 1
|
||||
assert summary.added == 1
|
||||
assert summary.errors == 0
|
||||
assert summary.summaries == 0
|
||||
# Phase 118 (A2): the one visible md IS summarized.
|
||||
assert summary.summaries == 1
|
||||
assert summary.summary_errors == 0
|
||||
|
||||
docs = db.scalars(select(Document)).all()
|
||||
@@ -73,8 +74,9 @@ def test_hidden_paths_not_indexed_by_default(db: Session, tmp_path: Path) -> Non
|
||||
for t in texts:
|
||||
assert "HIDDEN-MD-CONTENT" not in t and "HIDDEN-YAML-VALUE" not in t
|
||||
assert "EXCLUDED-CONTENT" not in t
|
||||
# The hidden yaml never reached the lite model.
|
||||
assert not llm.chat_calls
|
||||
# Only the visible md reached the lite model (phase 118, A2) — the
|
||||
# hidden files were never walked.
|
||||
assert len(llm.chat_calls) == 1
|
||||
_reset(db)
|
||||
|
||||
|
||||
@@ -105,14 +107,14 @@ def test_hidden_paths_indexed_when_flag_on(db: Session, tmp_path: Path) -> None:
|
||||
assert not any(d.path == ".venv/junk.md" for d in docs)
|
||||
chunks = db.scalars(select(Chunk)).all()
|
||||
assert not any("EXCLUDED-CONTENT" in c.content for c in chunks)
|
||||
# The hidden md was embedded like any visible md (no summary — the
|
||||
# markdown path skips the lite model).
|
||||
# The hidden md was embedded like any visible md — AND summarized
|
||||
# like any other doc (phase 118, A2: markdown included).
|
||||
note = db.scalar(select(Document).where(Document.path == ".hidden/note.md"))
|
||||
assert note is not None and note.summary is None
|
||||
assert note is not None and note.summary is not None
|
||||
assert any("HIDDEN-MD-CONTENT" in c.content for c in chunks)
|
||||
# The hidden yaml went through the FULL non-markdown path (phase 30):
|
||||
# a stored summary plus one embedded is_summary chunk on top of the
|
||||
# content chunks.
|
||||
# The hidden yaml went through the full doc path (phase 30): a stored
|
||||
# summary plus one embedded is_summary chunk on top of the content
|
||||
# chunks.
|
||||
yaml_doc = db.scalar(select(Document).where(Document.path == ".hidden/data.yaml"))
|
||||
assert yaml_doc is not None and yaml_doc.summary is not None
|
||||
yaml_chunks = [
|
||||
@@ -121,11 +123,16 @@ def test_hidden_paths_indexed_when_flag_on(db: Session, tmp_path: Path) -> None:
|
||||
assert any(c.is_summary for c in yaml_chunks)
|
||||
assert any(not c.is_summary for c in yaml_chunks)
|
||||
assert any("HIDDEN-YAML-VALUE" in c.content for c in yaml_chunks)
|
||||
# Only the yaml reached the lite model.
|
||||
assert summary.summaries == 1
|
||||
assert len(llm.chat_calls) == 1
|
||||
user = next(m["content"] for m in llm.chat_calls[0] if m["role"] == "user")
|
||||
assert "HIDDEN-YAML-VALUE" in user
|
||||
# EVERY doc reached the lite model (phase 118, A2: markdown too).
|
||||
assert summary.summaries == 3
|
||||
assert len(llm.chat_calls) == 3
|
||||
users = [
|
||||
next(m["content"] for m in msgs if m["role"] == "user")
|
||||
for msgs in llm.chat_calls
|
||||
]
|
||||
assert any("HIDDEN-YAML-VALUE" in u for u in users)
|
||||
assert any("HIDDEN-MD-CONTENT" in u for u in users)
|
||||
assert any("visible body" in u for u in users)
|
||||
_reset(db)
|
||||
|
||||
|
||||
|
||||
@@ -579,7 +579,8 @@ def test_limit_run_skips_folder_generation(
|
||||
capsys: pytest.CaptureFixture[str],
|
||||
) -> None:
|
||||
"""A ``--limit`` debug run (an incomplete walk) never generates —
|
||||
no ``lite`` call at all, no rows, the line says ``skipped``."""
|
||||
no overview/folder ``lite`` call, no rows, the line says
|
||||
``skipped``."""
|
||||
llm = FakeEmbedder()
|
||||
rc, out = _run_main(monkeypatch, llm, ["--source", str(src), "--limit", "2"], capsys)
|
||||
assert rc == 0
|
||||
@@ -587,7 +588,10 @@ def test_limit_run_skips_folder_generation(
|
||||
assert out.rstrip().endswith(
|
||||
"overview=skipped sources_version=skipped folder_summaries=skipped"
|
||||
)
|
||||
assert llm.chat_calls == [] # no lite call, any mode
|
||||
# Phase 118 (A2): the two walked docs still burn their document
|
||||
# summaries — but the folder generation never fires (and neither
|
||||
# does the overview).
|
||||
assert len(llm.chat_calls) == 2 # doc summaries only, any other mode absent
|
||||
assert _rows(db) == {} # an incomplete walk never writes rows
|
||||
|
||||
|
||||
|
||||
+169
-66
@@ -221,50 +221,71 @@ def test_agent_tools_names_and_parameters() -> None:
|
||||
"'homelab/active'). Omit it to list every source."
|
||||
)
|
||||
read = by_name["read"]["function"]
|
||||
# Tool-calling fast loop (2026-09-04, controlled fixture gate):
|
||||
# the do-not-read rule is FRONT-LOADED — the controlled gate's
|
||||
# telemetry showed the `lite` model obeying the user's "open it /
|
||||
# read it" and reading seed-context documents the <documents>
|
||||
# section already carries (every refusal of a 12-call run was
|
||||
# ALREADY_IN_CONTEXT); the rule now leads the description instead
|
||||
# of sitting mid-paragraph, and the tool is framed as "only for
|
||||
# documents NOT already in <documents>". Phase 95 (task 01): the
|
||||
# read-truncation sentence is inserted before the one-call-at-a-
|
||||
# time discipline clause (the discipline rule stays last, as in the
|
||||
# other two tools) — a capped read carries the TRUNCATED notice and
|
||||
# the `grep` follow-up (the pinned copy).
|
||||
# Phase 118 (task 04, A6): the read description is rewritten for
|
||||
# the summary-seed mode — the <documents> section shows the
|
||||
# top-ranked documents' SUMMARIES (their full texts are NOT in
|
||||
# the prompt yet); read adds one of them (or any other document)
|
||||
# to the context; only an already-read document is refused. The
|
||||
# combined-identity handoff to the `ls` output / the <documents>
|
||||
# blocks stays; the truncation-notice paragraph (phase 95) and
|
||||
# the one-call-at-a-time discipline clause (task 05) survive the
|
||||
# rewrite byte-identical.
|
||||
assert read["description"] == (
|
||||
"Do not call this tool for a document already shown in "
|
||||
"the <documents> section, even when the user asks you to "
|
||||
"open or read it — its full text is already in your "
|
||||
"prompt; answer directly from it. Use it only to add a "
|
||||
"document NOT already in <documents> to your context, "
|
||||
"by its combined `source/path` string. Very large "
|
||||
"documents are truncated: you receive the first part "
|
||||
"plus a TRUNCATED notice naming how many more characters "
|
||||
"exist — the notice is authoritative, the document did "
|
||||
"NOT end where it stopped. Follow it and use `grep` "
|
||||
"(pattern) to locate the rest — it searches the whole "
|
||||
"document. Call one tool at a time — wait for this "
|
||||
"result before your next call."
|
||||
"The <documents> section shows the SUMMARIES of the "
|
||||
"top-ranked documents — their full texts are NOT in "
|
||||
"your prompt yet. Use this tool to add one of them (or "
|
||||
"any other document) to your context, by its combined "
|
||||
"`source/path` string, exactly as shown in the `ls` "
|
||||
"output or the <documents> blocks. Do not re-read a "
|
||||
"document you have already read — its full text is "
|
||||
"already in your prompt. Very large documents are "
|
||||
"truncated: you receive the first part plus a TRUNCATED "
|
||||
"notice naming how many more characters exist — the "
|
||||
"notice is authoritative, the document did NOT end "
|
||||
"where it stopped. Follow it and use `grep` (pattern) "
|
||||
"to locate the rest — it searches the whole document. "
|
||||
"Call one tool at a time — wait for this result before "
|
||||
"your next call."
|
||||
)
|
||||
# The byte-preserved contracts inside the rewrite, pinned as
|
||||
# substrings: the combined source/path identity, the
|
||||
# truncation-notice paragraph, the one-call-at-a-time
|
||||
# discipline sentence.
|
||||
assert (
|
||||
"by its combined `source/path` string, exactly as shown in "
|
||||
"the `ls` output or the <documents> blocks"
|
||||
) in read["description"]
|
||||
assert (
|
||||
"Very large documents are truncated: you receive the first "
|
||||
"part plus a TRUNCATED notice naming how many more "
|
||||
"characters exist — the notice is authoritative, the "
|
||||
"document did NOT end where it stopped. Follow it and use "
|
||||
"`grep` (pattern) to locate the rest — it searches the "
|
||||
"whole document."
|
||||
) in read["description"]
|
||||
assert (
|
||||
"Call one tool at a time — wait for this result before "
|
||||
"your next call."
|
||||
) in read["description"]
|
||||
read_params = read["parameters"]
|
||||
assert read_params["type"] == "object"
|
||||
assert read_params["required"] == ["path"]
|
||||
assert set(read_params["properties"]) == {"path"}
|
||||
assert read_params["properties"]["path"]["type"] == "string"
|
||||
# The combined source/path string is the canonical document identity
|
||||
# (phase 70) — the description pins it with a worked example. Phase
|
||||
# 72 (task 02): the bare-path contract is stated up front; task 05
|
||||
# (live gate iteration 1): the do-not-re-read clause (the dedupe
|
||||
# refusal's prevention at the prompt).
|
||||
# (phase 70) — the description pins it with a worked example; the
|
||||
# bare-path contract is stated up front (phase 72, task 02). Phase
|
||||
# 118 (task 04, A6): the trailing clause is the do-not-RE-READ
|
||||
# teaching — the seeds are summary blocks, not full text, so a
|
||||
# first read of a suggested document succeeds and only an
|
||||
# already-read document is refused.
|
||||
assert read_params["properties"]["path"]["description"] == (
|
||||
"The document to add to your context, as the combined "
|
||||
"`source/path` string exactly as shown in the `ls` output (e.g. "
|
||||
"'homelab/active/container_caddy/caddy.md'). A bare document "
|
||||
"path (without the source name) will not resolve. Only pass a "
|
||||
"document NOT already shown in the <documents> section — it is "
|
||||
"already in your context; do not re-read it."
|
||||
"path (without the source name) will not resolve. Do not "
|
||||
"re-read a document you have already read — its full "
|
||||
"text is already in your prompt."
|
||||
)
|
||||
grep = by_name["grep"]["function"]
|
||||
# Task 05 (live gate iterations 2-6, refined in the 2026-09-03
|
||||
@@ -1455,15 +1476,17 @@ def test_read_bare_filename_without_slash_keeps_no_db_refusal(
|
||||
assert llm.requests[1][1] == AGENT_TOOLS
|
||||
|
||||
|
||||
def test_read_bare_path_of_seed_doc_gets_suggestion_then_dedupe(
|
||||
def test_read_bare_path_of_suggested_doc_gets_suggestion_then_succeeds(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""Dedupe precedence: the in-context dedupe fires on the SPLIT pair
|
||||
of the argument — the bare path of an in-context document
|
||||
(``read('app/rag/importer.py')`` with ``sample/app/rag/importer.py``
|
||||
seeded) is NOT that pair, so it is not a dedupe: it gets the
|
||||
suggestion line naming the combined identity, and the model's next,
|
||||
correctly-formed call is then deduped as ALREADY_IN_CONTEXT."""
|
||||
"""Dedupe precedence (phase 118): the in-context dedupe fires on
|
||||
the SPLIT pair of a document already READ — the bare path of a
|
||||
SUGGESTED (seeded) document (``read('app/rag/importer.py')`` with
|
||||
``sample/app/rag/importer.py`` seeded) is NOT that pair, so it is
|
||||
not a dedupe: it gets the suggestion line naming the combined
|
||||
identity, and the model's next, correctly-formed call ADDS the
|
||||
suggested document's full text (A6: the seeds are summary blocks,
|
||||
not full text — a first read of a suggested document succeeds)."""
|
||||
seed = [_doc("sample", "app/rag/importer.py", "Importer", "IMPORTER")]
|
||||
|
||||
def _find(db: Any, source: str, path: str) -> Document | None:
|
||||
@@ -1482,8 +1505,9 @@ def test_read_bare_path_of_seed_doc_gets_suggestion_then_dedupe(
|
||||
],
|
||||
[
|
||||
# Round 2: the corrected call (the suggested combined
|
||||
# identity) — the seed document is already in context, so it
|
||||
# dedupes.
|
||||
# identity) — the suggested document's full text is NOT in
|
||||
# the prompt (only its summary is), so this read succeeds
|
||||
# and adds the full content.
|
||||
ToolCallPiece(
|
||||
id="call_2",
|
||||
name="read",
|
||||
@@ -1493,12 +1517,18 @@ def test_read_bare_path_of_seed_doc_gets_suggestion_then_dedupe(
|
||||
[StreamPiece("content", "ans")],
|
||||
)
|
||||
asyncio.run(_run(llm, holder, _settings(), seed_docs=seed))
|
||||
assert holder.read_docs == [] and holder.tool_calls == 0 # both refused
|
||||
# Round 1 refused (the teaching), round 2 executed (A6).
|
||||
assert holder.read_docs == [seed[0]]
|
||||
assert holder.tool_calls == 1
|
||||
assert llm.requests[1][0][3]["content"] == (
|
||||
"No document at 'app/rag/importer.py' — "
|
||||
"did you mean 'sample/app/rag/importer.py'?"
|
||||
)
|
||||
assert llm.requests[2][0][5]["content"] == agent.ALREADY_IN_CONTEXT
|
||||
assert llm.requests[2][0][5]["content"] == (
|
||||
"Document sample/app/rag/importer.py:\n"
|
||||
"date: 2024-06-15\n"
|
||||
"IMPORTER"
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
@@ -1529,16 +1559,23 @@ def test_read_missing_arguments_refused(
|
||||
assert llm.requests[1][1] == AGENT_TOOLS
|
||||
|
||||
|
||||
def test_reading_a_seed_doc_is_already_in_context(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""The combined identity of a seeded document: its split pair is in
|
||||
the known set → ALREADY_IN_CONTEXT with no DB lookup (the dedupe
|
||||
check precedes the resolve)."""
|
||||
def test_reading_a_suggested_seed_doc_adds_its_full_text(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""Phase 118 (A6): a FIRST read of a suggested (seeded) document
|
||||
SUCCEEDS — the seed is a summary block in the prompt, not the full
|
||||
text, so it falls out of the dedupe set (``holder.read_docs``
|
||||
only): the read goes through the existing path unchanged — the
|
||||
full-content result (header + date line), appended to
|
||||
``holder.read_docs``, counted in ``holder.tool_calls``."""
|
||||
seed = [_doc("Homelab", "kubernetes.md", "Kubernetes", "K8S-CONTENT")]
|
||||
|
||||
def _boom(*_a: Any, **_k: Any) -> None:
|
||||
raise AssertionError("find_document must not be called for a seeded doc")
|
||||
|
||||
monkeypatch.setattr(agent, "find_document", _boom)
|
||||
monkeypatch.setattr(
|
||||
agent,
|
||||
"find_document",
|
||||
lambda db, source, path: seed[0]
|
||||
if (source, path) == ("Homelab", "kubernetes.md")
|
||||
else None,
|
||||
)
|
||||
holder = AgentHolder()
|
||||
llm = ScriptedLLM(
|
||||
[
|
||||
@@ -1551,11 +1588,58 @@ def test_reading_a_seed_doc_is_already_in_context(monkeypatch: pytest.MonkeyPatc
|
||||
[StreamPiece("content", "ans")],
|
||||
)
|
||||
asyncio.run(_run(llm, holder, _settings(), seed_docs=seed))
|
||||
assert holder.read_docs == [] and holder.tool_calls == 0
|
||||
assert llm.requests[1][0][3]["content"] == agent.ALREADY_IN_CONTEXT
|
||||
assert holder.read_docs == [seed[0]]
|
||||
assert holder.tool_calls == 1
|
||||
assert llm.requests[1][0][3]["content"] == (
|
||||
"Document Homelab/kubernetes.md:\ndate: 2024-06-15\nK8S-CONTENT"
|
||||
)
|
||||
assert llm.requests[1][1] == AGENT_TOOLS
|
||||
|
||||
|
||||
def test_re_reading_a_suggested_doc_is_deduped(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""Phase 118 (A6): once the suggested document's full text has been
|
||||
read into context, a SECOND read of the same combined identity is
|
||||
refused with the byte-identical ``ALREADY_IN_CONTEXT`` line — the
|
||||
dedupe set is ``holder.read_docs`` only, so the refusal fires with
|
||||
NO DB lookup (the dedupe check precedes the resolve) and the
|
||||
counters stay untouched."""
|
||||
seed = [_doc("Homelab", "kubernetes.md", "Kubernetes", "K8S-CONTENT")]
|
||||
lookups: list[tuple[str, str]] = []
|
||||
|
||||
def _find(db: Any, source: str, path: str) -> Document | None:
|
||||
lookups.append((source, path))
|
||||
return seed[0] if (source, path) == ("Homelab", "kubernetes.md") else None
|
||||
|
||||
monkeypatch.setattr(agent, "find_document", _find)
|
||||
holder = AgentHolder()
|
||||
llm = ScriptedLLM(
|
||||
[
|
||||
ToolCallPiece(
|
||||
id="call_1",
|
||||
name="read",
|
||||
arguments={"path": "Homelab/kubernetes.md"},
|
||||
)
|
||||
],
|
||||
[
|
||||
ToolCallPiece(
|
||||
id="call_2",
|
||||
name="read",
|
||||
arguments={"path": "Homelab/kubernetes.md"},
|
||||
)
|
||||
],
|
||||
[StreamPiece("content", "ans")],
|
||||
)
|
||||
asyncio.run(_run(llm, holder, _settings(), seed_docs=seed))
|
||||
assert holder.read_docs == [seed[0]] # appended exactly once
|
||||
assert holder.tool_calls == 1 # the re-read counts nothing
|
||||
assert lookups == [("Homelab", "kubernetes.md")] # the re-read deduped pre-lookup
|
||||
assert llm.requests[1][0][3]["content"] == (
|
||||
"Document Homelab/kubernetes.md:\ndate: 2024-06-15\nK8S-CONTENT"
|
||||
)
|
||||
assert llm.requests[2][0][5]["content"] == agent.ALREADY_IN_CONTEXT
|
||||
# Rejected → the tools are still offered on the next request (the
|
||||
# round cap is the only bound).
|
||||
assert llm.requests[1][1] == AGENT_TOOLS
|
||||
assert llm.requests[2][1] == AGENT_TOOLS
|
||||
|
||||
|
||||
def test_reading_an_already_read_doc_is_deduped(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
@@ -1718,28 +1802,47 @@ def test_run_agent_short_read_yields_no_tool_result_piece(
|
||||
assert isinstance(out[1], StreamPiece)
|
||||
|
||||
|
||||
def test_read_truncation_does_not_touch_refusal_paths(
|
||||
def test_read_truncation_paths_with_suggested_seed(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""Phase 95: the read refusal paths are untouched by the cap — a SEED
|
||||
document that is over the cap is still refused with
|
||||
``ALREADY_IN_CONTEXT`` (not truncated, nothing recorded, nothing
|
||||
counted), and an unknown path is still the no-document refusal (no
|
||||
"""Phase 118 (A6) re-target of the phase-95 pins ("seed" →
|
||||
"suggested"): a SUGGESTED (seeded) document that is over the cap is
|
||||
read through the ordinary cap path — cut at the cap + the shared
|
||||
marker + the pinned notice, the truncation recorded on the holder,
|
||||
and the call still successful (the seed is a summary, not full
|
||||
text) — and an unknown path is still the no-document refusal (no
|
||||
content is read, so no truncation either)."""
|
||||
big = "B" * 5000 # far over the tiny cap below
|
||||
seed = _doc("S", "seed.md", "Seed", big)
|
||||
monkeypatch.setattr(agent, "find_document", lambda db, source, path: None)
|
||||
monkeypatch.setattr(agent, "all_documents", lambda db: [])
|
||||
# (a) Reading the (over-cap) seed doc → ALREADY_IN_CONTEXT (refusal).
|
||||
monkeypatch.setattr(
|
||||
agent,
|
||||
"find_document",
|
||||
lambda db, source, path: seed
|
||||
if (source, path) == ("S", "seed.md")
|
||||
else None,
|
||||
)
|
||||
monkeypatch.setattr(agent, "all_documents", lambda db: [seed])
|
||||
# (a) Reading the (over-cap) SUGGESTED doc → the capped read: cut at
|
||||
# the cap, marker + pinned notice, holder entry, still counted.
|
||||
holder = AgentHolder()
|
||||
llm = ScriptedLLM(
|
||||
[ToolCallPiece(id="call_1", name="read", arguments={"path": "S/seed.md"})],
|
||||
[StreamPiece("content", "ans")],
|
||||
)
|
||||
asyncio.run(_run(llm, holder, _settings(read_max_chars=100), seed_docs=[seed]))
|
||||
assert llm.requests[1][0][3]["content"] == agent.ALREADY_IN_CONTEXT
|
||||
assert holder.read_truncations == []
|
||||
assert holder.tool_calls == 0 and holder.read_docs == []
|
||||
assert llm.requests[1][0][3]["content"] == (
|
||||
"Document S/seed.md:\n"
|
||||
"date: 2024-06-15\n"
|
||||
+ big[:100]
|
||||
+ "\n"
|
||||
+ TRUNCATION_MARKER
|
||||
+ "\n"
|
||||
+ READ_TRUNCATION_NOTICE.format(shown=100, total=5000)
|
||||
)
|
||||
# (argument, cap, total) — the raw argument, the cap kept, the
|
||||
# true length.
|
||||
assert holder.read_truncations == [("S/seed.md", 100, 5000)]
|
||||
assert holder.tool_calls == 1 and holder.read_docs == [seed]
|
||||
# (b) An unknown path → the no-document refusal (argument echoed),
|
||||
# even though a big doc could have truncated — no content is read.
|
||||
holder2 = AgentHolder()
|
||||
|
||||
+264
-120
@@ -8,6 +8,7 @@ session, and the LLM all faked, so the whole deflection contract
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import uuid
|
||||
from collections.abc import Iterator
|
||||
@@ -23,6 +24,7 @@ from app.main import app as fastapi_app
|
||||
from app.models import Document, KbOverview, QueryLog
|
||||
from app.rag.agent import AGENT_TOOLS
|
||||
from app.rag.llm import StreamPiece
|
||||
from app.rag.prompts import build_deflect_prompt
|
||||
from app.rag.retriever import RetrievedChunk, weak_hit_titles
|
||||
from app.rag.suggestions import MAX_SUGGESTIONS, derive_suggestions
|
||||
from tests.conftest import ADMIN_PASSWORD
|
||||
@@ -46,7 +48,9 @@ def _settings(threshold: float = 0.30, floor: float | None = None) -> Settings:
|
||||
)
|
||||
|
||||
|
||||
def _doc(title: str, content: str) -> Document:
|
||||
def _doc(
|
||||
title: str, content: str, summary: str | None = None
|
||||
) -> Document:
|
||||
return Document(
|
||||
id=uuid.uuid4(),
|
||||
source="Homelab",
|
||||
@@ -54,6 +58,10 @@ def _doc(title: str, content: str) -> Document:
|
||||
full_path="/tmp/doc.md",
|
||||
title=title,
|
||||
content=content,
|
||||
# Phase 30/118: the stored lite-model summary — the HIGH block's
|
||||
# BODY (task 03). ``None`` exercises the A5 preview fallback
|
||||
# (the first ``suggestion_preview_chars`` content chars).
|
||||
summary=summary,
|
||||
content_hash="0" * 64,
|
||||
# Phase 106, D5: the HIGH block formats the row's created_at
|
||||
# UTC date part — the detached fixture carries it (the NOT NULL
|
||||
@@ -179,11 +187,12 @@ def test_gate_lexical_only_chunk_does_not_inflate_cosine() -> None:
|
||||
plan = chat_api.plan_turn(chunks, _settings(threshold=0.30))
|
||||
assert plan.top_score == pytest.approx(0.55)
|
||||
assert plan.deflected is False # 0.55 >= 0.30 anyway
|
||||
# Phase 113 (the usefulness bar): Beta's doc ranks first by fused
|
||||
# score, but a lexical-only doc (cosine 0.0 by construction) cannot
|
||||
# clear the bar — it lands in the RELATED tier, never the cited one.
|
||||
assert plan.docs[0].title == "Alpha"
|
||||
assert plan.related_docs[0].title == "Beta"
|
||||
# Phase 118 (A3): the suggestion tier has NO floor — the lexical-
|
||||
# only doc (cosine 0.0 by construction) is SUGGESTED when it ranks.
|
||||
# Rank order is the fused score, so Beta (0.90) leads Alpha (0.50);
|
||||
# with two docs nothing is left for the related tier (rank 6+).
|
||||
assert [d.title for d in plan.suggested_docs] == ["Beta", "Alpha"]
|
||||
assert plan.related_docs == []
|
||||
|
||||
|
||||
# ---------- lexical support floor (A8 revised 2026-09-14) ----------
|
||||
@@ -373,178 +382,225 @@ def test_gate_zero_chunks_deflects_with_fallback_chips() -> None:
|
||||
assert 2 <= len(plan.suggestions) <= MAX_SUGGESTIONS
|
||||
|
||||
|
||||
# ---------- usefulness bar tiering (phase 113, LOCKED A2/A4) ----------
|
||||
# ---------- summary-seed tiering (phase 118, LOCKED A3/A4/A6) ----------
|
||||
|
||||
|
||||
def _bar_settings(
|
||||
def _tier_settings(
|
||||
threshold: float = 0.62,
|
||||
lex_floor: float = 0.35,
|
||||
source_floor: float = 0.35,
|
||||
suggested_cap: int = 5,
|
||||
related_cap: int = 2,
|
||||
top_n: int = 2,
|
||||
) -> Settings:
|
||||
"""Explicit code defaults (production calibration) — the env's mock-
|
||||
calibrated floor (tests/conftest.py) is overridden per test."""
|
||||
calibrated values (tests/conftest.py) are overridden per test.
|
||||
``source_usefulness_floor`` / ``top_n_docs`` are deliberately left
|
||||
at their code defaults: phase 118 retired their seeding role (A6) —
|
||||
``plan_turn`` never consults them (pinned in
|
||||
``test_plan_turn_does_not_consult_retired_seeding_settings``)."""
|
||||
return Settings(
|
||||
_env_file=None, # pyright: ignore[reportCallIssue]
|
||||
relevance_threshold=threshold,
|
||||
lexical_support_floor=lex_floor,
|
||||
source_usefulness_floor=source_floor,
|
||||
suggested_docs=suggested_cap,
|
||||
related_max_docs=related_cap,
|
||||
top_n_docs=top_n,
|
||||
)
|
||||
|
||||
|
||||
def test_plan_turn_high_tiers_strong_plus_weak() -> None:
|
||||
"""Grounded turn: the bar-clearing doc is cited (and in the prompt),
|
||||
the weak 2nd doc loses its citation slot and lands in related_docs —
|
||||
the recurring incident's fix at the plan level."""
|
||||
def _seven_docs_with_summaries() -> list[Document]:
|
||||
"""The 7-doc fixture (rank 1–7 by fused score): every document has a
|
||||
stored SUMMARY (distinct sentinel) and a distinct FULL-CONTENT
|
||||
sentinel that must never reach the prompt (A6: the ``read`` tool is
|
||||
the only full-text path)."""
|
||||
return [
|
||||
_doc(
|
||||
f"Doc {i}",
|
||||
f"FULL_CONTENT_SENTINEL_{i}_SHOULD_NEVER_REACH_THE_PROMPT",
|
||||
summary=f"SUMMARY_TEXT_{i}",
|
||||
)
|
||||
for i in range(7)
|
||||
]
|
||||
|
||||
|
||||
def test_plan_turn_high_seeds_top5_suggested_related_is_rank6plus() -> None:
|
||||
"""The phase-118 core pin (LOCKED A3/A6): the HIGH prompt seeds
|
||||
exactly the top-5 suggested documents' SUMMARY text and NONE of
|
||||
their full content; the related tier is rank 6+ (docs 6–7, capped
|
||||
by ``related_max_docs``)."""
|
||||
docs_in = _seven_docs_with_summaries()
|
||||
chunks = [
|
||||
_chunk(d, 0.9 - 0.1 * i, cosine=0.8 - 0.05 * i) for i, d in enumerate(docs_in)
|
||||
]
|
||||
plan = chat_api.plan_turn(chunks, _tier_settings())
|
||||
assert plan.deflected is False
|
||||
assert [d.title for d in plan.suggested_docs] == [f"Doc {i}" for i in range(5)]
|
||||
assert [d.title for d in plan.related_docs] == ["Doc 5", "Doc 6"]
|
||||
# The prompt seeds the five summaries …
|
||||
for i in range(5):
|
||||
assert f"SUMMARY_TEXT_{i}" in plan.system_prompt
|
||||
# … and NONE of the seven documents' full content (suggested OR
|
||||
# related) reaches the LLM (A6).
|
||||
for i in range(7):
|
||||
assert f"FULL_CONTENT_SENTINEL_{i}" not in plan.system_prompt
|
||||
|
||||
|
||||
def test_plan_turn_high_single_strong_doc_yields_one_suggested() -> None:
|
||||
"""The suggested cap is a CEILING, not a quota: one doc ⇒ one
|
||||
suggested doc, an empty related tier (nothing beyond rank 1)."""
|
||||
strong = _doc("Kubernetes Homelab Cluster", "STRONG_DOC_CONTENT")
|
||||
plan = chat_api.plan_turn([_chunk(strong, 0.90, cosine=0.80)], _tier_settings())
|
||||
assert plan.deflected is False
|
||||
assert [d.title for d in plan.suggested_docs] == ["Kubernetes Homelab Cluster"]
|
||||
assert plan.related_docs == []
|
||||
# No stored summary (the fixture default) → the A5 preview fallback
|
||||
# carries the short content whole (under the 400-char cap).
|
||||
assert "STRONG_DOC_CONTENT" in plan.system_prompt
|
||||
|
||||
|
||||
def test_plan_turn_high_suggests_strong_and_weak_no_floor() -> None:
|
||||
"""The recurring incident under phase 118 (A3): the weak 2nd doc no
|
||||
longer loses a citation slot to a bar — the floor never filters, so
|
||||
BOTH docs are suggested (rank order) and ride the citation surface
|
||||
(A4); the HIGH prompt seeds both summaries (the A5 fallback carries
|
||||
the short fixture content whole)."""
|
||||
strong = _doc("Kubernetes Homelab Cluster", "STRONG_DOC_CONTENT")
|
||||
weak = _doc("Backup Strategy", "WEAK_DOC_CONTENT")
|
||||
chunks = [_chunk(strong, 0.90, cosine=0.80), _chunk(weak, 0.80, cosine=0.20)]
|
||||
plan = chat_api.plan_turn(chunks, _bar_settings())
|
||||
plan = chat_api.plan_turn(chunks, _tier_settings())
|
||||
assert plan.deflected is False
|
||||
assert [d.title for d in plan.docs] == ["Kubernetes Homelab Cluster"]
|
||||
assert [d.title for d in plan.related_docs] == ["Backup Strategy"]
|
||||
# The HIGH prompt carries the cited doc's content only.
|
||||
assert "STRONG_DOC_CONTENT" in plan.system_prompt
|
||||
assert "WEAK_DOC_CONTENT" not in plan.system_prompt
|
||||
|
||||
|
||||
def test_plan_turn_high_single_strong_doc_yields_one_cited() -> None:
|
||||
"""top_n_docs is a CEILING, not a quota: one strong doc ⇒ one cited doc,
|
||||
an empty related tier (LOCKED A2)."""
|
||||
strong = _doc("Kubernetes Homelab Cluster", "STRONG_DOC_CONTENT")
|
||||
plan = chat_api.plan_turn([_chunk(strong, 0.90, cosine=0.80)], _bar_settings())
|
||||
assert plan.deflected is False
|
||||
assert [d.title for d in plan.docs] == ["Kubernetes Homelab Cluster"]
|
||||
assert [d.title for d in plan.suggested_docs] == [
|
||||
"Kubernetes Homelab Cluster",
|
||||
"Backup Strategy",
|
||||
]
|
||||
assert plan.related_docs == []
|
||||
assert "STRONG_DOC_CONTENT" in plan.system_prompt
|
||||
assert "WEAK_DOC_CONTENT" in plan.system_prompt
|
||||
|
||||
|
||||
def test_plan_turn_low_weak_hits_fall_to_related() -> None:
|
||||
"""Deflected turn: nothing clears the bar ⇒ the cited tier is empty
|
||||
and the weak hits fall to related_docs (the done frame's home for
|
||||
their visibility). The LOW prompt is unchanged (titles only)."""
|
||||
def test_plan_turn_low_weak_hits_are_suggested_record() -> None:
|
||||
"""Deflected turn: the weak hits are SUGGESTED too (no floor, A3) —
|
||||
the TurnPlan carries suggested + related for the durable record —
|
||||
while the LOW prompt itself stays byte-identical (weak-hit titles
|
||||
only, never content)."""
|
||||
a = _doc("Alpha", "ALPHA_DOC_NEVER_SENT")
|
||||
b = _doc("Beta", "BETA_DOC_NEVER_SENT")
|
||||
chunks = [_chunk(a, 0.30, cosine=0.20), _chunk(b, 0.20, cosine=0.15)]
|
||||
plan = chat_api.plan_turn(chunks, _bar_settings())
|
||||
plan = chat_api.plan_turn(chunks, _tier_settings())
|
||||
assert plan.deflected is True
|
||||
assert plan.docs == [] # no citation slot below the bar
|
||||
assert [d.title for d in plan.related_docs] == ["Alpha", "Beta"] # rank order
|
||||
assert [d.title for d in plan.suggested_docs] == ["Alpha", "Beta"] # rank order
|
||||
assert plan.related_docs == [] # nothing beyond rank 2 for 2 docs
|
||||
assert "ALPHA_DOC_NEVER_SENT" not in plan.system_prompt
|
||||
assert "Beta" in plan.system_prompt # weak-hit titles still carried
|
||||
assert plan.suggestions # chips unchanged
|
||||
|
||||
|
||||
def test_plan_turn_related_cap_zero_kills_the_related_tier() -> None:
|
||||
"""related_max_docs=0 is the kill switch: weak docs are scored but
|
||||
neither cited nor related (the pre-phase-113 visibility, minus the
|
||||
false citation — a deflected turn cites nothing)."""
|
||||
a = _doc("Alpha", "AAA")
|
||||
b = _doc("Beta", "BBB")
|
||||
chunks = [_chunk(a, 0.30, cosine=0.20), _chunk(b, 0.20, cosine=0.15)]
|
||||
plan = chat_api.plan_turn(chunks, _bar_settings(related_cap=0))
|
||||
assert plan.deflected is True
|
||||
assert plan.docs == []
|
||||
assert plan.related_docs == []
|
||||
|
||||
|
||||
def test_plan_turn_floor_zero_keeps_legacy_cited_docs() -> None:
|
||||
"""source_usefulness_floor=0 disables the bar: plan.docs is the legacy
|
||||
rank-ordered top-N (any cosine, incl. 0.0 lexical-only) and the
|
||||
related tier is empty."""
|
||||
a = _doc("Alpha", "AAA")
|
||||
b = _doc("Beta", "BBB")
|
||||
chunks = [
|
||||
_chunk(a, 0.90, cosine=0.0, fts_hit=True), # lexical-only, rank 1
|
||||
_chunk(b, 0.80, cosine=0.10),
|
||||
]
|
||||
plan = chat_api.plan_turn(
|
||||
chunks, _bar_settings(source_floor=0.0, lex_floor=0.05)
|
||||
)
|
||||
assert plan.deflected is False # 0.10 + the fts hit clears the 0.05 lex floor
|
||||
assert [d.title for d in plan.docs] == ["Alpha", "Beta"] # legacy order
|
||||
assert plan.related_docs == []
|
||||
|
||||
|
||||
def test_plan_turn_lexically_grounded_below_source_floor_has_no_cited_docs() -> None:
|
||||
"""The degenerate operator config (citation bar STRICTER than the
|
||||
grounding bar): a turn grounded by a corroborated-lexical hit whose
|
||||
cosine sits between the two floors has an EMPTY cited tier — the HIGH
|
||||
prompt carries no document content (the tools remain the escape
|
||||
hatch). The bar is a citation filter, not a gate input."""
|
||||
doc = _doc("Static DNS", "DNS_DOC_CONTENT")
|
||||
plan = chat_api.plan_turn(
|
||||
[_chunk(doc, 0.50, cosine=0.35, fts_hit=True)],
|
||||
_bar_settings(threshold=0.62, lex_floor=0.30, source_floor=0.50),
|
||||
)
|
||||
assert plan.deflected is False # 0.35 >= lex floor 0.30, fts fired
|
||||
assert plan.docs == [] # 0.35 < source floor 0.50 — no citation slot
|
||||
assert "DNS_DOC_CONTENT" not in plan.system_prompt
|
||||
# The doc still SCORED — it rides the related tier (the "nearby docs"
|
||||
# row), it is not invisible.
|
||||
assert [d.title for d in plan.related_docs] == ["Static DNS"]
|
||||
|
||||
|
||||
def test_plan_turn_related_tier_capped_in_rank_order() -> None:
|
||||
"""Grounded turn, four bar-clearing docs, ceiling 2: cited = the top-2
|
||||
in rank order; related = the next two (the ceiling overflow, any
|
||||
cosine), capped at related_max_docs."""
|
||||
docs_in = [
|
||||
_doc(f"Doc {i}", f"DOC_CONTENT_{i}") for i in range(4)
|
||||
]
|
||||
"""related_max_docs=0 is the kill switch: rank-6+ docs are scored
|
||||
and suggested-adjacent but neither suggested nor related — the
|
||||
done frame's row stays empty."""
|
||||
docs_in = _seven_docs_with_summaries()
|
||||
chunks = [
|
||||
_chunk(d, 0.9 - 0.1 * i, cosine=0.8 - 0.05 * i) for i, d in enumerate(docs_in)
|
||||
]
|
||||
plan = chat_api.plan_turn(chunks, _bar_settings(top_n=2, related_cap=2))
|
||||
plan = chat_api.plan_turn(chunks, _tier_settings(related_cap=0))
|
||||
assert plan.deflected is False
|
||||
assert [d.title for d in plan.docs] == ["Doc 0", "Doc 1"]
|
||||
assert [d.title for d in plan.related_docs] == ["Doc 2", "Doc 3"]
|
||||
assert len(plan.suggested_docs) == 5 # the suggestion tier is untouched
|
||||
assert plan.related_docs == []
|
||||
# The cap restores the rank-6+ row (rank order, capped).
|
||||
wide = chat_api.plan_turn(chunks, _tier_settings(related_cap=3))
|
||||
assert [d.title for d in wide.related_docs] == ["Doc 5", "Doc 6"]
|
||||
|
||||
|
||||
def test_plan_turn_suggested_docs_setting_caps_the_suggested_tier() -> None:
|
||||
"""``suggested_docs`` (``BOR_SUGGESTED_DOCS``) is honored as the
|
||||
suggestion cap: 3 here ⇒ the top-3 rank-ordered docs are suggested
|
||||
and rank 4+ falls to the related tier (capped at ``related_max_docs``)."""
|
||||
docs_in = _seven_docs_with_summaries()
|
||||
chunks = [
|
||||
_chunk(d, 0.9 - 0.1 * i, cosine=0.8 - 0.05 * i) for i, d in enumerate(docs_in)
|
||||
]
|
||||
plan = chat_api.plan_turn(chunks, _tier_settings(suggested_cap=3))
|
||||
assert [d.title for d in plan.suggested_docs] == ["Doc 0", "Doc 1", "Doc 2"]
|
||||
assert [d.title for d in plan.related_docs] == ["Doc 3", "Doc 4"] # cap 2
|
||||
|
||||
|
||||
def test_plan_turn_does_not_consult_retired_seeding_settings() -> None:
|
||||
"""Phase 118 (A6): ``top_n_docs`` and ``source_usefulness_floor``
|
||||
lost their seeding role — ``plan_turn`` never consults them. A
|
||||
degenerate config (the maximum legal bar — 0.62, above every chunk's
|
||||
cosine of 0.50 — and a top-N of 1) changes nothing: the suggested
|
||||
tier is still the no-floor top-5 in rank order and the related tier
|
||||
is still rank 6+.
|
||||
(Behavioral pin — the settings themselves stay, env back-compat.)"""
|
||||
docs_in = _seven_docs_with_summaries()
|
||||
chunks = [
|
||||
_chunk(docs_in[0], 0.9, cosine=0.50, fts_hit=True),
|
||||
] + [
|
||||
_chunk(d, 0.9 - 0.1 * i, cosine=0.50)
|
||||
for i, d in enumerate(docs_in[1:], start=1)
|
||||
]
|
||||
settings = Settings(
|
||||
_env_file=None, # pyright: ignore[reportCallIssue]
|
||||
relevance_threshold=0.62,
|
||||
lexical_support_floor=0.35,
|
||||
top_n_docs=1, # retired: the old full-text seeding ceiling
|
||||
# retired: the maximum legal bar (== threshold) — above every
|
||||
# cosine here (0.50), so the OLD tiering would cite nothing.
|
||||
source_usefulness_floor=0.62,
|
||||
related_max_docs=2,
|
||||
)
|
||||
plan = chat_api.plan_turn(chunks, settings)
|
||||
assert plan.deflected is False # 0.50 >= the 0.35 lex floor, fts fired
|
||||
assert [d.title for d in plan.suggested_docs] == [f"Doc {i}" for i in range(5)]
|
||||
assert [d.title for d in plan.related_docs] == ["Doc 5", "Doc 6"]
|
||||
|
||||
|
||||
# ---------- summary hits (phase 30: summary → full source document) ----------
|
||||
|
||||
|
||||
def test_summary_hit_on_selected_top_doc_counts() -> None:
|
||||
"""HIGH branch: the top document was hit via its summary chunk ⇒ 1.
|
||||
|
||||
Context assembly is unchanged (A7 revised): the *source* document's
|
||||
full content lands in the prompt, not the summary text alone.
|
||||
"""
|
||||
def test_summary_hit_on_suggested_doc_counts() -> None:
|
||||
"""HIGH branch: a suggested (rank-1) document hit via its summary
|
||||
chunk ⇒ 1. Phase 118 (A6): the *summary* is what the LLM sees in
|
||||
the prompt (no stored summary here → the A5 preview fallback carries
|
||||
the short fixture content whole)."""
|
||||
a = _doc("Alpha", "ALPHA_FULL_SOURCE_CONTENT")
|
||||
b = _doc("Beta", "BETA_FULL_SOURCE_CONTENT")
|
||||
chunks = [
|
||||
_chunk(a, 0.90, is_summary=True), # top doc reached through its summary
|
||||
_chunk(a, 0.90, is_summary=True), # suggested doc reached through its summary
|
||||
_chunk(b, 0.50),
|
||||
]
|
||||
plan = chat_api.plan_turn(chunks, _settings(threshold=0.30))
|
||||
assert plan.deflected is False
|
||||
assert plan.summary_hits == 1
|
||||
# The full source document is what the LLM sees (phase 24 contract).
|
||||
# The A5 preview fallback (short content, under the 400-char cap)
|
||||
# carries the content whole — the block body, never more.
|
||||
assert "ALPHA_FULL_SOURCE_CONTENT" in plan.system_prompt
|
||||
|
||||
|
||||
def test_summary_hit_outside_top_n_selection_not_counted() -> None:
|
||||
"""A summary chunk on a document outside the top-N (default 2) selection
|
||||
does not count — only hits that landed in the selected context do."""
|
||||
a = _doc("Alpha", "ALPHA_CONTENT")
|
||||
b = _doc("Beta", "BETA_CONTENT")
|
||||
c = _doc("Gamma", "GAMMA_CONTENT")
|
||||
def test_summary_hits_counts_suggested_parent_only() -> None:
|
||||
"""Phase 118 redefinition (redefined from the phase-113 cited set):
|
||||
a summary chunk counts ONLY when its parent document is in the
|
||||
SUGGESTED set — a rank-1 (suggested) parent counts, a rank-6
|
||||
(related-tier-only) parent does not."""
|
||||
docs_in = [_doc(f"Doc {i}", f"CONTENT_{i}") for i in range(7)]
|
||||
chunks = [
|
||||
_chunk(a, 0.90),
|
||||
_chunk(b, 0.80),
|
||||
_chunk(c, 0.70, is_summary=True), # 3rd-ranked doc — not selected
|
||||
_chunk(docs_in[0], 0.90, is_summary=True), # suggested parent — counts
|
||||
_chunk(docs_in[1], 0.80),
|
||||
_chunk(docs_in[2], 0.70),
|
||||
_chunk(docs_in[3], 0.60),
|
||||
_chunk(docs_in[4], 0.50),
|
||||
_chunk(docs_in[5], 0.40, is_summary=True), # related-only parent — does not
|
||||
_chunk(docs_in[6], 0.30),
|
||||
]
|
||||
plan = chat_api.plan_turn(chunks, _settings(threshold=0.30))
|
||||
assert plan.deflected is False
|
||||
assert [d.title for d in plan.docs] == ["Alpha", "Beta"]
|
||||
assert plan.summary_hits == 0
|
||||
assert [d.title for d in plan.suggested_docs] == [f"Doc {i}" for i in range(5)]
|
||||
assert [d.title for d in plan.related_docs] == ["Doc 5", "Doc 6"]
|
||||
assert plan.summary_hits == 1
|
||||
|
||||
|
||||
def test_low_branch_counts_summary_hit_on_selected_doc() -> None:
|
||||
def test_low_branch_counts_summary_hit_on_suggested_doc() -> None:
|
||||
"""LOW (deflected) branch records ``summary_hits`` too: the weak hit's
|
||||
parent is still the selected (weak-hit) document."""
|
||||
parent is still the suggested (weak-hit) document (no floor, A3)."""
|
||||
a = _doc("Gamma", "GAMMA_DOC_CONTENT")
|
||||
b = _doc("Delta", "DELTA_DOC_CONTENT")
|
||||
chunks = [
|
||||
@@ -683,7 +739,43 @@ def test_high_path_unaffected() -> None:
|
||||
assert "Reply in plain text only" not in plan.system_prompt
|
||||
assert "ALPHA_DOC_CONTENT" in plan.system_prompt
|
||||
assert "BETA_DOC_CONTENT" in plan.system_prompt
|
||||
assert [d.title for d in plan.docs] == ["Kubernetes Homelab Cluster", "Backup Strategy"]
|
||||
# Phase 118 (A3): both docs are suggested (no floor) — the durable-
|
||||
# record input — with nothing left for the related tier (rank 6+).
|
||||
assert [d.title for d in plan.suggested_docs] == [
|
||||
"Kubernetes Homelab Cluster",
|
||||
"Backup Strategy",
|
||||
]
|
||||
assert plan.related_docs == []
|
||||
|
||||
|
||||
# ---------- deflected branch: byte-identical (A8 gate untouched) ----------
|
||||
|
||||
|
||||
def test_low_prompt_byte_identical_to_pre_task_sha_pin() -> None:
|
||||
"""Phase 118 (A8 gate untouched): the deflected prompt is
|
||||
BYTE-IDENTICAL to pre-task on the same chunks — it is exactly
|
||||
``build_deflect_prompt(weak_hit_titles(…))`` (the tiering feeds the
|
||||
TurnPlan's durable-record fields, never the LOW prompt). The sha256
|
||||
pin makes any future LOW-body drift loud; the suggestions/chips and
|
||||
the deflected flag are unchanged."""
|
||||
a = _doc("Kubernetes Homelab Cluster", "ALPHA_DOC_CONTENT")
|
||||
b = _doc("Backup Strategy", "BETA_DOC_CONTENT")
|
||||
chunks = [_chunk(b, 0.10), _chunk(a, 0.20)]
|
||||
plan = chat_api.plan_turn(chunks, _settings())
|
||||
assert plan.deflected is True
|
||||
expected = build_deflect_prompt(weak_hit_titles(chunks))
|
||||
assert plan.system_prompt == expected # byte-identical to the pre-task build
|
||||
assert (
|
||||
hashlib.sha256(plan.system_prompt.encode("utf-8")).hexdigest()
|
||||
== "603395e013c97be8a13837eda533c0b7bd5da4f7a0806b0ae8ad7d6c52420913"
|
||||
)
|
||||
assert plan.suggestions # the "Maybe try" chips are unchanged
|
||||
# Both tiers still ride the plan (the durable-record inputs, A3).
|
||||
assert [d.title for d in plan.suggested_docs] == [
|
||||
"Kubernetes Homelab Cluster",
|
||||
"Backup Strategy",
|
||||
]
|
||||
assert plan.related_docs == []
|
||||
|
||||
|
||||
# ---------- weak_hit_titles (fake retriever mapping) ----------
|
||||
@@ -1059,3 +1151,55 @@ def test_endpoint_no_kb_row_prompt_unchanged(
|
||||
assert "<knowledge_base>" not in system["content"]
|
||||
log_lines = [r.getMessage() for r in caplog.records if "question=" in r.getMessage()]
|
||||
assert log_lines and "kb_chars=0" in log_lines[-1]
|
||||
|
||||
|
||||
def test_endpoint_log_line_records_suggested_after_summary_hits(
|
||||
client: TestClient,
|
||||
gate_env: tuple[_FakeSession, _CannedLLM],
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
caplog: pytest.LogCaptureFixture,
|
||||
) -> None:
|
||||
"""Phase 118 (PLAN §9 extension): the per-turn log line gains
|
||||
``suggested=N`` immediately AFTER ``summary_hits=N`` — the field
|
||||
order of every existing field is untouched (the phase-114
|
||||
``retries=N scaffold_stripped=N`` tail stays last). The value is the
|
||||
seeded suggestion tier's size (``len(plan.suggested_docs)``)."""
|
||||
_session, _llm = gate_env
|
||||
doc = _doc("Kubernetes Homelab Cluster", "TALOS_DOC_SENT")
|
||||
chunks = [_chunk(doc, 0.30), _chunk(doc, 0.20, is_summary=True)]
|
||||
monkeypatch.setattr(chat_api, "retrieve", _fake_retriever(chunks))
|
||||
|
||||
with caplog.at_level("INFO", logger="app.chat"):
|
||||
_ask(client, "How is my Kubernetes cluster set up?")
|
||||
|
||||
log_lines = [r.getMessage() for r in caplog.records if "question=" in r.getMessage()]
|
||||
assert log_lines
|
||||
line = log_lines[-1]
|
||||
# The new slot: suggested=N right after summary_hits=N (one suggested
|
||||
# doc — the single fixture doc — and one summary hit on it).
|
||||
assert "summary_hits=1 suggested=1" in line
|
||||
# The full field order (the phase-114 tail stays last).
|
||||
order = (
|
||||
"question=",
|
||||
"embed_ms=",
|
||||
"top_score=",
|
||||
"fts_hits=",
|
||||
"summary_hits=",
|
||||
"suggested=",
|
||||
"tuning=",
|
||||
"kb_chars=",
|
||||
"history_msgs=",
|
||||
"threshold=",
|
||||
"deflected=",
|
||||
"sources=",
|
||||
"thinking_chars=",
|
||||
"tool_calls=",
|
||||
"total_ms=",
|
||||
"retries=",
|
||||
"scaffold_stripped=",
|
||||
)
|
||||
idx = -1
|
||||
for field in order:
|
||||
pos = line.find(field)
|
||||
assert pos > idx, f"{field} out of order in the per-turn log line"
|
||||
idx = pos
|
||||
|
||||
@@ -282,6 +282,59 @@ def test_related_max_docs_rejects_negative(
|
||||
_settings()
|
||||
|
||||
|
||||
def test_suggested_docs_default_and_env_override(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""Phase 118 (LOCKED A3): the start-here suggestion tier cap —
|
||||
default 5 (the owner directive, TODO L3), env-tunable."""
|
||||
monkeypatch.delenv("BOR_SUGGESTED_DOCS", raising=False)
|
||||
assert _settings().suggested_docs == 5
|
||||
monkeypatch.setenv("BOR_SUGGESTED_DOCS", "3")
|
||||
assert _settings().suggested_docs == 3
|
||||
|
||||
|
||||
def test_suggested_docs_rejects_zero_and_negative(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""``0`` (no starting points) and a negative cap are typos (the
|
||||
``agent_max_rounds`` pattern); ``1`` is the minimum legal value."""
|
||||
monkeypatch.setenv("BOR_SUGGESTED_DOCS", "0")
|
||||
with pytest.raises(ValidationError, match="suggested_docs"):
|
||||
_settings()
|
||||
monkeypatch.setenv("BOR_SUGGESTED_DOCS", "-2")
|
||||
with pytest.raises(ValidationError, match="suggested_docs"):
|
||||
_settings()
|
||||
monkeypatch.setenv("BOR_SUGGESTED_DOCS", "1")
|
||||
assert _settings().suggested_docs == 1
|
||||
|
||||
|
||||
def test_suggestion_preview_chars_default_and_env_override(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""Phase 118, task 03 (LOCKED A5): the NULL-summary preview cap —
|
||||
default 400, env-tunable via ``BOR_SUGGESTION_PREVIEW_CHARS``."""
|
||||
monkeypatch.delenv("BOR_SUGGESTION_PREVIEW_CHARS", raising=False)
|
||||
assert _settings().suggestion_preview_chars == 400
|
||||
monkeypatch.setenv("BOR_SUGGESTION_PREVIEW_CHARS", "800")
|
||||
assert _settings().suggestion_preview_chars == 800
|
||||
|
||||
|
||||
def test_suggestion_preview_chars_rejects_zero_and_negative(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""``0``/negative would preview an empty/absent prefix — typos that
|
||||
fail loudly (the ``agent_max_rounds`` pattern); ``1`` is the minimum
|
||||
legal value."""
|
||||
monkeypatch.setenv("BOR_SUGGESTION_PREVIEW_CHARS", "0")
|
||||
with pytest.raises(ValidationError, match="suggestion_preview_chars"):
|
||||
_settings()
|
||||
monkeypatch.setenv("BOR_SUGGESTION_PREVIEW_CHARS", "-5")
|
||||
with pytest.raises(ValidationError, match="suggestion_preview_chars"):
|
||||
_settings()
|
||||
monkeypatch.setenv("BOR_SUGGESTION_PREVIEW_CHARS", "1")
|
||||
assert _settings().suggestion_preview_chars == 1
|
||||
|
||||
|
||||
def test_read_max_chars_default_and_env_override(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
|
||||
+272
-22
@@ -4,8 +4,10 @@ The walk tests are pure filesystem (``tmp_path``); the delta and summary
|
||||
tests run against the local compose Postgres (preferred — a real vector
|
||||
table), skipping with clear instructions when the stack is not up.
|
||||
|
||||
Summaries (phase 30): non-markdown files get a ``lite``-model summary via
|
||||
the fake's deterministic ``chat`` (``"Summary of <first token>"``); the
|
||||
Summaries (phase 30; phase 118, A2: every file, markdown included):
|
||||
every file gets a ``lite``-model summary via the fake's deterministic
|
||||
``chat`` (``"Summary of <first token>"``); an unchanged doc whose summary
|
||||
is NULL is backfilled on the next run (``summary_backfilled``); the
|
||||
sentinel word ``SUMMARY-BLOWUP`` makes ``chat`` raise :class:`LLMError`
|
||||
for the fail-soft path.
|
||||
"""
|
||||
@@ -13,10 +15,12 @@ from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
from sqlalchemy import func, select
|
||||
from sqlalchemy.orm import Session
|
||||
|
||||
import app.rag.importer as importer
|
||||
from app.config import Settings
|
||||
@@ -29,7 +33,7 @@ from app.rag.importer import (
|
||||
iter_importable_files,
|
||||
match_extension,
|
||||
)
|
||||
from app.rag.llm import EmbeddingError
|
||||
from app.rag.llm import EmbeddingError, LLMError
|
||||
from tests.fakes import FakeEmbedder
|
||||
|
||||
#: The original seven A9 formats as dotted suffixes (pre-phase-47 default
|
||||
@@ -47,6 +51,15 @@ class _PoisonEmbedder(FakeEmbedder):
|
||||
return await super().embed(texts)
|
||||
|
||||
|
||||
class _FailingChatEmbedder(FakeEmbedder):
|
||||
"""A ``lite`` model that always fails (drives the summary fail-soft
|
||||
path — including the phase-118 backfill)."""
|
||||
|
||||
async def chat(self, messages: list[dict[str, str]], model: str | None = None) -> str:
|
||||
self.chat_calls.append(list(messages))
|
||||
raise LLMError("simulated lite-model failure (test sentinel)")
|
||||
|
||||
|
||||
class _CapEmbedder(FakeEmbedder):
|
||||
"""Simulates the endpoint's ~1024-token input cap at ~1.1 chars/token:
|
||||
any single text over 1000 chars is rejected (URL-dense worst case)."""
|
||||
@@ -266,16 +279,19 @@ def test_added_then_unchanged_then_updated_then_pruned(db, tmp_path: Path) -> No
|
||||
try:
|
||||
s1 = asyncio.run(import_sources([root], llm, session=db))
|
||||
assert (s1.files, s1.added, s1.unchanged, s1.updated, s1.pruned) == (2, 2, 0, 0, 0)
|
||||
# a.md has two sections (2 chunks), b.md one (1 chunk).
|
||||
# a.md has two sections (2 chunks), b.md one (1 chunk) — content
|
||||
# chunks only; the ``is_summary`` chunks live in ``summaries``.
|
||||
assert s1.chunks == 3
|
||||
# Embeddings are stored with the configured dimension.
|
||||
assert s1.summaries == 2 # phase 118 (A2): markdown is summarized too
|
||||
# Embeddings are stored with the configured dimension: 3 content
|
||||
# chunks + 2 ``is_summary`` chunks (one per doc, phase 118 A2).
|
||||
n = db.scalar(
|
||||
select(func.count())
|
||||
.select_from(Chunk)
|
||||
.join(Document, Document.id == Chunk.document_id)
|
||||
.where(Document.source == root.name)
|
||||
)
|
||||
assert n == 3
|
||||
assert n == 5
|
||||
for c in db.scalars(
|
||||
select(Chunk)
|
||||
.join(Document, Document.id == Chunk.document_id)
|
||||
@@ -300,14 +316,15 @@ def test_added_then_unchanged_then_updated_then_pruned(db, tmp_path: Path) -> No
|
||||
assert db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "a.md")
|
||||
) is None
|
||||
# Chunks of the pruned document are gone (FK cascade).
|
||||
# Chunks of the pruned document are gone (FK cascade); b.md's
|
||||
# content chunk + its ``is_summary`` chunk survive (phase 118 A2).
|
||||
n_after = db.scalar(
|
||||
select(func.count())
|
||||
.select_from(Chunk)
|
||||
.join(Document, Document.id == Chunk.document_id)
|
||||
.where(Document.source == root.name)
|
||||
)
|
||||
assert n_after == 1
|
||||
assert n_after == 2
|
||||
finally:
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
@@ -438,8 +455,11 @@ def test_chunk_positions_and_titles(db, tmp_path: Path) -> None:
|
||||
select(Document).where(Document.source == root.name, Document.path == "multi.md")
|
||||
)
|
||||
assert doc is not None
|
||||
positions = sorted(c.position for c in doc.chunks)
|
||||
assert positions == list(range(len(doc.chunks))) and len(doc.chunks) >= 2
|
||||
# 0-based CONTENT positions (the ``is_summary`` chunk sits at −1,
|
||||
# phase 30/118).
|
||||
content = [c for c in doc.chunks if not c.is_summary]
|
||||
positions = sorted(c.position for c in content)
|
||||
assert positions == list(range(len(content))) and len(content) >= 2
|
||||
finally:
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
@@ -535,7 +555,7 @@ def test_quadlet_and_j2_files_get_stem_titles_and_per_format_counts(
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
|
||||
# ---------- phase 30: lite-model summaries for non-markdown files ----------
|
||||
# ---------- phase 30: lite-model summaries (phase 118: every file) ----------
|
||||
|
||||
|
||||
def test_non_markdown_file_gets_stored_and_indexed_summary(db, tmp_path: Path) -> None:
|
||||
@@ -571,9 +591,11 @@ def test_non_markdown_file_gets_stored_and_indexed_summary(db, tmp_path: Path) -
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
|
||||
def test_markdown_file_never_gets_summary(db, tmp_path: Path) -> None:
|
||||
"""Markdown is already natural language: no summary, no ``is_summary``
|
||||
chunk, and the ``lite`` model is never called."""
|
||||
def test_markdown_file_gets_stored_summary(db, tmp_path: Path) -> None:
|
||||
"""Phase 118 (A2): markdown is summarized too (the phase-30 exclusion
|
||||
is retired) — ``documents.summary`` is set and one ``is_summary``
|
||||
chunk (position −1, embedded) is indexed alongside the content
|
||||
chunks."""
|
||||
root = tmp_path / "mdsrc"
|
||||
root.mkdir()
|
||||
(root / "note.md").write_text("# Note\n\nmarkdown body\n")
|
||||
@@ -581,14 +603,240 @@ def test_markdown_file_never_gets_summary(db, tmp_path: Path) -> None:
|
||||
try:
|
||||
summary = asyncio.run(import_sources([root], llm, session=db))
|
||||
assert summary.added == 1
|
||||
assert summary.summaries == 0 and summary.summary_errors == 0
|
||||
assert llm.chat_calls == [] # the model was never asked
|
||||
assert summary.summaries == 1 and summary.summary_errors == 0
|
||||
assert llm.chat_calls # the lite model WAS asked (phase 118)
|
||||
doc = db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "note.md")
|
||||
)
|
||||
assert doc is not None
|
||||
assert doc.summary is None
|
||||
assert doc.chunks and all(not c.is_summary for c in doc.chunks)
|
||||
assert doc.summary is not None
|
||||
# Deterministic fake reply + the code-appended pointer line.
|
||||
assert doc.summary.startswith("Summary of")
|
||||
assert doc.summary.endswith(f"Source: {root.name}/note.md")
|
||||
schunks = [c for c in doc.chunks if c.is_summary]
|
||||
assert len(schunks) == 1
|
||||
assert schunks[0].position == -1
|
||||
assert schunks[0].content == doc.summary
|
||||
assert schunks[0].embedding is not None and len(schunks[0].embedding) == 768
|
||||
# Content chunks stay 0-based and are never flagged as summaries.
|
||||
content = [c for c in doc.chunks if not c.is_summary]
|
||||
assert sorted(c.position for c in content) == list(range(len(content)))
|
||||
finally:
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
|
||||
# ---------- phase 118 (A2): NULL-summary backfill on the unchanged path ----------
|
||||
|
||||
|
||||
def _clear_stored_summary(db: Session, doc: Document) -> None:
|
||||
"""Simulate a NULL-summary row (a pre-phase-30 row, or a cleared
|
||||
summary): the content stays, only the summary + its chunk go away."""
|
||||
doc.summary = None
|
||||
for c in [c for c in doc.chunks if c.is_summary]:
|
||||
doc.chunks.remove(c)
|
||||
db.commit()
|
||||
|
||||
|
||||
def test_unchanged_doc_with_null_summary_is_backfilled(db, tmp_path: Path) -> None:
|
||||
"""Phase 118 (A2): an unchanged doc whose summary is NULL gets a
|
||||
summary-only backfill on the next sync: summary stored + one embedded
|
||||
``is_summary`` chunk, counted ``summary_backfilled`` — never
|
||||
``summaries``, never added/updated/pruned, no content re-embed."""
|
||||
root = tmp_path / "bfill"
|
||||
root.mkdir()
|
||||
(root / "note.md").write_text("# Note\n\nmarkdown body\n")
|
||||
llm = FakeEmbedder()
|
||||
try:
|
||||
first = asyncio.run(import_sources([root], llm, session=db))
|
||||
assert (first.added, first.summaries) == (1, 1)
|
||||
doc = db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "note.md")
|
||||
)
|
||||
assert doc is not None and doc.summary is not None
|
||||
_clear_stored_summary(db, doc) # the NULL-summary row the backfill targets
|
||||
|
||||
embed_before = len(llm.calls)
|
||||
second = asyncio.run(import_sources([root], llm, session=db))
|
||||
assert (second.added, second.updated, second.pruned) == (0, 0, 0)
|
||||
assert second.unchanged == 1
|
||||
assert second.summary_backfilled == 1
|
||||
assert second.summaries == 0 and second.summary_errors == 0
|
||||
# No content re-embed: exactly one new embed batch, the summary
|
||||
# text only.
|
||||
assert len(llm.calls) == embed_before + 1
|
||||
|
||||
db.expire_all()
|
||||
doc = db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "note.md")
|
||||
)
|
||||
assert doc is not None
|
||||
assert doc.summary is not None
|
||||
assert llm.calls[-1] == [doc.summary] # only the backfilled summary
|
||||
schunks = [c for c in doc.chunks if c.is_summary]
|
||||
assert len(schunks) == 1
|
||||
assert schunks[0].position == -1
|
||||
assert schunks[0].content == doc.summary
|
||||
assert schunks[0].embedding is not None and len(schunks[0].embedding) == 768
|
||||
# The content chunk is untouched.
|
||||
content = [c for c in doc.chunks if not c.is_summary]
|
||||
assert len(content) == 1 and content[0].embedding is not None
|
||||
finally:
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
|
||||
def test_unchanged_doc_with_stored_summary_never_resummarizes(db, tmp_path: Path) -> None:
|
||||
"""Phase 118 (A2): an unchanged doc that ALREADY has a summary (the
|
||||
third sync of the lifecycle) makes no summary LLM call at all and
|
||||
gains no chunks — owner-edited (non-NULL) summaries are never
|
||||
touched."""
|
||||
root = tmp_path / "noref"
|
||||
root.mkdir()
|
||||
(root / "note.md").write_text("# Note\n\nmarkdown body\n")
|
||||
llm = FakeEmbedder()
|
||||
try:
|
||||
asyncio.run(import_sources([root], llm, session=db))
|
||||
chat_before = len(llm.chat_calls)
|
||||
embed_before = len(llm.calls)
|
||||
chunk_before = db.scalar(
|
||||
select(func.count())
|
||||
.select_from(Chunk)
|
||||
.join(Document, Document.id == Chunk.document_id)
|
||||
.where(Document.source == root.name)
|
||||
)
|
||||
second = asyncio.run(import_sources([root], llm, session=db))
|
||||
assert second.unchanged == 1
|
||||
assert second.summary_backfilled == 0 and second.summaries == 0
|
||||
assert second.summary_errors == 0
|
||||
assert len(llm.chat_calls) == chat_before # the model was never asked
|
||||
assert len(llm.calls) == embed_before # no embedding of any kind
|
||||
chunk_after = db.scalar(
|
||||
select(func.count())
|
||||
.select_from(Chunk)
|
||||
.join(Document, Document.id == Chunk.document_id)
|
||||
.where(Document.source == root.name)
|
||||
)
|
||||
assert chunk_after == chunk_before # no new chunk of any kind
|
||||
finally:
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
|
||||
def test_unchanged_doc_with_empty_string_summary_is_never_backfilled(
|
||||
db, tmp_path: Path
|
||||
) -> None:
|
||||
"""Phase 118 (A2, strict ``is None``): an empty-string summary is
|
||||
owner-set (phase 57) — the backfill skips it, the ``lite`` model is
|
||||
never called, and the value stays byte-identical."""
|
||||
root = tmp_path / "emptysum"
|
||||
root.mkdir()
|
||||
(root / "note.md").write_text("# Note\n\nmarkdown body\n")
|
||||
llm = FakeEmbedder()
|
||||
try:
|
||||
asyncio.run(import_sources([root], llm, session=db))
|
||||
doc = db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "note.md")
|
||||
)
|
||||
assert doc is not None
|
||||
doc.summary = "" # the owner-set empty string (never NULL)
|
||||
db.commit()
|
||||
|
||||
chat_before = len(llm.chat_calls)
|
||||
second = asyncio.run(import_sources([root], llm, session=db))
|
||||
assert second.unchanged == 1
|
||||
assert second.summary_backfilled == 0 and second.summaries == 0
|
||||
assert second.summary_errors == 0
|
||||
assert len(llm.chat_calls) == chat_before # the model was never asked
|
||||
|
||||
db.expire_all()
|
||||
doc = db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "note.md")
|
||||
)
|
||||
assert doc is not None
|
||||
assert doc.summary == "" # byte-identical — never overwritten
|
||||
finally:
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
|
||||
def test_backfill_runs_on_manually_dated_doc_without_touching_the_date(
|
||||
db, tmp_path: Path
|
||||
) -> None:
|
||||
"""Phase 118 (A2, assumption 7): ``created_at_manual`` protects the
|
||||
DATE only (phase 106, D1) — a manually-dated, NULL-summary doc still
|
||||
gets its backfilled summary, and the stored date stays byte-untouched
|
||||
even though a refresh was due."""
|
||||
root = tmp_path / "manualdate"
|
||||
root.mkdir()
|
||||
(root / "note.md").write_text("# Note\n\nmarkdown body\n")
|
||||
llm = FakeEmbedder()
|
||||
try:
|
||||
asyncio.run(import_sources([root], llm, session=db))
|
||||
doc = db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "note.md")
|
||||
)
|
||||
assert doc is not None
|
||||
manual = datetime(2020, 5, 4, 12, 0, 0, tzinfo=UTC)
|
||||
doc.created_at = manual
|
||||
doc.created_at_manual = True
|
||||
_clear_stored_summary(db, doc) # the NULL-summary row the backfill targets
|
||||
|
||||
second = asyncio.run(import_sources([root], llm, session=db))
|
||||
assert second.unchanged == 1
|
||||
assert second.summary_backfilled == 1 and second.summary_errors == 0
|
||||
# A date refresh WAS due (the mtime differs from the 2020
|
||||
# correction) but the manual flag withheld it — the backfill
|
||||
# never touches the date either.
|
||||
assert second.dates_updated == 0
|
||||
|
||||
db.expire_all()
|
||||
doc = db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "note.md")
|
||||
)
|
||||
assert doc is not None
|
||||
assert doc.summary is not None # the backfill landed
|
||||
assert doc.created_at == manual # byte-untouched
|
||||
assert doc.created_at_manual is True
|
||||
finally:
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
|
||||
def test_backfill_failure_is_fail_soft_and_date_still_refreshes(
|
||||
db, tmp_path: Path
|
||||
) -> None:
|
||||
"""Phase 118 (A2): a backfill whose ``lite`` call fails rolls back
|
||||
its own session work only — ``summary_errors=1``, the doc row
|
||||
untouched — while the UNCHANGED path's date refresh still runs
|
||||
afterwards."""
|
||||
root = tmp_path / "bfillfail"
|
||||
root.mkdir()
|
||||
(root / "note.md").write_text("# Note\n\nmarkdown body\n")
|
||||
llm = FakeEmbedder()
|
||||
try:
|
||||
asyncio.run(import_sources([root], llm, session=db))
|
||||
doc = db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "note.md")
|
||||
)
|
||||
assert doc is not None
|
||||
_clear_stored_summary(db, doc) # the NULL-summary row the backfill targets
|
||||
# Force a date drift so the refresh is DUE on this run.
|
||||
doc.created_at = datetime(2020, 1, 1, tzinfo=UTC)
|
||||
db.commit()
|
||||
|
||||
second = asyncio.run(import_sources([root], _FailingChatEmbedder(), session=db))
|
||||
assert second.unchanged == 1
|
||||
assert (second.added, second.updated, second.pruned) == (0, 0, 0)
|
||||
assert second.summary_errors == 1
|
||||
assert second.summary_backfilled == 0 and second.summaries == 0
|
||||
|
||||
db.expire_all()
|
||||
doc = db.scalar(
|
||||
select(Document).where(Document.source == root.name, Document.path == "note.md")
|
||||
)
|
||||
assert doc is not None
|
||||
assert doc.summary is None # the failed backfill left the row untouched
|
||||
assert not any(c.is_summary for c in doc.chunks)
|
||||
# …but the date refresh ran (the failure only rolled back the
|
||||
# summary's own session work).
|
||||
assert second.dates_updated == 1
|
||||
assert doc.created_at != datetime(2020, 1, 1, tzinfo=UTC)
|
||||
finally:
|
||||
_cleanup_source(db, root.name)
|
||||
|
||||
@@ -720,11 +968,13 @@ def test_import_summary_log_line_includes_summary_counters(
|
||||
caplog: pytest.LogCaptureFixture,
|
||||
) -> None:
|
||||
"""PLAN §9 summary line: the phase-30 counters sit between
|
||||
``embed_batches`` and ``formats``; the phase-106 date-refresh
|
||||
counter sits between ``summary_errors`` and ``formats``."""
|
||||
``embed_batches`` and ``formats``; the phase-118 backfill counter
|
||||
sits between ``summary_errors`` and ``dates_updated``; the
|
||||
phase-106 date-refresh counter sits before ``formats``."""
|
||||
s = ImportSummary()
|
||||
s.files, s.added, s.chunks, s.embed_batches = 3, 3, 5, 4
|
||||
s.summaries, s.summary_errors = 2, 1
|
||||
s.summary_backfilled = 1
|
||||
s.dates_updated = 0
|
||||
s.formats = {"md": 1, "yaml": 2}
|
||||
with caplog.at_level(logging.INFO, logger="app.importer"):
|
||||
@@ -732,8 +982,8 @@ def test_import_summary_log_line_includes_summary_counters(
|
||||
line = caplog.records[-1].getMessage()
|
||||
assert line == (
|
||||
"import: summary files=3 added=3 updated=0 unchanged=0 pruned=0 errors=0 "
|
||||
"chunks=5 embed_batches=4 summaries=2 summary_errors=1 dates_updated=0 "
|
||||
"formats=yaml:2,md:1"
|
||||
"chunks=5 embed_batches=4 summaries=2 summary_errors=1 summary_backfilled=1 "
|
||||
"dates_updated=0 formats=yaml:2,md:1"
|
||||
)
|
||||
|
||||
|
||||
|
||||
@@ -16,6 +16,7 @@ import hashlib
|
||||
|
||||
from app.rag.prompts import (
|
||||
PERSONA,
|
||||
SUGGEST_INTRO,
|
||||
TOOLS_SECTION,
|
||||
_base,
|
||||
build_deflect_prompt,
|
||||
@@ -67,9 +68,16 @@ def test_high_and_low_bases_byte_locked() -> None:
|
||||
|
||||
# ---------- TOOLS_SECTION (the HIGH prompt's locked ``<tools>`` copy) ----------
|
||||
|
||||
#: Pre-phase-112 anchors for ``TOOLS_SECTION``.
|
||||
TOOLS_SECTION_SHA256 = "b834cbe368055e65da82ae3e37a91e6c658c713954703fc79b849a6ebdf4aa53"
|
||||
TOOLS_SECTION_LEN = 2273
|
||||
#: Anchors for ``TOOLS_SECTION`` — pre-phase-112 values, RE-CUT for
|
||||
#: phase 118 (task 04, A6, owner directive 2026-09-15): the ``read``
|
||||
#: clause was rewritten for the summary-seed mode (the ``<documents>``
|
||||
#: section holds SUMMARIES — ``read`` adds the full text; only an
|
||||
#: already-read document is refused). Only that clause moved — the
|
||||
#: prefix (the ``ls``-clause opening) and the suffix (the
|
||||
#: discipline-rules ending) survived byte-identical, so they are the
|
||||
#: same anchors as pre-phase-118.
|
||||
TOOLS_SECTION_SHA256 = "87ee80faf0170da6ab1de518177313def6460785613fa074312a2a5f9f071750"
|
||||
TOOLS_SECTION_LEN = 2465
|
||||
TOOLS_SECTION_PREFIX = (
|
||||
"<tools>\n"
|
||||
"You may extend your context with three tools. `ls` lists the "
|
||||
@@ -120,3 +128,51 @@ def test_deflect_body_byte_locked() -> None:
|
||||
assert prompt.index("DEFLECT_MODE") < prompt.index(
|
||||
"Reply in plain text only"
|
||||
) # the marker precedes the plain-text line
|
||||
|
||||
|
||||
#: The full LOW prompt build on the canonical fixture titles — sha-pinned
|
||||
#: (phase 118, task 03): the deflection path is UNTOUCHED by the
|
||||
#: summary-seed re-revision (LOCKED A8) — the same inputs must produce
|
||||
#: the pre-phase bytes, so this anchor is a pre-phase-118 value.
|
||||
LOW_PROMPT_SHA256 = "726eddb4eb3bcc26c840011f6f8635d6af55aa09d4d064bbb9e256caa9665837"
|
||||
LOW_PROMPT_LEN = 968
|
||||
|
||||
|
||||
def test_low_prompt_build_byte_identical_to_pre_phase() -> None:
|
||||
"""Phase 118 contract: the LOW prompt output is byte-identical to
|
||||
pre-phase for identical inputs — the summary seeding (and the new
|
||||
``SUGGEST_INTRO`` line) never leaks into the deflection path."""
|
||||
prompt = build_deflect_prompt(["T1", "T2"])
|
||||
assert len(prompt) == LOW_PROMPT_LEN
|
||||
assert _sha256(prompt) == LOW_PROMPT_SHA256
|
||||
assert "SUGGEST_INTRO" not in prompt and "<documents>" not in prompt
|
||||
|
||||
|
||||
# ---------- SUGGEST_INTRO (phase 118, task 03 — the start-here framing) ----------
|
||||
|
||||
#: Phase-118 anchors for ``SUGGEST_INTRO`` — the ``<documents>``
|
||||
#: section's intro line (the owner's "start here if these summaries seem
|
||||
#: right to you" framing, TODO L3). The E2E mock's ``_document_block``
|
||||
#: parser is regex-based over the block markup (which stays byte-stable
|
||||
#: around the intro), so this constant is a prompt-copy lock, pinned the
|
||||
#: way ``TOOLS_SECTION`` is: sha256 + prefix + total length.
|
||||
SUGGEST_INTRO_SHA256 = "7b14d2dedc6ebc4e440d32dd1edb979a7461c94034b9541c9f37b9042f394d3a"
|
||||
SUGGEST_INTRO_LEN = 323
|
||||
SUGGEST_INTRO_PREFIX = (
|
||||
"The blocks below are the summaries of the top-ranked documents for "
|
||||
"your question — start here if one seems right to you: "
|
||||
)
|
||||
|
||||
|
||||
def test_suggest_intro_byte_locked() -> None:
|
||||
"""The start-here framing is LOCKED copy (phase 118): sha256 + exact
|
||||
prefix + total length; the three contracts it must carry (summaries
|
||||
are the starting points; ``read`` adds the full text, which is NOT
|
||||
in the prompt until read; cite by path) are pinned as substrings."""
|
||||
assert len(SUGGEST_INTRO) == SUGGEST_INTRO_LEN
|
||||
assert _sha256(SUGGEST_INTRO) == SUGGEST_INTRO_SHA256
|
||||
assert SUGGEST_INTRO.startswith(SUGGEST_INTRO_PREFIX)
|
||||
assert "call `read`" in SUGGEST_INTRO
|
||||
assert "combined `source/path`" in SUGGEST_INTRO
|
||||
assert "its full text is not in the prompt until you read it" in SUGGEST_INTRO
|
||||
assert "Cite the document(s) you used, by path." in SUGGEST_INTRO
|
||||
|
||||
+303
-24
@@ -17,9 +17,13 @@ stated up front — the combined ``source/path`` identity for
|
||||
the drill-down tree contract — one level per call, sources at the
|
||||
top, folders + files below, ``grep`` as the without-listing locator —
|
||||
while the ``read``/``grep`` clauses and the discipline rules are
|
||||
byte-identical): the teaching refusals in :mod:`app.rag.agent`
|
||||
re-state the same contract; the ``<tools>`` marker keying (HIGH
|
||||
only) is unchanged.
|
||||
byte-identical; phase 118, task 04: the ``read`` clause rewritten
|
||||
for the summary-seed mode — the ``<documents>`` section holds
|
||||
SUMMARIES, ``read`` adds the full text, and only an already-read
|
||||
document is refused — while the ``ls``/``grep`` clauses and the
|
||||
discipline rules stay byte-identical): the teaching refusals in
|
||||
:mod:`app.rag.agent` re-state the same contract; the ``<tools>``
|
||||
marker keying (HIGH only) is unchanged.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
@@ -32,6 +36,7 @@ from app.config import Settings
|
||||
from app.models import Document
|
||||
from app.rag.prompts import (
|
||||
PERSONA,
|
||||
SUGGEST_INTRO,
|
||||
TOOLS_SECTION,
|
||||
_base,
|
||||
build_deflect_prompt,
|
||||
@@ -55,7 +60,12 @@ KB_INTRO = "The basic categories of everything in this knowledge base (generated
|
||||
_FIXTURE_CREATED_AT = datetime(2024, 6, 15, 12, 0, 0, tzinfo=UTC)
|
||||
|
||||
|
||||
def _doc(path: str, content: str, title: str) -> Document:
|
||||
def _doc(
|
||||
path: str,
|
||||
content: str,
|
||||
title: str,
|
||||
summary: str | None = None,
|
||||
) -> Document:
|
||||
return Document(
|
||||
id=uuid.uuid4(),
|
||||
source="Homelab",
|
||||
@@ -65,6 +75,7 @@ def _doc(path: str, content: str, title: str) -> Document:
|
||||
content=content,
|
||||
content_hash="0" * 64,
|
||||
created_at=_FIXTURE_CREATED_AT,
|
||||
summary=summary,
|
||||
)
|
||||
|
||||
|
||||
@@ -93,23 +104,64 @@ def test_persona_owner_edits_are_preserved() -> None:
|
||||
assert "HONESTY GATE" in PERSONA # the gate itself is intact
|
||||
|
||||
|
||||
def test_high_prompt_carries_relevance_marker_and_full_documents() -> None:
|
||||
doc = _doc("kubernetes.md", "Talos Linux on three nodes.", "Kubernetes Homelab Cluster")
|
||||
def test_high_prompt_carries_relevance_marker_and_summary_block() -> None:
|
||||
"""Phase 118 (LOCKED A6): the grounded turn seeds the document's
|
||||
SUMMARY — the full content never reaches the prompt (the ``read``
|
||||
tool is the only full-text path)."""
|
||||
doc = _doc(
|
||||
"kubernetes.md",
|
||||
"Talos Linux on three nodes. FULL_CONTENT_SENTINEL_987654",
|
||||
"Kubernetes Homelab Cluster",
|
||||
summary="A Talos Linux cluster on three nodes.",
|
||||
)
|
||||
prompt = build_high_prompt([doc])
|
||||
assert "<relevance>HIGH</relevance>" in prompt
|
||||
assert "DEFLECT_MODE" not in prompt
|
||||
assert "<documents>" in prompt and "</documents>" in prompt
|
||||
assert 'path="kubernetes.md"' in prompt
|
||||
assert "Talos Linux on three nodes." in prompt
|
||||
assert "A Talos Linux cluster on three nodes." in prompt # the summary
|
||||
assert "FULL_CONTENT_SENTINEL_987654" not in prompt # never the content
|
||||
assert "Talos Linux on three nodes." not in prompt # nor the full sentence
|
||||
assert "HONESTY GATE" in prompt # persona intact
|
||||
|
||||
|
||||
def test_high_prompt_lists_multiple_documents_in_order() -> None:
|
||||
a = _doc("a.md", "CONTENT_A", "Title A")
|
||||
b = _doc("b.md", "CONTENT_B", "Title B")
|
||||
a = _doc("a.md", "CONTENT_A", "Title A", summary="Summary A.")
|
||||
b = _doc("b.md", "CONTENT_B", "Title B", summary="Summary B.")
|
||||
prompt = build_high_prompt([a, b])
|
||||
assert prompt.index("CONTENT_A") < prompt.index("CONTENT_B")
|
||||
assert prompt.index("Summary A.") < prompt.index("Summary B.")
|
||||
assert 'title="Title B"' in prompt
|
||||
assert "CONTENT_A" not in prompt and "CONTENT_B" not in prompt
|
||||
|
||||
|
||||
def test_high_prompt_seeds_summaries_of_all_suggested_docs() -> None:
|
||||
"""The phase-118 completion pin: a grounded prompt built from five
|
||||
summary-bearing documents contains ALL FIVE summaries + the intro,
|
||||
and ZERO full-content characters (the full texts are not in the
|
||||
prompt at all)."""
|
||||
docs = [
|
||||
_doc(
|
||||
f"doc{i}.md",
|
||||
f"FULL_CONTENT_SENTINEL_{i} " + "x" * 100,
|
||||
f"Title {i}",
|
||||
summary=f"SUMMARY_{i} of the document.",
|
||||
)
|
||||
for i in range(5)
|
||||
]
|
||||
prompt = build_high_prompt(docs)
|
||||
assert SUGGEST_INTRO in prompt
|
||||
for i in range(5):
|
||||
assert f"SUMMARY_{i} of the document.\n" in prompt
|
||||
assert f"FULL_CONTENT_SENTINEL_{i}" not in prompt
|
||||
assert f'path="doc{i}.md"' in prompt
|
||||
# One block per document, in the given order.
|
||||
assert prompt.count("<document ") == 5
|
||||
positions = [prompt.index(f"SUMMARY_{i}") for i in range(5)]
|
||||
assert positions == sorted(positions)
|
||||
# The intro leads the section: <documents> → intro → first block.
|
||||
i_open = prompt.index("<documents>")
|
||||
i_block = prompt.index("<document ")
|
||||
assert prompt[i_open:i_block] == f"<documents>\n{SUGGEST_INTRO}\n\n"
|
||||
|
||||
|
||||
def test_high_prompt_without_documents_stays_honest() -> None:
|
||||
@@ -143,7 +195,9 @@ def test_zero_note_prompt_is_byte_identical_to_pre_steering() -> None:
|
||||
"""Phase 15 contract: with no steering notes the prompt is exactly what
|
||||
it was before the <tuning> section existed. (Phase 37: the HIGH prompt
|
||||
additionally carries the ``<tools>`` section after the mode body — the
|
||||
fixtures account for it; the LOW prompt is untouched.)"""
|
||||
fixtures account for it; phase 118: the ``<documents>`` section leads
|
||||
with the ``SUGGEST_INTRO`` line — the fixture accounts for it; the
|
||||
LOW prompt is untouched.)"""
|
||||
doc = _doc("kubernetes.md", "Talos Linux on three nodes.", "Kubernetes Homelab Cluster")
|
||||
block = (
|
||||
'<document source="Homelab" path="kubernetes.md" '
|
||||
@@ -152,7 +206,14 @@ def test_zero_note_prompt_is_byte_identical_to_pre_steering() -> None:
|
||||
"</document>"
|
||||
)
|
||||
assert build_high_prompt([doc]) == (
|
||||
_base("HIGH") + "\n<documents>\n" + block + "\n</documents>" + "\n" + TOOLS_SECTION
|
||||
_base("HIGH")
|
||||
+ "\n<documents>\n"
|
||||
+ SUGGEST_INTRO
|
||||
+ "\n\n"
|
||||
+ block
|
||||
+ "\n</documents>"
|
||||
+ "\n"
|
||||
+ TOOLS_SECTION
|
||||
)
|
||||
# Phase 71: the LOW prompt carries the owner-permitted plain-text
|
||||
# line after the DEFLECT_MODE sentence (the marker-keying contract
|
||||
@@ -257,6 +318,86 @@ def test_tools_section_phase95_read_truncation_clause() -> None:
|
||||
assert "Very large documents are capped" in build_high_prompt([doc])
|
||||
|
||||
|
||||
def test_tools_section_phase118_summary_seed_read_clause() -> None:
|
||||
"""Phase 118 (task 04, A6): the ``read`` clause is rewritten for
|
||||
the summary-seed mode — the ``<documents>`` section holds
|
||||
SUMMARIES (a suggested document's full text is not in the prompt
|
||||
until ``read`` adds it); do not re-read an already-read document
|
||||
(answer from the text already in the prompt); if the user asks to
|
||||
open or read a suggested document, ``read`` it. The phase-72 "do
|
||||
not call ``read`` for a ``<documents>`` document" copy is retired.
|
||||
The ``ls`` clause (phase 94 drill-down contract), the ``grep``
|
||||
clause, and the discipline rules stay byte-identical — the
|
||||
prompt lock's prefix/suffix anchors survive (see
|
||||
``test_prompt_lock``)."""
|
||||
# The new summary-seed copy (pinned byte-for-byte).
|
||||
assert (
|
||||
"The <documents> section holds SUMMARIES — the full text of a "
|
||||
"suggested document is not in your prompt until you `read` it. "
|
||||
"Do not re-read a document you have already read — its full "
|
||||
"text is already in your prompt; answer directly from it. If "
|
||||
"the user asks you to open or read a suggested document, "
|
||||
"`read` it — that is the point of the section."
|
||||
) in TOOLS_SECTION
|
||||
# The read identity handoff now also names the <documents> summary
|
||||
# blocks (the combined identity is shown there), keeping the
|
||||
# "including the source name" contract.
|
||||
assert (
|
||||
"exactly as shown in the `ls` output — including the source "
|
||||
"name — or in the <documents> summary blocks — adding its "
|
||||
"full content to your context"
|
||||
) in TOOLS_SECTION
|
||||
# The retired pre-phase-118 copy is gone.
|
||||
assert "Do not call `read` for a document already shown in" not in TOOLS_SECTION
|
||||
assert "even when the user asks you to open or read it" not in TOOLS_SECTION
|
||||
# The ls clause (phase 94 drill-down contract) — byte-identical.
|
||||
assert (
|
||||
"`ls` lists the knowledge base as a tree, one level at a "
|
||||
"time: with no `path` it lists every synced source with its "
|
||||
"document count and a summary of its contents; with a source "
|
||||
"name (e.g. 'homelab') it lists that source's top-level "
|
||||
"folders and files; with a `source/folder` path it drills one "
|
||||
"level deeper. A listing shows only that level's subfolders "
|
||||
"and its own files — never the whole knowledge base in one "
|
||||
"call — and each folder line's summary says what the folder "
|
||||
"contains before you drill into it. File lines are `source: X "
|
||||
"| path: Y | title: Z`; to find one specific document without "
|
||||
"listing, use `grep`."
|
||||
) in TOOLS_SECTION
|
||||
# The grep clause — byte-identical.
|
||||
assert (
|
||||
"`grep` locates an exact string (case-insensitive) in the "
|
||||
"indexed documents and returns up to 20 matching "
|
||||
"`source/path:line: text` lines — a locator, not a "
|
||||
"context-adder: read the winner with `read`. A grep pattern "
|
||||
"is a plain substring, NEVER a regex — '.*' and '\\.' are "
|
||||
"literal text there; if such a pattern returns no matches, "
|
||||
"retry with the plain text you expect to see. For a normal "
|
||||
"search pass only `pattern` — its optional `path` argument "
|
||||
"limits the search to one document you already know, by the "
|
||||
"same combined `source/path` string; never a source name — a "
|
||||
"bare document path (without the source name) will not "
|
||||
"resolve there either."
|
||||
) in TOOLS_SECTION
|
||||
# The discipline rules — byte-identical.
|
||||
assert (
|
||||
"Make exactly one tool call per reply — a reply carrying two "
|
||||
"tool calls runs only the first, the second is discarded — "
|
||||
"and wait for the result before the next call. Never repeat a "
|
||||
"call that was refused or already succeeded — the refusal "
|
||||
"already told you the correct form. Answer as soon as you "
|
||||
"have what you need."
|
||||
) in TOOLS_SECTION
|
||||
# The new read clause rides the built HIGH prompt and never the
|
||||
# LOW (deflected) prompt.
|
||||
doc = _doc("kubernetes.md", "Talos Linux on three nodes.", "Kubernetes", summary="K8S")
|
||||
high = build_high_prompt([doc])
|
||||
assert "The <documents> section holds SUMMARIES" in high
|
||||
assert "`read` it — that is the point of the section." in high
|
||||
low = build_deflect_prompt(["T1"])
|
||||
assert "The <documents> section holds SUMMARIES" not in low
|
||||
|
||||
|
||||
def test_tools_section_phase72_clauses_in_high_prompt_not_low() -> None:
|
||||
"""Phase 72/94: the contract clauses ride the HIGH prompt with the
|
||||
rest of the section and never leak into the LOW/deflection prompt
|
||||
@@ -273,20 +414,37 @@ def test_tools_section_phase72_clauses_in_high_prompt_not_low() -> None:
|
||||
assert "including the source name" not in low
|
||||
|
||||
|
||||
def test_documents_section_has_no_leading_intro() -> None:
|
||||
"""Phase 72, task 05 (gate iterations 2-3, reverted): the
|
||||
``<documents>`` section must NOT lead with an in-context reminder
|
||||
or name the ``<document>`` blocks — the live telemetry showed that
|
||||
copy primed the model to latch the seed documents' paths as
|
||||
``ls`` scopes (the incident turn regressed to a cap-reached loop
|
||||
on run 2 and re-trapped on run 5), and the reminder never flipped
|
||||
the seed-doc ``read``s (15/15 across gate runs 1-5). The section
|
||||
is exactly the document blocks again."""
|
||||
def test_documents_section_leads_with_the_suggest_intro() -> None:
|
||||
"""Phase 118, task 03: the ``<documents>`` section leads with the
|
||||
start-here :data:`SUGGEST_INTRO` line BEFORE the first block (the
|
||||
phase-15 ``_STEERING_INTRO`` / phase-31 ``_KB_INTRO`` precedent) —
|
||||
only when at least one block is present. This is the summary-as-
|
||||
starting-point framing, not the reverted phase-72 do-not-read
|
||||
reminder (that copy taught the seed texts as already-read context;
|
||||
with A6's summary seeding the blocks are starting points the model
|
||||
may ``read`` through to full text)."""
|
||||
doc = _doc("kubernetes.md", "Talos Linux on three nodes.", "Kubernetes Homelab Cluster")
|
||||
high = build_high_prompt([doc])
|
||||
i_open = high.index("<documents>")
|
||||
i_block = high.index('<document source="Homelab"')
|
||||
assert high[i_open : i_block] == "<documents>\n" # no intro line
|
||||
assert high[i_open : i_block] == f"<documents>\n{SUGGEST_INTRO}\n\n"
|
||||
# The intro occurs exactly once and sits inside the section.
|
||||
assert high.count(SUGGEST_INTRO) == 1
|
||||
assert high.index(SUGGEST_INTRO) > high.index("<documents>")
|
||||
assert high.index(SUGGEST_INTRO) < high.index("</documents>")
|
||||
|
||||
|
||||
def test_documents_section_without_blocks_has_no_intro() -> None:
|
||||
"""The intro rides only on present blocks: an empty ``<documents>``
|
||||
section is exactly the fallback line again (no intro, no blocks)."""
|
||||
high = build_high_prompt([])
|
||||
assert SUGGEST_INTRO not in high
|
||||
i_open = high.index("<documents>")
|
||||
i_close = high.index("</documents>")
|
||||
assert (
|
||||
high[i_open : i_close]
|
||||
== "<documents>\n(no documents matched — do not invent specifics)\n"
|
||||
)
|
||||
|
||||
|
||||
def test_tools_section_old_names_and_budget_copy_gone() -> None:
|
||||
@@ -314,6 +472,117 @@ def test_high_prompt_still_ends_with_tools_section() -> None:
|
||||
assert old not in prompt
|
||||
|
||||
|
||||
# ---------- phase 118 (task 03): the NULL-summary preview fallback (A5) ----------
|
||||
|
||||
|
||||
def _content_only_doc(content: str) -> Document:
|
||||
return _doc("kubernetes.md", content, "Kubernetes Homelab Cluster")
|
||||
|
||||
|
||||
def test_null_summary_falls_back_to_preview_plus_marker() -> None:
|
||||
"""A doc whose ``summary`` is None (a fail-soft import miss) seeds the
|
||||
first ``suggestion_preview_chars`` (default 400) content characters +
|
||||
the shared ``[…truncated…]`` marker on its own line — never the full
|
||||
content, never an LLM call."""
|
||||
content = "A" * 400 + "B" * 600 # 1000 chars
|
||||
prompt = build_high_prompt([_content_only_doc(content)])
|
||||
body_start = prompt.index('date="2024-06-15">\n') + len('date="2024-06-15">\n')
|
||||
body_end = prompt.index("\n</document>", body_start)
|
||||
assert prompt[body_start:body_end] == content[:400] + "\n" + TRUNCATION_MARKER
|
||||
assert "B" * 40 not in prompt # beyond the 400-char cut
|
||||
assert prompt.count(TRUNCATION_MARKER) == 1 # only the block's marker
|
||||
|
||||
|
||||
def test_whitespace_summary_uses_the_same_preview_fallback() -> None:
|
||||
for blank in ("", " ", "\n \t "):
|
||||
doc = _doc("kubernetes.md", "A" * 500 + "B" * 500, "T", summary=blank)
|
||||
prompt = build_high_prompt([doc])
|
||||
assert "A" * 400 in prompt
|
||||
assert "B" * 40 not in prompt
|
||||
assert TRUNCATION_MARKER in prompt
|
||||
|
||||
|
||||
def test_short_content_preview_is_the_whole_content_unmarked() -> None:
|
||||
"""Content at or under the cap rides whole — nothing was cut, so no
|
||||
marker (the marker signals truncation, not the fallback)."""
|
||||
for content in ("short body", "C" * 399, "D" * 400):
|
||||
prompt = build_high_prompt([_content_only_doc(content)])
|
||||
assert content + "\n</document>" in prompt
|
||||
assert TRUNCATION_MARKER not in prompt
|
||||
|
||||
|
||||
def test_summary_stripped_and_never_truncated_by_the_preview_cap() -> None:
|
||||
"""The summary body is the stripped ``doc.summary`` — even when it is
|
||||
longer than the preview cap (the cap bounds only the FALLBACK; a
|
||||
stored summary is trusted context, LOCKED A5)."""
|
||||
summary = "S" * 900
|
||||
content = "CONTENT_SENTINEL " + "x" * 50
|
||||
prompt = build_high_prompt([_doc("kubernetes.md", content, "T", summary=f" {summary}\n")])
|
||||
assert summary in prompt
|
||||
assert content not in prompt
|
||||
assert TRUNCATION_MARKER not in prompt
|
||||
|
||||
|
||||
def test_preview_cap_setting_is_honored_on_the_fallback_path(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""A non-default ``suggestion_preview_chars`` (``BOR_SUGGESTION_PREVIEW_CHARS``
|
||||
— the env mapping is pinned in :mod:`tests.unit.test_config`) cuts the
|
||||
preview at the setting's cap on the fallback path."""
|
||||
from app.rag import prompts as prompts_mod
|
||||
|
||||
monkeypatch.setattr(
|
||||
prompts_mod,
|
||||
"get_settings",
|
||||
lambda: Settings(_env_file=None, suggestion_preview_chars=10), # pyright: ignore[reportCallIssue]
|
||||
)
|
||||
content = "A" * 100 + "B" * 100
|
||||
prompt = build_high_prompt([_content_only_doc(content)])
|
||||
assert "A" * 10 in prompt
|
||||
assert "A" * 11 not in prompt
|
||||
assert TRUNCATION_MARKER in prompt
|
||||
|
||||
|
||||
def test_preview_fallback_never_reads_settings_for_summarized_docs(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""House pattern (``build_steering_section`` / ``build_kb_section``):
|
||||
the fallback cap is read from settings on the fallback path ONLY — a
|
||||
prompt built from summary-bearing docs makes no settings call for
|
||||
it (a dead ``get_settings`` must not break such a build)."""
|
||||
from app.rag import prompts as prompts_mod
|
||||
|
||||
doc = _doc("kubernetes.md", "CONTENT_SENTINEL", "T", summary="the summary")
|
||||
|
||||
def _dead() -> None: # pragma: no cover - must never be called
|
||||
raise AssertionError("get_settings() called for a summarized doc")
|
||||
|
||||
monkeypatch.setattr(prompts_mod, "get_settings", _dead)
|
||||
prompt = build_high_prompt([doc])
|
||||
assert "the summary\n</document>" in prompt
|
||||
assert "CONTENT_SENTINEL" not in prompt
|
||||
|
||||
|
||||
# ---------- phase 118: the LOW prompt stays byte-identical ----------
|
||||
|
||||
|
||||
def test_low_prompt_byte_identical_to_pre_task() -> None:
|
||||
"""LOCKED A8 surface: the deflection prompt is untouched by the
|
||||
summary seeding — byte-identical build on the same inputs (the full
|
||||
sha pin lives in :mod:`tests.unit.test_prompt_lock`)."""
|
||||
for titles, notes, kb in (
|
||||
(["T1", "T2"], None, None),
|
||||
(["T1"], ["be concise"], None),
|
||||
([], None, OVERVIEW),
|
||||
(["T1", "T2"], ["be concise"], OVERVIEW),
|
||||
):
|
||||
prompt = build_deflect_prompt(titles, notes=notes, kb_overview=kb)
|
||||
assert "DEFLECT_MODE" in prompt
|
||||
assert "<documents>" not in prompt
|
||||
assert SUGGEST_INTRO not in prompt
|
||||
assert "<tools>" not in prompt
|
||||
|
||||
|
||||
def test_relevance_placeholder_rejected_for_garbage() -> None:
|
||||
with pytest.raises(ValueError, match="HIGH or LOW"):
|
||||
_base("MEDIUM")
|
||||
@@ -401,7 +670,9 @@ def test_no_overview_prompt_is_byte_identical_to_pre_phase() -> None:
|
||||
every prompt is exactly what it was before the ``<knowledge_base>``
|
||||
section existed — with or without steering notes. (Phase 37: the HIGH
|
||||
prompt additionally carries the ``<tools>`` section after the mode
|
||||
body — the fixtures account for it; the LOW prompt is untouched.)"""
|
||||
body; phase 118: the ``<documents>`` section leads with the
|
||||
``SUGGEST_INTRO`` line — the fixtures account for both; the LOW
|
||||
prompt is untouched.)"""
|
||||
doc = _doc("kubernetes.md", "Talos Linux on three nodes.", "Kubernetes Homelab Cluster")
|
||||
block = (
|
||||
'<document source="Homelab" path="kubernetes.md" '
|
||||
@@ -409,7 +680,15 @@ def test_no_overview_prompt_is_byte_identical_to_pre_phase() -> None:
|
||||
"Talos Linux on three nodes.\n"
|
||||
"</document>"
|
||||
)
|
||||
docs_block = "\n<documents>\n" + block + "\n</documents>" + "\n" + TOOLS_SECTION
|
||||
docs_block = (
|
||||
"\n<documents>\n"
|
||||
+ SUGGEST_INTRO
|
||||
+ "\n\n"
|
||||
+ block
|
||||
+ "\n</documents>"
|
||||
+ "\n"
|
||||
+ TOOLS_SECTION
|
||||
)
|
||||
high_plain = _base("HIGH") + docs_block
|
||||
high_steered = _base("HIGH") + "\n" + build_steering_section(["be concise"]) + docs_block
|
||||
# Phase 71: the owner-permitted plain-text line is part of the
|
||||
|
||||
@@ -58,6 +58,7 @@ def _doc(
|
||||
title: str = "T",
|
||||
content: str = "CONTENT",
|
||||
created_at: datetime = CREATED_AT,
|
||||
summary: str | None = None,
|
||||
) -> Document:
|
||||
return Document(
|
||||
id=uuid.uuid4(),
|
||||
@@ -68,6 +69,7 @@ def _doc(
|
||||
content=content,
|
||||
content_hash="0" * 64,
|
||||
created_at=created_at,
|
||||
summary=summary,
|
||||
)
|
||||
|
||||
|
||||
@@ -171,6 +173,32 @@ def test_high_block_date_always_present_for_every_document() -> None:
|
||||
assert prompt.count("<document ") == prompt.count(' date="')
|
||||
|
||||
|
||||
def test_high_block_date_survives_the_summary_body_change() -> None:
|
||||
"""Phase 118 (task 03): the block BODY became the document's summary
|
||||
(never the full content) — the D5 identity attributes, including
|
||||
``date`` after ``title``, survive byte-identical around the new
|
||||
body (the E2E mock's block parser keys off exactly these)."""
|
||||
doc = _doc(
|
||||
source="S",
|
||||
path="P",
|
||||
title="T",
|
||||
content="FULL_CONTENT_SENTINEL_42",
|
||||
created_at=CREATED_AT,
|
||||
summary="The stored summary.",
|
||||
)
|
||||
prompt = build_high_prompt([doc])
|
||||
block = (
|
||||
f'<document source="S" path="P" title="T" date="{DATE}">\n'
|
||||
"The stored summary.\n"
|
||||
"</document>"
|
||||
)
|
||||
assert block in prompt
|
||||
# Attribute order pinned: date directly after title, body after.
|
||||
assert f'title="T" date="{DATE}">' in prompt
|
||||
# The full content stays out (A6) — only the summary rides the block.
|
||||
assert "FULL_CONTENT_SENTINEL_42" not in prompt
|
||||
|
||||
|
||||
# --------------------------------------------------------------------
|
||||
# The deflection prompt — byte-identical to the pre-phase text (A8)
|
||||
# --------------------------------------------------------------------
|
||||
|
||||
@@ -11,12 +11,16 @@ from types import SimpleNamespace
|
||||
|
||||
import pytest
|
||||
|
||||
from app.config import Settings
|
||||
from app.models import Document
|
||||
from app.rag import retriever
|
||||
from app.rag.retriever import (
|
||||
TRUNCATION_MARKER,
|
||||
RetrievedChunk,
|
||||
select_documents,
|
||||
select_documents_tiered,
|
||||
select_related,
|
||||
select_suggested,
|
||||
)
|
||||
|
||||
|
||||
@@ -258,6 +262,231 @@ def test_select_documents_wrapper_is_legacy_tiering() -> None:
|
||||
)[0]
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Phase 118 — select_suggested: the top-N "start here" tier, NO floor (A3)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_suggested_rank_order_by_first_seen_chunk() -> None:
|
||||
"""The SAME stable walk as ``select_documents_tiered``: a document's
|
||||
rank is fixed by its FIRST seen chunk in score-descending order — a
|
||||
doc whose best chunk appears later in the input list still ranks
|
||||
where that chunk falls."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
c = _doc("c.md", "C" * 50)
|
||||
chunks = [
|
||||
_chunk(a, 0.4, position=0), # a's weak chunk comes first
|
||||
_chunk(b, 0.8),
|
||||
_chunk(a, 0.9, position=2), # a's best chunk comes last
|
||||
_chunk(c, 0.5),
|
||||
]
|
||||
out = select_suggested(chunks, n=5)
|
||||
assert [d.path for d in out] == ["a.md", "b.md", "c.md"]
|
||||
# The rows carry the full content byte-identical (A6: the content is
|
||||
# what ``read`` serves later — never truncated).
|
||||
assert out[0].content == "A" * 50
|
||||
assert TRUNCATION_MARKER not in out[0].content
|
||||
|
||||
|
||||
def test_suggested_dedupes_multiple_chunks_to_one_row() -> None:
|
||||
"""Multiple hit chunks of one document collapse to a single row."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
chunks = [
|
||||
_chunk(a, 0.2),
|
||||
_chunk(b, 0.7),
|
||||
_chunk(a, 0.9, position=2),
|
||||
_chunk(a, 0.5),
|
||||
]
|
||||
out = select_suggested(chunks, n=5)
|
||||
assert [d.path for d in out] == ["a.md", "b.md"] # one row per document
|
||||
assert out[0] is a
|
||||
|
||||
|
||||
def test_suggested_caps_at_n_in_rank_order() -> None:
|
||||
docs_in = [_doc(f"d{i}.md", "X" * 20) for i in range(7)]
|
||||
chunks = [_chunk(d, 0.9 - 0.1 * i) for i, d in enumerate(docs_in)]
|
||||
out = select_suggested(chunks, n=3)
|
||||
assert [d.path for d in out] == ["d0.md", "d1.md", "d2.md"]
|
||||
|
||||
|
||||
def test_suggested_default_cap_is_the_settings_value(
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""``n`` omitted → ``BOR_SUGGESTED_DOCS`` caps the walk — default 5
|
||||
(LOCKED A3, the top-5 "start here" directive, TODO L3) — and the cap
|
||||
is the setting's LIVE value, not a frozen constant."""
|
||||
docs_in = [_doc(f"d{i}.md", "X" * 20) for i in range(7)]
|
||||
chunks = [_chunk(d, 0.9 - 0.1 * i) for i, d in enumerate(docs_in)]
|
||||
settings = Settings(_env_file=None) # pyright: ignore[reportCallIssue]
|
||||
assert settings.suggested_docs == 5 # the production default
|
||||
monkeypatch.setattr(retriever, "get_settings", lambda: settings)
|
||||
assert [d.path for d in select_suggested(chunks)] == [
|
||||
f"d{i}.md" for i in range(5)
|
||||
]
|
||||
small = Settings(_env_file=None, suggested_docs=2) # pyright: ignore[reportCallIssue]
|
||||
monkeypatch.setattr(retriever, "get_settings", lambda: small)
|
||||
assert [d.path for d in select_suggested(chunks)] == ["d0.md", "d1.md"]
|
||||
|
||||
|
||||
def test_suggested_never_filters_on_cosine_floor() -> None:
|
||||
"""NO floor (LOCKED A3): a lexical-only hit (cosine 0.0 by
|
||||
construction) is a suggestion when it ranks — the contrast pin
|
||||
against ``select_documents_tiered``'s floored cited tier on the SAME
|
||||
input, which demotes it to the related tier."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
lexical_only = RetrievedChunk(
|
||||
chunk_id=uuid.uuid4(),
|
||||
position=0,
|
||||
content="X" * 10,
|
||||
score=0.9, # top fused rank (the FTS hit)
|
||||
document=a,
|
||||
cosine=0.0, # no vector rank — lexical-only
|
||||
fts_hit=True,
|
||||
)
|
||||
chunks = [lexical_only, _cos_chunk(b, 0.8, 0.5)]
|
||||
# Suggested: the floor never filters — a leads, b follows in rank order.
|
||||
assert [d.path for d in select_suggested(chunks, n=5)] == ["a.md", "b.md"]
|
||||
# Contrast: the same input through the phase-113 cited tier — the
|
||||
# 0.35 usefulness bar demotes the lexical-only doc to related.
|
||||
cited, related = select_documents_tiered(chunks, n=2, floor=0.35, related_cap=2)
|
||||
assert [d.path for d in cited] == ["b.md"]
|
||||
assert [d.path for d in related] == ["a.md"]
|
||||
|
||||
|
||||
def test_suggested_tie_break_inherited_from_fused_order() -> None:
|
||||
"""Equal fused scores keep the input (fused) order — the stable
|
||||
score-only walk inherits ``fuse()``'s (-score, -cosine, path,
|
||||
position) tie-break; the selector never re-sorts it away."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
chunks = [_cos_chunk(a, 0.7, 0.5), _cos_chunk(b, 0.7, 0.4)] # a wins on cosine
|
||||
assert [d.path for d in select_suggested(chunks, n=5)] == ["a.md", "b.md"]
|
||||
# A FULL tie (score AND cosine): the input position — ``fuse()``'s
|
||||
# path/position tie-break already applied — decides. ``m.md`` sorts
|
||||
# AFTER ``b2.md`` alphabetically, so any re-sort by path would flip
|
||||
# the order; the fused input order must win.
|
||||
m = _doc("m.md", "M" * 50)
|
||||
b2 = _doc("b2.md", "B" * 50)
|
||||
chunks = [_cos_chunk(m, 0.7, 0.4), _cos_chunk(b2, 0.7, 0.4)]
|
||||
assert [d.path for d in select_suggested(chunks, n=5)] == ["m.md", "b2.md"]
|
||||
|
||||
|
||||
def test_suggested_empty_chunks_yield_no_documents() -> None:
|
||||
assert select_suggested([], n=5) == []
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Phase 118, task 05 — select_related: the rank-6+ tier after the suggested set
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_related_walk_order_after_excluded_set() -> None:
|
||||
"""The SAME stable score-descending walk as ``select_suggested``:
|
||||
a document's rank is fixed by its FIRST seen chunk; documents in
|
||||
*excluded_ids* (the suggested set) are skipped and the rest come
|
||||
back in rank order."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
c = _doc("c.md", "C" * 50)
|
||||
d = _doc("d.md", "D" * 50)
|
||||
chunks = [
|
||||
_chunk(a, 0.4, position=0), # a's weak chunk comes first
|
||||
_chunk(b, 0.8),
|
||||
_chunk(a, 0.9, position=2), # a's best chunk last — a ranks first
|
||||
_chunk(c, 0.5),
|
||||
_chunk(d, 0.3),
|
||||
]
|
||||
out = select_related(chunks, {a.id, b.id}, cap=2)
|
||||
assert [x.path for x in out] == ["c.md", "d.md"]
|
||||
|
||||
|
||||
def test_related_skips_excluded_documents() -> None:
|
||||
"""Every document in *excluded_ids* is skipped, even when it would
|
||||
rank inside the cap — the suggested set never rides the related row."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
c = _doc("c.md", "C" * 50)
|
||||
chunks = [_chunk(a, 0.9), _chunk(b, 0.8), _chunk(c, 0.5)]
|
||||
out = select_related(chunks, {a.id, b.id}, cap=5)
|
||||
assert out == [c]
|
||||
# Every doc excluded → empty, even with room left in the cap.
|
||||
assert select_related(chunks, {a.id, b.id, c.id}, cap=5) == []
|
||||
|
||||
|
||||
def test_related_caps_at_cap_in_rank_order() -> None:
|
||||
"""The phase-118 turn wiring on 9 docs: suggested = the top 5,
|
||||
related = rank 6–7 (capped at 2), disjoint from the suggested set."""
|
||||
docs_in = [_doc(f"d{i}.md", "X" * 20) for i in range(9)]
|
||||
chunks = [_chunk(d, 0.9 - 0.1 * i) for i, d in enumerate(docs_in)]
|
||||
suggested = select_suggested(chunks, n=5)
|
||||
out = select_related(chunks, {d.id for d in suggested}, cap=2)
|
||||
assert [x.path for x in out] == ["d5.md", "d6.md"] # rank 6–7, capped
|
||||
suggested_paths = {d.path for d in suggested}
|
||||
assert suggested_paths.isdisjoint({x.path for x in out})
|
||||
|
||||
|
||||
def test_related_cap_zero_yields_empty() -> None:
|
||||
"""cap=0 is the kill switch (related_max_docs=0): no related docs,
|
||||
the pre-phase-113 visibility."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
c = _doc("c.md", "C" * 50)
|
||||
chunks = [_chunk(a, 0.9), _chunk(b, 0.8), _chunk(c, 0.5)]
|
||||
assert select_related(chunks, {a.id}, cap=0) == []
|
||||
|
||||
|
||||
def test_related_never_filters_on_cosine_floor() -> None:
|
||||
"""NO floor: a lexical-only (cosine 0.0) doc is related when it
|
||||
ranks after the excluded set — the related tier is visibility, not
|
||||
citation (phase 118 applies no cosine floor to it)."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
lexical_only = RetrievedChunk(
|
||||
chunk_id=uuid.uuid4(),
|
||||
position=0,
|
||||
content="X" * 10,
|
||||
score=0.9, # top fused rank (the FTS hit)
|
||||
document=a,
|
||||
cosine=0.0, # no vector rank — lexical-only
|
||||
fts_hit=True,
|
||||
)
|
||||
chunks = [lexical_only, _cos_chunk(b, 0.8, 0.5)]
|
||||
out = select_related(chunks, set(), cap=5)
|
||||
assert [x.path for x in out] == ["a.md", "b.md"]
|
||||
|
||||
|
||||
def test_related_dedupes_multiple_chunks_to_one_row() -> None:
|
||||
"""Multiple hit chunks of one document collapse to a single row
|
||||
(first-seen-chunk rank, dedupe by document.id — the shared walk)."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
chunks = [_chunk(a, 0.2), _chunk(b, 0.7), _chunk(a, 0.9, position=2)]
|
||||
out = select_related(chunks, set(), cap=5)
|
||||
assert [x.path for x in out] == ["a.md", "b.md"] # one row per document
|
||||
assert out[0] is a
|
||||
|
||||
|
||||
def test_related_tie_break_inherited_from_fused_order() -> None:
|
||||
"""Equal fused scores keep the input (fused) order — the stable
|
||||
score-only walk inherits ``fuse()``'s (-score, -cosine, path,
|
||||
position) tie-break; the selector never re-sorts it away."""
|
||||
a = _doc("a.md", "A" * 50)
|
||||
b = _doc("b.md", "B" * 50)
|
||||
chunks = [_cos_chunk(a, 0.7, 0.5), _cos_chunk(b, 0.7, 0.4)] # a wins on cosine
|
||||
assert [x.path for x in select_related(chunks, set(), cap=5)] == ["a.md", "b.md"]
|
||||
m = _doc("m.md", "M" * 50)
|
||||
b2 = _doc("b2.md", "B" * 50)
|
||||
chunks = [_cos_chunk(m, 0.7, 0.4), _cos_chunk(b2, 0.7, 0.4)] # full tie
|
||||
assert [x.path for x in select_related(chunks, set(), cap=5)] == ["m.md", "b2.md"]
|
||||
|
||||
|
||||
def test_related_empty_chunks_yield_no_documents() -> None:
|
||||
assert select_related([], set(), cap=5) == []
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Hybrid retrieval (A7): RRF fusion + lexical tsquery
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@@ -331,7 +331,16 @@ def test_extra_keys_still_forbidden() -> None:
|
||||
def test_minimal_message_still_validates() -> None:
|
||||
"""Optional keys may be ABSENT exactly as pre-phase-83."""
|
||||
msg = ChatMessage.model_validate({"who": "user", "text": "hi"})
|
||||
assert (msg.sources, msg.deflected, msg.suggestions, msg.thinking, msg.tools, msg.stopped) == (
|
||||
assert (
|
||||
msg.sources,
|
||||
msg.related,
|
||||
msg.deflected,
|
||||
msg.suggestions,
|
||||
msg.thinking,
|
||||
msg.tools,
|
||||
msg.stopped,
|
||||
) == (
|
||||
None,
|
||||
None,
|
||||
None,
|
||||
None,
|
||||
@@ -358,6 +367,7 @@ def test_realistic_bor_chat_v1_payload_round_trips() -> None:
|
||||
"who": "user",
|
||||
"text": "How did I install k3s on the new node?",
|
||||
"sources": None,
|
||||
"related": None,
|
||||
"deflected": None,
|
||||
"suggestions": None,
|
||||
"thinking": None,
|
||||
@@ -371,6 +381,14 @@ def test_realistic_bor_chat_v1_payload_round_trips() -> None:
|
||||
{"source": "Homelab", "path": "kubernetes.md", "title": "Kubernetes Cluster"},
|
||||
{"source": "Deployments", "path": "k3s-install.md", "title": "k3s Install Notes"},
|
||||
],
|
||||
# Phase 113 related-doc tier — an ACCEPTED key (the phase-113
|
||||
# omission of this field made extra="forbid" 422 every
|
||||
# done-time auto-save carrying it, so grounded turns' brain
|
||||
# messages never persisted — the A2 quiet failure swallowed
|
||||
# the 422). Round-trips like sources.
|
||||
"related": [
|
||||
{"source": "Homelab", "path": "traefik.md", "title": "Traefik Notes"}
|
||||
],
|
||||
"deflected": False,
|
||||
"suggestions": None,
|
||||
"thinking": "The kubernetes doc covers the cluster layout…",
|
||||
@@ -398,6 +416,7 @@ def test_realistic_bor_chat_v1_payload_round_trips() -> None:
|
||||
"who": "user",
|
||||
"text": "And what ports does Traefik expose?",
|
||||
"sources": None,
|
||||
"related": None,
|
||||
"deflected": None,
|
||||
"suggestions": None,
|
||||
"thinking": None,
|
||||
@@ -408,6 +427,7 @@ def test_realistic_bor_chat_v1_payload_round_trips() -> None:
|
||||
"who": "brain",
|
||||
"text": "Traefik exposes 80/443 on every node.",
|
||||
"sources": None,
|
||||
"related": None,
|
||||
"deflected": None,
|
||||
"suggestions": ["What is the Traefik dashboard password?"],
|
||||
"thinking": None,
|
||||
@@ -433,6 +453,7 @@ def test_realistic_payload_round_trips_through_update_model() -> None:
|
||||
"who": "brain",
|
||||
"text": "answer",
|
||||
"sources": [_source_ref()],
|
||||
"related": None,
|
||||
"deflected": None,
|
||||
"suggestions": ["follow-up?"],
|
||||
"thinking": "scratchpad",
|
||||
|
||||
@@ -1,13 +1,15 @@
|
||||
"""Unit: the phase-113 source-chip-quality contract (TODO L5 + L2c —
|
||||
"the 2nd chip is often noise the answer never used").
|
||||
|
||||
Phase 113 demotes sub-floor hits out of the citation surface: the
|
||||
done frame carries the cited tier in ``sources`` (rendered by
|
||||
``appendSources`` as ``.source-chip`` pills, UNCHANGED) and the
|
||||
related tier in ``related`` (rendered by the NEW ``appendRelated`` as
|
||||
the de-emphasized labeled row — ``.related-doc`` links, never
|
||||
``.source-chip``). A deflected turn carries ``sources: []`` → zero
|
||||
chips; its weak hits live in the related row only.
|
||||
Phase 118 re-tiers the same frame (LOCKED A3/A4): ``sources`` carries
|
||||
the suggested tier (top-5, NO floor) + the agent-read docs (deduped)
|
||||
— rendered by ``appendSources`` as ``.source-chip`` pills, UNCHANGED —
|
||||
and ``related`` carries rank 6+ after the suggested set (rendered by
|
||||
``appendRelated`` as the de-emphasized labeled row — ``.related-doc``
|
||||
links, never ``.source-chip``). A deflected turn carries
|
||||
``sources: []`` → zero chips; its weak hits are the suggested tier
|
||||
(the durable record), and with ≤5 retrieved docs the related row is
|
||||
empty.
|
||||
|
||||
This module pins the STATIC SOURCES the UI contract stands on, in the
|
||||
house source-pin pattern (the test_chip_sizing_question_cap.py
|
||||
@@ -26,24 +28,23 @@ house source-pin pattern (the test_chip_sizing_question_cap.py
|
||||
|
||||
1. **both docs weak** ("What is the capital of Mongolia?" →
|
||||
``Trooper_Nagraz.pl`` + ``Trooper_Begzei.pl``, both unrelated) —
|
||||
cited tier empty, the weak hits ride the related tier, capped at
|
||||
``related_max_docs``; the FTS hit without vector corroboration
|
||||
stays LOW (the A8-revised "Mongolia" case);
|
||||
the weak hits are SUGGESTED (no floor, A3; the durable record),
|
||||
nothing left for the related tier; the FTS hit without vector
|
||||
corroboration stays LOW (the A8-revised "Mongolia" case);
|
||||
2. **one strong + one weak** (the phase-gate question answered from
|
||||
``brain-of-reese/.agents/validate.sh``; the 2nd chip
|
||||
``ServMon/README.md`` unused) — exactly ONE cited ref, the weak
|
||||
doc in ``related``;
|
||||
``brain-of-reese/.agents/validate.sh``; ``ServMon/README.md``
|
||||
alongside) — BOTH suggested (no floor) ⇒ both cited refs (A4),
|
||||
no related tier;
|
||||
3. **the Nagraz case** (``Trooper_Nagraz.pl`` strong,
|
||||
``Trooper_Byzin.pl`` weak — same shape, different fixtures);
|
||||
4. **the meta/history question** (no doc clears the bar, the agent
|
||||
reads nothing — chips ``app/api/suggestions.py`` +
|
||||
``108_history_wire_check/00_phase.md``, neither used) — pinned on
|
||||
the DONE FRAME (endpoint-level, fake retriever/LLM/session): the
|
||||
frame is row-only — ``sources: []`` (the UI's chip list — zero
|
||||
chips) + the weak hits in ``related``;
|
||||
5. **the agent-read exemption** (LOCKED A2): a below-floor doc the
|
||||
agent ``read`` via the tool joins ``sources`` (cited, last) and is
|
||||
excluded from ``related``.
|
||||
4. **the meta/history question** (the agent reads nothing —
|
||||
``app/api/suggestions.py`` + ``108_history_wire_check/00_phase.md``
|
||||
alongside) — pinned on the DONE FRAME (endpoint-level, fake
|
||||
retriever/LLM/session): ``sources: []`` (zero chips) and an EMPTY
|
||||
related row (both weak docs are suggested, ≤5 docs retrieved);
|
||||
5. **the agent-read exemption** (LOCKED A4): a rank-6+ (related-
|
||||
tier) doc the agent ``read`` via the tool joins ``sources``
|
||||
(cited, last) and is excluded from ``related``.
|
||||
|
||||
The browser behavior (chip counts on a single-source question, zero
|
||||
chips on a deflected turn) is E2E-gated by
|
||||
@@ -411,14 +412,18 @@ def _shape_settings() -> Settings:
|
||||
"""The PRODUCTION calibration (the code defaults, explicit) — the
|
||||
four shapes were observed LIVE under this threshold/floor pair.
|
||||
``_env_file=None`` keeps the mock-calibrated values from
|
||||
``tests/conftest.py`` (and any local ``.env``) out of the pin."""
|
||||
``tests/conftest.py`` (and any local ``.env``) out of the pin.
|
||||
``source_usefulness_floor`` / ``top_n_docs`` are legacy phase-113
|
||||
settings — phase 118 retired their seeding role (A6; ``plan_turn``
|
||||
never consults them), they are carried here for completeness."""
|
||||
return Settings(
|
||||
_env_file=None, # pyright: ignore[reportCallIssue]
|
||||
relevance_threshold=0.62,
|
||||
lexical_support_floor=0.35,
|
||||
source_usefulness_floor=0.35, # LOCKED A2 default
|
||||
related_max_docs=2, # LOCKED A4 default
|
||||
top_n_docs=2, # the ceiling — never a quota (LOCKED A2)
|
||||
source_usefulness_floor=0.35, # retired by phase 118 (A6) — not consulted
|
||||
related_max_docs=2, # the rank-6+ row cap (LOCKED A4)
|
||||
top_n_docs=2, # retired by phase 118 (A6) — not consulted
|
||||
suggested_docs=5, # the "start here" cap (LOCKED A3)
|
||||
)
|
||||
|
||||
|
||||
@@ -461,12 +466,12 @@ def _chunk(
|
||||
def test_shape_1_mongolia_both_docs_weak_cite_nothing() -> None:
|
||||
"""Observed shape 1 (TODO L110–113): "What is the capital of
|
||||
Mongolia?" → chips ``Trooper_Nagraz.pl`` + ``Trooper_Begzei.pl``,
|
||||
BOTH unrelated. Both below the bar: the cited tier is EMPTY (zero
|
||||
citation chips) and the weak hits ride the related tier — in rank
|
||||
order, capped at ``related_max_docs`` (the 3rd weak doc drops out).
|
||||
The FTS hit without vector corroboration (0.20 < the 0.35 lexical
|
||||
floor) stays LOW — the A8-revised "Mongolia" case; the weak content
|
||||
never reaches the LLM."""
|
||||
BOTH unrelated. Phase 118 (A3): the floor never filters — the weak
|
||||
hits are the SUGGESTED tier (the durable record's input), and with
|
||||
three retrieved docs nothing is left for the related tier (rank
|
||||
6+). The FTS hit without vector corroboration (0.20 < the 0.35
|
||||
lexical floor) stays LOW — the A8-revised "Mongolia" case; the weak
|
||||
content never reaches the LLM (titles only)."""
|
||||
nagraz = _doc("scripts", "scripts/Trooper_Nagraz.pl", "Trooper_Nagraz.pl",
|
||||
"NAGRAZ_PL_CONTENT")
|
||||
begzei = _doc("scripts", "scripts/Trooper_Begzei.pl", "Trooper_Begzei.pl",
|
||||
@@ -476,64 +481,76 @@ def test_shape_1_mongolia_both_docs_weak_cite_nothing() -> None:
|
||||
chunks = [
|
||||
_chunk(nagraz, 0.033, cosine=0.20, fts_hit=True), # rank 1, lexical hit
|
||||
_chunk(begzei, 0.031, cosine=0.12),
|
||||
_chunk(third, 0.030, cosine=0.10), # below the cap — related drops it
|
||||
_chunk(third, 0.030, cosine=0.10),
|
||||
]
|
||||
plan = chat_api.plan_turn(chunks, _shape_settings())
|
||||
assert plan.deflected is True # 0.20 < 0.62 AND 0.20 < the 0.35 lex floor
|
||||
assert plan.docs == [] # NO citation slot below the bar
|
||||
assert [d.title for d in plan.related_docs] == [
|
||||
# No floor (A3): the weak hits are suggested, in rank order (≤5).
|
||||
assert [d.title for d in plan.suggested_docs] == [
|
||||
"Trooper_Nagraz.pl",
|
||||
"Trooper_Begzei.pl",
|
||||
] # rank order, capped at related_max_docs (2)
|
||||
assert len(plan.related_docs) <= 2
|
||||
"Trooper_Third.pl",
|
||||
]
|
||||
assert plan.related_docs == [] # no rank-6+ doc among 3 retrieved
|
||||
# The LOW prompt is titles only — none of the weak content is sent.
|
||||
assert "NAGRAZ_PL_CONTENT" not in plan.system_prompt
|
||||
assert "Trooper_Nagraz.pl" in plan.system_prompt # weak-hit titles carried
|
||||
assert plan.suggestions # the "Maybe try" chips are unchanged
|
||||
|
||||
|
||||
def test_shape_2_validate_sh_strong_plus_unused_second_chip() -> None:
|
||||
def test_shape_2_validate_sh_strong_plus_weak_second_suggested() -> None:
|
||||
"""Observed shape 2 (TODO L114–116): the phase-gate question is
|
||||
answered from ``brain-of-reese/.agents/validate.sh`` — the 2nd chip
|
||||
``ServMon/README.md`` was NEVER used. The strong doc clears the bar
|
||||
and takes the only cited slot (top_n_docs is a ceiling, not a
|
||||
quota); the weak 2nd doc demotes to related — never a citation.
|
||||
The HIGH prompt carries the cited content only."""
|
||||
answered from ``brain-of-reese/.agents/validate.sh`` with
|
||||
``ServMon/README.md`` retrieved alongside (weak cosine). Phase 118
|
||||
(A3): the floor never filters — the weak 2nd doc is SUGGESTED too
|
||||
(both docs seed the HIGH prompt as summaries; the A5 fallback
|
||||
carries the short fixture content whole), and A4 makes both
|
||||
citation refs on the done frame — the "unused 2nd chip" is the
|
||||
phase-113 shape, retired by the owner directive."""
|
||||
validate = _doc("brain-of-reese", ".agents/validate.sh", "validate.sh",
|
||||
"VALIDATE_SH_CONTENT")
|
||||
servmon = _doc("ServMon", "README.md", "ServMon README",
|
||||
"SERVMON_README_CONTENT")
|
||||
chunks = [
|
||||
_chunk(validate, 0.90, cosine=0.70), # clears threshold AND bar
|
||||
_chunk(validate, 0.90, cosine=0.70), # clears the threshold
|
||||
_chunk(servmon, 0.80, cosine=0.20), # high fused rank, weak cosine
|
||||
]
|
||||
plan = chat_api.plan_turn(chunks, _shape_settings())
|
||||
assert plan.deflected is False # 0.70 >= 0.62
|
||||
assert [d.title for d in plan.docs] == ["validate.sh"] # exactly ONE cited
|
||||
assert [d.title for d in plan.related_docs] == ["ServMon README"]
|
||||
assert "VALIDATE_SH_CONTENT" in plan.system_prompt
|
||||
assert "SERVMON_README_CONTENT" not in plan.system_prompt
|
||||
assert [d.title for d in plan.suggested_docs] == [
|
||||
"validate.sh",
|
||||
"ServMon README",
|
||||
] # both suggested (no floor), rank order
|
||||
assert plan.related_docs == [] # nothing beyond rank 2 for 2 docs
|
||||
assert "VALIDATE_SH_CONTENT" in plan.system_prompt # A5 preview fallback
|
||||
assert "SERVMON_README_CONTENT" in plan.system_prompt # ditto
|
||||
|
||||
|
||||
def test_shape_3_nagraz_answered_by_own_doc_byzin_uncited() -> None:
|
||||
def test_shape_3_nagraz_answered_by_own_doc_byzin_suggested() -> None:
|
||||
"""Observed shape 3 (TODO L117–119): the Trooper_Nagraz question is
|
||||
answered from ``Trooper_Nagraz.pl`` — the 2nd chip
|
||||
``Trooper_Byzin.pl`` uncited. The SAME shape as shape 2 with
|
||||
different fixtures — the bar filters the 2nd chip; it is not a
|
||||
coincidence of the validate.sh pair."""
|
||||
answered from ``Trooper_Nagraz.pl`` with ``Trooper_Byzin.pl``
|
||||
retrieved alongside (weak cosine). The SAME shape as shape 2 with
|
||||
different fixtures — phase 118's no-floor tiering suggests BOTH
|
||||
(the phase-113 "bar filters the 2nd chip" story is retired); the
|
||||
HIGH prompt seeds both summaries (A5 fallback for the short
|
||||
fixture content)."""
|
||||
nagraz = _doc("scripts", "scripts/Trooper_Nagraz.pl", "Trooper_Nagraz.pl",
|
||||
"NAGRAZ_PL_CONTENT")
|
||||
byzin = _doc("scripts", "scripts/Trooper_Byzin.pl", "Trooper_Byzin.pl",
|
||||
"BYZIN_PL_CONTENT")
|
||||
chunks = [
|
||||
_chunk(nagraz, 0.85, cosine=0.70),
|
||||
_chunk(byzin, 0.75, cosine=0.15), # below the bar
|
||||
_chunk(byzin, 0.75, cosine=0.15), # weak cosine — still suggested (A3)
|
||||
]
|
||||
plan = chat_api.plan_turn(chunks, _shape_settings())
|
||||
assert plan.deflected is False
|
||||
assert [d.title for d in plan.docs] == ["Trooper_Nagraz.pl"] # 1 cited
|
||||
assert [d.title for d in plan.related_docs] == ["Trooper_Byzin.pl"] # 1 related
|
||||
assert "BYZIN_PL_CONTENT" not in plan.system_prompt
|
||||
assert [d.title for d in plan.suggested_docs] == [
|
||||
"Trooper_Nagraz.pl",
|
||||
"Trooper_Byzin.pl",
|
||||
] # both suggested (no floor), rank order
|
||||
assert plan.related_docs == []
|
||||
assert "NAGRAZ_PL_CONTENT" in plan.system_prompt # A5 preview fallback
|
||||
assert "BYZIN_PL_CONTENT" in plan.system_prompt # ditto
|
||||
|
||||
|
||||
# ---------- task 03: done-frame wire (endpoint-level fakes, no stack) ----------
|
||||
@@ -664,21 +681,22 @@ def _fake_retriever(chunks: list[RetrievedChunk]) -> Any:
|
||||
return retrieve
|
||||
|
||||
|
||||
def test_shape_4_meta_question_deflected_frame_is_row_only(
|
||||
def test_shape_4_meta_question_deflected_frame_has_no_chips_or_row(
|
||||
client: TestClient,
|
||||
chip_env: tuple[_FakeSession, _CannedLLM],
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""Observed shape 4 (TODO L120–123), pinned on the DONE FRAME: a
|
||||
meta question about the conversation's own history → chips
|
||||
meta question about the conversation's own history → weak hits
|
||||
``app/api/suggestions.py`` + ``108_history_wire_check/00_phase.md``,
|
||||
neither used. No doc clears the bar and the agent reads nothing —
|
||||
the frame is ROW-ONLY: ``sources: []`` (the UI chips every source
|
||||
entry — zero chips) with the weak hits in ``related`` (rank order,
|
||||
≤ ``related_max_docs``) — the de-emphasized row's links (the row's
|
||||
rendering itself is pinned by task 02's source tests + the E2E).
|
||||
The weak retrieval stays durably recorded (LOCKED A3); the weak
|
||||
content never reaches the LLM (LOW prompt, titles only)."""
|
||||
neither used. Phase 118: the agent reads nothing, the weak hits
|
||||
are the SUGGESTED tier (no floor, A3 — the durable record's input)
|
||||
and, with only two retrieved docs, nothing reaches rank 6+ — the
|
||||
frame carries ``sources: []`` (zero chips — a deflected answer
|
||||
cites nothing) AND an empty ``related`` row (the row's rendering
|
||||
itself is pinned by task 02's source tests + the E2E). The weak
|
||||
retrieval stays durably recorded (LOCKED A3); the weak content
|
||||
never reaches the LLM (LOW prompt, titles only)."""
|
||||
session, llm = chip_env
|
||||
suggestions = _doc("brain-of-reese", "app/api/suggestions.py",
|
||||
"suggestions.py", "SUGGESTIONS_PY_CONTENT")
|
||||
@@ -701,13 +719,9 @@ def test_shape_4_meta_question_deflected_frame_is_row_only(
|
||||
assert done["type"] == "done"
|
||||
assert done["deflected"] is True
|
||||
assert done["sources"] == [] # zero citation chips on the wire
|
||||
related = done["related"]
|
||||
assert [(s["source"], s["path"]) for s in related] == [
|
||||
("brain-of-reese", "app/api/suggestions.py"),
|
||||
("brain-of-reese", ".agents/108_history_wire_check/00_phase.md"),
|
||||
] # rank order
|
||||
assert len(related) <= 2 # related_max_docs
|
||||
assert all(s["title"] for s in related) # the row's links carry the identity
|
||||
# Phase 118 (A3): both weak docs are suggested (≤5, no floor) —
|
||||
# nothing reaches rank 6+, so the related row is empty.
|
||||
assert done["related"] == []
|
||||
assert done["suggestions"] # the "Maybe try" chips are unchanged
|
||||
|
||||
(system, _user) = llm.seen[0][0], llm.seen[0][1]
|
||||
@@ -722,16 +736,17 @@ def test_shape_4_meta_question_deflected_frame_is_row_only(
|
||||
assert "108_history_wire_check/00_phase.md" in row.sources
|
||||
|
||||
|
||||
def test_done_frame_single_cited_ref_strong_plus_weak(
|
||||
def test_done_frame_carries_suggested_refs_strong_plus_weak(
|
||||
client: TestClient,
|
||||
chip_env: tuple[_FakeSession, _CannedLLM],
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""Shape 2 on the wire — the single-document question's input to
|
||||
"exactly one citation chip" (the E2E asserts the rendered chip):
|
||||
the bar-clearing doc is the ONLY ``sources`` ref; the weak 2nd doc
|
||||
rides ``related``; the tiers are disjoint (the done frame's dedupe).
|
||||
The durable record keeps the FULL retrieval (LOCKED A3)."""
|
||||
"""Shape 2 on the wire under phase 118 (LOCKED A4): the citation
|
||||
surface is the suggested tier + the agent's reads (deduped) —
|
||||
with two retrieved docs and no read, BOTH docs are ``sources``
|
||||
refs (no floor — A3); nothing reaches rank 6+, so ``related`` is
|
||||
empty; the tiers stay disjoint (the done frame's dedupe). The
|
||||
durable record keeps the FULL retrieval (LOCKED A3)."""
|
||||
session, _llm = chip_env
|
||||
validate = _doc("brain-of-reese", ".agents/validate.sh", "validate.sh",
|
||||
"VALIDATE_SH_CONTENT")
|
||||
@@ -753,10 +768,9 @@ def test_done_frame_single_cited_ref_strong_plus_weak(
|
||||
assert done["deflected"] is False
|
||||
assert [(s["source"], s["path"]) for s in done["sources"]] == [
|
||||
("brain-of-reese", ".agents/validate.sh"),
|
||||
] # EXACTLY one citation chip on the wire
|
||||
assert [(s["source"], s["path"]) for s in done["related"]] == [
|
||||
("ServMon", "README.md"),
|
||||
("ServMon", "README.md"), # A4: suggested + read — both suggested (A3)
|
||||
]
|
||||
assert done["related"] == [] # nothing reaches rank 6+ for 2 docs
|
||||
cited = {(s["source"], s["path"]) for s in done["sources"]}
|
||||
related = {(s["source"], s["path"]) for s in done["related"]}
|
||||
assert cited.isdisjoint(related)
|
||||
@@ -769,20 +783,25 @@ def test_done_frame_single_cited_ref_strong_plus_weak(
|
||||
assert "ServMon/README.md" in row.sources
|
||||
|
||||
|
||||
def test_agent_read_below_floor_doc_joins_sources(
|
||||
def test_agent_read_related_doc_is_cited_not_related(
|
||||
client: TestClient,
|
||||
chip_env: tuple[_FakeSession, _CannedLLM],
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
) -> None:
|
||||
"""The agent-read exemption (LOCKED A2): a doc UNDER the bar that
|
||||
the agent ``read`` via the tool is cited by definition — the model
|
||||
read it, so it was used. It joins ``sources`` (after the retrieved
|
||||
cited docs, deduped) and is EXCLUDED from ``related`` (a used doc
|
||||
must never read as "nearby"); the other below-floor doc stays in
|
||||
the tier. The read content reached the model (the tool result in
|
||||
the follow-up request)."""
|
||||
"""The agent-read exemption (LOCKED A4, phase-118 tiering): a
|
||||
rank-6+ doc — the related tier ("nearby docs") — that the agent
|
||||
``read`` via the tool is cited by definition: the model read it, so
|
||||
it was used. It joins ``sources`` (after the suggested docs — it
|
||||
was not suggested, so the read appends it last, deduped) and is
|
||||
EXCLUDED from ``related`` (a used doc must never read as "nearby");
|
||||
the other rank-6+ doc stays in the tier. The read content reached
|
||||
the model (the tool result in the follow-up request)."""
|
||||
session, _default_llm = chip_env
|
||||
strong = _doc("docs", "strong.md", "Strong", "STRONG_DOC_CONTENT")
|
||||
fillers = [
|
||||
_doc("docs", f"filler{i}.md", f"Filler {i}", f"FILLER_{i}_CONTENT")
|
||||
for i in range(1, 5) # ranks 2–5 — fill the suggested tier
|
||||
]
|
||||
weak_b = _doc("docs", "weak-b.md", "Weak B", "WEAK_B_READ_BY_AGENT")
|
||||
weak_c = _doc("docs", "weak-c.md", "Weak C", "WEAK_C_CONTENT")
|
||||
monkeypatch.setattr(
|
||||
@@ -790,9 +809,13 @@ def test_agent_read_below_floor_doc_joins_sources(
|
||||
"retrieve",
|
||||
_fake_retriever(
|
||||
[
|
||||
_chunk(strong, 0.90, cosine=0.70), # clears the bar
|
||||
_chunk(weak_b, 0.80, cosine=0.20), # below the bar — read by the agent
|
||||
_chunk(weak_c, 0.70, cosine=0.10), # below the bar — nobody reads it
|
||||
_chunk(strong, 0.90, cosine=0.70), # clears the threshold (rank 1)
|
||||
_chunk(fillers[0], 0.85, cosine=0.30), # ranks 2–5: suggested
|
||||
_chunk(fillers[1], 0.80, cosine=0.30),
|
||||
_chunk(fillers[2], 0.75, cosine=0.30),
|
||||
_chunk(fillers[3], 0.72, cosine=0.30),
|
||||
_chunk(weak_b, 0.70, cosine=0.20), # rank 6 — related; read by the agent
|
||||
_chunk(weak_c, 0.65, cosine=0.10), # rank 7 — related; nobody reads it
|
||||
]
|
||||
),
|
||||
)
|
||||
@@ -815,7 +838,10 @@ def test_agent_read_below_floor_doc_joins_sources(
|
||||
assert "WEAK_B_READ_BY_AGENT" in tool_msgs[0]["content"]
|
||||
|
||||
sources = [(s["source"], s["path"]) for s in done["sources"]]
|
||||
assert sources == [("docs", "strong.md"), ("docs", "weak-b.md")] # read ⇒ cited, last
|
||||
# A4: suggested (5) + the read doc (last — it was not suggested).
|
||||
assert sources[-1] == ("docs", "weak-b.md") # read ⇒ cited, last
|
||||
assert len(sources) == 6
|
||||
assert ("docs", "weak-c.md") not in sources # never suggested, never read
|
||||
related = [(s["source"], s["path"]) for s in done["related"]]
|
||||
assert related == [("docs", "weak-c.md")] # the read doc is not "nearby"
|
||||
assert set(sources).isdisjoint(set(related))
|
||||
@@ -823,3 +849,4 @@ def test_agent_read_below_floor_doc_joins_sources(
|
||||
(row,) = session.added
|
||||
assert isinstance(row, QueryLog)
|
||||
assert "weak-b.md" in row.sources # the full retrieval is recorded (A3)
|
||||
assert "weak-c.md" in row.sources # … rank 6+ included (suggested + related + read)
|
||||
|
||||
Reference in New Issue
Block a user