phase: 119_name_signal_read_chips
Build and Push Containers / build-and-push-app (push) Successful in 2m1s
Build and Push Containers / build-and-push-db (push) Successful in 18s

All verification complete. Final report:

**Phase 119 final verification pass — all criteria verified, one stale pin fixed.**
- Verified implementation of all 6 tasks: D1 component name-hit rule (`name_hit` flag, titles never matched, retired length tie-break), D2 `BOR_NAME_HIT_BONUS` (0.005 default, 0 = byte-identical kill switch, negative fails startup, selection-layer only, `eval_retrieval` `suggested:` line), D3 suggested-folder lines (after `SUGGEST_INTRO`, before first block), D4 cite-discipline `SUGGEST_INTRO` sentence (PERSONA/LOW/`TOOLS_SECTION` byte-pins intact), D5 `done.sources` = read docs only (frontend no-op on empty confirmed), D6 mock `repeat your folder map` echo + new suite + telemetry.
- Battery (replica restored per skill, fingerprint docs=1000/chunks=8866 verified, `eval_retrieval --from-file tests/fixtures/retrieval_battery.txt` re-run): **GATE PASS** — gitea README #4 in suggested top-5, forgejo 5/5 (README #1), gateway README in top-5 (#4), qwen3.8-27b quadlets top-5, Mongolia HIGH/fts=5 unchanged.
- New E2E in isolation: `4 passed` ×2 (deterministic). All 27 modified E2E suites in isolation: 26 green; **1 stale pin fixed** — `test_source_chip_quality.py` durable-record order pin pre-dated the D1 re-rank (`aliases` stem sub-component name-hits `ssh_aliases.txt`, deterministically lifting `backups.md` over `kubernetes.md`; probe-verified 0.016277 vs 0.016036, 4/4 stable) — re-pinned with the phase-119 rationale; suite green ×2.
- Gates: `uv run pytest --cov=app --cov-report=term-missing` → **2547 passed, app coverage 99%** (>90%); `uv run ruff check .` → All checks passed; `uv run pyright` → 0 errors.
- Completion criteria: 1 ✅ (battery, recorded), 2 ✅ (folder lines; block/LOW byte-identical pins green), 3 ✅ (read-only chips, zero-read chips nothing, related row + durable record untouched — unit+E2E agree), 4 ✅ (all green), 5 → commit/phase-move left to the harness per pass rules (nothing committed).
- Deviations: battery output + real-model telemetry recorded in `.agents/reports/119_name_signal_read_chips/task06_battery_and_e2e.md` and `TOOL_CALLING_TESTING.md` §11 (task files in `complete/` are immutable to this pass); gateway canonical doc at #4 vs overview's #3 was already documented at task 06 (containment gate met).
- Next pending phase: **none** — `todo/` holds only phase 119.
This commit is contained in:
2026-09-16 15:50:48 -04:00
parent 795fb56425
commit a5b63f83ad
89 changed files with 4377 additions and 655 deletions
+23 -14
View File
@@ -60,12 +60,15 @@ Test → story mapping (Playwright Mapping Rule):
``tool`` frames (``ls``, then ``read`` with the combined path, ahead
of any delta), the UI shows the transient calling-tool status while a
tool runs, the bubble shows both tool lines, the final answer quotes
the read document, and the source chips include the read document
(viewer link).
the read document, and the source chips are EXACTLY the read
document (phase 119, LOCKED A1 — chips cite read docs only; viewer
link).
2. ``test_tool_lines_re_render_after_reload`` — the persisted record
(phase 14) re-renders the tool lines.
3. ``test_plain_grounded_question_has_no_tool_frames`` — no marker → no
``tool`` frames, the answer renders exactly as today (regression
``tool`` frames, the answer renders exactly as today, and (phase 119,
LOCKED A1) the zero-read grounded turn chips NOTHING (the retrieval
doc is cited only in the durable record, never as a chip) (regression
inside the story file).
4. ``test_deflected_question_has_no_tool_frames`` — the tools are
grounded-only: a deflected turn runs none.
@@ -444,8 +447,12 @@ def test_marker_question_lists_reads_and_quotes(
)
done = next(f for f in frames if f.get("type") == "done")
assert done["deflected"] is False
# Phase 119 (LOCKED A1): done.sources = the READ docs only —
# exactly the scripted read; the retrieval (seed) doc is suggested
# context, never a citation chip (the retired phase-118 A4
# suggested+read union is gone). The durable record below still
# carries both (LOCKED A3, untouched).
assert [(s["source"], s["path"]) for s in done["sources"]] == [
(SEED_SOURCE, SEED_PATH),
(READ_SOURCE, READ_PATH),
]
@@ -465,16 +472,17 @@ def test_marker_question_lists_reads_and_quotes(
expect(bubble).to_contain_text(ANSWER_PREFIX)
expect(bubble).to_contain_text(ANSWER_QUOTE)
# Source chips: the retrieval doc AND the read doc (deduped, in
# order) — the read chip links to the viewer.
# Source chips: the read doc ONLY (phase 119, LOCKED A1 — the
# retrieval doc was never read, so it never chips; the retired
# phase-118 A4 union is gone) — the read chip links to the viewer.
chips = page.locator(".msg.brain .source-chip")
expect(chips).to_have_count(2)
expect(chips.nth(0)).to_contain_text(SEED_SP)
expect(chips).to_have_count(1)
chip_read = page.locator(".msg.brain .source-chip", has_text=READ_PATH)
expect(chip_read).to_have_count(1)
expect(chip_read.first).to_have_attribute("href", READ_CHIP_HREF)
# Durable record: grounded, both sources logged (retrieval + read).
# Durable record: grounded, both sources logged (suggested + read —
# LOCKED A3, untouched by phase 119 A1).
row = _last_query_log()
assert row.question == MARKER_QUESTION
assert row.deflected is False
@@ -540,14 +548,15 @@ def test_plain_grounded_question_has_no_tool_frames(
assert _tool_frames(_frames(page)) == []
expect(page.locator(".tool-call")).to_have_count(0)
# The standard grounded answer, citing the retrieval doc only — the
# referenced JSON stays OUT of the sources (it was never read).
# The standard grounded answer, with ZERO citation chips — the
# turn read nothing, so (phase 119, LOCKED A1) the chip row is
# empty: the retrieval doc is cited in the durable record only
# (the retired phase-118 A4 union is gone); the referenced JSON
# stays out of the record too (it was never read).
bubble = page.locator(".msg.brain .bubble").last
expect(bubble).to_contain_text(PLAIN_QUESTION)
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER)
chips = page.locator(".msg.brain .source-chip")
expect(chips).to_have_count(1)
expect(chips.first).to_contain_text(SEED_SP)
expect(page.locator(".msg.brain .source-chip")).to_have_count(0)
row = _last_query_log()
assert row.question == PLAIN_QUESTION