phase: 119_name_signal_read_chips
All verification complete. Final report: **Phase 119 final verification pass — all criteria verified, one stale pin fixed.** - Verified implementation of all 6 tasks: D1 component name-hit rule (`name_hit` flag, titles never matched, retired length tie-break), D2 `BOR_NAME_HIT_BONUS` (0.005 default, 0 = byte-identical kill switch, negative fails startup, selection-layer only, `eval_retrieval` `suggested:` line), D3 suggested-folder lines (after `SUGGEST_INTRO`, before first block), D4 cite-discipline `SUGGEST_INTRO` sentence (PERSONA/LOW/`TOOLS_SECTION` byte-pins intact), D5 `done.sources` = read docs only (frontend no-op on empty confirmed), D6 mock `repeat your folder map` echo + new suite + telemetry. - Battery (replica restored per skill, fingerprint docs=1000/chunks=8866 verified, `eval_retrieval --from-file tests/fixtures/retrieval_battery.txt` re-run): **GATE PASS** — gitea README #4 in suggested top-5, forgejo 5/5 (README #1), gateway README in top-5 (#4), qwen3.8-27b quadlets top-5, Mongolia HIGH/fts=5 unchanged. - New E2E in isolation: `4 passed` ×2 (deterministic). All 27 modified E2E suites in isolation: 26 green; **1 stale pin fixed** — `test_source_chip_quality.py` durable-record order pin pre-dated the D1 re-rank (`aliases` stem sub-component name-hits `ssh_aliases.txt`, deterministically lifting `backups.md` over `kubernetes.md`; probe-verified 0.016277 vs 0.016036, 4/4 stable) — re-pinned with the phase-119 rationale; suite green ×2. - Gates: `uv run pytest --cov=app --cov-report=term-missing` → **2547 passed, app coverage 99%** (>90%); `uv run ruff check .` → All checks passed; `uv run pyright` → 0 errors. - Completion criteria: 1 ✅ (battery, recorded), 2 ✅ (folder lines; block/LOW byte-identical pins green), 3 ✅ (read-only chips, zero-read chips nothing, related row + durable record untouched — unit+E2E agree), 4 ✅ (all green), 5 → commit/phase-move left to the harness per pass rules (nothing committed). - Deviations: battery output + real-model telemetry recorded in `.agents/reports/119_name_signal_read_chips/task06_battery_and_e2e.md` and `TOOL_CALLING_TESTING.md` §11 (task files in `complete/` are immutable to this pass); gateway canonical doc at #4 vs overview's #3 was already documented at task 06 (containment gate met). - Next pending phase: **none** — `todo/` holds only phase 119.
This commit is contained in:
@@ -722,3 +722,84 @@ within the ~20 % band. The summary-seed behavior is proven against the
|
||||
real configured chat model — full text enters the context only through
|
||||
the capped `read` tool, and a summary-only answer is the intended fast
|
||||
path.
|
||||
|
||||
## 11. Phase 119 — name-signal + read-chips TELEMETRY, 2026-09-16
|
||||
(Telemetry-only — NOT a re-triggered gate.) Phase 119 changed
|
||||
retrieval (the component name-hit rule + the bounded name-hit bonus),
|
||||
the grounded prompt (the suggested-folder context lines + the
|
||||
cite-discipline sentence), and the citation surface (`done.sources` =
|
||||
read docs only) — but the `AGENT_TOOLS` / `TOOLS_SECTION` tool copy is
|
||||
BYTE-IDENTICAL (the phase-117 copy is untouched, an invariant of every
|
||||
phase-119 task), so the tool-copy gate is NOT re-triggered. The
|
||||
fixture battery was re-run against the real configured chat model as
|
||||
TELEMETRY for the owner (task 06, D6); the four conditions are read
|
||||
under the phase-118 A7 semantics (1/2/4 gated, 3 reported).
|
||||
|
||||
**The verdict run.** `uv run python -m scripts.agent_realmodel_check
|
||||
--restore --mode fixture` — the configured chat model (`turbo` per
|
||||
`.env`), the fixture KB restored from the tracked dump, the real
|
||||
grounded path, the full 10-question battery:
|
||||
|
||||
```
|
||||
$ uv run python -m scripts.agent_realmodel_check --restore --mode fixture
|
||||
restore: ok in 0.04s (8 docs, 2 sources)
|
||||
turn 01 | emitted=2 executed=2 cap=no defl=no | 17.10s | List the files in this directory.
|
||||
turn 02 | emitted=4 executed=4 cap=no defl=no | 14.33s | List the documents you have in the …
|
||||
turn 03 | emitted=9 executed=9 cap=no defl=no | 30.92s | List every document you have indexed.
|
||||
turn 04 | emitted=1 executed=1 cap=no defl=no | 12.16s | Open the document …
|
||||
turn 05 | emitted=1 executed=1 cap=no defl=no | 10.20s | Read …
|
||||
turn 06 | emitted=1 executed=1 cap=no defl=no | 7.61s | Open the document …
|
||||
turn 07 | emitted=2 executed=2 cap=no defl=no | 9.63s | Find the exact string "rbm-8842" in …
|
||||
turn 08 | emitted=1 executed=1 cap=no defl=no | 15.02s | Which document has the title "Lab …"
|
||||
turn 09 | emitted=1 executed=1 cap=no defl=no | 13.00s | What do you know about the qwen 3.8 …
|
||||
turn 10 | emitted=3 executed=3 cap=no defl=no | 11.71s | List the files in the deployments …
|
||||
gate: turbo PASS turns=10 answered=10 caps=0 tool-turns=10 calls 25/25 executed (100%) contract 25/25 (100%) 2026-09-16 (wall 141.8s)
|
||||
```
|
||||
|
||||
**The four conditions (read under the phase-118 A7 semantics) and
|
||||
metrics.**
|
||||
|
||||
| condition | gated under A7 | result |
|
||||
|---|---|---|
|
||||
| 1. all turns answer | yes | 10/10 answered — **GREEN** |
|
||||
| 2. zero round-cap hits | yes | caps=0 — **GREEN** |
|
||||
| 3. ≥6/10 turns emit ≥1 tool call | **no (reported)** | 10/10 tool-turns |
|
||||
| 4. contract accuracy ≥ 0.90 | yes | 25/25 (100%) — **GREEN** |
|
||||
| executed / emitted (reported) | no | 25/25 (100%) |
|
||||
|
||||
Contract line: **contract 25/25 (100%)**. Wall time: **141.8 s**.
|
||||
Model: **turbo** (the configured chat model).
|
||||
|
||||
**Per-turn reading (the phase's intended latency effect, highlighted).**
|
||||
All four designed read turns (04/05/06/09) each emitted exactly one
|
||||
contract-correct `read` **in round 1** — no `ls` drill-downs at all
|
||||
before the read: the suggested-folder context lines (phase 119, D3)
|
||||
put the target files' names in the grounded prompt, which is exactly
|
||||
the live-turn failure the phase fixes (the 2026-09-16 owner report:
|
||||
the pre-phase gitea turn walked three `ls` levels — `deploy` →
|
||||
`reeseapps` → `gitea` — before it could `read` the canonical README).
|
||||
On this fixture battery the read targets were already seed-suggested,
|
||||
so round-1 reads held from the phase-118 run — the drill-down
|
||||
reduction shows instead on the real product-name questions (the
|
||||
eval battery, recorded in the phase's task 06). Turn 08 (the title
|
||||
lookup) flipped from the phase-118 run's zero-call summary answer to
|
||||
a single round-1 `read` of `deployments/ansible/lab-inventory.md` —
|
||||
a within-contract choice (the cite-discipline sentence, D4, licenses
|
||||
citing a read suggested doc). The listing turns drilled slightly
|
||||
deeper than the phase-118 run (01: 2 calls vs 1; 03: 9 vs 8; 07:
|
||||
grep + a follow-up `read` = 2 vs 1) — +4 calls total, no caps.
|
||||
|
||||
**Baseline comparison.** Against the phase-118 `turbo` run
|
||||
(§10: 118.1 s wall, 21/21 calls, tool-turns 9/10): contract and
|
||||
executed stay at 100 %, caps remain 0, tool-turns rise to 10/10, and
|
||||
the wall time is 141.8 s (+20.1 %) — at the edge of the ~20 % band
|
||||
the phase-94 gate used as its slowdown tripwire and just above the
|
||||
97–135 s `turbo` range recorded in §3 (endpoint-load variance). As
|
||||
telemetry this is within the normal band; no copy iteration was
|
||||
triggered (the tool copy did not change).
|
||||
|
||||
**Conclusion (telemetry).** All four conditions GREEN/reported-green
|
||||
under the A7 semantics, contract 100 %, zero caps — the phase-119
|
||||
retrieval/prompt/citation changes did not degrade the real model's
|
||||
tool-calling behavior, and the read turns confirm the D3
|
||||
intended effect: files named in the prompt are read in round 1.
|
||||
|
||||
Reference in New Issue
Block a user