phase: 119_name_signal_read_chips
Build and Push Containers / build-and-push-app (push) Successful in 2m1s
Build and Push Containers / build-and-push-db (push) Successful in 18s

All verification complete. Final report:

**Phase 119 final verification pass — all criteria verified, one stale pin fixed.**
- Verified implementation of all 6 tasks: D1 component name-hit rule (`name_hit` flag, titles never matched, retired length tie-break), D2 `BOR_NAME_HIT_BONUS` (0.005 default, 0 = byte-identical kill switch, negative fails startup, selection-layer only, `eval_retrieval` `suggested:` line), D3 suggested-folder lines (after `SUGGEST_INTRO`, before first block), D4 cite-discipline `SUGGEST_INTRO` sentence (PERSONA/LOW/`TOOLS_SECTION` byte-pins intact), D5 `done.sources` = read docs only (frontend no-op on empty confirmed), D6 mock `repeat your folder map` echo + new suite + telemetry.
- Battery (replica restored per skill, fingerprint docs=1000/chunks=8866 verified, `eval_retrieval --from-file tests/fixtures/retrieval_battery.txt` re-run): **GATE PASS** — gitea README #4 in suggested top-5, forgejo 5/5 (README #1), gateway README in top-5 (#4), qwen3.8-27b quadlets top-5, Mongolia HIGH/fts=5 unchanged.
- New E2E in isolation: `4 passed` ×2 (deterministic). All 27 modified E2E suites in isolation: 26 green; **1 stale pin fixed** — `test_source_chip_quality.py` durable-record order pin pre-dated the D1 re-rank (`aliases` stem sub-component name-hits `ssh_aliases.txt`, deterministically lifting `backups.md` over `kubernetes.md`; probe-verified 0.016277 vs 0.016036, 4/4 stable) — re-pinned with the phase-119 rationale; suite green ×2.
- Gates: `uv run pytest --cov=app --cov-report=term-missing` → **2547 passed, app coverage 99%** (>90%); `uv run ruff check .` → All checks passed; `uv run pyright` → 0 errors.
- Completion criteria: 1 ✅ (battery, recorded), 2 ✅ (folder lines; block/LOW byte-identical pins green), 3 ✅ (read-only chips, zero-read chips nothing, related row + durable record untouched — unit+E2E agree), 4 ✅ (all green), 5 → commit/phase-move left to the harness per pass rules (nothing committed).
- Deviations: battery output + real-model telemetry recorded in `.agents/reports/119_name_signal_read_chips/task06_battery_and_e2e.md` and `TOOL_CALLING_TESTING.md` §11 (task files in `complete/` are immutable to this pass); gateway canonical doc at #4 vs overview's #3 was already documented at task 06 (containment gate met).
- Next pending phase: **none** — `todo/` holds only phase 119.
This commit is contained in:
2026-09-16 15:50:48 -04:00
parent 795fb56425
commit a5b63f83ad
89 changed files with 4377 additions and 655 deletions
+81
View File
@@ -722,3 +722,84 @@ within the ~20 % band. The summary-seed behavior is proven against the
real configured chat model — full text enters the context only through
the capped `read` tool, and a summary-only answer is the intended fast
path.
## 11. Phase 119 — name-signal + read-chips TELEMETRY, 2026-09-16
(Telemetry-only — NOT a re-triggered gate.) Phase 119 changed
retrieval (the component name-hit rule + the bounded name-hit bonus),
the grounded prompt (the suggested-folder context lines + the
cite-discipline sentence), and the citation surface (`done.sources` =
read docs only) — but the `AGENT_TOOLS` / `TOOLS_SECTION` tool copy is
BYTE-IDENTICAL (the phase-117 copy is untouched, an invariant of every
phase-119 task), so the tool-copy gate is NOT re-triggered. The
fixture battery was re-run against the real configured chat model as
TELEMETRY for the owner (task 06, D6); the four conditions are read
under the phase-118 A7 semantics (1/2/4 gated, 3 reported).
**The verdict run.** `uv run python -m scripts.agent_realmodel_check
--restore --mode fixture` — the configured chat model (`turbo` per
`.env`), the fixture KB restored from the tracked dump, the real
grounded path, the full 10-question battery:
```
$ uv run python -m scripts.agent_realmodel_check --restore --mode fixture
restore: ok in 0.04s (8 docs, 2 sources)
turn 01 | emitted=2 executed=2 cap=no defl=no | 17.10s | List the files in this directory.
turn 02 | emitted=4 executed=4 cap=no defl=no | 14.33s | List the documents you have in the …
turn 03 | emitted=9 executed=9 cap=no defl=no | 30.92s | List every document you have indexed.
turn 04 | emitted=1 executed=1 cap=no defl=no | 12.16s | Open the document …
turn 05 | emitted=1 executed=1 cap=no defl=no | 10.20s | Read …
turn 06 | emitted=1 executed=1 cap=no defl=no | 7.61s | Open the document …
turn 07 | emitted=2 executed=2 cap=no defl=no | 9.63s | Find the exact string "rbm-8842" in …
turn 08 | emitted=1 executed=1 cap=no defl=no | 15.02s | Which document has the title "Lab …"
turn 09 | emitted=1 executed=1 cap=no defl=no | 13.00s | What do you know about the qwen 3.8 …
turn 10 | emitted=3 executed=3 cap=no defl=no | 11.71s | List the files in the deployments …
gate: turbo PASS turns=10 answered=10 caps=0 tool-turns=10 calls 25/25 executed (100%) contract 25/25 (100%) 2026-09-16 (wall 141.8s)
```
**The four conditions (read under the phase-118 A7 semantics) and
metrics.**
| condition | gated under A7 | result |
|---|---|---|
| 1. all turns answer | yes | 10/10 answered — **GREEN** |
| 2. zero round-cap hits | yes | caps=0 — **GREEN** |
| 3. ≥6/10 turns emit ≥1 tool call | **no (reported)** | 10/10 tool-turns |
| 4. contract accuracy ≥ 0.90 | yes | 25/25 (100%) — **GREEN** |
| executed / emitted (reported) | no | 25/25 (100%) |
Contract line: **contract 25/25 (100%)**. Wall time: **141.8 s**.
Model: **turbo** (the configured chat model).
**Per-turn reading (the phase's intended latency effect, highlighted).**
All four designed read turns (04/05/06/09) each emitted exactly one
contract-correct `read` **in round 1** — no `ls` drill-downs at all
before the read: the suggested-folder context lines (phase 119, D3)
put the target files' names in the grounded prompt, which is exactly
the live-turn failure the phase fixes (the 2026-09-16 owner report:
the pre-phase gitea turn walked three `ls` levels — `deploy` →
`reeseapps` → `gitea` — before it could `read` the canonical README).
On this fixture battery the read targets were already seed-suggested,
so round-1 reads held from the phase-118 run — the drill-down
reduction shows instead on the real product-name questions (the
eval battery, recorded in the phase's task 06). Turn 08 (the title
lookup) flipped from the phase-118 run's zero-call summary answer to
a single round-1 `read` of `deployments/ansible/lab-inventory.md` —
a within-contract choice (the cite-discipline sentence, D4, licenses
citing a read suggested doc). The listing turns drilled slightly
deeper than the phase-118 run (01: 2 calls vs 1; 03: 9 vs 8; 07:
grep + a follow-up `read` = 2 vs 1) — +4 calls total, no caps.
**Baseline comparison.** Against the phase-118 `turbo` run
(§10: 118.1 s wall, 21/21 calls, tool-turns 9/10): contract and
executed stay at 100 %, caps remain 0, tool-turns rise to 10/10, and
the wall time is 141.8 s (+20.1 %) — at the edge of the ~20 % band
the phase-94 gate used as its slowdown tripwire and just above the
97–135 s `turbo` range recorded in §3 (endpoint-load variance). As
telemetry this is within the normal band; no copy iteration was
triggered (the tool copy did not change).
**Conclusion (telemetry).** All four conditions GREEN/reported-green
under the A7 semantics, contract 100 %, zero caps — the phase-119
retrieval/prompt/citation changes did not degrade the real model's
tool-calling behavior, and the read turns confirm the D3
intended effect: files named in the prompt are read in round 1.