2026-09-06 fixture runs: contract 100 %, executed 100 %, wall ~113 s (2 runs). Derived battery: FAIL only on usage floor (5/10 tool-turns) — answers seeded questions from context, which is ideal grounded behavior. Wall time ~2.8× lite (113 s vs 40 s). Model is clean.
21 KiB
Tool-Calling Testing Methodology (controlled KB + one-command fast loop)
How to test, measure, and iterate on the agent's tool calling
(ls / read / grep) against the real configured chat model
(lite per .env) — fast enough to iterate on, controlled enough to
trust.
This methodology was set up on 2026-09-04 after phase 72 spent a long iteration cycle on an uncontrolled database (clear → git-clone the homelab repo → re-import 38–51 documents → re-embed → re-generate the KB overview → run → repeat). The old loop took many minutes per iteration and every run measured a different knowledge base, so the numbers never converged. The fix: a hand-written, unguessable, fixed-size knowledge base, snapshotted to a SQL dump, restored in ~0.03 s, and a fixed 10-question battery with one unambiguously correct tool behavior per question.
1. The fast loop (one command)
podman compose up -d db # once
uv run python -m scripts.agent_realmodel_check --restore --mode fixture
That is the whole loop:
- restore the fixture KB from
tests/fixtures/test_kb.dump.sql(one transaction — no git clone, no re-embedding, nolitecalls; ~0.03 s hot), - run the 10-question fixture battery through the real grounded
path — the exact mirror of
app/api/chat.py: embed → hybrid retrieval → the honesty gate (plan_turn) → the real prompt (persona + KB overview +<documents>+<tools>) →run_agentagainst the live endpoint, - print one line per turn plus the verdict.
Measured timings (2026-09-04, this machine):
| step | time |
|---|---|
| restore fixture KB | 0.03 s (0.2 s first run — psycopg connect) |
3-turn micro-loop (--turns 3) |
~12 s end-to-end (incl. ~1 s uv/python startup) |
| full 10-turn fixture loop | ~43–51 s wall |
| one-off KB rebuild (real embeddings, 9 chunks) | ~1–2 s |
Iteration workflow. When tuning the copy levers (§4), do not run the full battery — run the micro-loop on the first three turns (the incident turn + both listing traps, the fastest signal):
uv run python -m scripts.agent_realmodel_check --restore --mode fixture --turns 3
~12 s per variant. Run the full 10-turn battery only when a variant looks good and you want the real verdict.
Timing is visible, by design: every turn line carries its wall seconds and the verdict line carries the run's total wall time, so a slow-down (endpoint load, a retry storm, a copy that makes the model ramble) is visible on the same line as the accuracy:
turn 01 | emitted=1 executed=1 cap=no defl=no | 4.06s | List the files in …
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 8/11 executed (73%) contract 11/11 (100%) 2026-09-04 (wall 43.4s)
Notes on speed, measured (not guessed):
--concurrency 2/--concurrency 3was tested and does not help: the aipi endpoint serializes generation server-side, so parallel turns finish in the same total wall time (45.5 s @ 3-way vs ~44 s sequential) with the same aggregates. Sequential stays the default for clean telemetry.- The LLM is ~95 % of the cost (1–3 model rounds per turn at ~2–6 s each). Database work per turn is milliseconds. Don't optimize it.
2. The controlled knowledge base
tests/fixtures/agent_kb/
├── deployments/
│ ├── ansible/lab-inventory.md
│ ├── ci/gitlab-runner.md
│ └── quadlet/mimir-service.md
└── homelab/
├── backups/restic-rack7.md
├── containers/qwen38-llamacpp.md
├── containers/uptime-kuma.md
├── networking/meridian-notes.md
└── networking/vela-bridges.md
8 hand-written markdown documents, 2 sources (source name =
directory basename, the importer's rule). Every document carries
specifics no model can guess: the rack7 cluster, 10.77.42.0/24
and the VLAN 130 lab-iot pool, PVE build 8.3.4-1-lab1, port 18443
(Uptime Kuma) and ntfy topic reese-uptime-7, restic machine ID
rbm-8842, the 17 2 * * * schedule, ghcr.io/reese/obsidian-bor:2026.7.14
on 127.0.0.1:18765, the Qwen 3.8 llama.cpp launch line, ansible-core
2.19.4, … If an answer contains those specifics, the model got them
from the KB (via retrieval or a tool call) — not from its weights.
Two deliberate design rules:
- Non-topical file names for the
readtargets.vela-bridges.md,meridian-notes.md,mimir-service.mdcarry no words their content repeats. Why: hybrid retrieval seeds the question's top-2 documents into the prompt's<documents>section; FTS is OR-matched, so any question that names a document's topic words seeds that document. If the document the user asks to "open" is already in context, the correct behavior becomes ambiguous (answer from context vs. read it) and the model's well-formed re-read gets the app's in-context dedupe refusal — a test artifact, not a capability signal. With non-topical names the read must actually happen, exactly once, in the combinedsource/pathform: unambiguous, and a real test ofread. - The grep token is unique.
rbm-8842occurs in exactly one document, so thegrepturn has a definite answer.
The KB is imported through the real pipeline (import_sources —
real chunking, real embed-model vectors) and the resulting database
state is snapshotted to tests/fixtures/test_kb.dump.sql — a
data-only SQL script (TRUNCATE + one multi-row INSERT per app table:
documents, chunks + embeddings, the git_sources local rows that make
the source registry self-contained, the static KB overview, the
sources version). Restoring it puts the whole known state back in one
transaction; the generated chunks.tsv column is recomputed by
Postgres. The dump is verified by round-trip at build time (restore +
per-table checksum compare — a serialization bug fails the build).
# Rebuild the KB + dump — only when the fixture documents, the
# chunker, or the embedding model change. NOT part of the loop.
uv run python -m scripts.load_test_kb
# restores the fixture KB standalone (what --restore runs inline)
uv run python -m scripts.restore_test_kb
The build script also prints a retrieval report — for every battery question, whether the real honesty gate grounds it and which two documents would seed the context. The battery's design contract is all 10 grounded (a deflected turn offers no tools at all — it wouldn't be a tool-calling turn) with the intended seed pattern; if a question deflects or seeds the wrong document, the build says so and the fixture content is adjusted until the report is right. That report is what makes the test design checkable in 2 seconds.
Another caveat: the dump bakes in the build machine's absolute paths
(documents.full_path, the git_sources local rows) — they are
display metadata only (the gate never walks disk), so a dump built on
one machine restores fine on another. If that ever matters, rebuild.
Caveat (measured): vector cosine in an 8-document KB sits at
~0.55–0.68 for generic questions, so one read target
(vela-bridges.md) is seeded by cosine even though no FTS token
hits it. That turn is then a discipline turn (target in context —
answer from it, don't re-read), not a read turn. The battery has three
guaranteed read turns; the fourth is what the embedding lottery makes
of it.
3. The battery and the metrics
The battery (locked for the methodology — don't swap in easier questions)
| # | question | tests | expected ideal |
|---|---|---|---|
| 1 | List the files in this directory. | the phase-72 incident; full listing needs ls (8 docs, 2 in seed) |
ls() |
| 2 | List the documents you have in the homelab source. | scoped ls by the correct source name |
ls(path='homelab') |
| 3 | List every document you have indexed. | no-arg listing | ls() |
| 4 | Open the document homelab/networking/vela-bridges.md … | read, combined form |
read('homelab/networking/vela-bridges.md') |
| 5 | Read deployments/quadlet/mimir-service.md and summarize it. | read, unseeded target |
read(…) (or ls first, then read) |
| 6 | Open the document homelab/networking/meridian-notes.md … | read, unseeded target |
read(…) |
| 7 | Find the exact string "rbm-8842" in your documents … | grep, pattern only |
grep(pattern='rbm-8842') — the match line alone answers it |
| 8 | Which document has the title "Lab Ansible Inventory"? Summarize it. | title lookup; target IS seeded | answer from context (or ls) |
| 9 | What do you know about the qwen 3.8 llama.cpp setup? … | topic lookup; target IS seeded | answer from context |
| 10 | List the files in the deployments directory. | source name phrased as a directory | ls(path='deployments') |
Questions 4–6 name the full combined identity (no bare-path trap — that is the job of the locked derived battery, §6). Questions 7–9 name content, so their target document is seeded; the correct behavior there is to not re-read what is already in the prompt.
The four pass conditions
- all 10 turns answer (no
LLMError/MalformedReplyError); - zero turns hit the round cap (the incident's loop signature);
- ≥6 of 10 turns emit ≥1 tool call (the model keeps using tools);
- the accuracy bar (the mode decides which one):
fixturemode — contract accuracy ≥ 0.90 (§5 below), with the executed ratio reported alongside;derivedmode (the phase-72 locked gate) — executed/emitted ≥ 0.90, byte-compatible with the phase-72 task file.
Current standing (2026-09-04, lite, fixture KB)
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 8/11 executed (73%) contract 11/11 (100%) (wall 43.4s)
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 8/13 executed (62%) contract 12/13 (92%) (wall 50.6s)
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 7/11 executed (64%) contract 11/11 (100%) (wall 46.8s)
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=9 calls 9/12 executed (75%) contract 11/12 (92%) 2026-09-06 (wall 40.4s)
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 9/14 executed (64%) contract 13/14 (93%) 2026-09-06 (wall 40.5s)
Contract accuracy ≥ 90 %: met (100 / 92 / 100 / 93 / 92 / 93). The executed ratio sits at 58–73 % for the reason documented in §5 — an app semantics choice, not a model defect, and the open design question in §7.
Model comparison — turbo (2026-09-05, same fixture KB, .env
chat model switched to turbo):
gate: turbo PASS turns=10 answered=10 caps=0 tool-turns=7 calls 9/9 executed (100%) contract 9/9 (100%) 2026-09-05 (wall 105.1s)
gate: turbo PASS turns=10 answered=10 caps=0 tool-turns=7 calls 7/7 executed (100%) contract 7/7 (100%) 2026-09-05 (wall 135.5s)
gate: turbo FAIL turns=10 answered=10 caps=0 tool-turns=5 calls 5/5 executed (100%) contract 5/5 (100%) 2026-09-05 (wall 77.1s) [derived battery — MISS: 5/10 tool-turn floor]
gate: turbo PASS turns=10 answered=10 caps=0 tool-turns=7 calls 9/9 executed (100%) contract 9/9 (100%) 2026-09-06 (wall 113.7s)
gate: turbo PASS turns=10 answered=10 caps=0 tool-turns=7 calls 7/7 executed (100%) contract 7/7 (100%) 2026-09-06 (wall 112.8s)
gate: turbo FAIL turns=10 answered=10 caps=0 tool-turns=5 calls 5/5 executed (100%) contract 5/5 (100%) 2026-09-06 (wall 97.2s) [derived battery — MISS: 5/10 tool-turn floor]
Reads: the re-read habit is model-specific. lite re-reads a seeded
named document ~100 % of the time (copy-invariant, §4); turbo
answered 7 of 8 seeded-target questions straight from the
<documents> context with zero tool calls — the exact
"don't re-read" behavior the copy levers could not buy from lite
(one re-read in the sample, 12 % vs ~100 %). Consequence:
turbo hits 100 % on both metrics on the fixture battery — the
executed ratio reaches 100 % naturally once the redundant reads
stop, which corroborates §7's framing (the block on lite is the
model's re-read habit, not a gate or app defect). On the locked
derived battery turbo fails only the usage floor condition (≥ 6/10
turns with ≥ 1 emitted call: 5/10) — it answers the seeded read-target
questions from context instead of making the (refusable) read call the
trap design expects; accuracy on every call it does make is still
100 %. The cost: ~2.5× slower wall time (97–114 s per full loop
vs 40–41 s for lite, with individual turns 7–20 s).
4. The copy levers (what you iterate)
All three are fixed-template constants with byte-pinned unit tests —
change the constant, update the pin, run uv run pytest tests/unit -q
(~10 s), then the micro-loop:
| lever | where | what it teaches |
|---|---|---|
| refusal templates | app/rag/agent.py (LS_PATH_NOT_A_SOURCE, NO_SOURCE_NOT_A_DIRECTORY, NO_DOCUMENT_DID_YOU_MEAN[_MAN], ALREADY_IN_CONTEXT, …) |
the correct form after a misuse — self-correction in one round |
| tool descriptions | app/rag/agent.py AGENT_TOOLS |
the contract at call time (the most local text the model reads) |
<tools> prompt section |
app/rag/prompts.py TOOLS_SECTION |
the contract up front, every grounded turn |
Unit pins to follow the constants: tests/unit/test_agent.py
(description + refusal pins), tests/unit/test_prompts.py
(TOOLS_SECTION substring pins — the listed substrings must survive
any rewording). The E2E mock keys off marker presence (<tools>,
DEFLECT_MODE), not wording — rewording is safe there.
What has been tried on this model (2026-09-03 → 04, all measured live) — so the next iteration doesn't repeat it:
| variant | re-reads of seeded docs | note |
|---|---|---|
phase-72: mid-paragraph do-not-read rule (TOOLS_SECTION + read description) |
15/15 (never flipped) | 9 runs, 38–51 doc KBs |
leading in-<documents>-section reminder naming the blocks |
0/15 flipped | reverted — primed seed paths as ls scopes (incident turn regressed to a cap loop) |
front-loaded do-not-read as the read description's first sentence |
no improvement | + one 6-emitted variance spike |
per-block note="…do not call read on it" attribute on each <document> header |
no improvement | reverted |
Conclusion: the re-read of a salient seeded document is
copy-invariant behavior of the lite model (it obeys the user's
"open it / read it" over every prompt-level rule tried). The levers
that do work on this model: the teaching refusals (bare-path
self-correction in exactly one round — 4/4 in the derived battery;
NO_DOCUMENT_DID_YOU_MEAN naming the combined identity), the
one-call-per-reply and never-repeat rules (no cap hits, no repeat
loops in any controlled run), and the grep pattern-only clause (the
source-scoped-grep misuse is gone).
Do not touch while iterating: the battery questions, the
thresholds, the fixture documents (that would be moving the goal
posts — if the battery needs changing, it is a methodology change,
say so), the refusal mechanics (a refusal is still a refusal,
counts in nothing, consumes a round — phase-72 locked decision), the
tool names/argument shapes (ls(path?) / read(path) /
grep(pattern, path?) — phase-70 locked surface).
5. The two metrics — read this before arguing about the numbers
The verdict carries both:
- contract accuracy = emitted calls that are well-formed and
target a resolvable entity ÷ emitted (
classify_callinscripts/agent_realmodel_check.py, mirroringapp/rag/agent._execute_tool's resolution rules gate-side). A call is a contract violation when the model aimed wrong: unknown tool, missing argument, a bare document path where the combinedsource/pathbelongs, a nonexistent document identity, a source name where a document belongs (ls(path='.'),ls(path='/'),grep(path='homelab')— the entire phase-72 incident class). - executed/emitted (the phase-72 locked metric) = calls the app
actually executed ÷ emitted. Every refusal class counts against it
— including
ALREADY_IN_CONTEXT, the app's dedupe refusal when the model reads a document whose full text is already in the<documents>context.
Why the fixture gate's accuracy bar is contract accuracy, and why this is honest rather than goalpost-moving:
- The re-read is a correct tool call — right tool, well-formed
arguments, a real document identity — that the app declines for
redundancy. The phase-72 incident the owner was frustrated by
(garbage scopes, loops, cap hits) is exactly the class contract
accuracy measures, and it is gone: 0 contract violations in 2 of
3 fixture runs, 3 in the third (one directory-scoped
grep('mimir-service', path='deployments/quadlet')exploration that self-corrected vialsin two rounds). - The executed ratio is blocked at 58–73 % by the re-reads alone — and §4 shows five independent copy variants failed to change that behavior even once. Gating the fast loop on a number no lever can move would make it permanently red and useless for iteration.
- Both numbers are always printed. Nothing is hidden; the executed ratio stays the pass bar for the locked derived gate.
The remaining question — should a redundant-but-correct read count as a failure at all? — is an app-semantics decision, not a copy lever (§7).
6. The derived gate (phase 72, locked)
--mode derived (the default) runs the phase-72 locked battery —
derived from the live catalog's first two documents, including the two
bare-path traps (read('ansible/lab-inventory.md') without the
source prefix, etc.) — with the phase-72 locked conditions, including
executed/emitted ≥ 0.90. Against the fixture KB (2026-09-04):
gate: lite FAIL turns=10 answered=10 caps=0 tool-turns=10 calls 5/15 executed (33%) contract 12/15 (80%) (wall 47.7s)
gate: lite FAIL turns=10 answered=10 caps=0 tool-turns=10 calls 5/14 executed (36%) contract 10/14 (71%) 2026-09-06 (wall 38.3s)
Reading that result: the teaching works — every bare-path trap
self-corrected in exactly one round (the did-you-mean refusal named
the combined identity, the model used it next round), zero cap hits,
10/10 answered. The executed bar fails because the corrected read then
hits ALREADY_IN_CONTEXT — the trap question names the document's
topic words, so the document is seeded, and the correct combined-form
read is dedupe-refused. Same wall as §5, now on the locked gate:
the ≥90 % executed bar is unreachable under the current refusal
semantics regardless of copy. The gate runs as-is, unchanged, and
reports it.
7. Open design question (for the owner)
The only thing standing between the lite model and a ≥90 %
executed ratio is one refusal's semantics: ALREADY_IN_CONTEXT.
Options, with trade-offs:
- Keep as-is (phase-72 locked): a redundant read is a refusal, counts in nothing. The model is taught not to re-read; the cost is that the executed metric can't reach 90 % while the model's copy-invariant re-read habit exists. Contract accuracy (the capability metric) is ~100 %.
- Count an in-context read as executed (return the document,
dedupe the context — the
holder.read_docsdedupe already makes a re-read a no-op content-wise). The executed metric would jump to ~100 %; the teaching signal weakens (the model never sees the refusal it is being taught by). - Hybrid: execute it, but mark the turn
redundant_reads=Nin the log line and the verdict, keeping the signal without the wall.
The controlled methodology makes this a 50-second experiment either
way: change the one branch in app/rag/agent.py::_execute_tool,
update its unit pins, run the full fixture loop.
8. Reproducing from scratch
# 0. Prereqs: the usual dev setup (AGENTS.md quick reference)
podman compose up -d db
cp .env.example .env # once; LLM endpoint + DB URL
uv run alembic upgrade head
# 1. Build the controlled KB + dump (one-off, ~2 s — real embeddings)
uv run python -m scripts.load_test_kb
# → prints the retrieval report (all 10 must be grounded) and
# verifies the dump by round-trip.
# 2. The loop
uv run python -m scripts.agent_realmodel_check --restore --mode fixture --turns 3 # ~12 s micro-loop
uv run python -m scripts.agent_realmodel_check --restore --mode fixture # ~45 s full gate
uv run python -m scripts.agent_realmodel_check --restore # phase-72 locked gate
# 3. After touching the copy levers
uv run pytest tests/unit -q # pins in sync?
uv run pytest --cov=app --cov-report=term-missing | tail -3 # >90 %
uv run ruff check . && uv run pyright
uv run pytest tests/e2e/test_tool_path_teaching.py -v --no-cov # E2E in isolation
Exit codes, both gate and restore/build: 0 pass/ok, 1 fail (with the per-condition breakdown — the MISS lines name the lever to iterate), 2 precondition (DB down, dump missing, schema not applied — each with the actionable fix on the same line).
Diagnosing a bad run: every call is logged by run_agent
(agent tool=… args=… round=…/…) — correlate the arguments with the
refusal templates in app/rag/agent.py to see which teaching line the
model hit, and which refusal class (contract violation vs. in-context
dedupe) the rejection was.