**Phase 118 final verification pass — complete.** All criteria verified; 4 pre-existing defects found and fixed.
- **Verified:** summary-seed wiring (`select_suggested` top-5 no-floor → summary blocks, no full text in HIGH prompt), all-doc markdown summaries + NULL backfill (`summary_backfilled`, no `sources_meta` bump), `read` adds full text with `read_docs`-only dedupe, `done.sources` = suggested+read / durable record = suggested+related+read + `suggested=N` log line (seen live in E2E), byte-locked PERSONA/LOW/TOOLS_SECTION, battery gate PASS recorded in `TOOL_CALLING_TESTING.md` §10 (turbo 2026-09-16: 1/2/4 GREEN, cond-3 reported 9/10 per A7, contract 21/21, caps 0).
- **Defects fixed (all pre-existing, none phase-118):** ① `ChatMessage` schema missing the phase-113 `related` key → `extra="forbid"` 422'd every done-time auto-save of grounded turns with a related tier, leaving `message_count=1` (root cause of `test_share_chat` 3F; browser-level instrumentation proved the PUT 422) — added the field + unit/integration pins; ② `test_theme_semantic_completion` pins stale vs phase-117 debox (border/chip removed) — re-targeted to assert border/chip *absence*; ③ `test_header_consistency` `<26`px pin red on 26.125px native date-input line — bound relaxed to `<34` (wrap-detection intent kept); ④ `test_navbar_refresh` bor.chat.v1 key set updated for `related`.
- **Test/lint/coverage:** `uv run pytest --cov=app --cov-report=term-missing` → **2506 passed, app/ 99%** (>90%); `uv run ruff check . && uv run pyright` → clean, 0 errors.
- **E2E:** new story suite in isolation → **2 passed**; full 103-suite matrix sweep (each isolated) → **all 103 green** after the fixes; `test_share_chat` 4 passed, `test_theme_semantic_completion` 8 passed, `test_header_consistency` 3 passed, `test_navbar_refresh` 7 passed.
- **Deviations:** none from LOCKED decisions. Note: orphaned diagnostic uvicorn processes briefly made E2E sessions exercise stale code — killed and re-verified; a sweep-regenerated tracked screenshot was restored. No commits made (harness commits).
- **Completion criteria:** all 7 ✅ (commit/phase-move is the harness's step).
- **Next pending phase:** none — `todo/` holds only this phase's overview pending the harness move.
873 lines
35 KiB
Python
873 lines
35 KiB
Python
"""The real-model tool-calling gate (live, the configured chat model).
|
|
|
|
The phase-72 pass condition (owner directive 2026-09-03 — "test with the
|
|
real lite model until tool calls work consistently; don't pass until a
|
|
sufficient number of tool calls succeed"): this script drives a fixed
|
|
question battery through the **real** grounded path — the exact mirror
|
|
of ``app.api.chat`` (embed → retrieve → the honesty gate via
|
|
``plan_turn`` → the steering notes + KB overview exactly as
|
|
``app.api.chat`` reads them → ``build_high_prompt`` / deflection prompt →
|
|
``run_agent`` with the configured chat model and the real Postgres KB, a
|
|
fresh ``AgentHolder`` per turn; a deflected turn runs the same
|
|
``tools=None`` + one-bounded-recovery stream the deflected API branch
|
|
runs) — and applies the four LOCKED pass conditions:
|
|
|
|
1. all turns answer (no ``LLMError`` / ``MalformedReplyError``);
|
|
2. zero turns hit the round cap (the incident's loop signature — hitting
|
|
the cap means the teaching did not end the loop);
|
|
3. >=6 of 10 turns emit >=1 tool call (the model keeps USING tools — it
|
|
does not abandon them and answer from seed context alone, the
|
|
incident's end state);
|
|
4. the accuracy bar across the whole run — ``derived`` mode (the
|
|
phase-72 LOCKED gate): executed / emitted >= 0.90 (refusals count in
|
|
nothing); ``fixture`` mode (the controlled methodology): contract
|
|
accuracy — well-formed calls targeting resolvable entities —
|
|
>= 0.90, with the executed ratio reported alongside (see
|
|
``classify_call`` and ``TOOL_CALLING_TESTING.md`` for why the two
|
|
metrics differ: the app's ALREADY_IN_CONTEXT dedupe refusal is an
|
|
app-semantics choice, not a tool-calling error).
|
|
|
|
Batteries (``--mode``):
|
|
|
|
* ``derived`` (default, the phase-72 LOCKED battery) — the fixed
|
|
10-question battery derived from the live catalog's first two documents
|
|
``D1 = (s1, p1, t1)`` / ``D2 = (s2, p2, t2)`` (catalog order); the
|
|
grep question's token is the first whitespace-split word of
|
|
``D2.content`` with length >= 6 (leading/trailing non-alphanumerics
|
|
stripped, lowercased), falling back to the first word of ``t2``.
|
|
* ``fixture`` (the controlled fast loop, 2026-09-04 owner directive) —
|
|
:data:`FIXTURE_BATTERY`, the curated 10 questions pinned to the
|
|
hand-written fixture KB (``tests/fixtures/agent_kb/``). Combine with
|
|
``--restore`` so the whole iteration is one command against a known,
|
|
unguessable, re-embed-free knowledge base (``TOOL_CALLING_TESTING.md``).
|
|
|
|
Speed levers (the fast loop): ``--restore`` (restore the fixture dump in
|
|
one transaction — no git clone, no re-embedding, sub-second); ``--turns
|
|
N`` (run only the first N questions — the micro-loop for copy
|
|
iteration; the verdict is then marked ``partial`` and condition 3 is
|
|
reported, not gated); every turn line carries its wall seconds and the
|
|
verdict line carries the run's total wall time, so a slow-down is
|
|
visible in the same line that carries the accuracy.
|
|
|
|
House probe pattern (``scripts/llm_probe.py``): ``uv run python -m
|
|
scripts.agent_realmodel_check`` — argparse, dotenv, plain module, no
|
|
debugpy. The script never modifies the KB (no commits, no query_log
|
|
rows — the only writes are the ``--restore`` snapshot restore, which is
|
|
explicit and transactional).
|
|
|
|
Preconditions (exit 2 with an actionable line on failure): the DB is
|
|
reachable; ``BOR_AGENT_MAX_ROUNDS`` is > 0 (the gate needs tools
|
|
enabled); ``derived`` mode — the catalog holds >=2 documents and the
|
|
FIRST TWO catalog documents' ``path``s each contain ``/`` (the
|
|
bare-path traps need nested paths); ``fixture`` mode — the fixture dump
|
|
exists (build it with ``uv run python -m scripts.load_test_kb``).
|
|
|
|
Exit codes: **0 PASS**, **1 FAIL** (the per-condition breakdown is
|
|
printed so the copy-lever iteration loop can target the right lever),
|
|
**2 precondition failure**. For refusal diagnosis, every call is
|
|
already logged by ``run_agent`` (``agent tool=… args=… round=…/…``) —
|
|
correlate the logged arguments with the refusal templates in
|
|
``app/rag/agent.py`` to see which teaching line the model hit.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import asyncio
|
|
import logging
|
|
import sys
|
|
import time
|
|
from dataclasses import dataclass, field
|
|
from datetime import date
|
|
from pathlib import Path
|
|
from typing import Any
|
|
|
|
from dotenv import load_dotenv
|
|
from sqlalchemy import select
|
|
|
|
from app.api.chat import plan_turn
|
|
from app.api.steering import load_steering_notes
|
|
from app.config import Settings, get_settings
|
|
from app.db import SessionLocal, db_available
|
|
from app.models import Document
|
|
from app.rag.agent import (
|
|
CORRECTION_INSTRUCTION,
|
|
AgentHolder,
|
|
MalformedReplyError,
|
|
find_document,
|
|
list_source_names,
|
|
run_agent,
|
|
)
|
|
from app.rag.llm import (
|
|
EmbeddingError,
|
|
LLMClient,
|
|
LLMError,
|
|
StreamPiece,
|
|
ToolCallPiece,
|
|
chat_stream_retried,
|
|
)
|
|
from app.rag.overview import load_kb_overview
|
|
from app.rag.retriever import retrieve
|
|
from app.rag.scaffolding import ScaffoldingFilter
|
|
|
|
logger = logging.getLogger("agent_realmodel_check")
|
|
|
|
#: Repo-relative fixture dump (written by scripts.load_test_kb).
|
|
DEFAULT_DUMP_PATH = Path("tests/fixtures/test_kb.dump.sql")
|
|
|
|
|
|
def _catalog(db) -> list[tuple[str, str, str]]:
|
|
"""The full catalog as ``(source, path, title)`` in ``(source, path)``
|
|
order — the battery builder's pre-phase-94 ``list_catalog`` contract,
|
|
kept local to this script (phase 94 deleted the agent accessor: the
|
|
drill-down ``ls`` lists one bounded source level at a time)."""
|
|
rows = db.execute(
|
|
select(Document.source, Document.path, Document.title).order_by(
|
|
Document.source, Document.path
|
|
)
|
|
).all()
|
|
return [(source, path, title) for source, path, title in rows]
|
|
|
|
#: The per-turn line truncates the question at this width (the locked
|
|
#: format prints ``turn NN | emitted=E executed=X cap=Y|N | <question>``).
|
|
QUESTION_DISPLAY_WIDTH = 40
|
|
|
|
#: The controlled fast-loop battery (2026-09-04, owner directive): the
|
|
#: curated 10 questions pinned to the fixture KB
|
|
#: (``tests/fixtures/agent_kb/`` — sources ``deployments`` / ``homelab``,
|
|
#: 8 hand-written documents whose specifics — ``rack7``,
|
|
#: ``10.77.42.0/24``, VLAN 130, port 18443, ntfy topic
|
|
#: ``reese-uptime-7``, machine ID ``rbm-8842``, the ``17 2 * * *``
|
|
#: schedule, image ``ghcr.io/reese/obsidian-bor:2026.7.14``, port
|
|
#: 18765 — are not guessable by any model).
|
|
#:
|
|
#: Design rule (the controlled-test property): every question has ONE
|
|
#: unambiguously correct tool behavior, verified by the load script's
|
|
#: retrieval report. The ``read`` targets (Q4/Q5/Q6) are named so the
|
|
#: question's tokens do NOT lexically seed the target document (the
|
|
#: filenames carry no topical words the content repeats) — the read
|
|
#: must actually happen, exactly once, in the combined form. The
|
|
#: discipline turns (Q7/Q8/Q9) name content that IS seeded — the
|
|
#: correct behavior there is to answer from the ``<documents>`` context
|
|
#: (or ``ls``), NOT to re-read. The bare-path traps are the job of the
|
|
#: locked derived battery, not this one.
|
|
#:
|
|
#: 1. the phase-72 incident ("list the files in this directory" — a
|
|
#: full listing needs ``ls``: 8 docs, 2 in seed context);
|
|
#: 2. a scoped ``ls`` by the correct source name;
|
|
#: 3. the no-arg listing;
|
|
#: 4. a ``read`` of an unseeded document (combined form, given whole);
|
|
#: 5. a second ``read`` of an unseeded document (explicit "Read …");
|
|
#: 6. a third ``read`` of an unseeded document;
|
|
#: 7. the ``grep`` turn (``rbm-8842`` — a string that occurs in exactly
|
|
#: one fixture document; the grep line alone answers "which ones");
|
|
#: 8. a title lookup — the target IS seeded; summarize from context;
|
|
#: 9. a topic lookup — the target IS seeded; answer from context;
|
|
#: 10. the source name phrased as a directory (the ``ls(path=…)`` scope
|
|
#: trap).
|
|
FIXTURE_BATTERY: list[str] = [
|
|
"List the files in this directory.",
|
|
"List the documents you have in the homelab source.",
|
|
"List every document you have indexed.",
|
|
"Open the document homelab/networking/vela-bridges.md and tell me what it covers.",
|
|
"Read deployments/quadlet/mimir-service.md and summarize it.",
|
|
"Open the document homelab/networking/meridian-notes.md and tell me what it covers.",
|
|
'Find the exact string "rbm-8842" in your documents and tell me which ones '
|
|
"contain it.",
|
|
'Which document has the title "Lab Ansible Inventory"? Summarize it.',
|
|
"What do you know about the qwen 3.8 llama.cpp setup? Give me the exact "
|
|
"launch arguments.",
|
|
"List the files in the deployments directory.",
|
|
]
|
|
|
|
|
|
def _alnum_edge(word: str) -> str:
|
|
"""Strip leading/trailing non-alphanumerics from *word*."""
|
|
start = 0
|
|
end = len(word)
|
|
while start < end and not word[start].isalnum():
|
|
start += 1
|
|
while end > start and not word[end - 1].isalnum():
|
|
end -= 1
|
|
return word[start:end]
|
|
|
|
|
|
def derive_token(content: str, title: str) -> str:
|
|
"""The derived battery's grep token (locked by the phase-72 task
|
|
file). The first whitespace-split word of *content* whose stripped
|
|
length is >= 6 (leading/trailing non-alphanumerics stripped,
|
|
lowercased); the fallback is the first word of *title* (the same
|
|
cleanup)."""
|
|
for word in content.split():
|
|
token = _alnum_edge(word).lower()
|
|
if len(token) >= 6:
|
|
return token
|
|
words = title.split()
|
|
return _alnum_edge(words[0]).lower() if words else "document"
|
|
|
|
|
|
def build_battery(
|
|
catalog: list[tuple[str, str, str]], d2_content: str
|
|
) -> list[str]:
|
|
"""The fixed 10-question battery (locked by the phase-72 task file —
|
|
do not swap in easier questions), derived from the live catalog's
|
|
first two documents ``D1 = (s1, p1, t1)`` / ``D2 = (s2, p2, t2)``
|
|
(catalog order):
|
|
|
|
1. the incident ("list the files in this directory" — the
|
|
harness-prior ``ls(path='.')`` misuse);
|
|
2. a scoped ``ls`` by the correct source name (``s1``);
|
|
3. the no-arg listing;
|
|
4. a bare-path ``read`` trap (``p1`` without its source prefix);
|
|
5. the combined form (the correct shape);
|
|
6. a second bare-path ``read`` trap (``p2``);
|
|
7. the ``grep`` turn (the derived token — guaranteed to occur in
|
|
D2's content);
|
|
8. a title lookup + read (``t2``);
|
|
9. a title lookup + open (``t1``);
|
|
10. the source name phrased as a directory (``s2`` — the
|
|
``ls(path=…)`` scope trap).
|
|
"""
|
|
(s1, p1, t1), (s2, p2, t2) = catalog[0], catalog[1]
|
|
token = derive_token(d2_content, t2)
|
|
return [
|
|
"List the files in this directory.",
|
|
f"List the documents you have in the {s1} source.",
|
|
"List every document you have indexed.",
|
|
f"What does the document {p1} contain? Open it and tell me.",
|
|
f"Read {s1}/{p1} and summarize it.",
|
|
f"Open the document {p2} and tell me what it covers.",
|
|
f'Find the exact string "{token}" in your documents and tell me '
|
|
"which ones contain it.",
|
|
f'Which document has the title "{t2}"? Read it and summarize.',
|
|
f"What do you know about {t1}? Open the relevant document and give "
|
|
"me specifics.",
|
|
f"List the files in the {s2} directory.",
|
|
]
|
|
|
|
|
|
def check_preconditions(
|
|
settings: Settings, mode: str, dump: Path
|
|
) -> int | None:
|
|
"""The locked preconditions — ``2`` on failure (an actionable line is
|
|
printed), ``None`` when all hold: the DB is reachable and
|
|
``agent_max_rounds`` > 0 (both modes); ``derived`` — the catalog
|
|
holds >=2 documents and the first two catalog documents' ``path``s
|
|
each contain ``/`` (the bare-path traps need nested paths);
|
|
``fixture`` — the fixture dump exists."""
|
|
if not db_available():
|
|
print(
|
|
"precondition failed: database unreachable — start Postgres "
|
|
"with `podman compose up -d db` and re-run"
|
|
)
|
|
return 2
|
|
if settings.agent_max_rounds <= 0:
|
|
print(
|
|
"precondition failed: BOR_AGENT_MAX_ROUNDS is "
|
|
f"{settings.agent_max_rounds} (the no-tools kill switch) — set "
|
|
"it to a positive value for the gate"
|
|
)
|
|
return 2
|
|
if mode == "fixture":
|
|
if not dump.is_file():
|
|
print(
|
|
"precondition failed: fixture dump missing "
|
|
f"({dump}) — build it once: "
|
|
"`uv run python -m scripts.load_test_kb`"
|
|
)
|
|
return 2
|
|
return None
|
|
with SessionLocal() as db:
|
|
catalog = _catalog(db)
|
|
if len(catalog) < 2:
|
|
print(
|
|
f"precondition failed: catalog holds {len(catalog)} document(s) "
|
|
"(need >= 2) — import a knowledge base first: "
|
|
"`uv run python -m scripts.import_docs`"
|
|
)
|
|
return 2
|
|
bad = [(source, path) for source, path, _ in catalog[:2] if "/" not in path]
|
|
if bad:
|
|
print(
|
|
"precondition failed: the first two catalog documents' paths must "
|
|
f"each contain '/' (the bare-path traps need nested paths) — got "
|
|
f"{bad!r}; import a source with a nested directory layout"
|
|
)
|
|
return 2
|
|
return None
|
|
|
|
|
|
@dataclass
|
|
class TurnResult:
|
|
"""One battery turn's measurements (from the consumed stream + the
|
|
holder — no app-code changes for measurement)."""
|
|
|
|
index: int
|
|
question: str
|
|
max_rounds: int # settings.agent_max_rounds for this run
|
|
emitted: int = 0 # ToolCallPieces the stream yielded
|
|
executed: int = 0 # holder.tool_calls (refusals count in nothing)
|
|
answered: bool = True # the stream finished without LLMError
|
|
error: str = "" # the terminal error (when not answered)
|
|
deflected: bool = False # the honesty gate deflected (no tools offered)
|
|
seconds: float = 0.0 # wall time for the whole turn (embed → settled)
|
|
calls: list[tuple[str, dict[str, Any]]] = field(
|
|
default_factory=list
|
|
) # every emitted (name, arguments) — the contract-accuracy input
|
|
|
|
@property
|
|
def cap_reached(self) -> bool:
|
|
"""Every capped round emitted a call, so the cap implies at
|
|
least ``max_rounds`` emissions — and never the reverse."""
|
|
return self.emitted >= self.max_rounds
|
|
|
|
def display(self) -> str:
|
|
"""The per-turn line: ``turn NN | emitted=E executed=X
|
|
cap=Y|N defl=Y|N | Ss | <question>`` (the locked core plus the
|
|
2026-09-04 additions — the deflection flag and the turn's wall
|
|
seconds — so a slow-down is visible on the same line)."""
|
|
text = self.question
|
|
if len(text) > QUESTION_DISPLAY_WIDTH:
|
|
cut = text[:QUESTION_DISPLAY_WIDTH]
|
|
text = cut.rsplit(" ", 1)[0].rstrip(" ,;:") + " …"
|
|
return (
|
|
f"turn {self.index:02d} | emitted={self.emitted} "
|
|
f"executed={self.executed} cap={'yes' if self.cap_reached else 'no'} "
|
|
f"defl={'yes' if self.deflected else 'no'} | {self.seconds:5.2f}s "
|
|
f"| {text}"
|
|
)
|
|
|
|
|
|
async def run_turn(
|
|
llm: LLMClient, settings: Settings, index: int, question: str
|
|
) -> TurnResult:
|
|
"""One battery question through the REAL grounded path — the exact
|
|
mirror of ``app.api.chat`` (same prompt the UI gets): embed the
|
|
question, retrieve, the honesty gate (``plan_turn``), read the
|
|
steering notes + KB overview the way chat.py reads them, then —
|
|
grounded — ``run_agent`` with a **fresh** :class:`AgentHolder`, or —
|
|
deflected — the ``tools=None`` stream with the one bounded
|
|
scaffolding recovery. Every yielded piece is consumed to the end;
|
|
``emitted`` counts the yielded :class:`ToolCallPiece` values,
|
|
``executed`` is the holder's executed-call count (refusals count in
|
|
nothing). A turn that dies with ``LLMError`` /
|
|
``MalformedReplyError`` (the latter subclasses the former) or an
|
|
``EmbeddingError`` is not ``answered``. ``seconds`` is the turn's
|
|
wall time (embed → settled).
|
|
"""
|
|
result = TurnResult(
|
|
index=index, question=question, max_rounds=settings.agent_max_rounds
|
|
)
|
|
started = time.monotonic()
|
|
try:
|
|
with SessionLocal() as db:
|
|
steering_notes = load_steering_notes(db)
|
|
kb_text = (load_kb_overview(db) or "").strip()
|
|
question_vec = await llm.embed_one(question)
|
|
chunks = retrieve(db, question, question_vec)
|
|
plan = plan_turn(chunks, settings, notes=steering_notes, kb_overview=kb_text)
|
|
if plan.deflected:
|
|
result.deflected = True
|
|
await _run_deflected(llm, settings, plan.system_prompt, question, result)
|
|
else:
|
|
holder = AgentHolder()
|
|
# SEC-14-04: pass a session factory, not the outer session
|
|
stream = run_agent(
|
|
llm,
|
|
lambda: SessionLocal(),
|
|
system_prompt=plan.system_prompt,
|
|
user_message=question,
|
|
seed_docs=plan.suggested_docs, # phase 118: the suggestion tier
|
|
settings=settings,
|
|
holder=holder,
|
|
)
|
|
async for piece in stream:
|
|
if isinstance(piece, ToolCallPiece):
|
|
result.emitted += 1
|
|
result.calls.append((piece.name, dict(piece.arguments)))
|
|
result.executed = holder.tool_calls
|
|
except (LLMError, EmbeddingError) as e:
|
|
result.answered = False
|
|
result.error = f"{type(e).__name__}: {e}"
|
|
result.seconds = time.monotonic() - started
|
|
return result
|
|
|
|
|
|
async def _run_deflected(
|
|
llm: LLMClient,
|
|
settings: Settings,
|
|
system_prompt: str,
|
|
question: str,
|
|
result: TurnResult,
|
|
) -> None:
|
|
"""The deflected mirror of ``app.api.chat``: one ``tools=None``
|
|
request through the retry primitive with a caller-owned
|
|
:class:`ScaffoldingFilter`, and — when the filter wiped the whole
|
|
reply — exactly ONE bounded recovery (``tools=None``,
|
|
:data:`CORRECTION_INSTRUCTION` folded into the single system
|
|
message, a fresh filter). A second empty reply raises
|
|
:class:`MalformedReplyError` (the turn is not ``answered``). No tool
|
|
call can be emitted on this path (``tools=None``)."""
|
|
messages = [
|
|
{"role": "system", "content": system_prompt},
|
|
{"role": "user", "content": question},
|
|
]
|
|
first_filter = ScaffoldingFilter()
|
|
content_chars = 0
|
|
stream = chat_stream_retried(
|
|
llm,
|
|
messages,
|
|
tools=None,
|
|
retries=settings.llm_retries,
|
|
delay=settings.llm_retry_delay,
|
|
scaffolding=first_filter,
|
|
)
|
|
try:
|
|
async for piece in stream:
|
|
if isinstance(piece, StreamPiece) and piece.kind == "content":
|
|
content_chars += len(piece.text)
|
|
finally:
|
|
await stream.aclose()
|
|
if content_chars == 0 and first_filter.stripped_chars > 0:
|
|
recovery_filter = ScaffoldingFilter()
|
|
recovered = chat_stream_retried(
|
|
llm,
|
|
[
|
|
{
|
|
"role": "system",
|
|
"content": system_prompt + "\n" + CORRECTION_INSTRUCTION,
|
|
},
|
|
*messages[1:],
|
|
],
|
|
tools=None,
|
|
retries=settings.llm_retries,
|
|
delay=settings.llm_retry_delay,
|
|
scaffolding=recovery_filter,
|
|
)
|
|
recovery_content = 0
|
|
try:
|
|
async for piece in recovered:
|
|
if isinstance(piece, StreamPiece) and piece.kind == "content":
|
|
recovery_content += len(piece.text)
|
|
finally:
|
|
await recovered.aclose()
|
|
if recovery_content == 0:
|
|
# Terminal, as in chat.py: the dedicated error, no done.
|
|
raise MalformedReplyError(
|
|
"the deflected model answered in raw tool-scaffolding twice "
|
|
"in a row — no clean answer"
|
|
)
|
|
|
|
|
|
def _ls_folder_prefixes(catalog: set[tuple[str, str]]) -> dict[str, set[str]]:
|
|
"""The existing ``ls`` folders per source (the phase-94 drill-down
|
|
contract): for each source, the set of source-relative folder
|
|
prefixes — every slash-boundary prefix of its indexed paths.
|
|
|
|
Mirrors the ``00_phase.md`` existence rule that
|
|
``app.rag.agent._execute_tool`` applies: a folder ``F`` exists ⟺
|
|
some indexed path of the source starts with ``F + "/"`` — a
|
|
document's OWN path is never a folder, so the document's path itself
|
|
is not in the set (``ls`` of a file path is a refusal, as is a bare
|
|
folder name missing its source prefix).
|
|
"""
|
|
folders: dict[str, set[str]] = {}
|
|
for source, path in catalog:
|
|
folder = path[: path.rfind("/")] if "/" in path else ""
|
|
while folder:
|
|
folders.setdefault(source, set()).add(folder)
|
|
folder = folder[: folder.rfind("/")] if "/" in folder else ""
|
|
return folders
|
|
|
|
|
|
def classify_call(
|
|
name: str,
|
|
args: dict[str, Any],
|
|
catalog: set[tuple[str, str]],
|
|
sources: set[str],
|
|
ls_folders: dict[str, set[str]],
|
|
) -> bool:
|
|
"""Contract correctness of ONE emitted call (the tool-calling
|
|
accuracy metric, 2026-09-04 controlled methodology).
|
|
|
|
A call is contract-correct when it uses a known tool with well-formed
|
|
required arguments that target a RESOLVABLE entity — the phase-72
|
|
incident class (unknown scopes, bare document paths, hallucinated
|
|
identities, ``ls(path='/')``-style misuse) is exactly what this
|
|
flags. An ``ALREADY_IN_CONTEXT`` re-read is NOT flagged: the call is
|
|
well-formed and names a real document — the app's context dedupe
|
|
refusing a redundant read is an app-semantics choice, not a
|
|
tool-calling error (the controlled gate's telemetry — the same
|
|
re-read 15/15 across copy variants — is documented in
|
|
``TOOL_CALLING_TESTING.md``). The classification mirrors
|
|
``app.rag.agent._execute_tool``'s resolution rules gate-side (no
|
|
app-code changes for measurement) — including the phase-94 ``ls``
|
|
revision (owner-permitted tool-surface change, ``00_phase.md``):
|
|
``ls(path)`` is contract-correct for ``""`` (the synced sources),
|
|
a registered source name (its root folder), or a ``source/folder``
|
|
path naming an EXISTING folder (``ls_folders`` — the drill-down);
|
|
an unknown first segment, a bare folder name (no source prefix), or
|
|
a folder matching no indexed prefix is the violation.
|
|
"""
|
|
if name == "ls":
|
|
raw = args.get("path")
|
|
scope = raw.strip() if isinstance(raw, str) else ""
|
|
if not scope:
|
|
return True # the top level (the synced sources)
|
|
source, _, rest = scope.partition("/")
|
|
if source not in sources:
|
|
return False # unknown first segment (the phase-72 incident class)
|
|
if not rest:
|
|
return True # a registered source's root folder
|
|
return rest in ls_folders.get(source, ())
|
|
if name == "read":
|
|
raw = args.get("path")
|
|
arg = raw.strip() if isinstance(raw, str) else ""
|
|
if "/" not in arg:
|
|
return False # a bare name can never be a document
|
|
source, _, path = arg.partition("/")
|
|
return (source, path) in catalog
|
|
if name == "grep":
|
|
raw_pattern = args.get("pattern")
|
|
if not (isinstance(raw_pattern, str) and raw_pattern.strip()):
|
|
return False
|
|
raw_path = args.get("path")
|
|
scope = raw_path.strip() if isinstance(raw_path, str) else ""
|
|
if scope:
|
|
if "/" not in scope:
|
|
return False
|
|
source, _, path = scope.partition("/")
|
|
return (source, path) in catalog
|
|
return True
|
|
return False # unknown tool
|
|
|
|
|
|
def score_contract(
|
|
turns: list[TurnResult],
|
|
catalog: set[tuple[str, str]],
|
|
sources: set[str],
|
|
) -> int:
|
|
"""Contract accuracy across the run: how many emitted calls are
|
|
contract-correct (:func:`classify_call`) against the run's catalog
|
|
and source names. Per-turn counts are attached on each
|
|
:class:`TurnResult` as ``_contract_ok`` (measurement state, not a
|
|
dataclass field — the display line stays the locked format)."""
|
|
ls_folders = _ls_folder_prefixes(catalog)
|
|
ok = 0
|
|
for turn in turns:
|
|
turn_ok = sum(
|
|
1
|
|
for name, args in turn.calls
|
|
if classify_call(name, args, catalog, sources, ls_folders)
|
|
)
|
|
turn._contract_ok = turn_ok # type: ignore[attr-defined]
|
|
ok += turn_ok
|
|
return ok
|
|
|
|
|
|
def evaluate(
|
|
turns: list[TurnResult], partial: bool = False, mode: str = "derived"
|
|
) -> tuple[bool, list[tuple[str, bool, str]]]:
|
|
"""The pass conditions → ``(passed, [(name, ok, detail)])``.
|
|
|
|
``contract_ok`` per turn was attached by :func:`score_contract`
|
|
before this call (0 while unattached — main always scores first).
|
|
|
|
1. all turns ``answered``; 2. zero ``cap_reached`` turns; 3. >=6 of
|
|
10 turns with ``emitted >= 1`` (on a ``partial`` run — ``--turns N``
|
|
with N < the battery length — the count is REPORTED but not gated:
|
|
a short micro-loop exists for copy iteration, not as the gate);
|
|
4. the accuracy bar — ``derived`` mode (the phase-72 LOCKED gate):
|
|
``executed / emitted >= 0.90``; ``fixture`` mode (the 2026-09-04
|
|
controlled methodology): **contract accuracy** (well-formed calls
|
|
targeting resolvable entities, :func:`classify_call`) >= 0.90 — with
|
|
the executed ratio REPORTED alongside (a run with zero emitted calls
|
|
fails condition 3 anyway, so an empty denominator does not sink
|
|
condition 4).
|
|
"""
|
|
n = len(turns)
|
|
failed = [t for t in turns if not t.answered]
|
|
caps = [t for t in turns if t.cap_reached]
|
|
tool_turns = sum(1 for t in turns if t.emitted >= 1)
|
|
emitted = sum(t.emitted for t in turns)
|
|
executed = sum(t.executed for t in turns)
|
|
deflected = sum(1 for t in turns if t.deflected)
|
|
contract_ok = sum(
|
|
getattr(t, "_contract_ok", 0) for t in turns # type: ignore[attr-defined]
|
|
)
|
|
failed_detail = f"{n - len(failed)}/{n}"
|
|
if failed:
|
|
failed_detail += "; " + "; ".join(
|
|
f"turn {t.index:02d}: {t.error}" for t in failed
|
|
)
|
|
conditions: list[tuple[str, bool, str]] = [
|
|
("all turns answered", not failed, failed_detail),
|
|
(
|
|
"zero cap-reached turns",
|
|
not caps,
|
|
"0" if not caps else f"{len(caps)} hit the round cap: "
|
|
+ ", ".join(f"turn {t.index:02d}" for t in caps),
|
|
),
|
|
(
|
|
">= 6 of 10 turns with >= 1 emitted tool call",
|
|
True if partial else tool_turns >= 6,
|
|
f"{tool_turns}/{n}"
|
|
+ (" (partial run — reported, not gated)" if partial else "")
|
|
+ (
|
|
f"; {deflected} deflected (no tools offered — the honesty gate)"
|
|
if deflected
|
|
else ""
|
|
),
|
|
),
|
|
]
|
|
if emitted == 0:
|
|
conditions.append(
|
|
(
|
|
"accuracy bar >= 0.90 across the run",
|
|
True, # no calls emitted — condition 3 already fails
|
|
"0/0 (no calls emitted — condition 3 fails)",
|
|
)
|
|
)
|
|
elif mode == "fixture":
|
|
# The controlled methodology's accuracy bar: contract accuracy.
|
|
# The executed ratio is reported right below it (not gated — it
|
|
# includes the app's ALREADY_IN_CONTEXT dedupe refusals, which
|
|
# the controlled telemetry shows are copy-invariant model
|
|
# behavior, not tool-calling errors).
|
|
contract_ratio = contract_ok / emitted
|
|
conditions.append(
|
|
(
|
|
"contract accuracy >= 0.90 (well-formed calls, resolvable targets)",
|
|
contract_ratio >= 0.90,
|
|
f"{contract_ok}/{emitted} ({round(100 * contract_ratio)}%)",
|
|
)
|
|
)
|
|
conditions.append(
|
|
(
|
|
"executed / emitted (reported — includes in-context dedupe refusals)",
|
|
True,
|
|
f"{executed}/{emitted} ({round(100 * executed / emitted)}%)",
|
|
)
|
|
)
|
|
else:
|
|
ratio = executed / emitted
|
|
conditions.append(
|
|
(
|
|
"executed / emitted >= 0.90 across the run (phase-72 locked)",
|
|
ratio >= 0.90,
|
|
f"{executed}/{emitted} ({round(100 * ratio)}%) — "
|
|
f"contract {contract_ok}/{emitted} ({round(100 * contract_ok / emitted)}%)",
|
|
)
|
|
)
|
|
return all(ok for _name, ok, _detail in conditions), conditions
|
|
|
|
|
|
def verdict_line(
|
|
model: str,
|
|
turns: list[TurnResult],
|
|
passed: bool,
|
|
wall_seconds: float,
|
|
partial: bool,
|
|
mode: str = "derived",
|
|
) -> str:
|
|
"""The single stable verdict line (model = the configured chat
|
|
model, date = run date, the run's total wall time since 2026-09-04,
|
|
both accuracy metrics since the 2026-09-04 controlled methodology)::
|
|
|
|
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=8 calls
|
|
7/12 executed (58%) contract 12/12 (100%) 2026-09-04 (wall 48.8s)
|
|
"""
|
|
answered = sum(1 for t in turns if t.answered)
|
|
caps = sum(1 for t in turns if t.cap_reached)
|
|
tool_turns = sum(1 for t in turns if t.emitted >= 1)
|
|
emitted = sum(t.emitted for t in turns)
|
|
executed = sum(t.executed for t in turns)
|
|
contract_ok = sum(
|
|
getattr(t, "_contract_ok", 0) for t in turns # type: ignore[attr-defined]
|
|
)
|
|
pct = round(100 * executed / emitted) if emitted else 0
|
|
cpct = round(100 * contract_ok / emitted) if emitted else 0
|
|
deflected = sum(1 for t in turns if t.deflected)
|
|
return (
|
|
f"gate: {model} {'PASS' if passed else 'FAIL'} turns={len(turns)}"
|
|
+ (" (partial)" if partial else "")
|
|
+ f" answered={answered} caps={caps} tool-turns={tool_turns}"
|
|
+ (f" deflected={deflected}" if deflected else "")
|
|
+ f" calls {executed}/{emitted} executed ({pct}%) "
|
|
f"contract {contract_ok}/{emitted} ({cpct}%) "
|
|
f"{date.today().isoformat()} (wall {wall_seconds:.1f}s)"
|
|
)
|
|
|
|
|
|
async def run_battery(
|
|
llm: LLMClient,
|
|
settings: Settings,
|
|
battery: list[str],
|
|
concurrency: int = 1,
|
|
) -> list[TurnResult]:
|
|
"""The whole battery, printing the per-turn line as each turn
|
|
settles. ``concurrency > 1`` runs that many turns at once against
|
|
the endpoint (the aggregate verdict is unchanged — the conditions
|
|
are run-wide sums; the per-turn lines may then print out of order)."""
|
|
turns: list[TurnResult] = []
|
|
if concurrency <= 1:
|
|
for index, question in enumerate(battery, start=1):
|
|
turns.append(await run_turn(llm, settings, index, question))
|
|
print(turns[-1].display())
|
|
return turns
|
|
semaphore = asyncio.Semaphore(concurrency)
|
|
|
|
async def one(index: int, question: str) -> TurnResult:
|
|
async with semaphore:
|
|
return await run_turn(llm, settings, index, question)
|
|
|
|
gather = [
|
|
asyncio.create_task(one(index, question))
|
|
for index, question in enumerate(battery, start=1)
|
|
]
|
|
for finished in asyncio.as_completed(gather):
|
|
result = await finished
|
|
turns.append(result)
|
|
print(result.display())
|
|
turns.sort(key=lambda t: t.index)
|
|
return turns
|
|
|
|
|
|
def main(argv: list[str] | None = None) -> int:
|
|
# CLI-only: pick up .env without side effects on import (the
|
|
# house probe pattern, cf. scripts/llm_probe.py).
|
|
load_dotenv()
|
|
parser = argparse.ArgumentParser(
|
|
description=(
|
|
"The real-model tool-calling gate: drive a fixed question "
|
|
"battery through the real grounded path against the live "
|
|
"endpoint with the configured chat model, and print the "
|
|
"PASS/FAIL verdict against the four locked conditions "
|
|
"(exit 0 PASS, 1 FAIL, 2 precondition failure). The fast "
|
|
"loop: --restore --mode fixture."
|
|
)
|
|
)
|
|
parser.add_argument(
|
|
"--mode",
|
|
choices=["derived", "fixture"],
|
|
default="derived",
|
|
help=(
|
|
"derived (default, the phase-72 locked battery from the live "
|
|
"catalog) or fixture (the curated battery pinned to the "
|
|
"fixture KB — pair with --restore)"
|
|
),
|
|
)
|
|
parser.add_argument(
|
|
"--restore",
|
|
action="store_true",
|
|
help="restore the fixture KB dump into the database first (one "
|
|
"transaction — no git clone, no re-embedding)",
|
|
)
|
|
parser.add_argument(
|
|
"--turns",
|
|
type=int,
|
|
default=None,
|
|
metavar="N",
|
|
help="run only the first N battery questions (the micro-loop for "
|
|
"copy iteration; the verdict is marked partial and condition 3 "
|
|
"is reported, not gated)",
|
|
)
|
|
parser.add_argument(
|
|
"--concurrency",
|
|
type=int,
|
|
default=1,
|
|
metavar="N",
|
|
help="run up to N turns at once (default 1 — sequential; the "
|
|
"aggregate verdict is unchanged)",
|
|
)
|
|
parser.add_argument(
|
|
"--dump",
|
|
type=Path,
|
|
default=None,
|
|
metavar="PATH",
|
|
help="the fixture dump for --restore / --mode fixture (default: "
|
|
f"{DEFAULT_DUMP_PATH})",
|
|
)
|
|
args = parser.parse_args(argv)
|
|
logging.basicConfig(
|
|
level=logging.INFO, format="%(levelname)s %(name)s: %(message)s"
|
|
)
|
|
|
|
settings = get_settings()
|
|
dump = args.dump or DEFAULT_DUMP_PATH
|
|
rc = check_preconditions(settings, args.mode, dump)
|
|
if rc is not None:
|
|
return rc
|
|
|
|
run_started = time.monotonic()
|
|
if args.restore:
|
|
from scripts.restore_test_kb import restore_dump
|
|
|
|
print(f"restore: {dump} …")
|
|
t0 = time.monotonic()
|
|
restored = restore_dump(dump)
|
|
print(
|
|
f"restore: ok in {time.monotonic() - t0:.2f}s "
|
|
f"({restored.docs} docs, {len(restored.sources)} sources)"
|
|
)
|
|
|
|
if args.mode == "fixture":
|
|
battery = list(FIXTURE_BATTERY)
|
|
logger.info(
|
|
"gate: model=%s mode=fixture battery=%d questions (fixture KB)",
|
|
settings.llm_chat_model,
|
|
len(battery),
|
|
)
|
|
else:
|
|
with SessionLocal() as db:
|
|
catalog = _catalog(db)
|
|
s2, p2, _t2 = catalog[1]
|
|
d2 = find_document(db, s2, p2)
|
|
d2_content = d2.content if d2 is not None else ""
|
|
battery = build_battery(catalog, d2_content)
|
|
logger.info(
|
|
"gate: model=%s kb_docs=%d battery=%d questions",
|
|
settings.llm_chat_model,
|
|
len(catalog),
|
|
len(battery),
|
|
)
|
|
|
|
if args.turns is not None:
|
|
if args.turns <= 0:
|
|
print("error: --turns must be >= 1")
|
|
return 2
|
|
battery = battery[: args.turns]
|
|
|
|
for number, question in enumerate(battery, start=1):
|
|
print(f" {number:02d}. {question}")
|
|
|
|
llm = LLMClient(settings)
|
|
turns = asyncio.run(
|
|
run_battery(llm, settings, battery, concurrency=max(1, args.concurrency))
|
|
)
|
|
wall = time.monotonic() - run_started
|
|
|
|
# The contract-accuracy classification needs the run's catalog +
|
|
# source names (the KB is static across the run — the restore, when
|
|
# any, happened before the battery).
|
|
with SessionLocal() as db:
|
|
catalog_set = set((s, p) for s, p, _t in _catalog(db))
|
|
sources_set = set(list_source_names(db))
|
|
score_contract(turns, catalog_set, sources_set)
|
|
passed, conditions = evaluate(
|
|
turns, partial=args.turns is not None, mode=args.mode
|
|
)
|
|
print(
|
|
verdict_line(
|
|
settings.llm_chat_model, turns, passed, wall, args.turns is not None, args.mode
|
|
)
|
|
)
|
|
if not passed:
|
|
print("conditions (the MISS(es) mark the copy lever to iterate):")
|
|
for name, ok, detail in conditions:
|
|
print(f" [{'ok ' if ok else 'MISS'}] {name}: {detail}")
|
|
return 0 if passed else 1
|
|
|
|
|
|
if __name__ == "__main__":
|
|
sys.exit(main())
|