"""The real-model tool-calling gate (live, the configured chat model). The phase-72 pass condition (owner directive 2026-09-03 — "test with the real lite model until tool calls work consistently; don't pass until a sufficient number of tool calls succeed"): this script drives a fixed question battery through the **real** grounded path — the exact mirror of ``app.api.chat`` (embed → retrieve → the honesty gate via ``plan_turn`` → the steering notes + KB overview exactly as ``app.api.chat`` reads them → ``build_high_prompt`` / deflection prompt → ``run_agent`` with the configured chat model and the real Postgres KB, a fresh ``AgentHolder`` per turn; a deflected turn runs the same ``tools=None`` + one-bounded-recovery stream the deflected API branch runs) — and applies the four LOCKED pass conditions: 1. all turns answer (no ``LLMError`` / ``MalformedReplyError``); 2. zero turns hit the round cap (the incident's loop signature — hitting the cap means the teaching did not end the loop); 3. >=6 of 10 turns emit >=1 tool call (the model keeps USING tools — it does not abandon them and answer from seed context alone, the incident's end state); 4. the accuracy bar across the whole run — ``derived`` mode (the phase-72 LOCKED gate): executed / emitted >= 0.90 (refusals count in nothing); ``fixture`` mode (the controlled methodology): contract accuracy — well-formed calls targeting resolvable entities — >= 0.90, with the executed ratio reported alongside (see ``classify_call`` and ``TOOL_CALLING_TESTING.md`` for why the two metrics differ: the app's ALREADY_IN_CONTEXT dedupe refusal is an app-semantics choice, not a tool-calling error). Batteries (``--mode``): * ``derived`` (default, the phase-72 LOCKED battery) — the fixed 10-question battery derived from the live catalog's first two documents ``D1 = (s1, p1, t1)`` / ``D2 = (s2, p2, t2)`` (catalog order); the grep question's token is the first whitespace-split word of ``D2.content`` with length >= 6 (leading/trailing non-alphanumerics stripped, lowercased), falling back to the first word of ``t2``. * ``fixture`` (the controlled fast loop, 2026-09-04 owner directive) — :data:`FIXTURE_BATTERY`, the curated 10 questions pinned to the hand-written fixture KB (``tests/fixtures/agent_kb/``). Combine with ``--restore`` so the whole iteration is one command against a known, unguessable, re-embed-free knowledge base (``TOOL_CALLING_TESTING.md``). Speed levers (the fast loop): ``--restore`` (restore the fixture dump in one transaction — no git clone, no re-embedding, sub-second); ``--turns N`` (run only the first N questions — the micro-loop for copy iteration; the verdict is then marked ``partial`` and condition 3 is reported, not gated); every turn line carries its wall seconds and the verdict line carries the run's total wall time, so a slow-down is visible in the same line that carries the accuracy. House probe pattern (``scripts/llm_probe.py``): ``uv run python -m scripts.agent_realmodel_check`` — argparse, dotenv, plain module, no debugpy. The script never modifies the KB (no commits, no query_log rows — the only writes are the ``--restore`` snapshot restore, which is explicit and transactional). Preconditions (exit 2 with an actionable line on failure): the DB is reachable; ``BOR_AGENT_MAX_ROUNDS`` is > 0 (the gate needs tools enabled); ``derived`` mode — the catalog holds >=2 documents and the FIRST TWO catalog documents' ``path``s each contain ``/`` (the bare-path traps need nested paths); ``fixture`` mode — the fixture dump exists (build it with ``uv run python -m scripts.load_test_kb``). Exit codes: **0 PASS**, **1 FAIL** (the per-condition breakdown is printed so the copy-lever iteration loop can target the right lever), **2 precondition failure**. For refusal diagnosis, every call is already logged by ``run_agent`` (``agent tool=… args=… round=…/…``) — correlate the logged arguments with the refusal templates in ``app/rag/agent.py`` to see which teaching line the model hit. """ from __future__ import annotations import argparse import asyncio import logging import sys import time from dataclasses import dataclass, field from datetime import date from pathlib import Path from typing import Any from dotenv import load_dotenv from sqlalchemy import select from app.api.chat import plan_turn from app.api.steering import load_steering_notes from app.config import Settings, get_settings from app.db import SessionLocal, db_available from app.models import Document from app.rag.agent import ( CORRECTION_INSTRUCTION, AgentHolder, MalformedReplyError, find_document, list_source_names, run_agent, ) from app.rag.llm import ( EmbeddingError, LLMClient, LLMError, StreamPiece, ToolCallPiece, chat_stream_retried, ) from app.rag.overview import load_kb_overview from app.rag.retriever import retrieve from app.rag.scaffolding import ScaffoldingFilter logger = logging.getLogger("agent_realmodel_check") #: Repo-relative fixture dump (written by scripts.load_test_kb). DEFAULT_DUMP_PATH = Path("tests/fixtures/test_kb.dump.sql") def _catalog(db) -> list[tuple[str, str, str]]: """The full catalog as ``(source, path, title)`` in ``(source, path)`` order — the battery builder's pre-phase-94 ``list_catalog`` contract, kept local to this script (phase 94 deleted the agent accessor: the drill-down ``ls`` lists one bounded source level at a time).""" rows = db.execute( select(Document.source, Document.path, Document.title).order_by( Document.source, Document.path ) ).all() return [(source, path, title) for source, path, title in rows] #: The per-turn line truncates the question at this width (the locked #: format prints ``turn NN | emitted=E executed=X cap=Y|N | ``). QUESTION_DISPLAY_WIDTH = 40 #: The controlled fast-loop battery (2026-09-04, owner directive): the #: curated 10 questions pinned to the fixture KB #: (``tests/fixtures/agent_kb/`` — sources ``deployments`` / ``homelab``, #: 8 hand-written documents whose specifics — ``rack7``, #: ``10.77.42.0/24``, VLAN 130, port 18443, ntfy topic #: ``reese-uptime-7``, machine ID ``rbm-8842``, the ``17 2 * * *`` #: schedule, image ``ghcr.io/reese/obsidian-bor:2026.7.14``, port #: 18765 — are not guessable by any model). #: #: Design rule (the controlled-test property): every question has ONE #: unambiguously correct tool behavior, verified by the load script's #: retrieval report. The ``read`` targets (Q4/Q5/Q6) are named so the #: question's tokens do NOT lexically seed the target document (the #: filenames carry no topical words the content repeats) — the read #: must actually happen, exactly once, in the combined form. The #: discipline turns (Q7/Q8/Q9) name content that IS seeded — the #: correct behavior there is to answer from the ```` context #: (or ``ls``), NOT to re-read. The bare-path traps are the job of the #: locked derived battery, not this one. #: #: 1. the phase-72 incident ("list the files in this directory" — a #: full listing needs ``ls``: 8 docs, 2 in seed context); #: 2. a scoped ``ls`` by the correct source name; #: 3. the no-arg listing; #: 4. a ``read`` of an unseeded document (combined form, given whole); #: 5. a second ``read`` of an unseeded document (explicit "Read …"); #: 6. a third ``read`` of an unseeded document; #: 7. the ``grep`` turn (``rbm-8842`` — a string that occurs in exactly #: one fixture document; the grep line alone answers "which ones"); #: 8. a title lookup — the target IS seeded; summarize from context; #: 9. a topic lookup — the target IS seeded; answer from context; #: 10. the source name phrased as a directory (the ``ls(path=…)`` scope #: trap). FIXTURE_BATTERY: list[str] = [ "List the files in this directory.", "List the documents you have in the homelab source.", "List every document you have indexed.", "Open the document homelab/networking/vela-bridges.md and tell me what it covers.", "Read deployments/quadlet/mimir-service.md and summarize it.", "Open the document homelab/networking/meridian-notes.md and tell me what it covers.", 'Find the exact string "rbm-8842" in your documents and tell me which ones ' "contain it.", 'Which document has the title "Lab Ansible Inventory"? Summarize it.', "What do you know about the qwen 3.8 llama.cpp setup? Give me the exact " "launch arguments.", "List the files in the deployments directory.", ] def _alnum_edge(word: str) -> str: """Strip leading/trailing non-alphanumerics from *word*.""" start = 0 end = len(word) while start < end and not word[start].isalnum(): start += 1 while end > start and not word[end - 1].isalnum(): end -= 1 return word[start:end] def derive_token(content: str, title: str) -> str: """The derived battery's grep token (locked by the phase-72 task file). The first whitespace-split word of *content* whose stripped length is >= 6 (leading/trailing non-alphanumerics stripped, lowercased); the fallback is the first word of *title* (the same cleanup).""" for word in content.split(): token = _alnum_edge(word).lower() if len(token) >= 6: return token words = title.split() return _alnum_edge(words[0]).lower() if words else "document" def build_battery( catalog: list[tuple[str, str, str]], d2_content: str ) -> list[str]: """The fixed 10-question battery (locked by the phase-72 task file — do not swap in easier questions), derived from the live catalog's first two documents ``D1 = (s1, p1, t1)`` / ``D2 = (s2, p2, t2)`` (catalog order): 1. the incident ("list the files in this directory" — the harness-prior ``ls(path='.')`` misuse); 2. a scoped ``ls`` by the correct source name (``s1``); 3. the no-arg listing; 4. a bare-path ``read`` trap (``p1`` without its source prefix); 5. the combined form (the correct shape); 6. a second bare-path ``read`` trap (``p2``); 7. the ``grep`` turn (the derived token — guaranteed to occur in D2's content); 8. a title lookup + read (``t2``); 9. a title lookup + open (``t1``); 10. the source name phrased as a directory (``s2`` — the ``ls(path=…)`` scope trap). """ (s1, p1, t1), (s2, p2, t2) = catalog[0], catalog[1] token = derive_token(d2_content, t2) return [ "List the files in this directory.", f"List the documents you have in the {s1} source.", "List every document you have indexed.", f"What does the document {p1} contain? Open it and tell me.", f"Read {s1}/{p1} and summarize it.", f"Open the document {p2} and tell me what it covers.", f'Find the exact string "{token}" in your documents and tell me ' "which ones contain it.", f'Which document has the title "{t2}"? Read it and summarize.', f"What do you know about {t1}? Open the relevant document and give " "me specifics.", f"List the files in the {s2} directory.", ] def check_preconditions( settings: Settings, mode: str, dump: Path ) -> int | None: """The locked preconditions — ``2`` on failure (an actionable line is printed), ``None`` when all hold: the DB is reachable and ``agent_max_rounds`` > 0 (both modes); ``derived`` — the catalog holds >=2 documents and the first two catalog documents' ``path``s each contain ``/`` (the bare-path traps need nested paths); ``fixture`` — the fixture dump exists.""" if not db_available(): print( "precondition failed: database unreachable — start Postgres " "with `podman compose up -d db` and re-run" ) return 2 if settings.agent_max_rounds <= 0: print( "precondition failed: BOR_AGENT_MAX_ROUNDS is " f"{settings.agent_max_rounds} (the no-tools kill switch) — set " "it to a positive value for the gate" ) return 2 if mode == "fixture": if not dump.is_file(): print( "precondition failed: fixture dump missing " f"({dump}) — build it once: " "`uv run python -m scripts.load_test_kb`" ) return 2 return None with SessionLocal() as db: catalog = _catalog(db) if len(catalog) < 2: print( f"precondition failed: catalog holds {len(catalog)} document(s) " "(need >= 2) — import a knowledge base first: " "`uv run python -m scripts.import_docs`" ) return 2 bad = [(source, path) for source, path, _ in catalog[:2] if "/" not in path] if bad: print( "precondition failed: the first two catalog documents' paths must " f"each contain '/' (the bare-path traps need nested paths) — got " f"{bad!r}; import a source with a nested directory layout" ) return 2 return None @dataclass class TurnResult: """One battery turn's measurements (from the consumed stream + the holder — no app-code changes for measurement).""" index: int question: str max_rounds: int # settings.agent_max_rounds for this run emitted: int = 0 # ToolCallPieces the stream yielded executed: int = 0 # holder.tool_calls (refusals count in nothing) answered: bool = True # the stream finished without LLMError error: str = "" # the terminal error (when not answered) deflected: bool = False # the honesty gate deflected (no tools offered) seconds: float = 0.0 # wall time for the whole turn (embed → settled) calls: list[tuple[str, dict[str, Any]]] = field( default_factory=list ) # every emitted (name, arguments) — the contract-accuracy input @property def cap_reached(self) -> bool: """Every capped round emitted a call, so the cap implies at least ``max_rounds`` emissions — and never the reverse.""" return self.emitted >= self.max_rounds def display(self) -> str: """The per-turn line: ``turn NN | emitted=E executed=X cap=Y|N defl=Y|N | Ss | `` (the locked core plus the 2026-09-04 additions — the deflection flag and the turn's wall seconds — so a slow-down is visible on the same line).""" text = self.question if len(text) > QUESTION_DISPLAY_WIDTH: cut = text[:QUESTION_DISPLAY_WIDTH] text = cut.rsplit(" ", 1)[0].rstrip(" ,;:") + " …" return ( f"turn {self.index:02d} | emitted={self.emitted} " f"executed={self.executed} cap={'yes' if self.cap_reached else 'no'} " f"defl={'yes' if self.deflected else 'no'} | {self.seconds:5.2f}s " f"| {text}" ) async def run_turn( llm: LLMClient, settings: Settings, index: int, question: str ) -> TurnResult: """One battery question through the REAL grounded path — the exact mirror of ``app.api.chat`` (same prompt the UI gets): embed the question, retrieve, the honesty gate (``plan_turn``), read the steering notes + KB overview the way chat.py reads them, then — grounded — ``run_agent`` with a **fresh** :class:`AgentHolder`, or — deflected — the ``tools=None`` stream with the one bounded scaffolding recovery. Every yielded piece is consumed to the end; ``emitted`` counts the yielded :class:`ToolCallPiece` values, ``executed`` is the holder's executed-call count (refusals count in nothing). A turn that dies with ``LLMError`` / ``MalformedReplyError`` (the latter subclasses the former) or an ``EmbeddingError`` is not ``answered``. ``seconds`` is the turn's wall time (embed → settled). """ result = TurnResult( index=index, question=question, max_rounds=settings.agent_max_rounds ) started = time.monotonic() try: with SessionLocal() as db: steering_notes = load_steering_notes(db) kb_text = (load_kb_overview(db) or "").strip() question_vec = await llm.embed_one(question) chunks = retrieve(db, question, question_vec) plan = plan_turn(chunks, settings, notes=steering_notes, kb_overview=kb_text) if plan.deflected: result.deflected = True await _run_deflected(llm, settings, plan.system_prompt, question, result) else: holder = AgentHolder() # SEC-14-04: pass a session factory, not the outer session stream = run_agent( llm, lambda: SessionLocal(), system_prompt=plan.system_prompt, user_message=question, seed_docs=plan.docs, settings=settings, holder=holder, ) async for piece in stream: if isinstance(piece, ToolCallPiece): result.emitted += 1 result.calls.append((piece.name, dict(piece.arguments))) result.executed = holder.tool_calls except (LLMError, EmbeddingError) as e: result.answered = False result.error = f"{type(e).__name__}: {e}" result.seconds = time.monotonic() - started return result async def _run_deflected( llm: LLMClient, settings: Settings, system_prompt: str, question: str, result: TurnResult, ) -> None: """The deflected mirror of ``app.api.chat``: one ``tools=None`` request through the retry primitive with a caller-owned :class:`ScaffoldingFilter`, and — when the filter wiped the whole reply — exactly ONE bounded recovery (``tools=None``, :data:`CORRECTION_INSTRUCTION` folded into the single system message, a fresh filter). A second empty reply raises :class:`MalformedReplyError` (the turn is not ``answered``). No tool call can be emitted on this path (``tools=None``).""" messages = [ {"role": "system", "content": system_prompt}, {"role": "user", "content": question}, ] first_filter = ScaffoldingFilter() content_chars = 0 stream = chat_stream_retried( llm, messages, tools=None, retries=settings.llm_retries, delay=settings.llm_retry_delay, scaffolding=first_filter, ) try: async for piece in stream: if isinstance(piece, StreamPiece) and piece.kind == "content": content_chars += len(piece.text) finally: await stream.aclose() if content_chars == 0 and first_filter.stripped_chars > 0: recovery_filter = ScaffoldingFilter() recovered = chat_stream_retried( llm, [ { "role": "system", "content": system_prompt + "\n" + CORRECTION_INSTRUCTION, }, *messages[1:], ], tools=None, retries=settings.llm_retries, delay=settings.llm_retry_delay, scaffolding=recovery_filter, ) recovery_content = 0 try: async for piece in recovered: if isinstance(piece, StreamPiece) and piece.kind == "content": recovery_content += len(piece.text) finally: await recovered.aclose() if recovery_content == 0: # Terminal, as in chat.py: the dedicated error, no done. raise MalformedReplyError( "the deflected model answered in raw tool-scaffolding twice " "in a row — no clean answer" ) def _ls_folder_prefixes(catalog: set[tuple[str, str]]) -> dict[str, set[str]]: """The existing ``ls`` folders per source (the phase-94 drill-down contract): for each source, the set of source-relative folder prefixes — every slash-boundary prefix of its indexed paths. Mirrors the ``00_phase.md`` existence rule that ``app.rag.agent._execute_tool`` applies: a folder ``F`` exists ⟺ some indexed path of the source starts with ``F + "/"`` — a document's OWN path is never a folder, so the document's path itself is not in the set (``ls`` of a file path is a refusal, as is a bare folder name missing its source prefix). """ folders: dict[str, set[str]] = {} for source, path in catalog: folder = path[: path.rfind("/")] if "/" in path else "" while folder: folders.setdefault(source, set()).add(folder) folder = folder[: folder.rfind("/")] if "/" in folder else "" return folders def classify_call( name: str, args: dict[str, Any], catalog: set[tuple[str, str]], sources: set[str], ls_folders: dict[str, set[str]], ) -> bool: """Contract correctness of ONE emitted call (the tool-calling accuracy metric, 2026-09-04 controlled methodology). A call is contract-correct when it uses a known tool with well-formed required arguments that target a RESOLVABLE entity — the phase-72 incident class (unknown scopes, bare document paths, hallucinated identities, ``ls(path='/')``-style misuse) is exactly what this flags. An ``ALREADY_IN_CONTEXT`` re-read is NOT flagged: the call is well-formed and names a real document — the app's context dedupe refusing a redundant read is an app-semantics choice, not a tool-calling error (the controlled gate's telemetry — the same re-read 15/15 across copy variants — is documented in ``TOOL_CALLING_TESTING.md``). The classification mirrors ``app.rag.agent._execute_tool``'s resolution rules gate-side (no app-code changes for measurement) — including the phase-94 ``ls`` revision (owner-permitted tool-surface change, ``00_phase.md``): ``ls(path)`` is contract-correct for ``""`` (the synced sources), a registered source name (its root folder), or a ``source/folder`` path naming an EXISTING folder (``ls_folders`` — the drill-down); an unknown first segment, a bare folder name (no source prefix), or a folder matching no indexed prefix is the violation. """ if name == "ls": raw = args.get("path") scope = raw.strip() if isinstance(raw, str) else "" if not scope: return True # the top level (the synced sources) source, _, rest = scope.partition("/") if source not in sources: return False # unknown first segment (the phase-72 incident class) if not rest: return True # a registered source's root folder return rest in ls_folders.get(source, ()) if name == "read": raw = args.get("path") arg = raw.strip() if isinstance(raw, str) else "" if "/" not in arg: return False # a bare name can never be a document source, _, path = arg.partition("/") return (source, path) in catalog if name == "grep": raw_pattern = args.get("pattern") if not (isinstance(raw_pattern, str) and raw_pattern.strip()): return False raw_path = args.get("path") scope = raw_path.strip() if isinstance(raw_path, str) else "" if scope: if "/" not in scope: return False source, _, path = scope.partition("/") return (source, path) in catalog return True return False # unknown tool def score_contract( turns: list[TurnResult], catalog: set[tuple[str, str]], sources: set[str], ) -> int: """Contract accuracy across the run: how many emitted calls are contract-correct (:func:`classify_call`) against the run's catalog and source names. Per-turn counts are attached on each :class:`TurnResult` as ``_contract_ok`` (measurement state, not a dataclass field — the display line stays the locked format).""" ls_folders = _ls_folder_prefixes(catalog) ok = 0 for turn in turns: turn_ok = sum( 1 for name, args in turn.calls if classify_call(name, args, catalog, sources, ls_folders) ) turn._contract_ok = turn_ok # type: ignore[attr-defined] ok += turn_ok return ok def evaluate( turns: list[TurnResult], partial: bool = False, mode: str = "derived" ) -> tuple[bool, list[tuple[str, bool, str]]]: """The pass conditions → ``(passed, [(name, ok, detail)])``. ``contract_ok`` per turn was attached by :func:`score_contract` before this call (0 while unattached — main always scores first). 1. all turns ``answered``; 2. zero ``cap_reached`` turns; 3. >=6 of 10 turns with ``emitted >= 1`` (on a ``partial`` run — ``--turns N`` with N < the battery length — the count is REPORTED but not gated: a short micro-loop exists for copy iteration, not as the gate); 4. the accuracy bar — ``derived`` mode (the phase-72 LOCKED gate): ``executed / emitted >= 0.90``; ``fixture`` mode (the 2026-09-04 controlled methodology): **contract accuracy** (well-formed calls targeting resolvable entities, :func:`classify_call`) >= 0.90 — with the executed ratio REPORTED alongside (a run with zero emitted calls fails condition 3 anyway, so an empty denominator does not sink condition 4). """ n = len(turns) failed = [t for t in turns if not t.answered] caps = [t for t in turns if t.cap_reached] tool_turns = sum(1 for t in turns if t.emitted >= 1) emitted = sum(t.emitted for t in turns) executed = sum(t.executed for t in turns) deflected = sum(1 for t in turns if t.deflected) contract_ok = sum( getattr(t, "_contract_ok", 0) for t in turns # type: ignore[attr-defined] ) failed_detail = f"{n - len(failed)}/{n}" if failed: failed_detail += "; " + "; ".join( f"turn {t.index:02d}: {t.error}" for t in failed ) conditions: list[tuple[str, bool, str]] = [ ("all turns answered", not failed, failed_detail), ( "zero cap-reached turns", not caps, "0" if not caps else f"{len(caps)} hit the round cap: " + ", ".join(f"turn {t.index:02d}" for t in caps), ), ( ">= 6 of 10 turns with >= 1 emitted tool call", True if partial else tool_turns >= 6, f"{tool_turns}/{n}" + (" (partial run — reported, not gated)" if partial else "") + ( f"; {deflected} deflected (no tools offered — the honesty gate)" if deflected else "" ), ), ] if emitted == 0: conditions.append( ( "accuracy bar >= 0.90 across the run", True, # no calls emitted — condition 3 already fails "0/0 (no calls emitted — condition 3 fails)", ) ) elif mode == "fixture": # The controlled methodology's accuracy bar: contract accuracy. # The executed ratio is reported right below it (not gated — it # includes the app's ALREADY_IN_CONTEXT dedupe refusals, which # the controlled telemetry shows are copy-invariant model # behavior, not tool-calling errors). contract_ratio = contract_ok / emitted conditions.append( ( "contract accuracy >= 0.90 (well-formed calls, resolvable targets)", contract_ratio >= 0.90, f"{contract_ok}/{emitted} ({round(100 * contract_ratio)}%)", ) ) conditions.append( ( "executed / emitted (reported — includes in-context dedupe refusals)", True, f"{executed}/{emitted} ({round(100 * executed / emitted)}%)", ) ) else: ratio = executed / emitted conditions.append( ( "executed / emitted >= 0.90 across the run (phase-72 locked)", ratio >= 0.90, f"{executed}/{emitted} ({round(100 * ratio)}%) — " f"contract {contract_ok}/{emitted} ({round(100 * contract_ok / emitted)}%)", ) ) return all(ok for _name, ok, _detail in conditions), conditions def verdict_line( model: str, turns: list[TurnResult], passed: bool, wall_seconds: float, partial: bool, mode: str = "derived", ) -> str: """The single stable verdict line (model = the configured chat model, date = run date, the run's total wall time since 2026-09-04, both accuracy metrics since the 2026-09-04 controlled methodology):: gate: lite PASS turns=10 answered=10 caps=0 tool-turns=8 calls 7/12 executed (58%) contract 12/12 (100%) 2026-09-04 (wall 48.8s) """ answered = sum(1 for t in turns if t.answered) caps = sum(1 for t in turns if t.cap_reached) tool_turns = sum(1 for t in turns if t.emitted >= 1) emitted = sum(t.emitted for t in turns) executed = sum(t.executed for t in turns) contract_ok = sum( getattr(t, "_contract_ok", 0) for t in turns # type: ignore[attr-defined] ) pct = round(100 * executed / emitted) if emitted else 0 cpct = round(100 * contract_ok / emitted) if emitted else 0 deflected = sum(1 for t in turns if t.deflected) return ( f"gate: {model} {'PASS' if passed else 'FAIL'} turns={len(turns)}" + (" (partial)" if partial else "") + f" answered={answered} caps={caps} tool-turns={tool_turns}" + (f" deflected={deflected}" if deflected else "") + f" calls {executed}/{emitted} executed ({pct}%) " f"contract {contract_ok}/{emitted} ({cpct}%) " f"{date.today().isoformat()} (wall {wall_seconds:.1f}s)" ) async def run_battery( llm: LLMClient, settings: Settings, battery: list[str], concurrency: int = 1, ) -> list[TurnResult]: """The whole battery, printing the per-turn line as each turn settles. ``concurrency > 1`` runs that many turns at once against the endpoint (the aggregate verdict is unchanged — the conditions are run-wide sums; the per-turn lines may then print out of order).""" turns: list[TurnResult] = [] if concurrency <= 1: for index, question in enumerate(battery, start=1): turns.append(await run_turn(llm, settings, index, question)) print(turns[-1].display()) return turns semaphore = asyncio.Semaphore(concurrency) async def one(index: int, question: str) -> TurnResult: async with semaphore: return await run_turn(llm, settings, index, question) gather = [ asyncio.create_task(one(index, question)) for index, question in enumerate(battery, start=1) ] for finished in asyncio.as_completed(gather): result = await finished turns.append(result) print(result.display()) turns.sort(key=lambda t: t.index) return turns def main(argv: list[str] | None = None) -> int: # CLI-only: pick up .env without side effects on import (the # house probe pattern, cf. scripts/llm_probe.py). load_dotenv() parser = argparse.ArgumentParser( description=( "The real-model tool-calling gate: drive a fixed question " "battery through the real grounded path against the live " "endpoint with the configured chat model, and print the " "PASS/FAIL verdict against the four locked conditions " "(exit 0 PASS, 1 FAIL, 2 precondition failure). The fast " "loop: --restore --mode fixture." ) ) parser.add_argument( "--mode", choices=["derived", "fixture"], default="derived", help=( "derived (default, the phase-72 locked battery from the live " "catalog) or fixture (the curated battery pinned to the " "fixture KB — pair with --restore)" ), ) parser.add_argument( "--restore", action="store_true", help="restore the fixture KB dump into the database first (one " "transaction — no git clone, no re-embedding)", ) parser.add_argument( "--turns", type=int, default=None, metavar="N", help="run only the first N battery questions (the micro-loop for " "copy iteration; the verdict is marked partial and condition 3 " "is reported, not gated)", ) parser.add_argument( "--concurrency", type=int, default=1, metavar="N", help="run up to N turns at once (default 1 — sequential; the " "aggregate verdict is unchanged)", ) parser.add_argument( "--dump", type=Path, default=None, metavar="PATH", help="the fixture dump for --restore / --mode fixture (default: " f"{DEFAULT_DUMP_PATH})", ) args = parser.parse_args(argv) logging.basicConfig( level=logging.INFO, format="%(levelname)s %(name)s: %(message)s" ) settings = get_settings() dump = args.dump or DEFAULT_DUMP_PATH rc = check_preconditions(settings, args.mode, dump) if rc is not None: return rc run_started = time.monotonic() if args.restore: from scripts.restore_test_kb import restore_dump print(f"restore: {dump} …") t0 = time.monotonic() restored = restore_dump(dump) print( f"restore: ok in {time.monotonic() - t0:.2f}s " f"({restored.docs} docs, {len(restored.sources)} sources)" ) if args.mode == "fixture": battery = list(FIXTURE_BATTERY) logger.info( "gate: model=%s mode=fixture battery=%d questions (fixture KB)", settings.llm_chat_model, len(battery), ) else: with SessionLocal() as db: catalog = _catalog(db) s2, p2, _t2 = catalog[1] d2 = find_document(db, s2, p2) d2_content = d2.content if d2 is not None else "" battery = build_battery(catalog, d2_content) logger.info( "gate: model=%s kb_docs=%d battery=%d questions", settings.llm_chat_model, len(catalog), len(battery), ) if args.turns is not None: if args.turns <= 0: print("error: --turns must be >= 1") return 2 battery = battery[: args.turns] for number, question in enumerate(battery, start=1): print(f" {number:02d}. {question}") llm = LLMClient(settings) turns = asyncio.run( run_battery(llm, settings, battery, concurrency=max(1, args.concurrency)) ) wall = time.monotonic() - run_started # The contract-accuracy classification needs the run's catalog + # source names (the KB is static across the run — the restore, when # any, happened before the battery). with SessionLocal() as db: catalog_set = set((s, p) for s, p, _t in _catalog(db)) sources_set = set(list_source_names(db)) score_contract(turns, catalog_set, sources_set) passed, conditions = evaluate( turns, partial=args.turns is not None, mode=args.mode ) print( verdict_line( settings.llm_chat_model, turns, passed, wall, args.turns is not None, args.mode ) ) if not passed: print("conditions (the MISS(es) mark the copy lever to iterate):") for name, ok, detail in conditions: print(f" [{'ok ' if ok else 'MISS'}] {name}: {detail}") return 0 if passed else 1 if __name__ == "__main__": sys.exit(main())