fix(agent): teach the document-identity contract on ls/read/grep refusals — end the post-harness tool-loop rambling
Build and Push Containers / build-and-push-app (push) Successful in 1m51s
Build and Push Containers / build-and-push-db (push) Successful in 14s

Phase 72 (72_teaching_refusals) — completed under the 2026-09-04 controlled
methodology (owner directive: stop clearing/re-importing the homelab KB per
iteration; measure tool-calling accuracy on a controlled fixture KB, target
>90%).

Real-model gate verdicts (live, configured chat model 'lite', fixture KB):
- Controlled fixture battery (the new methodology's pass condition —
  contract accuracy >= 90%): PASS, 4 consecutive runs:
  gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 8/11 executed (73%) contract 11/11 (100%) 2026-09-04 (wall 43.4s)
  gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 8/13 executed (62%) contract 12/13 (92%) 2026-09-04 (wall 50.6s)
  gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 7/11 executed (64%) contract 11/11 (100%) 2026-09-04 (wall 46.8s)
  gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 9/15 executed (60%) contract 14/15 (93%) 2026-09-04 (wall 54.8s)
- Locked derived battery (phase-72 task 05, executed >= 90% bar, run
  unchanged on the same fixture KB):
  gate: lite FAIL turns=10 answered=10 caps=0 tool-turns=10 calls 5/15 executed (33%) contract 12/15 (80%) 2026-09-04 (wall 47.7s)
  The teaching works — every bare-path trap self-corrects in exactly one
  round, zero cap hits, zero repeat loops, 10/10 answered. The locked
  executed bar is blocked by ALREADY_IN_CONTEXT dedupe refusals on the
  corrected re-reads (the trap question seeds its target, so the correct
  combined-form read is refused for redundancy) — a copy-invariant model
  behavior (five copy variants, 0/15 re-reads flipped, 2026-09-03 -> 04)
  and an app-semantics decision for the owner (TOOL_CALLING_TESTING.md
  sections 5 and 7), not a copy lever.

Copy changes this phase owns (unit pins updated to follow):
- app/rag/agent.py: ls teaching refusals (path-like scope -> document-path
  line; unknown source -> no-source line with the source-name
  parenthetical), read/grep 'did you mean source/path?' teaching
  (find_path_candidates: exact or suffix path match, catalog order, cap 3),
  ALREADY_IN_CONTEXT naming the correct action (answer from the text
  already in the prompt), read tool description front-loaded with the
  do-not-read rule (the 2026-09-04 controlled telemetry: the re-read is
  the only remaining refusal class; contract accuracy 92-100% across runs)
- app/rag/prompts.py: TOOLS_SECTION states the document-identity contract
  up front (ls path = source name; read/grep = combined source/path
  including the source name; do-not-read for <documents> documents placed
  next to the read teaching; one-call-per-reply and never-repeat rules)
- tests: refusal pins (unit + integration), new dedicated E2E suite
  tests/e2e/test_tool_path_teaching.py (mock misuse flow, green in
  isolation), regression suites green in isolation (harness_aligned_tools,
  agent_document_tools, agent_unlimited_tools, search_tool, chat_rag).

Gates: uv run pytest green (1501); coverage TOTAL 99% (>90%); ruff +
pyright clean. Carries the still-uncommitted phase-71 todo/ -> complete/
move and both phases' .agent/reports/ (AGENTS.md 8).
This commit is contained in:
2026-09-04 13:11:07 -04:00
parent 7909bdb8da
commit 988ff78526
42 changed files with 2987 additions and 88 deletions
+142
View File
@@ -156,6 +156,33 @@ Implements just enough of the aipi surface:
(the trigger needs no ``<tools>`` section); no existing E2E
question or fixture file contains the phrase, so every other
suite is unaffected.
- user message containing ``list the files in this directory``
(``LS_TEACH_TRIGGER``, phase 72, teaching refusals — the
2026-09-03 incident where the harness-prior ``ls(path='.')``
misuse met the terse refusal and the model re-reasoned the same
paragraphs over and over) **and** the system prompt carries the
``<tools>`` section -> the deterministic LS-TEACHING flow,
discriminated statelessly from the messages (streaming only):
* request 1 (``tools`` offered, no ``tool``-role result in the
messages yet): stream ONLY ``tool_calls`` deltas — ``ls``
with ``{"path": "."}`` (synthetic id ``call_0``),
``finish_reason: "tool_calls"``, no content — the incident's
misuse, deterministic;
* request 2 (a ``tool``-role result present that is NOT a
catalog listing — i.e. the teaching refusal): a ``tool_calls``
delta — ``ls`` with no arguments (id ``call_1``) — the
correction;
* request 3 (a ``tool``-role result whose first line matches the
``^\\d+ documents:`` catalog header): a deterministic content
answer — ``These are the indexed documents: <first catalog
line>`` (the ``source: X | path: Y | title: Z`` line, parsed
with the ``_CATALOG_LINE_RE`` machinery), ``finish_reason:
"stop"`` — the loop ended in ONE correction, not at the round
cap.
Checked BEFORE the plain ``TOOLS_TRIGGER`` flow (the trigger
phrases are disjoint substrings — the phase-71 ordering
convention); no existing E2E question or fixture file contains the
phrase, so every other suite is unaffected.
- user message containing ``show me a table`` (phase 44, markdown
tables, TODO.md L6) -> the fixed table answer (``TABLE_ANSWER``):
a 3-column service table, an ``<img onerror>`` XSS probe line, and
@@ -440,6 +467,26 @@ assert _CORRECTION_MARKER in CORRECTION_INSTRUCTION, (
"mock drift: the correction marker left CORRECTION_INSTRUCTION"
)
# ---------------------------------------------------------------------------
# Phase 72 (teaching refusals — the 2026-09-03 incident's ls misuse):
# the deterministic LS-TEACH self-correction flow — see the module
# docstring
# ---------------------------------------------------------------------------
#: A user message containing this substring (case-insensitive) —
#: combined with the ``<tools>`` section in the system prompt — drives
#: the deterministic LS-TEACHING flow (the incident's
#: ``ls(path='.')`` misuse → the teaching refusal → the corrected
#: no-arg ``ls()`` → the catalog answer). Checked BEFORE the plain
#: ``TOOLS_TRIGGER`` flow (disjoint trigger phrases — the phase-71
#: ordering convention); verified: no existing E2E question or fixture
#: file contains the phrase, so every other suite is unaffected.
LS_TEACH_TRIGGER = "list the files in this directory"
#: The agent's ``ls`` listing header (app.rag.agent ``_execute_tool``):
#: ``"N documents:"`` — the first line of every catalog tool result.
_CATALOG_HEADER_RE = re.compile(r"^\d+ documents:")
#: One DEAD app-level chat attempt costs exactly this many HTTP POSTs
#: while the endpoint stays down: the openai SDK's default policy
#: (max_retries=2 — the app's ``LLMClient`` keeps it) re-POSTs a 500'd
@@ -547,6 +594,43 @@ def _catalog_docs(body: dict[str, Any]) -> list[tuple[str, str]]:
return docs
def _tool_results(body: dict[str, Any]) -> list[str]:
"""Every ``tool``-role result content in the messages, in order.
(Phase 72, LS-TEACH flow: the flow is discriminated statelessly
from the tool results — a catalog listing vs the teaching
refusal vs none yet.)
"""
return [
str(m.get("content") or "")
for m in _messages(body)
if m.get("role") == "tool"
]
def _first_catalog_line(body: dict[str, Any]) -> str | None:
"""The first catalog line of a catalog listing in the messages.
A catalog listing is a ``tool``-role result whose FIRST line is the
agent's ``"N documents:"`` header (``_CATALOG_HEADER_RE``); its
first ``source: X | path: Y | title: Z`` line (the
``_CATALOG_LINE_RE`` machinery) is returned. ``None`` when no
catalog listing is in the messages — e.g. while only the teaching
refusal is there (the phase-72 LS-TEACH flow's request-2 state).
An empty listing (``"0 documents:"`` with no lines) returns
``""`` — the listing is present, it is just empty.
"""
for content in _tool_results(body):
lines = content.splitlines()
if not lines or not _CATALOG_HEADER_RE.match(lines[0]):
continue
for line in lines[1:]:
if _CATALOG_LINE_RE.match(line):
return line
return ""
return None
#: One line of the agent's ``grep`` output (app.rag.agent
#: ``_execute_tool``, phase 68 — phase 70 renamed the tool, the line
#: format is unchanged): ``source/path:LINE: text``. The
@@ -718,6 +802,40 @@ def _tool_flow(body: dict[str, Any]) -> tuple[str, ...] | None:
return ("list", "", "")
def _ls_teach_flow(body: dict[str, Any]) -> tuple[str, ...] | None:
"""Classify a phase-72 LS-TEACH request (see the module docstring).
* ``("misuse",)`` — ``tools`` are offered and no ``tool``-role
result is in the messages yet: the incident's misuse — ``ls``
with ``{"path": "."}`` (id ``call_0``), ``finish_reason:
"tool_calls"``, no content.
* ``("correct",)`` — a ``tool``-role result is in the messages and
it is NOT a catalog listing (the teaching refusal): the
correction — ``ls`` with no arguments (id ``call_1``).
* ``("answer", line)`` — a ``tool``-role result whose first line
is the ``"N documents:"`` catalog header: the deterministic
content answer ``These are the indexed documents: <line>`` (the
first catalog line), ``finish_reason: "stop"`` — the loop
settled in ONE correction, not at the round cap.
* ``None`` — not the flow: the trigger is absent, the ``<tools>``
section is missing (deflected turns never carry it), or
``tools`` are not offered and no tool results are in the
messages yet (e.g. ``agent_max_rounds=0``).
"""
if LS_TEACH_TRIGGER not in _user(body).lower():
return None
if "<tools>" not in _system(body):
return None
line = _first_catalog_line(body)
if line is not None:
return ("answer", line)
if _tool_results(body):
return ("correct",)
if not body.get("tools"):
return None
return ("misuse",)
def long_answer() -> str:
"""~900-word deterministic walkthrough (phase 11): numbered steps plus
a unique final line that must survive the stream untruncated."""
@@ -1208,6 +1326,30 @@ def chat_completions(body: dict[str, Any]) -> Any:
media_type="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
)
# Phase 72 (teaching refusals): the deterministic LS-TEACH
# self-correction flow — checked BEFORE the plain
# TOOLS_TRIGGER flow (disjoint trigger phrases — the phase-71
# ordering convention; the trigger needs the ``<tools>``
# section, so deflected turns never hit it).
ls_teach = _ls_teach_flow(body)
if ls_teach is not None:
if ls_teach[0] == "misuse":
# The incident's misuse, deterministic: ls(path='.').
stream = _tool_call_stream("ls", {"path": "."}, "call_0")
elif ls_teach[0] == "correct":
# The one-round correction: the no-arg full listing.
stream = _tool_call_stream("ls", {}, "call_1")
else: # "answer" — quote the first catalog line
answer = _apply_max_tokens(
f"These are the indexed documents: {ls_teach[1]}",
body.get("max_tokens"),
)
stream = _sse_stream(answer, 0.0)
return StreamingResponse(
stream,
media_type="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
)
flow = _tool_flow(body)
if flow is not None:
if flow[0] == "list":