Phase 72 (72_teaching_refusals) — completed under the 2026-09-04 controlled methodology (owner directive: stop clearing/re-importing the homelab KB per iteration; measure tool-calling accuracy on a controlled fixture KB, target >90%). Real-model gate verdicts (live, configured chat model 'lite', fixture KB): - Controlled fixture battery (the new methodology's pass condition — contract accuracy >= 90%): PASS, 4 consecutive runs: gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 8/11 executed (73%) contract 11/11 (100%) 2026-09-04 (wall 43.4s) gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 8/13 executed (62%) contract 12/13 (92%) 2026-09-04 (wall 50.6s) gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 7/11 executed (64%) contract 11/11 (100%) 2026-09-04 (wall 46.8s) gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 9/15 executed (60%) contract 14/15 (93%) 2026-09-04 (wall 54.8s) - Locked derived battery (phase-72 task 05, executed >= 90% bar, run unchanged on the same fixture KB): gate: lite FAIL turns=10 answered=10 caps=0 tool-turns=10 calls 5/15 executed (33%) contract 12/15 (80%) 2026-09-04 (wall 47.7s) The teaching works — every bare-path trap self-corrects in exactly one round, zero cap hits, zero repeat loops, 10/10 answered. The locked executed bar is blocked by ALREADY_IN_CONTEXT dedupe refusals on the corrected re-reads (the trap question seeds its target, so the correct combined-form read is refused for redundancy) — a copy-invariant model behavior (five copy variants, 0/15 re-reads flipped, 2026-09-03 -> 04) and an app-semantics decision for the owner (TOOL_CALLING_TESTING.md sections 5 and 7), not a copy lever. Copy changes this phase owns (unit pins updated to follow): - app/rag/agent.py: ls teaching refusals (path-like scope -> document-path line; unknown source -> no-source line with the source-name parenthetical), read/grep 'did you mean source/path?' teaching (find_path_candidates: exact or suffix path match, catalog order, cap 3), ALREADY_IN_CONTEXT naming the correct action (answer from the text already in the prompt), read tool description front-loaded with the do-not-read rule (the 2026-09-04 controlled telemetry: the re-read is the only remaining refusal class; contract accuracy 92-100% across runs) - app/rag/prompts.py: TOOLS_SECTION states the document-identity contract up front (ls path = source name; read/grep = combined source/path including the source name; do-not-read for <documents> documents placed next to the read teaching; one-call-per-reply and never-repeat rules) - tests: refusal pins (unit + integration), new dedicated E2E suite tests/e2e/test_tool_path_teaching.py (mock misuse flow, green in isolation), regression suites green in isolation (harness_aligned_tools, agent_document_tools, agent_unlimited_tools, search_tool, chat_rag). Gates: uv run pytest green (1501); coverage TOTAL 99% (>90%); ruff + pyright clean. Carries the still-uncommitted phase-71 todo/ -> complete/ move and both phases' .agent/reports/ (AGENTS.md 8).
1428 lines
66 KiB
Python
1428 lines
66 KiB
Python
"""Deterministic OpenAI-compatible mock for E2E tests (aipi stand-in).
|
|
|
|
Implements just enough of the aipi surface:
|
|
|
|
* ``GET /v1/models``
|
|
* ``POST /v1/embeddings`` — real bag-of-words vectors (768-dim, L2-normed).
|
|
Because similarity is *genuine token overlap*, the relevance threshold
|
|
behaves the same way it will in production: related questions score high,
|
|
unrelated ones score low and trigger honest deflection.
|
|
* ``POST /v1/chat/completions`` — streaming (SSE) or not. The content keys
|
|
off markers in the system prompt:
|
|
- user message containing ``write a long answer`` -> a ~900-word
|
|
deterministic numbered answer (long-answers story, phase 11)
|
|
- ``SUMMARY_MODE`` -> the deterministic summary digest: the first 24
|
|
tokens of the user message (the summarizer puts the capped document
|
|
content there) — byte-stable for a given fixture (document summaries,
|
|
phase 30)
|
|
- ``KB_OVERVIEW_MODE`` -> the deterministic outline: the first 8 tokens
|
|
of the user message (the generator puts the document list there) —
|
|
byte-stable for a given KB (KB overview, phase 31)
|
|
- ``DEFLECT_MODE`` -> honest "I haven't done anything like that" answer
|
|
- otherwise -> upbeat answer quoting the provided document context
|
|
- user message containing ``pretend to think slowly`` -> 3s warm-up delay
|
|
(used by the loading-feedback story).
|
|
- user message containing ``think out loud`` -> the answer is preceded by
|
|
~2 700 chars of deterministic ``reasoning_content`` chunks (the
|
|
thinking-display story, phase 17; lengthened in phase 21 so the
|
|
rendered scratchpad overflows the 320px ``.thinking-text`` window)
|
|
- user message containing ``think out loud then hesitate`` -> the
|
|
``think out loud`` stream, then a 4s pause before the first content
|
|
frame (the sources-midstream story, phase 20 — a deterministic
|
|
"leave during pure thinking" navigation window).
|
|
- user message containing ``think in paragraphs`` (``THINK_PARAS_TRIGGER``)
|
|
-> the ``think out loud`` scratchpad WITH REAL paragraph breaks
|
|
("\n\n"), streamed at 60-char frames (vs the mock's 12-char default).
|
|
One frame renders several lines — a real-model-sized delta, the
|
|
condition under which a POST-render pin-state reading (the old
|
|
app.js) measured the chunk's height instead of the user's position
|
|
and the think-window follow died at the first 2-newline gap. The
|
|
regression pin for the pre-render capture in app.js (2026-08-29,
|
|
owner report). Checked BEFORE ``think out loud`` (it is the more
|
|
specific phrase); existing E2E questions carry neither, so every
|
|
other suite is unaffected.
|
|
- system prompt containing ``<tuning>`` (phase 15, steering notes) ->
|
|
the composed answer ends with `` (tuning: <first note line>)`` —
|
|
makes prompt injection observable in the UI deterministically.
|
|
- system prompt containing ``<knowledge_base>`` (phase 31, KB overview)
|
|
-> the composed answer ends with `` (kb: <first bullet line>)`` —
|
|
the same echo convention for the overview's prompt injection.
|
|
- user message containing ``show the end of your notes`` (phase 24,
|
|
whole-document context) -> the answer quotes the **last 160 chars of
|
|
the ``<documents>`` block** — a tail echo, byte-stable across runs, so
|
|
a sentinel placed at the *end* of a document appears in the rendered
|
|
answer iff the whole document was in the prompt. (Phase 37: the HIGH
|
|
prompt now ends with a ``<tools>`` section after ``</documents>``, so
|
|
the echo targets the block itself; its tail still includes the
|
|
closing tag — same sentinel semantics.)
|
|
- user message containing ``use your tools`` (phase 37, agent document
|
|
tools; phase 70: the flow emits the harness-aligned names — ``ls``
|
|
/ ``read`` with the combined ``source/path`` identity) **and** the
|
|
system prompt carries the ``<tools>`` section -> the deterministic
|
|
SINGLE-READ tool flow, discriminated statelessly from the messages
|
|
(the ``tools`` parameter gates the list/read steps — a no-tools
|
|
request with no tool results is not the flow):
|
|
* request 1 (``tools`` offered, no tool results yet): stream ONLY
|
|
``tool_calls`` deltas — ``ls`` (synthetic id ``call_0``, no
|
|
arguments), ``finish_reason: "tool_calls"``, no content;
|
|
* request 2 (a ``tool``-role catalog result in the messages):
|
|
parse the FIRST catalog line (``source: X | path: Y | title: Z``
|
|
— the labeled ``source:`` / ``path:`` fields, phase 63) and
|
|
stream a ``tool_calls`` delta calling ``read`` on the JOINED
|
|
combined ``source/path`` (the mock joins the two labeled fields
|
|
— the catalog format is unchanged, so this join is the only
|
|
parse change, phase 70) (id ``call_1``);
|
|
* request 3 (a ``tool``-role read result in the messages —
|
|
content starting with the agent's ``"Document <source/path>:"``
|
|
header): a content answer, deterministic: ``Read
|
|
<source/path>. <first 80 chars of the read document's
|
|
content>`` — so a suite can assert the read document reached
|
|
the model and landed in the answer. Reached regardless of the
|
|
``tools`` parameter (phase 45 keeps the tools offered until the
|
|
round cap).
|
|
The single-read flow stops at ONE read result; the MULTI-READ
|
|
variant below reads two.
|
|
- user message containing BOTH ``use your tools`` AND ``read two
|
|
documents`` (``MULTI_READ_TRIGGER``, phase 45 task 02) **and** the
|
|
system prompt carries the ``<tools>`` section -> the deterministic
|
|
MULTI-READ flow (list → read #1 → read #2 → answer), classified by
|
|
the COUNT of ``tool``-role read results (content starting with the
|
|
agent's ``"Document <source/path>:"`` prefix); phase 70: the same
|
|
flow on the harness-aligned names — ``ls``, then ``read`` on the
|
|
JOINED combined ``source/path`` of each catalog line:
|
|
* 0 read results, no catalog yet: ``ls`` (id ``call_0``);
|
|
* 0 read results, catalog present: ``read`` on the JOINED
|
|
combined ``source/path`` of the FIRST catalog line
|
|
(id ``call_1``);
|
|
* 1 read result: ``read`` on the JOINED combined ``source/path``
|
|
of the SECOND catalog line — the first listing line whose
|
|
``source/path`` differs from the one already read (id
|
|
``call_2``); a one-document catalog degenerates to the
|
|
single-read answer (nothing second to read);
|
|
* 2 read results: the forced answer, byte-stable: the single-read
|
|
shape quoting the FIRST read result, plus the line ``I read
|
|
<sp1> and <sp2>.`` naming both read paths in read order — so a
|
|
suite can assert the model used BOTH documents.
|
|
All other requests (including the marker without a ``<tools>``
|
|
section, or with the tool conversation not yet started and no tools
|
|
offered — e.g. ``agent_max_rounds=0``) behave exactly as today.
|
|
``E2E_REAL_LLM=1`` ignores the mock entirely (the real model does
|
|
what it does).
|
|
- user message containing ``search your documents``
|
|
(``SEARCH_TRIGGER``, phase 68 search tool — renamed to the
|
|
harness-aligned ``grep`` in phase 70, same match/output contract)
|
|
**and** the system prompt carries the ``<tools>`` section -> the
|
|
deterministic SEARCH tool flow, discriminated statelessly from the
|
|
messages (streaming only):
|
|
* request 1 (``tools`` offered, no search result yet): stream
|
|
ONLY ``tool_calls`` deltas — ``grep`` with
|
|
``{"pattern": SEARCH_PATTERN}`` (id ``call_0``);
|
|
* request 2 (a ``tool``-role search result in the messages —
|
|
recognizable by its ``source/path:line: text`` match lines or
|
|
the sentinel in its content): the content answer, deterministic:
|
|
``Found <first matched line's content up to 80 chars>`` — so a
|
|
suite can assert the search result reached the model and landed
|
|
in the answer.
|
|
Checked BEFORE the plain ``use your tools`` flow (it is the more
|
|
specific phrase — same convention as ``think in paragraphs``); no
|
|
existing E2E question or fixture file contains the trigger, so
|
|
every other suite is unaffected.
|
|
- user message containing ``emit raw tool markup``
|
|
(``SCAFFOLD_TRIGGER``, phase 71, tool-scaffolding guardrails — the
|
|
2026-09-03 incident where a deflected round streamed the model's
|
|
raw ``<|tool_call_start|>…<|tool_call_end|>`` markup into the UI)
|
|
**or** ``always emit raw tool markup``
|
|
(``SCAFFOLD_ALWAYS_TRIGGER``, checked FIRST — it contains the
|
|
former phrase) -> the deterministic SCAFFOLDING flow, independent
|
|
of the ``<tools>`` marker (both grounded and deflected turns hit
|
|
it):
|
|
* ``SCAFFOLD_ALWAYS_TRIGGER``: EVERY request (the one bounded
|
|
recovery included) streams ONLY ``delta.content`` chunks
|
|
carrying the incident span ``SCAFFOLD_SPAN`` —
|
|
``<|tool_call_start|>[read(path='search_docs/reese-notes.md')]
|
|
<|tool_call_end|>`` — split across the mock's 12-char chunks
|
|
(the filter's boundary path), ``finish_reason: "stop"``, no
|
|
structured ``tool_calls``, no reasoning — the terminal
|
|
malformed-reply path.
|
|
* ``SCAFFOLD_TRIGGER``: request 1 (no ``CORRECTION_INSTRUCTION``
|
|
in the system prompt) streams the same scaffolding-only span;
|
|
request 2 (the system prompt carries the harness constant — a
|
|
stable substring of ``app.rag.agent.CORRECTION_INSTRUCTION``,
|
|
IMPORTED into this module so the mock can never drift from
|
|
it: the one bounded recovery, ``tools=None`` with the
|
|
correction folded into the single system prompt) streams the
|
|
clean ``SCAFFOLD_RECOVERY_ANSWER`` — the recovery path.
|
|
Checked BEFORE the ``SEARCH_TRIGGER`` / ``TOOLS_TRIGGER`` flows
|
|
(the trigger needs no ``<tools>`` section); no existing E2E
|
|
question or fixture file contains the phrase, so every other
|
|
suite is unaffected.
|
|
- user message containing ``list the files in this directory``
|
|
(``LS_TEACH_TRIGGER``, phase 72, teaching refusals — the
|
|
2026-09-03 incident where the harness-prior ``ls(path='.')``
|
|
misuse met the terse refusal and the model re-reasoned the same
|
|
paragraphs over and over) **and** the system prompt carries the
|
|
``<tools>`` section -> the deterministic LS-TEACHING flow,
|
|
discriminated statelessly from the messages (streaming only):
|
|
* request 1 (``tools`` offered, no ``tool``-role result in the
|
|
messages yet): stream ONLY ``tool_calls`` deltas — ``ls``
|
|
with ``{"path": "."}`` (synthetic id ``call_0``),
|
|
``finish_reason: "tool_calls"``, no content — the incident's
|
|
misuse, deterministic;
|
|
* request 2 (a ``tool``-role result present that is NOT a
|
|
catalog listing — i.e. the teaching refusal): a ``tool_calls``
|
|
delta — ``ls`` with no arguments (id ``call_1``) — the
|
|
correction;
|
|
* request 3 (a ``tool``-role result whose first line matches the
|
|
``^\\d+ documents:`` catalog header): a deterministic content
|
|
answer — ``These are the indexed documents: <first catalog
|
|
line>`` (the ``source: X | path: Y | title: Z`` line, parsed
|
|
with the ``_CATALOG_LINE_RE`` machinery), ``finish_reason:
|
|
"stop"`` — the loop ended in ONE correction, not at the round
|
|
cap.
|
|
Checked BEFORE the plain ``TOOLS_TRIGGER`` flow (the trigger
|
|
phrases are disjoint substrings — the phase-71 ordering
|
|
convention); no existing E2E question or fixture file contains the
|
|
phrase, so every other suite is unaffected.
|
|
- user message containing ``show me a table`` (phase 44, markdown
|
|
tables, TODO.md L6) -> the fixed table answer (``TABLE_ANSWER``):
|
|
a 3-column service table, an ``<img onerror>`` XSS probe line, and
|
|
a deliberately wide 5-column table — byte-stable, so the story E2E
|
|
can assert the rendered ``<table class="md-table">`` shape, the
|
|
escaped XSS line, and the wrapper's horizontal scroll inside the
|
|
46rem column. Checked BEFORE the ``DEFLECT_MODE`` branch (a
|
|
deflection prompt never carries the marker, same reasoning as
|
|
``SUMMARY_MODE``), so a marker question always gets the table
|
|
answer; the E2E asks it against an on-topic fixture (HIGH gate) and
|
|
asserts non-deflection.
|
|
|
|
Failure injection (phase 67, LLM retry, TODO.md L3) — deterministic
|
|
dead-endpoint behavior for the retry E2E suite (``tests/e2e/
|
|
test_llm_retry.py``). The mock is single-conversation per e2e server, so
|
|
the sequences are driven by module-level counters that reset per
|
|
trigger phrase after the success they guard (a second question with
|
|
the same trigger re-drives the sequence from zero):
|
|
- user message containing ``fail then answer`` (``RETRY_TRIGGER``):
|
|
the first ``RETRY_DEAD_ATTEMPTS`` (2) app-level streaming attempts
|
|
respond 500 (JSON body, like a dead proxy) and the third streams
|
|
the normal composed answer — 2 = 1 original attempt + 1 retry under
|
|
the default ``BOR_LLM_RETRIES=3``, so a suite exercises a real
|
|
retry without waiting for the 4-attempt exhaustion. Counted in
|
|
APP-LEVEL attempts, not raw HTTP POSTs: while the endpoint stays
|
|
dead, the openai SDK's default policy (max_retries=2 — the app's
|
|
``LLMClient`` keeps it) re-POSTs a 500'd streaming request twice
|
|
before surfacing the error, so each dead attempt costs exactly 3
|
|
POSTs (``_HTTPS_PER_DEAD_ATTEMPT``).
|
|
- user message containing ``always fail``
|
|
(``ALWAYS_FAIL_TRIGGER``): EVERY streaming chat/completions request
|
|
responds 500 — the retry-budget exhaustion path (the terminal
|
|
error banner in the UI).
|
|
- embeddings request whose input contains ``embed fail once``
|
|
(``EMBED_FAIL_TRIGGER``): the FIRST such request responds 500, the
|
|
next returns the normal bag-of-words vector — the endpoint's
|
|
pre-stream embedding retry loop. Raw httpx on the client side (no
|
|
SDK-level retries), so one POST per attempt: the counter is per
|
|
POST here, unlike the chat counter above.
|
|
Non-streaming requests (document summaries, KB overview) never 500 —
|
|
the retry scope is the chat turn only (owner-locked A1).
|
|
|
|
``max_tokens`` is honored deterministically (token ≈ whitespace word),
|
|
like a real endpoint: an answer longer than the cap is truncated. This
|
|
is what makes the phase-11 truncation regression observable.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
import hashlib
|
|
import math
|
|
import os
|
|
import re
|
|
import signal
|
|
import threading
|
|
import time
|
|
import uuid
|
|
from typing import Any
|
|
|
|
from fastapi import FastAPI
|
|
from fastapi.responses import JSONResponse, StreamingResponse
|
|
|
|
from app.rag.agent import CORRECTION_INSTRUCTION # phase 71: the harness constant
|
|
|
|
app = FastAPI()
|
|
|
|
DIM = 768
|
|
TOKEN_RE = re.compile(r"[a-z0-9]+")
|
|
|
|
|
|
def embed_text(text: str) -> list[float]:
|
|
vec = [0.0] * DIM
|
|
for tok in TOKEN_RE.findall(text.lower()):
|
|
idx = int(hashlib.md5(tok.encode()).hexdigest(), 16) % DIM
|
|
vec[idx] += 1.0
|
|
norm = math.sqrt(sum(v * v for v in vec)) or 1.0
|
|
return [v / norm for v in vec]
|
|
|
|
|
|
def _messages(body: dict[str, Any]) -> list[dict[str, str]]:
|
|
return body.get("messages", [])
|
|
|
|
|
|
def _system(body: dict[str, Any]) -> str:
|
|
return " ".join(m.get("content", "") for m in _messages(body) if m.get("role") == "system")
|
|
|
|
|
|
def _user(body: dict[str, Any]) -> str:
|
|
parts = [m.get("content", "") for m in _messages(body) if m.get("role") == "user"]
|
|
return parts[-1] if parts else ""
|
|
|
|
|
|
def _context(body: dict[str, Any]) -> str:
|
|
"""The document context is the longest system/user message in practice."""
|
|
msgs = _messages(body)
|
|
return max((m.get("content", "") for m in msgs), key=len)
|
|
|
|
|
|
LONG_ANSWER_TRIGGER = "write a long answer"
|
|
#: ~920 words — comfortably past the old hard 700-token cap (where the
|
|
#: tail would be cut) yet short enough to stream in ~8s at the mock's
|
|
#: per-chunk pacing.
|
|
LONG_ANSWER_LINES = 40
|
|
LONG_ANSWER_END = "LONG-ANSWER-END"
|
|
|
|
#: Phase 17 (thinking-display story): a user message containing this
|
|
#: substring (case-insensitive) is answered with a deterministic
|
|
#: ``reasoning_content`` stream ahead of the content — same convention as
|
|
#: the other user-message triggers above. Existing E2E questions do not
|
|
#: contain the substring, so every other suite is unaffected.
|
|
THINKING_TRIGGER = "think out loud"
|
|
|
|
#: Regression pin (2026-08-29, owner report): a user message containing
|
|
#: this substring (case-insensitive) gets the phase-17 scratchpad WITH
|
|
#: REAL paragraph breaks ("\n\n"), streamed at ``THINK_PARAS_CHUNK``
|
|
#: chars/frame — a single frame renders several lines (a real-model-sized
|
|
#: delta), which is the condition under which the old POST-render pin
|
|
#: reading in app.js died at the first 2-newline gap. Existing E2E
|
|
#: questions do not contain the phrase, so every other suite is
|
|
#: unaffected (checked before ``THINKING_TRIGGER`` — the more specific
|
|
#: phrase wins).
|
|
THINK_PARAS_TRIGGER = "think in paragraphs"
|
|
#: 60-char thinking frames for the paragraph trigger (the mock default is
|
|
#: 12 — a 12-char frame renders at most one line ≈ 22px, always inside
|
|
#: the 32px think-window band, which is why the old code passed the
|
|
#: 12-char suites while the real model's larger deltas killed the pin).
|
|
THINK_PARAS_CHUNK = 60
|
|
|
|
#: Phase 20 (sources-midstream bug): a user message containing this
|
|
#: substring (case-insensitive) gets the phase-17 thinking stream followed
|
|
#: by a multi-second pause before the FIRST content frame — the
|
|
#: navigation window for the "leave during pure thinking" scenario
|
|
#: (owner-confirmed A1.2: nothing brain-side may be persisted then).
|
|
#: Strictly longer than ``THINKING_TRIGGER``, so the phase-17 suite's
|
|
#: questions are unaffected.
|
|
SLOW_PRETOKEN_TRIGGER = "think out loud then hesitate"
|
|
PRE_CONTENT_PAUSE_S = 4.0
|
|
|
|
#: Phase 24 (whole-document-context story): a user message containing this
|
|
#: substring (case-insensitive) gets an answer quoting the TAIL of the
|
|
#: document context (see the module docstring). Verified 2026-08-24: no
|
|
#: existing E2E question or fixture file contains the phrase, so every
|
|
#: other suite is unaffected.
|
|
END_OF_NOTES_TRIGGER = "show the end of your notes"
|
|
|
|
#: The ``<documents>`` block of the system prompt (phase 37: the HIGH
|
|
#: prompt ends with the ``<tools>`` section after ``</documents>``, so the
|
|
#: phase-24 tail echo targets the block, not the raw message tail).
|
|
_DOCUMENTS_BLOCK_RE = re.compile(r"<documents>.*?</documents>", re.S)
|
|
|
|
#: Phase 37 (agent-document-tools story; phase 70: the flow emits the
|
|
#: harness-aligned names): a user message containing this substring
|
|
#: (case-insensitive) — combined with the ``<tools>`` section in the
|
|
#: system prompt — drives the deterministic tool flow documented in the
|
|
#: module docstring (ls → read on the first catalog line's combined
|
|
#: ``source/path`` → the quoted answer). Existing E2E questions do not
|
|
#: contain the phrase, so every other suite is unaffected.
|
|
TOOLS_TRIGGER = "use your tools"
|
|
|
|
#: Phase 45 (agent-unlimited-tools story, task 02): a user message
|
|
#: containing BOTH ``TOOLS_TRIGGER`` and this substring (case-insensitive
|
|
#: — the check lowercases the user message) drives the deterministic
|
|
#: MULTI-READ tool flow (list → read #1 → read #2 → the forced answer
|
|
#: naming both read paths) — see the module docstring. The existing
|
|
#: phase-37 E2E question carries ``TOOLS_TRIGGER`` but not this phrase,
|
|
#: so the 3-step flow is untouched.
|
|
MULTI_READ_TRIGGER = "read two documents"
|
|
|
|
#: Phase 68 (search tool, TODO.md L4; phase 70: renamed to the
|
|
#: harness-aligned ``grep``): a user message containing this substring
|
|
#: (case-insensitive) — combined with the ``<tools>`` section in the
|
|
#: system prompt — drives the deterministic SEARCH tool flow (grep for
|
|
#: ``SEARCH_PATTERN`` → the "Found …" answer), documented in the module
|
|
#: docstring. Checked BEFORE ``TOOLS_TRIGGER``
|
|
#: (the more specific phrase wins — the same convention as
|
|
#: ``THINK_PARAS_TRIGGER``); verified 2026-09-01: no existing E2E
|
|
#: question or fixture file contains the phrase, so every other suite
|
|
#: is unaffected.
|
|
SEARCH_TRIGGER = "search your documents"
|
|
|
|
#: The sentinel the search flow greps for: the e2e fixture document
|
|
#: (``tests/fixtures/search_docs/reese-notes.md``) carries exactly one
|
|
#: line containing it, so the search result — and the "Found …" answer
|
|
#: that quotes its first matched line — is byte-stable (the sentinel
|
|
#: convention of ``END_OF_NOTES_TRIGGER``).
|
|
SEARCH_PATTERN = "reese-sentinel-42"
|
|
|
|
#: Phase 44 (markdown-tables story, TODO.md L6): a user message
|
|
#: containing this substring (case-insensitive) gets the fixed table
|
|
#: answer (``TABLE_ANSWER`` below) — a 3-column table, an XSS probe
|
|
#: line, and a deliberately wide table (see the module docstring).
|
|
#: Existing E2E questions do not contain the phrase, so every other
|
|
#: suite is unaffected.
|
|
TABLE_TRIGGER = "show me a table"
|
|
|
|
#: The fixed table answer (phase 44) — byte-stable on purpose: the story
|
|
#: E2E asserts the rendered table shape, the escaped ``<img onerror>``
|
|
#: line (the XSS payload must survive the mock byte-for-byte), and the
|
|
#: wide table's ``scrollWidth > clientWidth`` inside the 46rem column.
|
|
TABLE_ANSWER = (
|
|
"Here's the shape, in a table:\n"
|
|
"\n"
|
|
"| Service | Port | Host |\n"
|
|
"|---|---|---|\n"
|
|
"| Caddy | 80 | homelab-gw |\n"
|
|
"| GitLab | 8929 | homelab-git |\n"
|
|
"| ntfy | 2087 | homelab-ntfy |\n"
|
|
"\n"
|
|
"<img src=x onerror=alert(1)>\n"
|
|
"\n"
|
|
"And the wide one:\n"
|
|
"\n"
|
|
"| A very long column header to force overflow | Second column with "
|
|
"some padding text | Third column | Fourth | Fifth |\n"
|
|
"|---|---|---|---|---|\n"
|
|
"| value-one | value-two | value-three | value-four | value-five |"
|
|
)
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Phase 67 (LLM retry, TODO.md L3): deterministic failure injection
|
|
# ---------------------------------------------------------------------------
|
|
|
|
#: A user message containing this substring (case-insensitive) gets
|
|
#: ``RETRY_DEAD_ATTEMPTS`` dead streaming attempts (500, JSON body) before
|
|
#: the normal composed answer streams — 1 original attempt + 1 retry under
|
|
#: the default ``BOR_LLM_RETRIES=3`` (see the module docstring).
|
|
RETRY_TRIGGER = "fail then answer"
|
|
#: App-level attempts the endpoint stays dead for before the answer.
|
|
RETRY_DEAD_ATTEMPTS = 2
|
|
|
|
#: A user message containing this substring (case-insensitive) makes
|
|
#: EVERY streaming chat/completions request respond 500 — the
|
|
#: retry-budget exhaustion path (the terminal error banner in the UI).
|
|
ALWAYS_FAIL_TRIGGER = "always fail"
|
|
|
|
#: An embeddings request whose input contains this substring
|
|
#: (case-insensitive) 500s on its FIRST POST; the next returns the normal
|
|
#: bag-of-words vector — the endpoint's pre-stream embedding retry loop.
|
|
EMBED_FAIL_TRIGGER = "embed fail once"
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Phase 71 (tool-scaffolding guardrails, 2026-09-03 incident):
|
|
# deterministic raw-markup flows — see the module docstring
|
|
# ---------------------------------------------------------------------------
|
|
|
|
#: A user message containing this substring (case-insensitive) drives
|
|
#: the scaffolding flow: request 1 streams ONLY the incident's raw tool
|
|
#: markup as ``delta.content``; the follow-up request carrying the
|
|
#: harness correction in the system prompt (the one bounded recovery)
|
|
#: streams the clean answer. Independent of the ``<tools>`` marker —
|
|
#: both grounded and deflected turns hit it. Existing E2E questions do
|
|
#: not contain the phrase, so every other suite is unaffected.
|
|
SCAFFOLD_TRIGGER = "emit raw tool markup"
|
|
|
|
#: A user message containing this substring (checked BEFORE
|
|
#: ``SCAFFOLD_TRIGGER`` — it contains that phrase) streams the
|
|
#: scaffolding-only span on EVERY request, recovery included — the
|
|
#: terminal malformed-reply path (the dedicated error frame, no done).
|
|
SCAFFOLD_ALWAYS_TRIGGER = "always emit raw tool markup"
|
|
|
|
#: The incident span (2026-09-03): the model's chat-template tool
|
|
#: syntax, emitted as plain ``delta.content`` although no tools were
|
|
#: offered. Streamed through the mock's 12-char chunking, so it always
|
|
#: spans ≥2 wire chunks (the filter's boundary path).
|
|
SCAFFOLD_SPAN = (
|
|
"<|tool_call_start|>[read(path='search_docs/reese-notes.md')]"
|
|
"<|tool_call_end|>"
|
|
)
|
|
|
|
#: The clean answer the one bounded recovery produces (byte-stable —
|
|
#: the dedicated E2E suite asserts the recovered bubble and the wire's
|
|
#: delta text against it).
|
|
SCAFFOLD_RECOVERY_ANSWER = "Here is the plain-text answer the recovery produced."
|
|
|
|
#: The stable substring of the harness-owned correction constant the
|
|
#: recovery request carries in its system prompt. Keyed on a substring
|
|
#: (not the whole constant) so a re-wrap of the constant cannot silently
|
|
#: re-route the mock; the module-level assert below fails loudly if the
|
|
#: substring ever leaves the constant (the mock must never drift from
|
|
#: ``app.rag.agent.CORRECTION_INSTRUCTION``).
|
|
_CORRECTION_MARKER = "no tool syntax"
|
|
assert _CORRECTION_MARKER in CORRECTION_INSTRUCTION, (
|
|
"mock drift: the correction marker left CORRECTION_INSTRUCTION"
|
|
)
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Phase 72 (teaching refusals — the 2026-09-03 incident's ls misuse):
|
|
# the deterministic LS-TEACH self-correction flow — see the module
|
|
# docstring
|
|
# ---------------------------------------------------------------------------
|
|
|
|
#: A user message containing this substring (case-insensitive) —
|
|
#: combined with the ``<tools>`` section in the system prompt — drives
|
|
#: the deterministic LS-TEACHING flow (the incident's
|
|
#: ``ls(path='.')`` misuse → the teaching refusal → the corrected
|
|
#: no-arg ``ls()`` → the catalog answer). Checked BEFORE the plain
|
|
#: ``TOOLS_TRIGGER`` flow (disjoint trigger phrases — the phase-71
|
|
#: ordering convention); verified: no existing E2E question or fixture
|
|
#: file contains the phrase, so every other suite is unaffected.
|
|
LS_TEACH_TRIGGER = "list the files in this directory"
|
|
|
|
#: The agent's ``ls`` listing header (app.rag.agent ``_execute_tool``):
|
|
#: ``"N documents:"`` — the first line of every catalog tool result.
|
|
_CATALOG_HEADER_RE = re.compile(r"^\d+ documents:")
|
|
|
|
#: One DEAD app-level chat attempt costs exactly this many HTTP POSTs
|
|
#: while the endpoint stays down: the openai SDK's default policy
|
|
#: (max_retries=2 — the app's ``LLMClient`` keeps it) re-POSTs a 500'd
|
|
#: streaming request twice before surfacing the error to
|
|
#: ``chat_stream_retried``. The failure counters below therefore count
|
|
#: app-level attempts (groups of this size), not raw POSTs — the visible
|
|
#: sequence (one SSE ``retry`` frame after each dead attempt, the answer
|
|
#: on the third) stays deterministic regardless of the SDK's internal
|
|
#: backoff pacing.
|
|
_HTTPS_PER_DEAD_ATTEMPT = 3
|
|
|
|
#: Module-level failure counters — the mock is single-conversation per
|
|
#: e2e server. Keyed by trigger phrase (reset per trigger): the number
|
|
#: of matching POSTs served so far. Each sequence resets after the
|
|
#: success it guards, so a second question carrying the same trigger
|
|
#: re-drives the failure sequence from zero.
|
|
_fail_posts: dict[str, int] = {}
|
|
|
|
|
|
def _llm_500(why: str) -> JSONResponse:
|
|
"""A dead-proxy 500 with a JSON error body (phase 67 injection)."""
|
|
return JSONResponse(
|
|
status_code=500,
|
|
content={
|
|
"error": {
|
|
"message": f"upstream connection reset ({why})",
|
|
"type": "proxy_error",
|
|
}
|
|
},
|
|
)
|
|
|
|
|
|
def _bump_fail(key: str) -> int:
|
|
n = _fail_posts.get(key, 0) + 1
|
|
_fail_posts[key] = n
|
|
return n
|
|
|
|
|
|
def _chat_dead(key: str, dead_attempts: int) -> bool:
|
|
"""Bump *key*'s counter; True while the endpoint stays dead.
|
|
|
|
Counted in app-level attempts (see ``_HTTPS_PER_DEAD_ATTEMPT``): the
|
|
first ``dead_attempts * _HTTPS_PER_DEAD_ATTEMPT`` POSTs 500 and the
|
|
next attempt's first POST streams (the caller resets the counter on
|
|
the success).
|
|
"""
|
|
return _bump_fail(key) <= dead_attempts * _HTTPS_PER_DEAD_ATTEMPT
|
|
|
|
|
|
#: The agent's ``read`` tool-result prefix (app.rag.agent
|
|
#: ``_execute_tool``): ``"Document <source/path>:\n<content>"``.
|
|
_READ_RESULT_PREFIX = "Document "
|
|
|
|
#: One line of the agent's ``ls`` output (app.rag.agent
|
|
#: ``_execute_tool``, phase 63): labeled, pipe-delimited fields —
|
|
#: ``source: X | path: Y | title: Z`` — unambiguous for LLM parsing even
|
|
#: when the path contains ``/`` characters.
|
|
_CATALOG_LINE_RE = re.compile(
|
|
r"^source: (?P<source>.+?) \| path: (?P<path>.+?) \| title: .+$"
|
|
)
|
|
|
|
|
|
def _read_results(body: dict[str, Any]) -> list[tuple[str, str]]:
|
|
"""The read results in the messages, in order: ``(source/path, content)``.
|
|
|
|
A read result is a ``tool``-role message whose content starts with
|
|
the agent's read-result prefix (``app.rag.agent`` ``_execute_tool``):
|
|
``"Document <source/path>:\n<content>"``. The header is stripped of
|
|
the prefix AND the trailing colon so the path stays clean.
|
|
"""
|
|
out: list[tuple[str, str]] = []
|
|
for m in _messages(body):
|
|
if m.get("role") != "tool":
|
|
continue
|
|
content = str(m.get("content") or "")
|
|
if content.startswith(_READ_RESULT_PREFIX):
|
|
header, _, doc_content = content.partition("\n")
|
|
sp = header[len(_READ_RESULT_PREFIX):].strip().removesuffix(":")
|
|
out.append((sp, doc_content))
|
|
return out
|
|
|
|
|
|
def _catalog_docs(body: dict[str, Any]) -> list[tuple[str, str]]:
|
|
"""Every ``(source, path)`` in the catalog tool result, in listing order.
|
|
|
|
Catalog lines are ``source: X | path: Y | title: Z`` (the agent's
|
|
``ls`` output — phase 63: labeled, pipe-delimited
|
|
fields, unambiguous even for paths full of ``/``): the line-level
|
|
regex recovers the ``source`` and ``path`` fields directly. The
|
|
``"N documents:"`` header line matches no line and is skipped;
|
|
read-result messages are full documents, not listings, and are
|
|
skipped too.
|
|
"""
|
|
docs: list[tuple[str, str]] = []
|
|
for m in _messages(body):
|
|
if m.get("role") != "tool":
|
|
continue
|
|
content = str(m.get("content") or "")
|
|
if content.startswith(_READ_RESULT_PREFIX):
|
|
continue
|
|
for line in content.splitlines():
|
|
match = _CATALOG_LINE_RE.match(line)
|
|
if match:
|
|
docs.append((match.group("source"), match.group("path")))
|
|
return docs
|
|
|
|
|
|
def _tool_results(body: dict[str, Any]) -> list[str]:
|
|
"""Every ``tool``-role result content in the messages, in order.
|
|
|
|
(Phase 72, LS-TEACH flow: the flow is discriminated statelessly
|
|
from the tool results — a catalog listing vs the teaching
|
|
refusal vs none yet.)
|
|
"""
|
|
return [
|
|
str(m.get("content") or "")
|
|
for m in _messages(body)
|
|
if m.get("role") == "tool"
|
|
]
|
|
|
|
|
|
def _first_catalog_line(body: dict[str, Any]) -> str | None:
|
|
"""The first catalog line of a catalog listing in the messages.
|
|
|
|
A catalog listing is a ``tool``-role result whose FIRST line is the
|
|
agent's ``"N documents:"`` header (``_CATALOG_HEADER_RE``); its
|
|
first ``source: X | path: Y | title: Z`` line (the
|
|
``_CATALOG_LINE_RE`` machinery) is returned. ``None`` when no
|
|
catalog listing is in the messages — e.g. while only the teaching
|
|
refusal is there (the phase-72 LS-TEACH flow's request-2 state).
|
|
An empty listing (``"0 documents:"`` with no lines) returns
|
|
``""`` — the listing is present, it is just empty.
|
|
"""
|
|
for content in _tool_results(body):
|
|
lines = content.splitlines()
|
|
if not lines or not _CATALOG_HEADER_RE.match(lines[0]):
|
|
continue
|
|
for line in lines[1:]:
|
|
if _CATALOG_LINE_RE.match(line):
|
|
return line
|
|
return ""
|
|
return None
|
|
|
|
|
|
#: One line of the agent's ``grep`` output (app.rag.agent
|
|
#: ``_execute_tool``, phase 68 — phase 70 renamed the tool, the line
|
|
#: format is unchanged): ``source/path:LINE: text``. The
|
|
#: non-greedy prefix keeps nested paths (``/`` in the path) intact.
|
|
_SEARCH_LINE_RE = re.compile(r"^(?P<sp>.+?):(?P<line>\d+): (?P<text>.*)$")
|
|
|
|
|
|
def _search_result_line(body: dict[str, Any]) -> str | None:
|
|
"""The first matched line's text of a search result in the messages.
|
|
|
|
A search result is a ``tool``-role message — never a read result
|
|
(those start with the agent's ``"Document "`` prefix) — that either
|
|
carries ``source/path:LINE: text`` match lines (the agent's
|
|
``grep`` output, phase 68) or the sentinel pattern
|
|
itself (its no-match line quotes the pattern). Returns the first
|
|
match line's ``text`` part (already 200-char-capped server-side),
|
|
or the message's first line in the sentinel-only shape, or ``None``
|
|
when no search result is in the messages yet.
|
|
"""
|
|
sentinel = SEARCH_PATTERN.lower()
|
|
for m in _messages(body):
|
|
if m.get("role") != "tool":
|
|
continue
|
|
content = str(m.get("content") or "")
|
|
if content.startswith(_READ_RESULT_PREFIX):
|
|
continue
|
|
for line in content.splitlines():
|
|
match = _SEARCH_LINE_RE.match(line)
|
|
if match:
|
|
return match.group("text")
|
|
if sentinel in content.lower():
|
|
lines = content.splitlines()
|
|
return lines[0] if lines else ""
|
|
return None
|
|
|
|
|
|
def _search_flow(body: dict[str, Any]) -> tuple[str, ...] | None:
|
|
"""Classify a SEARCH_TRIGGER request into a step of the search flow.
|
|
|
|
* ``("search",)`` — ``tools`` are offered and no search result is
|
|
in the messages yet: the model greps the whole KB for
|
|
``SEARCH_PATTERN`` (id ``call_0``).
|
|
* ``("found", first_line)`` — a ``tool``-role search result is in
|
|
the messages: the model answers, quoting the first matched line
|
|
(``Found <first matched line's content up to 80 chars>``). Reached
|
|
regardless of the ``tools`` parameter (phase 45 keeps the tools
|
|
offered until the round cap).
|
|
* ``None`` — not the search flow: the trigger is absent, the
|
|
``<tools>`` section is missing (deflected turns never carry it),
|
|
or ``tools`` are not offered and no search result is in the
|
|
messages yet (e.g. ``agent_max_rounds=0``).
|
|
"""
|
|
if SEARCH_TRIGGER not in _user(body).lower():
|
|
return None
|
|
if "<tools>" not in _system(body):
|
|
return None
|
|
first_line = _search_result_line(body)
|
|
if first_line is not None:
|
|
return ("found", first_line)
|
|
if not body.get("tools"):
|
|
return None
|
|
return ("search",)
|
|
|
|
|
|
def _scaffold_flow(body: dict[str, Any]) -> str | None:
|
|
"""Classify a phase-71 scaffolding request (see the module docstring).
|
|
|
|
* ``"scaffold"`` — stream ONLY the incident span
|
|
(``SCAFFOLD_SPAN``) as ``delta.content`` chunks: ``finish_reason:
|
|
"stop"``, no structured ``tool_calls``, no reasoning. EVERY
|
|
request for ``SCAFFOLD_ALWAYS_TRIGGER`` (the recovery included),
|
|
and the FIRST request of ``SCAFFOLD_TRIGGER`` (no correction in
|
|
the system prompt yet).
|
|
* ``"recovery"`` — ``SCAFFOLD_TRIGGER`` whose system prompt carries
|
|
the harness correction (the one bounded recovery: ``tools=None``,
|
|
the constant folded into the single system prompt by
|
|
``app.api.chat`` / ``app.rag.agent``): stream the clean
|
|
``SCAFFOLD_RECOVERY_ANSWER``.
|
|
* ``None`` — not the scaffolding flow. The discrimination is
|
|
stateless, like the other marker flows: the trigger phrase in
|
|
the user message plus the correction's presence in the system
|
|
prompt.
|
|
"""
|
|
user = _user(body).lower()
|
|
if SCAFFOLD_ALWAYS_TRIGGER in user: # checked FIRST — it contains SCAFFOLD_TRIGGER
|
|
return "scaffold"
|
|
if SCAFFOLD_TRIGGER in user:
|
|
if _CORRECTION_MARKER in _system(body):
|
|
return "recovery"
|
|
return "scaffold"
|
|
return None
|
|
|
|
|
|
def _tool_flow(body: dict[str, Any]) -> tuple[str, ...] | None:
|
|
"""Classify a marker request into one step of the tool flow.
|
|
|
|
Single-read (phase 37 — the user message carries ``TOOLS_TRIGGER``
|
|
only):
|
|
|
|
* ``("list", "", "")`` — ``tools`` are offered and no tool results
|
|
are in the messages yet: the model lists the catalog.
|
|
* ``("read", source, path, "call_1")`` — a ``tool``-role catalog
|
|
result is in the messages: the model reads its FIRST
|
|
``source: X | path: Y | title: Z`` line (the labeled
|
|
``source:`` / ``path:`` fields, phase 63), emitted as ``read`` on
|
|
the JOINED combined ``source/path`` (phase 70: the mock joins
|
|
the two fields — the canonical document identity).
|
|
* ``("answer", "source/path", content)`` — a ``tool``-role read
|
|
result (``"Document <source/path>:\n<content>"``) is in the
|
|
messages: the model answers, quoting the read document. Reached
|
|
regardless of the ``tools`` parameter (phase 45 keeps the tools
|
|
offered until the round cap).
|
|
|
|
Multi-read (phase 45 task 02 — the user message carries BOTH
|
|
``TOOLS_TRIGGER`` and ``MULTI_READ_TRIGGER``), classified by the
|
|
count of ``tool``-role read results:
|
|
|
|
* 0 read results: ``("list", "", "")`` (no catalog yet) or
|
|
``("read", source, path, "call_1")`` on the FIRST catalog doc.
|
|
* 1 read result: ``("read", source, path, "call_2")`` on the SECOND
|
|
catalog doc — the first listing line whose ``source/path``
|
|
differs from the one already read. A one-document catalog
|
|
degenerates to the single-read ``("answer", ...)`` shape (nothing
|
|
second to read).
|
|
* 2 read results: ``("multi_answer", "", text)`` — the forced
|
|
answer, byte-stable: the single-read shape quoting the FIRST read
|
|
result, plus ``I read <sp1> and <sp2>.`` (both read paths, read
|
|
order). The second element is unused.
|
|
|
|
* ``None`` — not the marker flow: the request behaves exactly as
|
|
before (marker absent, no ``<tools>`` section, or a no-tools
|
|
request with no tool results — e.g. ``agent_max_rounds=0``).
|
|
"""
|
|
user = _user(body).lower()
|
|
if TOOLS_TRIGGER not in user:
|
|
return None
|
|
if "<tools>" not in _system(body):
|
|
return None
|
|
reads = _read_results(body)
|
|
if MULTI_READ_TRIGGER in user:
|
|
if not reads:
|
|
if not body.get("tools"):
|
|
return None
|
|
docs = _catalog_docs(body)
|
|
if not docs:
|
|
return ("list", "", "")
|
|
return ("read", docs[0][0], docs[0][1], "call_1")
|
|
if len(reads) == 1:
|
|
skip = reads[0][0]
|
|
second = next(
|
|
(d for d in _catalog_docs(body) if f"{d[0]}/{d[1]}" != skip), None
|
|
)
|
|
if second is None:
|
|
# One-document catalog: nothing second to read — the
|
|
# single-read answer shape (deterministic degenerate).
|
|
return ("answer", reads[0][0], reads[0][1])
|
|
return ("read", second[0], second[1], "call_2")
|
|
(sp1, c1), (sp2, _c2) = reads[0], reads[1]
|
|
answer = f"Read {sp1}. {c1[:80]} I read {sp1} and {sp2}."
|
|
return ("multi_answer", "", answer)
|
|
# Phase-37 single-read flow — byte-identical to the original.
|
|
if reads:
|
|
return ("answer", reads[0][0], reads[0][1])
|
|
if not body.get("tools"):
|
|
return None
|
|
docs = _catalog_docs(body)
|
|
if docs:
|
|
return ("read", docs[0][0], docs[0][1], "call_1")
|
|
return ("list", "", "")
|
|
|
|
|
|
def _ls_teach_flow(body: dict[str, Any]) -> tuple[str, ...] | None:
|
|
"""Classify a phase-72 LS-TEACH request (see the module docstring).
|
|
|
|
* ``("misuse",)`` — ``tools`` are offered and no ``tool``-role
|
|
result is in the messages yet: the incident's misuse — ``ls``
|
|
with ``{"path": "."}`` (id ``call_0``), ``finish_reason:
|
|
"tool_calls"``, no content.
|
|
* ``("correct",)`` — a ``tool``-role result is in the messages and
|
|
it is NOT a catalog listing (the teaching refusal): the
|
|
correction — ``ls`` with no arguments (id ``call_1``).
|
|
* ``("answer", line)`` — a ``tool``-role result whose first line
|
|
is the ``"N documents:"`` catalog header: the deterministic
|
|
content answer ``These are the indexed documents: <line>`` (the
|
|
first catalog line), ``finish_reason: "stop"`` — the loop
|
|
settled in ONE correction, not at the round cap.
|
|
* ``None`` — not the flow: the trigger is absent, the ``<tools>``
|
|
section is missing (deflected turns never carry it), or
|
|
``tools`` are not offered and no tool results are in the
|
|
messages yet (e.g. ``agent_max_rounds=0``).
|
|
"""
|
|
if LS_TEACH_TRIGGER not in _user(body).lower():
|
|
return None
|
|
if "<tools>" not in _system(body):
|
|
return None
|
|
line = _first_catalog_line(body)
|
|
if line is not None:
|
|
return ("answer", line)
|
|
if _tool_results(body):
|
|
return ("correct",)
|
|
if not body.get("tools"):
|
|
return None
|
|
return ("misuse",)
|
|
|
|
|
|
def long_answer() -> str:
|
|
"""~900-word deterministic walkthrough (phase 11): numbered steps plus
|
|
a unique final line that must survive the stream untruncated."""
|
|
lines = [
|
|
f"{i}. Step {i}: configure node-{i} with the homelab defaults and "
|
|
f"verify that step {i} of the long walkthrough is complete before moving on."
|
|
for i in range(1, LONG_ANSWER_LINES + 1)
|
|
]
|
|
lines.append(LONG_ANSWER_END)
|
|
return "\n".join(lines)
|
|
|
|
|
|
#: First numbered note line of a ``<tuning>`` section (phase 15).
|
|
_TUNING_BLOCK_RE = re.compile(r"<tuning>\n(.*?)\n</tuning>", re.S)
|
|
_NOTE_LINE_RE = re.compile(r"^\d+\.\s*(.+)$")
|
|
|
|
|
|
def first_tuning_note(system: str) -> str | None:
|
|
"""The first steering note in the system prompt, or ``None``.
|
|
|
|
The prompt numbers notes 1..N oldest-first (see
|
|
``app.rag.prompts.build_steering_section``); the mock echoes the first
|
|
one into its answer so prompt injection is observable in the UI.
|
|
"""
|
|
block = _TUNING_BLOCK_RE.search(system)
|
|
if not block:
|
|
return None
|
|
for line in block.group(1).splitlines():
|
|
m = _NOTE_LINE_RE.match(line.strip())
|
|
if m:
|
|
return m.group(1).strip()
|
|
return None
|
|
|
|
|
|
#: First ``-`` bullet line of a ``<knowledge_base>`` section (phase 31).
|
|
_KB_BLOCK_RE = re.compile(r"<knowledge_base>\n(.*?)\n</knowledge_base>", re.S)
|
|
_KB_BULLET_RE = re.compile(r"^-(?:\s+(.*))?$")
|
|
|
|
|
|
def first_kb_bullet(system: str) -> str | None:
|
|
"""The first outline bullet in the system prompt, or ``None``.
|
|
|
|
The stored outline (phase 31) is ``-`` bullet lines (see
|
|
``app.rag.overview.OVERVIEW_INSTRUCTION``); the mock echoes the first
|
|
one into its answer as `` (kb: <bullet>)`` — the exact
|
|
:func:`first_tuning_note` convention, so prompt injection of the
|
|
``<knowledge_base>`` section is observable in the UI.
|
|
"""
|
|
block = _KB_BLOCK_RE.search(system)
|
|
if not block:
|
|
return None
|
|
for line in block.group(1).splitlines():
|
|
m = _KB_BULLET_RE.match(line.strip())
|
|
if m:
|
|
return (m.group(1) or "").strip()
|
|
return None
|
|
|
|
|
|
def compose_answer(body: dict[str, Any]) -> str:
|
|
system = _system(body)
|
|
user = _user(body)
|
|
if LONG_ANSWER_TRIGGER in user.lower():
|
|
answer = long_answer()
|
|
elif "SUMMARY_MODE" in system:
|
|
# Document summaries (phase 30): the ``lite`` stand-in returns a
|
|
# deterministic digest — the first 24 tokens of the user message
|
|
# (the summarizer puts the capped document content there). Byte-
|
|
# stable for a given fixture, so the summary chunk's retrieval
|
|
# rank is a pure function of the fixture text. Checked BEFORE the
|
|
# DEFLECT_MODE branch (task 06) so a deflection prompt that ever
|
|
# carries the marker cannot shadow the summary call.
|
|
answer = (
|
|
f"This document covers "
|
|
f"{' '.join(TOKEN_RE.findall(user.lower())[:24])}."
|
|
)
|
|
elif "KB_OVERVIEW_MODE" in system:
|
|
# KB overview (phase 31): the ``lite`` stand-in returns the
|
|
# deterministic outline — the first 8 tokens of the user message
|
|
# (the generator puts the document list there). Byte-stable for a
|
|
# given KB, so the stored row is a pure function of the fixture.
|
|
# Checked BEFORE the DEFLECT_MODE branch, like SUMMARY_MODE, so a
|
|
# prompt that ever carries both markers cannot shadow the
|
|
# overview call.
|
|
answer = "Knowledge base outline:\n- " + " ".join(
|
|
TOKEN_RE.findall(user.lower())[:8]
|
|
)
|
|
elif TABLE_TRIGGER in user.lower():
|
|
# Markdown tables (phase 44, TODO.md L6): the story E2E's
|
|
# deterministic table answer — a 3-column table, the
|
|
# <img onerror> XSS probe line (it must survive the mock
|
|
# byte-for-byte so the E2E can prove the renderer neutralizes
|
|
# it), and a wide 5-column table (guarantees scrollWidth >
|
|
# clientWidth inside the 46rem column). Byte-stable. Checked
|
|
# BEFORE the DEFLECT_MODE branch: a deflection prompt never
|
|
# carries the marker (it lives in the user message, same
|
|
# reasoning as SUMMARY_MODE), so a marker question always gets
|
|
# the table answer, whatever the gate says; the E2E asks it
|
|
# against an on-topic fixture, where the gate is HIGH, and
|
|
# asserts non-deflection as part of the table test.
|
|
answer = TABLE_ANSWER
|
|
elif "DEFLECT_MODE" in system:
|
|
answer = (
|
|
"Ah — I haven't done anything like that, so I don't want to make stuff up! "
|
|
"You're thinking bigger than my notes for a second. Try asking about "
|
|
"kubernetes, backups, or deploying a new service — I know those inside out. "
|
|
"You've got this!"
|
|
)
|
|
elif END_OF_NOTES_TRIGGER in user.lower():
|
|
# Whole-document-context story (phase 24): echo the tail of the
|
|
# document context. Byte-stable across runs — a sentinel on the
|
|
# document's last line appears in the answer iff the whole
|
|
# document was in the prompt. (The tail includes the closing
|
|
# </documents> — harmless for the E2E sentinel assertions.)
|
|
# Phase 37: the HIGH prompt now ends with the <tools> section
|
|
# after </documents>, so the echo targets the <documents> block
|
|
# itself — the sentinel semantics are unchanged.
|
|
block = _DOCUMENTS_BLOCK_RE.search(_system(body))
|
|
tail_source = block.group(0) if block else _context(body)
|
|
answer = (
|
|
f"…and the very end of my notes reads: “{tail_source[-160:]}” "
|
|
"(Deterministic mock answer for E2E.)"
|
|
)
|
|
else:
|
|
ctx = _context(body)
|
|
snippet = ctx[:220].replace("\n", " ").strip()
|
|
answer = (
|
|
f"Great question — you've absolutely got this! Here's what my notes say about "
|
|
f"“{user.strip()[:80]}”: {snippet}… That's the gist from the docs; happy to "
|
|
"dig into any of it. (Deterministic mock answer for E2E.)"
|
|
)
|
|
# Steering (phase 15): when the system prompt carries <tuning>, the
|
|
# answer ends with the first note — deterministically observable.
|
|
note = first_tuning_note(system)
|
|
if note:
|
|
answer = f"{answer} (tuning: {note})"
|
|
# KB overview (phase 31): when the system prompt carries
|
|
# <knowledge_base>, the answer ends with the first outline bullet —
|
|
# mirrors the steering echo exactly (appended after it, so the kb
|
|
# suffix is the last thing rendered).
|
|
bullet = first_kb_bullet(system)
|
|
if bullet:
|
|
answer = f"{answer} (kb: {bullet})"
|
|
return answer
|
|
|
|
|
|
def compose_thinking(body: dict[str, Any]) -> str:
|
|
"""Deterministic reasoning scratchpad (thinking-display story, phase 17;
|
|
lengthened in phase 21).
|
|
|
|
A fixed "Step 1… Step 4" template interleaved with a "Scratch" deep-dive
|
|
block, quoting the first ~60 chars of the user question: unique per
|
|
question, byte-stable across runs, ~2 700 chars total (≈ 230 frames at
|
|
the mock's 12-char/0.02s pacing). The length is deliberate (phase 21,
|
|
thinking-no-scroll story): rendered in the 320px ``.thinking-text``
|
|
window it overflows by ~2x, so the live-tail clip and the no-user-scroll
|
|
contract are observable in E2E. The ``Step 2: Check my notes`` line
|
|
fragment (phase 17) and the ``nothing is invented`` tail (phase 20's
|
|
THINKING_TAIL) are what the E2E assertions key off — both are preserved.
|
|
"""
|
|
q = _user(body).strip()[:60]
|
|
return _thinking_template(q)
|
|
|
|
|
|
def _thinking_template(q: str) -> str:
|
|
"""The fixed Step/Scratch scratchpad (``compose_thinking`` and its
|
|
paragraph variant share the exact same text — only the line
|
|
separators differ)."""
|
|
return (
|
|
f"Step 1: Read the question carefully — “{q}” — and figure out what kind of "
|
|
"answer it wants (a how-to, a lookup, or a design decision) before touching "
|
|
"the docs, so I don't over- or under-answer.\n"
|
|
"Step 2: Check my notes for the closest match. The homelab kubernetes file "
|
|
"is the obvious candidate, but I should also consider whether a deployments "
|
|
"note covers the same ground better.\n"
|
|
"Scratch 1: the kubernetes file is organized by component — control plane, "
|
|
"worker nodes, ingress, storage — so I can map each part of the question to "
|
|
"a section instead of summarizing the whole file at once, and keep the "
|
|
"answer anchored to the structure the notes actually use.\n"
|
|
"Scratch 2: I should check whether the deployments note duplicates any of "
|
|
"that ground; if it does, I will prefer the homelab file because the "
|
|
"question is phrased around the cluster itself, and I will say which file "
|
|
"each fact came from so the citation is honest.\n"
|
|
"Scratch 3: versions and ports are the facts most likely to be stale in my "
|
|
"memory — the etcd backup schedule, the ingress controller port, the "
|
|
"registry mirror address — so I will re-read those lines verbatim before "
|
|
"writing a single one of them into the answer.\n"
|
|
"Scratch 4: if the answer needs a sequence, for example how a node joins the "
|
|
"cluster or how the load balancer fronts the control plane, I will keep the "
|
|
"order exactly as the notes write it rather than re-deriving it from general "
|
|
"kubernetes knowledge that may not match this setup.\n"
|
|
"Scratch 5: anything I cannot find in the notes — a host I do not recognize, "
|
|
"a version I am not sure about, a schedule I cannot place — gets left out of "
|
|
"the answer instead of guessed, because the honesty rule beats a longer "
|
|
"answer every single time.\n"
|
|
"Scratch 6: one more pass over the question wording to make sure I am "
|
|
"answering the cluster setup, not some other homelab topic that shares the "
|
|
"same vocabulary, and I will stay on the specific the question asked about.\n"
|
|
"Scratch 7: I will also verify that the file describes the current setup — "
|
|
"if the notes mention a migration from an older cluster, I should answer "
|
|
"from the post-migration section and not mix in the old host names or the "
|
|
"old port numbers that no longer apply.\n"
|
|
"Scratch 8: final shape check before I commit — short paragraphs, a few "
|
|
"bullets at most, the document path cited where the fact came from, and no "
|
|
"invented facts anywhere in the draft.\n"
|
|
"Step 3: Re-read the relevant sections top to bottom so every specific — "
|
|
"hosts, versions, ports, schedules — is exact as written rather than "
|
|
"remembered, and note which document each fact comes from.\n"
|
|
"Step 4: Draft the answer around those specifics, keep it tight with short "
|
|
"paragraphs and bullets where it helps, cite the documents by path, and "
|
|
"double-check that nothing is invented."
|
|
)
|
|
|
|
|
|
def compose_thinking_paragraphs(body: dict[str, Any]) -> str:
|
|
"""The phase-17 scratchpad with REAL paragraph breaks (\n\n, 2026-08-29
|
|
regression pin): the same deterministic text as ``compose_thinking``,
|
|
with a blank line inserted after scratchpad lines 2 and 6 (0-based) —
|
|
two genuine \"2-newline gaps\" in the rendered scratchpad. Unique per
|
|
question, byte-stable across runs (same length contract + 2 chars)."""
|
|
base = compose_thinking(body)
|
|
lines = base.split("\n")
|
|
out: list[str] = []
|
|
for i, line in enumerate(lines):
|
|
out.append(line)
|
|
if i in (2, 6):
|
|
out.append("") # blank line -> a real \"\n\n\" gap
|
|
return "\n".join(out)
|
|
|
|
|
|
@app.post("/__shutdown__")
|
|
def shutdown() -> dict[str, Any]:
|
|
"""Test hook (loading-feedback story): terminate this mock process to
|
|
simulate an LLM outage. The E2E fixture restores a fresh instance on
|
|
the same port afterwards, so the rest of the session keeps working."""
|
|
|
|
def _die() -> None:
|
|
time.sleep(0.1) # let the HTTP response flush before we exit
|
|
os.kill(os.getpid(), signal.SIGTERM)
|
|
|
|
threading.Thread(target=_die, daemon=True).start()
|
|
return {"status": "shutting down"}
|
|
|
|
|
|
@app.get("/v1/models")
|
|
def models() -> dict[str, Any]:
|
|
return {
|
|
"object": "list",
|
|
"data": [
|
|
{"id": "turbo", "object": "model"},
|
|
{"id": "embed", "object": "model"},
|
|
{"id": "lite", "object": "model"},
|
|
],
|
|
}
|
|
|
|
|
|
@app.post("/v1/embeddings")
|
|
def embeddings(body: dict[str, Any]) -> Any: # dict, or a 500 (phase 67)
|
|
raw = body.get("input")
|
|
if isinstance(raw, str):
|
|
raw = [raw]
|
|
inputs: list[Any] = list(raw) if isinstance(raw, list) else []
|
|
# Phase 67 (embedding retry): the first embeddings request whose
|
|
# input carries the marker 500s; the next returns the normal
|
|
# bag-of-words vector (see the module docstring). Raw httpx on the
|
|
# client side — no SDK-level retries — so one POST per app attempt:
|
|
# the counter is per POST here (unlike the chat counter below).
|
|
joined = " ".join(str(t) for t in inputs if isinstance(t, str)).lower()
|
|
if EMBED_FAIL_TRIGGER in joined:
|
|
n = _bump_fail(EMBED_FAIL_TRIGGER)
|
|
if n == 1:
|
|
return _llm_500(EMBED_FAIL_TRIGGER)
|
|
_fail_posts[EMBED_FAIL_TRIGGER] = 0 # the vector went out — restart
|
|
data = [
|
|
{"object": "embedding", "index": i, "embedding": embed_text(t)}
|
|
for i, t in enumerate(inputs)
|
|
]
|
|
return {
|
|
"object": "list",
|
|
"data": data,
|
|
"model": body.get("model", "embed"),
|
|
"usage": {"prompt_tokens": 8, "total_tokens": 8},
|
|
}
|
|
|
|
|
|
def _sse_stream(
|
|
answer: str,
|
|
delay: float,
|
|
thinking: str = "",
|
|
pre_content_delay: float = 0.0,
|
|
chunk: int = 12,
|
|
) -> Any:
|
|
"""SSE frames for one chat completion (phase 17: + reasoning).
|
|
|
|
``chunk`` (default 12) is the slice size for BOTH the thinking and
|
|
the content frames — the ``think in paragraphs`` trigger raises it
|
|
to ``THINK_PARAS_CHUNK`` (60) so a single frame renders past the
|
|
32px think-window band (see ``THINK_PARAS_TRIGGER``). At 12 the
|
|
output is byte-identical to the original.
|
|
|
|
When ``thinking`` is non-empty its ``chunk``-sized slices go out FIRST as
|
|
``delta.reasoning_content`` frames — same 0.02s cadence and envelope
|
|
as the content frames, the aipi wire convention (reasoning before
|
|
content). Without ``thinking`` the output is byte-identical to the
|
|
content-only stream, so the other story suites are unaffected.
|
|
|
|
``pre_content_delay`` (phase 20) inserts a silence gap between the end
|
|
of the thinking stream and the first content frame — the client stays
|
|
in its pre-token "thinking" state the whole time (0.02s cadence and
|
|
frame shapes are unchanged, so 0.0 is byte-identical to before).
|
|
"""
|
|
model = "turbo"
|
|
chunk_id = f"chatcmpl-{uuid.uuid4()}"
|
|
if delay:
|
|
time.sleep(delay)
|
|
for piece in re.findall(rf".{{1,{chunk}}}", thinking, re.S):
|
|
payload = {
|
|
"id": chunk_id,
|
|
"object": "chat.completion.chunk",
|
|
"created": int(time.time()),
|
|
"model": model,
|
|
"choices": [
|
|
{"index": 0, "delta": {"reasoning_content": piece}, "finish_reason": None}
|
|
],
|
|
}
|
|
yield f"data: {json_dumps(payload)}\n\n"
|
|
time.sleep(0.02)
|
|
if pre_content_delay:
|
|
time.sleep(pre_content_delay)
|
|
for piece in re.findall(rf".{{1,{chunk}}}", answer, re.S):
|
|
payload = {
|
|
"id": chunk_id,
|
|
"object": "chat.completion.chunk",
|
|
"created": int(time.time()),
|
|
"model": model,
|
|
"choices": [{"index": 0, "delta": {"content": piece}, "finish_reason": None}],
|
|
}
|
|
yield f"data: {json_dumps(payload)}\n\n"
|
|
time.sleep(0.02)
|
|
yield (
|
|
"data: "
|
|
+ json_dumps(
|
|
{
|
|
"id": chunk_id,
|
|
"object": "chat.completion.chunk",
|
|
"created": int(time.time()),
|
|
"model": model,
|
|
"choices": [{"index": 0, "delta": {}, "finish_reason": "stop"}],
|
|
}
|
|
)
|
|
+ "\n\n"
|
|
)
|
|
yield "data: [DONE]\n\n"
|
|
|
|
|
|
def json_dumps(obj: dict[str, Any]) -> str:
|
|
import json
|
|
|
|
return json.dumps(obj)
|
|
|
|
|
|
def _apply_max_tokens(answer: str, max_tokens: Any) -> str:
|
|
"""Deterministic stand-in for the endpoint's output cap: one token ≈
|
|
one whitespace-separated word. Answers within the cap pass through
|
|
byte-identical, so existing (short) answers are unaffected."""
|
|
if not isinstance(max_tokens, int) or max_tokens <= 0:
|
|
return answer
|
|
words = answer.split()
|
|
if len(words) <= max_tokens:
|
|
return answer
|
|
return " ".join(words[:max_tokens])
|
|
|
|
|
|
def _tool_call_stream(name: str, arguments: dict[str, Any], call_id: str) -> Any:
|
|
"""SSE frames for one tool-call-only chat completion (phase 37).
|
|
|
|
The OpenAI wire convention the app accumulates (``app/rag/llm.py``):
|
|
the first partial of index 0 carries ``id`` + ``type`` +
|
|
``function.name`` plus the first ``function.arguments`` fragment;
|
|
the remaining fragments (deterministic 16-char split — so the
|
|
multi-fragment accumulation path is exercised) arrive on later
|
|
chunks; the final chunk carries ``finish_reason: "tool_calls"``.
|
|
No ``content`` / ``reasoning_content`` frames — the turn asked for a
|
|
tool instead of answering.
|
|
|
|
Pacing: 0.1 s per frame — deliberately SLOWER than the content
|
|
stream's 0.02 s, so the UI's transient "calling tool" state (held
|
|
from the first ``tool`` frame until the first answer ``delta``) is a
|
|
comfortable observation window for the story E2E (~1 s across the
|
|
two tool requests).
|
|
"""
|
|
model = "turbo"
|
|
chunk_id = f"chatcmpl-{uuid.uuid4()}"
|
|
raw_args = json_dumps(arguments) if arguments else "{}"
|
|
frags = [raw_args[i : i + 16] for i in range(0, len(raw_args), 16)] or ["{}"]
|
|
for i, frag in enumerate(frags):
|
|
tc: dict[str, Any] = {"index": 0, "function": {"arguments": frag}}
|
|
delta: dict[str, Any] = {"tool_calls": [tc]}
|
|
if i == 0:
|
|
tc = {
|
|
"index": 0,
|
|
"id": call_id,
|
|
"type": "function",
|
|
"function": {"name": name, "arguments": frag},
|
|
}
|
|
delta = {"role": "assistant", "tool_calls": [tc]}
|
|
payload = {
|
|
"id": chunk_id,
|
|
"object": "chat.completion.chunk",
|
|
"created": int(time.time()),
|
|
"model": model,
|
|
"choices": [{"index": 0, "delta": delta, "finish_reason": None}],
|
|
}
|
|
yield f"data: {json_dumps(payload)}\n\n"
|
|
time.sleep(0.1)
|
|
yield (
|
|
"data: "
|
|
+ json_dumps(
|
|
{
|
|
"id": chunk_id,
|
|
"object": "chat.completion.chunk",
|
|
"created": int(time.time()),
|
|
"model": model,
|
|
"choices": [{"index": 0, "delta": {}, "finish_reason": "tool_calls"}],
|
|
}
|
|
)
|
|
+ "\n\n"
|
|
)
|
|
yield "data: [DONE]\n\n"
|
|
|
|
|
|
@app.post("/v1/chat/completions")
|
|
def chat_completions(body: dict[str, Any]) -> Any:
|
|
user_lower = _user(body).lower()
|
|
# Phase 37 (agent document tools): the deterministic marker flow.
|
|
# The app's chat path is the only streaming consumer of this mock, so
|
|
# the flow handles streaming requests; a non-streaming marker request
|
|
# (never issued by the app) falls through to the regular answer.
|
|
if body.get("stream"):
|
|
# Phase 67 (LLM retry): deterministic failure injection — see
|
|
# the module docstring. Checked before the marker tool flow: the
|
|
# injection markers never combine with the tool-flow markers in
|
|
# any suite, and a dead endpoint answers nothing (no flow).
|
|
if ALWAYS_FAIL_TRIGGER in user_lower:
|
|
return _llm_500(ALWAYS_FAIL_TRIGGER)
|
|
if RETRY_TRIGGER in user_lower:
|
|
if _chat_dead(RETRY_TRIGGER, RETRY_DEAD_ATTEMPTS):
|
|
return _llm_500(RETRY_TRIGGER)
|
|
_fail_posts[RETRY_TRIGGER] = 0 # the answer streamed — restart
|
|
# Phase 71 (tool-scaffolding guardrails): the deterministic raw-
|
|
# markup flow — checked BEFORE the search/tool marker flows (the
|
|
# trigger is independent of the ``<tools>`` marker, so both
|
|
# grounded and deflected turns hit it; SCAFFOLD_ALWAYS_TRIGGER
|
|
# is checked first inside the classifier — the more specific
|
|
# phrase wins, same convention as THINK_PARAS_TRIGGER).
|
|
scaffold_flow = _scaffold_flow(body)
|
|
if scaffold_flow is not None:
|
|
# Request 1 (or EVERY request for the ALWAYS trigger): the
|
|
# incident span as plain delta.content, 12-char chunks
|
|
# (the span always spans ≥2 chunks — the filter's boundary
|
|
# path), finish_reason "stop", no tool_calls, no reasoning.
|
|
# Request 2 of the recovery trigger: the clean answer.
|
|
answer = (
|
|
SCAFFOLD_RECOVERY_ANSWER
|
|
if scaffold_flow == "recovery"
|
|
else SCAFFOLD_SPAN
|
|
)
|
|
return StreamingResponse(
|
|
_sse_stream(answer, 0.0),
|
|
media_type="text/event-stream",
|
|
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
|
|
)
|
|
# Phase 68 (search tool): the deterministic search marker flow —
|
|
# checked BEFORE the phase-37 tool flow (the more specific
|
|
# trigger phrase wins, same convention as THINK_PARAS_TRIGGER).
|
|
search_flow = _search_flow(body)
|
|
if search_flow is not None:
|
|
if search_flow[0] == "search":
|
|
stream = _tool_call_stream(
|
|
"grep", {"pattern": SEARCH_PATTERN}, "call_0"
|
|
)
|
|
else: # "found" — quote the first matched line (80 chars)
|
|
answer = _apply_max_tokens(
|
|
f"Found {search_flow[1][:80]}", body.get("max_tokens")
|
|
)
|
|
stream = _sse_stream(answer, 0.0)
|
|
return StreamingResponse(
|
|
stream,
|
|
media_type="text/event-stream",
|
|
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
|
|
)
|
|
# Phase 72 (teaching refusals): the deterministic LS-TEACH
|
|
# self-correction flow — checked BEFORE the plain
|
|
# TOOLS_TRIGGER flow (disjoint trigger phrases — the phase-71
|
|
# ordering convention; the trigger needs the ``<tools>``
|
|
# section, so deflected turns never hit it).
|
|
ls_teach = _ls_teach_flow(body)
|
|
if ls_teach is not None:
|
|
if ls_teach[0] == "misuse":
|
|
# The incident's misuse, deterministic: ls(path='.').
|
|
stream = _tool_call_stream("ls", {"path": "."}, "call_0")
|
|
elif ls_teach[0] == "correct":
|
|
# The one-round correction: the no-arg full listing.
|
|
stream = _tool_call_stream("ls", {}, "call_1")
|
|
else: # "answer" — quote the first catalog line
|
|
answer = _apply_max_tokens(
|
|
f"These are the indexed documents: {ls_teach[1]}",
|
|
body.get("max_tokens"),
|
|
)
|
|
stream = _sse_stream(answer, 0.0)
|
|
return StreamingResponse(
|
|
stream,
|
|
media_type="text/event-stream",
|
|
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
|
|
)
|
|
flow = _tool_flow(body)
|
|
if flow is not None:
|
|
if flow[0] == "list":
|
|
stream = _tool_call_stream("ls", {}, "call_0")
|
|
elif flow[0] == "read":
|
|
# flow[3] is the synthetic call id — "call_1" for the
|
|
# single-read flow and the multi-read first read,
|
|
# "call_2" for the multi-read second read (phase 45,
|
|
# task 02). Phase 70: the harness-aligned shape — one
|
|
# combined ``source/path`` argument (the mock joins the
|
|
# two catalog fields; the catalog format is unchanged).
|
|
stream = _tool_call_stream(
|
|
"read", {"path": f"{flow[1]}/{flow[2]}"}, flow[3]
|
|
)
|
|
elif flow[0] == "multi_answer":
|
|
# Phase 45 (task 02): the multi-read forced answer —
|
|
# computed in _tool_flow, byte-stable.
|
|
stream = _sse_stream(_apply_max_tokens(flow[2], body.get("max_tokens")), 0.0)
|
|
else: # "answer" — quote the read document (first 80 chars)
|
|
answer = _apply_max_tokens(
|
|
f"Read {flow[1]}. {flow[2][:80]}", body.get("max_tokens")
|
|
)
|
|
stream = _sse_stream(answer, 0.0)
|
|
return StreamingResponse(
|
|
stream,
|
|
media_type="text/event-stream",
|
|
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
|
|
)
|
|
|
|
answer = _apply_max_tokens(compose_answer(body), body.get("max_tokens"))
|
|
delay = 3.0 if "pretend to think slowly" in _user(body) else 0.0
|
|
# ``think in paragraphs`` wins over ``think out loud`` (more specific):
|
|
# the same scratchpad WITH real "\n\n" paragraph breaks, at 60-char
|
|
# frames (real-model-sized deltas — the 32px-band regression pin).
|
|
if THINK_PARAS_TRIGGER in user_lower:
|
|
thinking = compose_thinking_paragraphs(body)
|
|
chunk = THINK_PARAS_CHUNK
|
|
elif THINKING_TRIGGER in user_lower:
|
|
thinking = compose_thinking(body)
|
|
chunk = 12
|
|
else:
|
|
thinking = ""
|
|
chunk = 12
|
|
pre_content = (
|
|
PRE_CONTENT_PAUSE_S if SLOW_PRETOKEN_TRIGGER in user_lower else 0.0
|
|
)
|
|
|
|
if not body.get("stream"):
|
|
message: dict[str, Any] = {"role": "assistant", "content": answer}
|
|
if thinking:
|
|
# Harmless future-proofing: the app only uses streaming, but a
|
|
# non-streaming client that reads the field gets the reasoning.
|
|
message["reasoning_content"] = thinking
|
|
return {
|
|
"id": f"chatcmpl-{uuid.uuid4()}",
|
|
"object": "chat.completion",
|
|
"created": int(time.time()),
|
|
"model": body.get("model", "turbo"),
|
|
"choices": [
|
|
{"index": 0, "message": message, "finish_reason": "stop"}
|
|
],
|
|
"usage": {"prompt_tokens": 100, "completion_tokens": 50, "total_tokens": 150},
|
|
}
|
|
|
|
return StreamingResponse(
|
|
_sse_stream(
|
|
answer,
|
|
delay,
|
|
thinking=thinking,
|
|
pre_content_delay=pre_content,
|
|
chunk=chunk,
|
|
),
|
|
media_type="text/event-stream",
|
|
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
|
|
)
|