feat(rag): agent document tools — list/read tools with env-tuned budgets, SSE tool events + "calling tool" UI

Grounded chat turns now run the agent loop (app/rag/agent.py) instead
of a bare chat_stream: while the per-turn budgets last
(BOR_AGENT_LIST_CALLS / BOR_AGENT_READ_CALLS, default 1 each) the model
gets list_documents (the indexed catalog, /api/docs order) and
read_document (full text, never truncated — A7-revised contract); once
both budgets are spent the tools key is dropped from the request and
the model must answer. Rejected calls (unknown tool, unknown/missing
path, document already in context, spent budget) consume no budget.
Budgets 0/0 make exactly one tools=None request — byte-identical to
the pre-phase path (budgets-as-kill-switch). Deflected turns keep the
direct chat_stream (A8 unchanged; the LOW prompt never carries the
<tools> section).

SSE contract gains {"type":"tool","name":...,"argument":
"source/path"|null} frames ahead of the answer deltas (PLAN §4
extension, owner permission 2026-08-26); done.sources, query_log.sources
and the per-turn log line (gains tool_calls=N) report the retrieval
docs + read docs, deduped. The UI shows a "calling tool"
button/label state and one visible .tool-call line per call above the
answer; the lines persist with the chat record and re-render on
reload. chat_stream passes tools through and accumulates streaming
tool_calls deltas into ToolCallPiece (tools=None stays byte-identical).

E2E: deterministic mock tool flow ("use your tools" + <tools> marker:
list -> read first catalog line -> quoted answer) plus the story suite
(marker flow, reload re-render, plain/deflected no-tool regressions).
Docs: .env.example + README (the two tools, the budgets, the SSE tool
frame, the "calling tool" UI state).

probe: turbo tool_calls=supported 2026-08-26 (uv run python -m
scripts.llm_probe --tools — non-streaming + streaming
finish_reason=tool_calls, indexed delta.tool_calls partials)
This commit is contained in:
2026-08-26 22:39:14 -04:00
parent 9efffcb428
commit 15c1272828
30 changed files with 3594 additions and 67 deletions
+39
View File
@@ -569,6 +569,45 @@ details.thinking .thinking-text {
details.thinking .thinking-text p,
details.thinking .thinking-text ul { margin: 0 0 0.5rem; }
/* Agent tool-call lines (phase 37): one visible "calling tool" row per
`tool` SSE frame — in the same wrap as the Thinking block, above the
answer, below the Thinking summary. Deliberately distinct from the
scratchpad: accent palette (--accent-ink) vs the brand-ink summary,
own icon, own accent left border. Contrast: --accent-ink on the row's
--surface ≈10.4:1 (11.6:1 on the page bg), and --ink on --brand-soft
in the path `code` ≈11.5:1 — all comfortably AA in the (single dark)
theme. Inline rows only: appending lines never shifts the 46rem chat
column (no new container), and the rows are not interactive — no
focus targets. */
.tool-calls {
display: flex;
flex-direction: column;
gap: 0.25rem;
}
.tool-call {
display: flex;
align-items: baseline;
gap: 0.45rem;
background: var(--surface);
border: 1px solid var(--line);
border-left: 3px solid var(--accent-line);
border-radius: var(--radius-sm);
padding: 0.3rem 0.75rem;
color: var(--accent-ink);
font-size: 0.8rem;
line-height: 1.4;
overflow-wrap: anywhere;
}
.tool-call code {
font-family: var(--mono);
font-size: 0.95em;
background: var(--brand-soft);
color: var(--ink);
padding: 0.05em 0.35em;
border-radius: 5px;
overflow-wrap: anywhere;
}
.msg-meta {
font-size: 0.75rem;
color: var(--ink-soft);