chore(agent): track .agent/ planning tree in git
Build and Push Containers / build-and-push-app (push) Successful in 12s
Build and Push Containers / build-and-push-db (push) Successful in 10s

Remove the blanket .agent/ gitignore so the phase roadmap, user
stories, reports, and PLAN.md are versioned with the code. Only
runtime artifacts (.agent/phase-sessions/, .agent/pipeline.log)
remain ignored. Update AGENTS.md git protocol rule to match.
This commit is contained in:
2026-09-01 10:18:22 -04:00
parent 5fa620fde5
commit 4971e2859d
818 changed files with 23964 additions and 4 deletions
@@ -0,0 +1,65 @@
# Story: Agent document tools (list + read)
**Phase:** `37_agent_document_tools` · **E2E:** `tests/e2e/test_agent_document_tools.py`
## Narrative
As **a user**, when my question depends on content the retriever did not put
in context — e.g. a document that *references* another file ("the exact JSON
shape is in example-record-file.json") — today I get a guess or an apology
("I don't want to invent it!"). I want the answer agent to have **two
tools**: list the available sources, and read the one extra document it
thinks is relevant — with the number of opportunities configured by
environment variables — and the UI to show a "calling tool" state in
addition to "thinking".
- **Given** a grounded (HIGH-gate) chat turn
- **When** the model calls `list_documents` (≤ `BOR_AGENT_LIST_CALLS`,
default 1) or `read_document(source, path)` (≤ `BOR_AGENT_READ_CALLS`,
default 1)
- **Then** the app executes the tool server-side (DB-only), streams a
`tool` SSE event the UI renders as a "calling tool" line + button state,
and — once both budgets are spent — the model is required to answer with
the augmented context; `done` sources include the read document.
## Acceptance criteria
1. Live probe (task 01): `scripts/llm_probe.py --tools` verifies
`turbo`'s tool-calling support (non-streaming + streaming) against the
aipi endpoint; the verdict is recorded in the `app/rag/agent.py`
docstring + the commit message (phase-17 "verified live" convention).
2. `app/rag/llm.py`: `chat_stream(messages, tools=None)` — `tools` passed
through to the API; streaming `tool_calls` deltas (partial `name` /
fragmented `arguments`, indexed) accumulated into
`ToolCallPiece(id, name, arguments)`; `tools=None` → byte-identical
request to today.
3. `app/rag/agent.py`: the loop — budgets from
`BOR_AGENT_LIST_CALLS` / `BOR_AGENT_READ_CALLS` (default 1/1); the
`list_documents` tool returns the DB catalog
(`source/path — title` lines, `/api/docs` order); `read_document`
appends the **full** document text (A7-revised: never truncated); after
both budgets are spent the tools are dropped and the model must answer;
safety cap on model rounds; the deflection (LOW) path is byte-identical.
4. `app/api/chat.py` + `app/schemas.py`: new SSE
`{"type":"tool","name":…,"argument":…}` event (PLAN §4 revision note,
owner permission 2026-08-26); `done.sources` + `query_log.sources`
include the read document (deduped); the per-turn log line gains
`tool_calls=N` (PLAN §9 note).
5. Frontend: "calling tool" button/label state + visible `.tool-call`
lines in the bubble (distinct from the Thinking block), aria-live,
tool lines persisted with the chat record (phase 14/17 convention);
UI Structure Check + no CDN.
6. `mock_llm.py` deterministic: marker-driven list → read → answer
sequence (the mock parses the catalog out of the tool result it was
given); the story E2E asserts the tool UI, the answer quoting the read
document, the read document as a source chip, and reload re-render;
existing suites stay green (no tool events without the marker).
7. Unit + integration green, `app/` coverage >90%, story E2E green in
isolation, ruff + pyright clean, one `--no-gpg-sign` commit.
## Playwright Mapping Rule
`tests/e2e/test_agent_document_tools.py` — one story, one file, run in
isolation: the marker question triggers two tool events; the UI shows the
"calling tool" state while tools run; the final answer quotes the read
document; the source chips include it; after a reload the tool lines
re-render; a plain question and a deflected question produce no tool
events.