chore(agent): phase roadmap from TODO.md — 8 phases (40–47), 24 tasks
Converts the 9 TODO items into an executable phase roadmap (Protocol B, appended after phase 39): - 40 tuning toggle anonymous flash (TODO L3) - 41 sync fail-fast + modal when a model is down (TODO L4) - 42 no reply autoscroll (TODO L5) - 43 thinking scroll back — user scroll + gated autoscroll (TODO L7) - 44 markdown tables (TODO L6) - 45 agent unlimited tool calls behind BOR_AGENT_MAX_ROUNDS (TODO L8) - 46 mobile hamburger nav (TODO L9) - 47 quadlet + jinja import formats, A9 revision (TODO L10–L11) Each phase carries a user story, a dedicated Playwright E2E suite plan, and owner-locked decisions (R1 A9 format extension, R2 phase-37 budget revision, A1–A5 scope decisions) confirmed 2026-08-27. Also records the completed phases 30–39 todo/ -> complete/ moves that were pending in the working tree. TODO.md is cleared (items now live in .agent/phases/todo/).
This commit is contained in:
@@ -0,0 +1,103 @@
|
||||
# Story: Agent makes as many tool calls as it wants
|
||||
|
||||
**Phase:** `45_agent_unlimited_tools` · **Source:** `TODO.md` L8 ·
|
||||
**E2E:** `tests/e2e/test_agent_unlimited_tools.py`
|
||||
|
||||
## Bug report (verbatim, `TODO.md` L8)
|
||||
|
||||
> "Allow the LLM to make as many tool calls as it wants, remove the
|
||||
> restrictions, they're causing problems getting correct answers"
|
||||
|
||||
## Narrative
|
||||
|
||||
As **the owner**, the phase-37 per-turn budgets — one `list_documents`
|
||||
call and one `read_document` call — cap answers: complex questions need
|
||||
several documents, and the "No reading budget left" refusal is where
|
||||
correct answers die. The per-tool restrictions are being **removed**:
|
||||
the model may call the tools as many times as it needs, within a single
|
||||
configurable **round cap** that exists only to stop a pathological
|
||||
infinite loop (and doubles as the no-tools kill switch).
|
||||
|
||||
- **Given** a grounded chat turn (retrieval found relevant docs)
|
||||
- **When** the model calls `list_documents` / `read_document`
|
||||
- **Then** every valid call is executed — re-lists included — until the
|
||||
model answers or the round cap is reached, at which point a final
|
||||
no-tools answer is forced.
|
||||
|
||||
## Acceptance criteria
|
||||
1. **No per-tool budgets:** `BOR_AGENT_LIST_CALLS` /
|
||||
`BOR_AGENT_READ_CALLS` are gone from `app/config.py`,
|
||||
`.env.example`, and the agent loop; the `LIST_EXHAUSTED` /
|
||||
`READ_EXHAUSTED` refusals no longer exist.
|
||||
2. **Round cap only:** new `agent_max_rounds`
|
||||
(`BOR_AGENT_MAX_ROUNDS`, default **10**) counts tool rounds; at the
|
||||
cap the loop forces one final `tools=None` answer. `0` disables the
|
||||
tools entirely — the request goes out with `tools=None`,
|
||||
byte-identical to the pre-phase-37 path (the kill switch survives,
|
||||
per the owner-locked revision).
|
||||
3. **Non-budget rejections kept:** unknown tool → `"Unknown tool."`,
|
||||
missing args → the `MISSING_READ_ARGS` refusal, already-in-context
|
||||
document → `"Already in your context."` — none of these consume a
|
||||
round's *budget* (there is none) but the round cap still bounds a
|
||||
stream that keeps emitting rejected calls.
|
||||
4. **Everything downstream unchanged:** the `tool` SSE event shape, the
|
||||
per-turn `tool_calls=N` log field (budget-consuming → now: executed),
|
||||
`done.sources` extension, and the UI tool lines are untouched.
|
||||
5. **PLAN recorded:** the phase-37 locked decision
|
||||
("budgets-as-kill-switch") is revised in `.agent/PLAN.md` with an
|
||||
owner-permission note (2026-08-27, `TODO.md` L8), following the
|
||||
established revision-note pattern (phases 16/19/24/37).
|
||||
|
||||
## Owner-confirmed (2026-08-27, roadmap R2)
|
||||
1. **Both budget env vars removed.** New single guard
|
||||
`BOR_AGENT_MAX_ROUNDS` (default **10** tool rounds, then forced final
|
||||
answer); **`0` = no-tools kill switch** (pre-phase-37 behavior).
|
||||
2. **Re-listing is allowed** (a second `list_documents` is a normal
|
||||
executed call — it even counts in `tool_calls=N`).
|
||||
|
||||
## UI Visualization & Structure
|
||||
- **Config** (`app/config.py`): delete `agent_list_calls` /
|
||||
`agent_read_calls`; add `agent_max_rounds: int = 10` with a docstring
|
||||
(0 disables the tools entirely — the loop makes exactly one request
|
||||
with `tools=None`). `.env.example`: the two `BOR_AGENT_*_CALLS` lines
|
||||
are replaced by `BOR_AGENT_MAX_ROUNDS=10` with an updated comment.
|
||||
- **Agent loop** (`app/rag/agent.py`): `run_agent` —
|
||||
`max_rounds = settings.agent_max_rounds`; `tools = AGENT_TOOLS if
|
||||
max_rounds > 0 else None`; the loop drops `list_left` / `read_left`
|
||||
and the budget-driven `tools = None` transition; after each executed
|
||||
call `rounds += 1` and at `rounds >= max_rounds` the existing forced
|
||||
final answer path runs (now the *only* exit besides "no calls").
|
||||
`_execute_tool` loses its budget parameters + the two exhaustion
|
||||
constants; `AGENT_TOOLS`' `read_document` description drops "exactly
|
||||
one more"; module/docstrings and the probe reference updated.
|
||||
`AgentHolder` unchanged (`tool_calls` still counts executed calls).
|
||||
- **Tests:** `tests/unit/test_agent.py` rewritten around the round cap
|
||||
(always-listing mock LLM: N rounds then forced answer; `max_rounds=0`
|
||||
→ single `tools=None` request; rejected-call spam bounded by the cap);
|
||||
`tests/unit/test_config.py` (default 10, env override, 0);
|
||||
`tests/integration/test_chat_api.py` budget fixtures →
|
||||
`agent_max_rounds`; `tests/e2e/mock_llm.py` `_tool_flow` extended: the
|
||||
existing deterministic 3-step flow stays, plus a multi-read variant
|
||||
triggered by a marker (e.g. the user message containing
|
||||
`"read two documents"`) — stateless classification by counting
|
||||
`Document …:` tool messages (list → read #1 → read #2 → answer).
|
||||
- **Non-goals:** no per-call cost cap, no streaming change, no new
|
||||
endpoint.
|
||||
|
||||
## Playwright Mapping Rule
|
||||
**Test Scenario → `tests/e2e/test_agent_unlimited_tools.py`** (mock
|
||||
LLM; DB up):
|
||||
1. `test_multi_read_turn` — a grounded question carrying both
|
||||
`TOOLS_TRIGGER` and the multi-read marker: the turn streams tool
|
||||
lines for `list_documents` **and two** `read_document` calls, then a
|
||||
final answer; the bubble is not deflected.
|
||||
2. `test_done_sources_include_reads` — the final `done` (observed via
|
||||
the source chips) lists the retrieved doc(s) plus **both** read
|
||||
documents, deduped.
|
||||
3. `test_relist_allowed` (unit-level via integration, plus E2E
|
||||
observable state) — a re-listed catalog does not produce a refusal
|
||||
line; the UI shows a tool line per executed call.
|
||||
4. `test_single_tool_flow_regression` (phase 37) — the original
|
||||
3-step flow still answers with exactly one read (runs against the
|
||||
unchanged `tests/e2e/test_agent_document_tools.py` in the
|
||||
regression pass, not duplicated here).
|
||||
Reference in New Issue
Block a user