Files
brain-of-reese/.agent/user_stories/agent-unlimited-tools.md
T
ducoterra 02c76ad328 chore(agent): phase roadmap from TODO.md — 8 phases (40–47), 24 tasks
Converts the 9 TODO items into an executable phase roadmap (Protocol B,
appended after phase 39):

- 40 tuning toggle anonymous flash (TODO L3)
- 41 sync fail-fast + modal when a model is down (TODO L4)
- 42 no reply autoscroll (TODO L5)
- 43 thinking scroll back — user scroll + gated autoscroll (TODO L7)
- 44 markdown tables (TODO L6)
- 45 agent unlimited tool calls behind BOR_AGENT_MAX_ROUNDS (TODO L8)
- 46 mobile hamburger nav (TODO L9)
- 47 quadlet + jinja import formats, A9 revision (TODO L10–L11)

Each phase carries a user story, a dedicated Playwright E2E suite plan,
and owner-locked decisions (R1 A9 format extension, R2 phase-37 budget
revision, A1–A5 scope decisions) confirmed 2026-08-27.

Also records the completed phases 30–39 todo/ -> complete/ moves that
were pending in the working tree. TODO.md is cleared (items now live in
.agent/phases/todo/).
2026-08-27 18:25:53 -04:00

5.5 KiB

Story: Agent makes as many tool calls as it wants

Phase: 45_agent_unlimited_tools · Source: TODO.md L8 · E2E: tests/e2e/test_agent_unlimited_tools.py

Bug report (verbatim, TODO.md L8)

"Allow the LLM to make as many tool calls as it wants, remove the restrictions, they're causing problems getting correct answers"

Narrative

As the owner, the phase-37 per-turn budgets — one list_documents call and one read_document call — cap answers: complex questions need several documents, and the "No reading budget left" refusal is where correct answers die. The per-tool restrictions are being removed: the model may call the tools as many times as it needs, within a single configurable round cap that exists only to stop a pathological infinite loop (and doubles as the no-tools kill switch).

  • Given a grounded chat turn (retrieval found relevant docs)
  • When the model calls list_documents / read_document
  • Then every valid call is executed — re-lists included — until the model answers or the round cap is reached, at which point a final no-tools answer is forced.

Acceptance criteria

  1. No per-tool budgets: BOR_AGENT_LIST_CALLS / BOR_AGENT_READ_CALLS are gone from app/config.py, .env.example, and the agent loop; the LIST_EXHAUSTED / READ_EXHAUSTED refusals no longer exist.
  2. Round cap only: new agent_max_rounds (BOR_AGENT_MAX_ROUNDS, default 10) counts tool rounds; at the cap the loop forces one final tools=None answer. 0 disables the tools entirely — the request goes out with tools=None, byte-identical to the pre-phase-37 path (the kill switch survives, per the owner-locked revision).
  3. Non-budget rejections kept: unknown tool → "Unknown tool.", missing args → the MISSING_READ_ARGS refusal, already-in-context document → "Already in your context." — none of these consume a round's budget (there is none) but the round cap still bounds a stream that keeps emitting rejected calls.
  4. Everything downstream unchanged: the tool SSE event shape, the per-turn tool_calls=N log field (budget-consuming → now: executed), done.sources extension, and the UI tool lines are untouched.
  5. PLAN recorded: the phase-37 locked decision ("budgets-as-kill-switch") is revised in .agent/PLAN.md with an owner-permission note (2026-08-27, TODO.md L8), following the established revision-note pattern (phases 16/19/24/37).

Owner-confirmed (2026-08-27, roadmap R2)

  1. Both budget env vars removed. New single guard BOR_AGENT_MAX_ROUNDS (default 10 tool rounds, then forced final answer); 0 = no-tools kill switch (pre-phase-37 behavior).
  2. Re-listing is allowed (a second list_documents is a normal executed call — it even counts in tool_calls=N).

UI Visualization & Structure

  • Config (app/config.py): delete agent_list_calls / agent_read_calls; add agent_max_rounds: int = 10 with a docstring (0 disables the tools entirely — the loop makes exactly one request with tools=None). .env.example: the two BOR_AGENT_*_CALLS lines are replaced by BOR_AGENT_MAX_ROUNDS=10 with an updated comment.
  • Agent loop (app/rag/agent.py): run_agent — max_rounds = settings.agent_max_rounds; tools = AGENT_TOOLS if max_rounds > 0 else None; the loop drops list_left / read_left and the budget-driven tools = None transition; after each executed call rounds += 1 and at rounds >= max_rounds the existing forced final answer path runs (now the only exit besides "no calls"). _execute_tool loses its budget parameters + the two exhaustion constants; AGENT_TOOLS' read_document description drops "exactly one more"; module/docstrings and the probe reference updated. AgentHolder unchanged (tool_calls still counts executed calls).
  • Tests: tests/unit/test_agent.py rewritten around the round cap (always-listing mock LLM: N rounds then forced answer; max_rounds=0 → single tools=None request; rejected-call spam bounded by the cap); tests/unit/test_config.py (default 10, env override, 0); tests/integration/test_chat_api.py budget fixtures → agent_max_rounds; tests/e2e/mock_llm.py _tool_flow extended: the existing deterministic 3-step flow stays, plus a multi-read variant triggered by a marker (e.g. the user message containing "read two documents") — stateless classification by counting Document …: tool messages (list → read #1 → read #2 → answer).
  • Non-goals: no per-call cost cap, no streaming change, no new endpoint.

Playwright Mapping Rule

Test Scenario → tests/e2e/test_agent_unlimited_tools.py (mock LLM; DB up):

  1. test_multi_read_turn — a grounded question carrying both TOOLS_TRIGGER and the multi-read marker: the turn streams tool lines for list_documents and two read_document calls, then a final answer; the bubble is not deflected.
  2. test_done_sources_include_reads — the final done (observed via the source chips) lists the retrieved doc(s) plus both read documents, deduped.
  3. test_relist_allowed (unit-level via integration, plus E2E observable state) — a re-listed catalog does not produce a refusal line; the UI shows a tool line per executed call.
  4. test_single_tool_flow_regression (phase 37) — the original 3-step flow still answers with exactly one read (runs against the unchanged tests/e2e/test_agent_document_tools.py in the regression pass, not duplicated here).