Files
brain-of-reese/.agents/phases/complete/45_agent_unlimited_tools/00_phase.md
T
ducoterra dbf2af26c6 refactor(agents): migrate .agent/ planning tree to .agents/
Standardize on the .agents/ directory (shared with project skills):
phases/, user_stories/, reports/, screenshots/, validate.sh, and
phase-sessions/ + pipeline.log all move to .agents/ (git mv preserves
history; runtime artifacts move alongside).

Updates every reference in AGENTS.md, README.md, .gitignore, app
docstrings, and test story headers. Historical KB content in data/
and the runtime pipeline.log transcript are left untouched.
2026-09-05 10:57:07 -04:00

4.6 KiB

Phase 45 — Agent makes as many tool calls as it wants

Source: TODO.md L8 — "Allow the LLM to make as many tool calls as it wants, remove the restrictions, they're causing problems getting correct answers" Story: .agents/user_stories/agent-unlimited-tools.md Context: Phase 37 shipped the grounded-turn agent loop (app/rag/agent.py::run_agent) with per-turn budgets — agent_list_calls / agent_read_calls (default 1 each, BOR_AGENT_LIST_CALLS / BOR_AGENT_READ_CALLS), "budgets-as-kill-switch" locked decision. The exhaustion refusals (LIST_EXHAUSTED / READ_EXHAUSTED) are where correct multi-document answers die. Owner direction (2026-08-27): remove both budgets; the loop keeps one guard — a configurable round cap that also doubles as the no-tools kill switch (0).

Objective

list_documents / read_document can be called as many times as the model needs (re-lists included), bounded only by BOR_AGENT_MAX_ROUNDS (default 10; 0 = no tools, byte-identical to the pre-phase-37 path).

Dependencies

  • 44_markdown_tables (todo) — sequential only (no shared files: this phase is app/ + tests + mock).
  • 37_agent_document_tools (complete) — the loop, the tool SSE event, the UI tool lines, the per-turn tool_calls=N log field, and the phase-37 locked decision being revised.
  • 31_kb_overview_prompt (complete) — the lite one-shot path is untouched by this phase.

Tasks

  1. 01_config_round_cap.md — the server core, atomically: agent_max_rounds replaces the budgets in app/config.py + app/rag/agent.py, unit + integration rewrites, .env.example (one task so the per-task gate stays green).
  2. 02_mock_multi_read_flow.md — the E2E mock's deterministic multi-read (list → read #1 → read #2 → answer) flow.
  3. 03_unlimited_tools_e2e_and_commit.md — story E2E + phase-37 regression + PLAN.md revision note + commit.

Testing & Quality

  • Unit: tests/unit/test_agent.py rewritten around the round cap (always-calling mock LLM: N tool rounds then a forced tools=None final answer; max_rounds=0 → exactly one request with tools=None; rejected-call spam — unknown tool / already-in-context — is bounded by the cap, not by budgets; re-lists execute and count in tool_calls); tests/unit/test_config.py (default 10, BOR_AGENT_MAX_ROUNDS override, 0, the budget env vars are gone).
  • Integration: tests/integration/test_chat_api.py — the agent_list_calls=0, agent_read_calls=0 fixtures become agent_max_rounds=0; the tool SSE event shape and the done.sources extension assertions stay.
  • Coverage: >90% on app/ — agent.py + config.py fully covered.
  • E2E (mandatory, A16): tests/e2e/test_agent_unlimited_tools.py, run in isolation.

Completion Criteria

  • BOR_AGENT_LIST_CALLS / BOR_AGENT_READ_CALLS are gone (config, .env.example, agent, tests); no exhaustion refusal remains.
  • A multi-read turn (list + 2 reads) streams three tool lines, answers non-deflected, and done.sources lists the retrieved doc(s) + both reads deduped.
  • agent_max_rounds=0 → single tools=None request (kill switch); at the cap the loop forces a final no-tools answer (log warning kept).
  • tool SSE event shape and tool_calls=N per-turn log field unchanged.
  • .agents/PLAN.md carries the phase-37 revision note (owner permission 2026-08-27, TODO.md L8) — the only PLAN edit in this phase.
  • uv run pytest green; uv run pytest --cov=app --cov-report=term-missing >90%.
  • uv run pytest tests/e2e/test_agent_unlimited_tools.py -v --no-cov green in isolation (DB up).
  • Regression E2E suites green in isolation: test_agent_document_tools.py, test_chat_rag.py, test_smoke.py.
  • uv run ruff check . && uv run pyright clean.
  • One --no-gpg-sign commit; phase dir moved to .agents/phases/complete/.

Locked decisions

  • Owner-locked revision (2026-08-27, roadmap R2): the phase-37 "budgets-as-kill-switch" decision is revised — both per-tool budgets removed; BOR_AGENT_MAX_ROUNDS (default 10) is the only loop guard and the kill switch (0). Recorded as a PLAN.md revision note (the established owner-permission pattern, like the A10/A7/A9/A15 notes) — a recorded revision, not a silent deviation (AGENTS.md rule 3).
  • A15 extension unchanged — the tool SSE event shape, the done shape, and the per-turn log line (tool_calls=N) are untouched; the revision note amends the phase-37 note's budget wording only.
  • Rejections kept: Unknown tool., MISSING_READ_ARGS, Already in your context. (non-budget rejections; the cap bounds their pathological repetition).
  • A17 honoured — one atomic commit.