Files
brain-of-reese/app/config.py
T
ducoterra b855d0aef9 feat(rag): unbounded agent tool calls behind a round cap (owner revision)
Phase 45 (owner permission 2026-08-27, TODO.md L8: "allow the LLM
to make as many tool calls as it wants"): the phase-37 per-turn tool
budgets (BOR_AGENT_LIST_CALLS / BOR_AGENT_READ_CALLS, default 1 each)
and their exhaustion refusals are removed — a grounded turn now offers
list_documents / read_document for the whole turn (re-lists included),
bounded only by the round cap:

- app/config.py: agent_max_rounds (BOR_AGENT_MAX_ROUNDS, default 10,
  negative rejected) replaces agent_list_calls / agent_read_calls;
  .env.example + README document the single knob; app/rag/prompts.py
  docstrings follow.
- app/rag/agent.py: the loop runs tools until the model answers or
  rounds >= max_rounds, at which point it forces one final no-tools
  answer (the cap is the only forced exit); 0 = no tools — exactly one
  tools=None request, byte-identical to the pre-phase-37 path (the
  kill switch). Rejected calls (unknown tool / missing args /
  already-in-context / unknown path) still consume a round, so
  pathological rejected-call streams are bounded by the cap. The
  per-call log line is now tool/args/round=N/M; the per-turn
  tool_calls=N field and the tool SSE event are unchanged.
- tests/e2e/mock_llm.py: MULTI_READ_TRIGGER ("read two documents") —
  the deterministic list -> read #1 -> read #2 -> forced-answer flow
  (byte-stable "I read <sp1> and <sp2>." line), classified by the
  count of tool-role read results; the phase-37 single-read flow stays
  byte-identical (unit-pinned in tests/unit/test_mock_tool_flow.py).
- tests/e2e/test_agent_unlimited_tools.py (new, story suite,
  mock-only): three tool frames/lines in order (one list, two reads —
  the second read is what the old read budget refused) + the
  both-named non-deflected answer; done.sources + chips = retrieval
  doc + both reads, deduped; no budget refusal rendered; the
  single-read marker flow regression (exactly one read, single tool
  pair).
- .agent/PLAN.md: the phase-45 SSE revision note (owner-locked, R2) —
  the only PLAN edit this phase; the phase-37 note's budget clause is
  marked removed.

Unit/integration rewrites (test_agent.py round-cap matrix incl. the
kill switch and rejected-call spam, test_config.py, test_chat_api.py
agent_max_rounds=0 fixtures) landed with the server core so every gate
stays green.

uv run pytest: 756 passed, app/ coverage 99%; ruff + pyright clean;
story E2E 4/4 in isolation (ran twice); regression E2E suites
(agent_document_tools unmodified, chat_rag, smoke) green in isolation.

Also records the 45_agent_unlimited_tools todo/ -> complete/ task-file
moves (00/01/02 pending in the working tree, task 03 moves on success).
2026-08-28 04:50:56 -04:00

208 lines
9.4 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Application settings.
Every setting can be overridden with an environment variable prefixed
``BOR_`` (or a local gitignored ``.env`` file — see ``.env.example``).
"""
from __future__ import annotations
import os
from functools import lru_cache
from pydantic import field_validator
from pydantic_settings import BaseSettings, SettingsConfigDict
#: The A9 import formats (PLAN anchor A9, revised 2026-08-21).
#: ``BOR_IMPORT_EXTENSIONS`` may narrow — but never widen — this set.
_ALLOWED_IMPORT_EXTENSIONS: frozenset[str] = frozenset(
{"md", "markdown", "txt", "yaml", "yml", "json", "py"}
)
class Settings(BaseSettings):
model_config = SettingsConfigDict(
env_file=".env",
env_file_encoding="utf-8",
env_prefix="BOR_",
extra="ignore",
)
# --- App ---
app_name: str = "Brain of Reese"
app_version: str = "0.1.0"
environment: str = "development"
log_level: str = "INFO"
static_dir: str = "frontend"
# --- Database (PostgreSQL 17 + pgvector) ---
database_url: str = "postgresql+psycopg://reese:reese@localhost:5432/brain_of_reese"
# --- LLM (self-hosted, OpenAI-compatible "aipi" endpoint) ---
llm_base_url: str = "https://aipi.reeseapps.com/v1"
llm_api_key: str = ""
llm_chat_model: str = "turbo"
llm_embed_model: str = "embed"
#: One-shot (non-streaming) completion model (A5 extended, phase 30):
#: document summaries at import time and the KB overview (phase 31).
#: Served by the same OpenAI-compatible endpoint — no new model
#: management. Called via ``LLMClient.chat()``.
llm_summary_model: str = "lite"
#: Operator kill-switch for the ``thinking`` SSE events (phase 17,
#: ``BOR_STREAM_THINKING``; ``0``/``false`` → off). When off, thinking
#: pieces are still counted for the per-turn log line but never
#: emitted — the answer stream itself is unchanged.
stream_thinking: bool = True
# --- RAG tuning ---
embedding_dim: int = 768 # verified against aipi /v1 (embed model)
top_n_docs: int = 2
# Honesty gate (A8, re-tuned 2026-08-21): the ``embed`` model's cosine
# scores compress into 0.41–0.84 on the real corpus, so the old 0.30
# default never discriminated. LOW only fires when best cosine < this
# AND no candidate chunk matches the question lexically (see A8).
relevance_threshold: float = 0.62
#: Maximum output tokens a chat answer may use (owner instruction
#: 2026-08-22: answers must run to their natural end — the old hard
#: 700-token cap cut long answers off mid-sentence).
max_output_tokens: int = 32_768
chunk_target_chars: int = 2_000
chunk_overlap_chars: int = 200
embed_batch_size: int = 16
#: Total char budget for the ``<tuning>`` section of the system prompt
#: (phase 15, steering notes). The newest-fitting notes are kept and the
#: overflow is replaced by the ``[…truncated…]`` marker.
steering_max_chars: int = 8_000
#: Cap on the document content sent to the ``lite`` summary model in one
#: call (phase 30, ``BOR_SUMMARY_MAX_CHARS``). Overflow is cut at the cap
#: and the shared ``[…truncated…]`` marker is appended (see
#: ``app.rag.summarizer``).
summary_max_chars: int = 12_000
#: Char budget for the ``<knowledge_base>`` section of the system prompt
#: (phase 31: lite-generated KB overview, ``app.rag.overview``). The
#: newest-fitting prefix of the stored outline is kept and the overflow
#: is replaced by the shared ``[…truncated…]`` marker (phase 15
#: convention — ``app.rag.prompts``).
kb_overview_max_chars: int = 4_000
#: Cap on the document list (source/path/title/first summary line per
#: row) sent to the ``lite`` overview model in one call (phase 31,
#: ``app.rag.overview``). Overflow is cut at the cap and the shared
#: ``[…truncated…]`` marker is appended (summarizer convention).
overview_input_max_chars: int = 40_000
#: Hard cap on the agent tool rounds per grounded turn (phase 45,
#: revising phase 37's per-tool budgets — owner permission
#: 2026-08-27, TODO L8: "allow the LLM to make as many tool calls
#: as it wants"). Every tool call the model emits consumes a
#: round; at the cap the loop forces one final no-tools answer.
#: ``0`` disables the tools entirely — the turn is a single
#: request with ``tools=None`` (the pre-phase-37 path — the kill
#: switch). Negative values are rejected at startup (validator).
agent_max_rounds: int = 10
# --- Hybrid retrieval (A7, revised 2026-08-21) ---
# cosine top-N ∪ Postgres FTS top-N, fused with Reciprocal Rank Fusion
# (score = Σ 1/(rrf_k + rank) over the lists a chunk appears in).
#
# The vector window is deliberately wider than the lexical one: a
# name-your-tool question's best *lexical* chunk (e.g. the "Install"
# section of gitlab.md) can sit far down the vector ranking because the
# question embeds close to generic templates. A 100-wide window is what
# lets such chunks double-hit (one RRF term per list) and outrank a
# template that owns vector rank 1 — measured 2026-08-22 against the
# live 2774-chunk KB for "How did I install gitlab?" (gitlab.md:1 at
# vrank 100 / lrank 3 → fused 0.0221 vs the template's 0.0164).
hybrid_vector_candidates: int = 100
hybrid_lexical_candidates: int = 30
rrf_k: int = 60
# --- Admin & sign-in (phase 16; A10 revised 2026-08-22) ---
# Single-admin auth via a signed session cookie (Starlette
# SessionMiddleware — no new services, no DB tables). Both secrets are
# REQUIRED at startup: ``create_app()`` refuses to boot when either is
# empty (``app.core.auth.ensure_admin_configured``). The password is
# plaintext on purpose (homelab scope, owner decision 2026-08-22);
# the session secret signs the cookie (``secrets.token_hex(32)``).
admin_password: str = ""
session_secret: str = ""
#: Signed-cookie lifetime in seconds (default 12 h, refreshed on
#: session writes — sliding for an active admin).
session_max_age: int = 43_200
session_cookie: str = "bor_session"
# --- Import scope (A9, revised 2026-08-21) ---
# Comma-separated list of lowercased file extensions (no dot) imported
# by ``scripts/import_docs.py``. Hidden (dot) path components are always
# skipped, plus the importer's exclusion list.
# Stored as a raw CSV string (env-native — no JSON) and parsed on demand
# via :py:meth:`import_extension_set`. ``mode="after"`` validation runs
# against the raw string so a typo fails loudly at startup.
import_extensions: str = "md,markdown,txt,yaml,yml,json,py"
#: List of git repo URLs to clone/pull into ``sources_dir`` before
#: indexing (phase 28); comma-separated, stored raw. Empty means no git
#: sources — ``import_docs`` then falls back to ``--source`` / the old
#: ``DEFAULT_SOURCES``.
git_sources: str = ""
#: Where ``import_docs`` clones/pulls the ``git_sources`` repos (phase
#: 28). Stored as a raw string — ``Path.expanduser()`` is applied in
#: the import script, not here.
sources_dir: str = "~/bor-sources"
@field_validator("import_extensions")
@classmethod
def _import_extensions_known(cls, v: str) -> str:
"""Reject unknown/empty formats loudly instead of silently importing
nothing (a typo like ``md,jsonn`` would otherwise walk zero files)."""
exts = {part.strip().lstrip(".").lower() for part in v.split(",") if part.strip()}
if not exts:
raise ValueError("import_extensions must name at least one format")
unknown = exts - _ALLOWED_IMPORT_EXTENSIONS
if unknown:
raise ValueError(
f"unknown import extension(s): {', '.join(sorted(unknown))} — "
f"allowed: {', '.join(sorted(_ALLOWED_IMPORT_EXTENSIONS))}"
)
return v
@field_validator("agent_max_rounds")
@classmethod
def _agent_max_rounds_non_negative(cls, v: int) -> int:
"""``0`` is the no-tools kill switch — a negative value is a typo."""
if v < 0:
raise ValueError("agent_max_rounds must be >= 0 (0 = no tools)")
return v
# Suggested questions (onboarding + empty state).
suggestions: list[str] = [
"How is my Kubernetes cluster set up?",
"What's my backup strategy?",
"How do I deploy a new service?",
"What's currently running in the homelab?",
]
@property
def import_extension_set(self) -> frozenset[str]:
"""Lowercased, dotted extension set (``.md``) for path filtering."""
return frozenset(
f".{part.strip().lstrip('.').lower()}"
for part in self.import_extensions.split(",")
if part.strip()
)
@property
def git_source_list(self) -> list[str]:
"""Non-empty, stripped git URLs from :py:attr:`git_sources` (phase 28).
Whitespace around each entry is trimmed and empty entries dropped;
an unset/empty value yields ``[]`` (the import script then uses its
legacy local-directory defaults).
"""
return [part.strip() for part in self.git_sources.split(",") if part.strip()]
@property
def effective_api_key(self) -> str:
"""API key for aipi: explicit setting, then $AIPI_KEY, then a placeholder."""
return self.llm_api_key or os.environ.get("AIPI_KEY", "") or "not-needed"
@lru_cache
def get_settings() -> Settings:
return Settings()