feat(rag): pass chat history with prior thinking to the LLM
Phase 74 (TODO.md L4): a follow-up question now reaches the model WITH the conversation so far — every prior user/brain turn and the prior thinking blocks on brain turns (preserve-thinking) — while POST /api/chat stays stateless (A10): the client provides the history in the request body and the server stores nothing new. Server (task 01): - ChatRequest.history: optional list[HistoryTurn] (who: user|brain, text, optional thinking) — absent/empty keeps the request byte-identical to pre-phase-74 (the two-message [system, user] request; the kill-switch semantics are pinned in the integration suite). - app.rag.prompts.history_to_messages: pure mapper — walks the turns newest-first against the settings budgets (history_max_turns=40 / history_max_chars=24000, BOR_HISTORY_MAX_TURNS / BOR_HISTORY_MAX_CHARS); a capped turn is dropped WHOLE (never cut mid-answer); the kept window is returned oldest-first; brain turns carry their thinking as reasoning_content (A4) only when non-empty. - Both branches feed it: the deflected path splices it between the system prompt and the current user message (the phase-71 recovery still rebuilds from messages[1:]), the grounded agent receives run_agent(..., history=hist); llm.py's message params widen to list[dict[str, Any]] (string-only messages stay byte-identical on the wire — the SDK passes message dicts through verbatim). - The per-turn log line (PLAN §9) gains history_msgs=N after kb_chars=N. - Pins: tests/unit/test_history.py (mapper: mapping, reasoning gating, both budgets, drop-whole, ordering, empty default), tests/unit/test_config.py (the two settings + env overrides), tests/unit/test_agent.py (the history splice + the default), tests/integration/test_chat_api.py (deflected AND grounded forward the history incl. reasoning_content, no-history byte-identity, 422 pins, the log field). Client (task 02): - runTurn — the single funnel for fresh send / phase-49 retry / phase-53 stale-regen — sends history = the conversation record minus the current question, with thinking only on brain records that streamed one (undefined drops the key from the JSON, the record's convention); the question is never duplicated into the history. Wire proof (task 03): - The mock's echo my history marker (HISTORY_TRIGGER) answers with the deterministic history echo — history: N prior messages; last answer tail: <last 24 chars>; thinking: yes|no — checked BEFORE the DEFLECT_MODE branch (like TABLE_TRIGGER), so it fires on both turn branches whatever the gate says; the module docstring records the user/assistant-only history invariant that keeps every existing (tool-result-classified) marker flow unaffected. - tests/e2e/test_llm_history.py (isolated): a grounded follow-up and a deflected follow-up both receive history: 2 prior messages + thinking: yes + the byte-exact tail of turn 1's answer (derived from the persisted bor.chat.v1 record — the same array the client maps into the body); a cold start receives history: 0 prior messages / last answer tail: none / thinking: no. - Regressions green in isolation: chat_rag, chat_history (phase 50), agent_document_tools, harness_aligned_tools, stop_generation, retry_answer, response_to_docs.
This commit is contained in:
@@ -147,6 +147,48 @@ def test_llm_retry_delay_rejects_negative(monkeypatch: pytest.MonkeyPatch) -> No
|
||||
_settings()
|
||||
|
||||
|
||||
def test_history_budget_defaults(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""Phase 74 (TODO L4): the client-provided history is trimmed
|
||||
newest-first against the newest 40 turns within a total of
|
||||
24 000 chars (text + prior thinking, owner-locked A3)."""
|
||||
monkeypatch.delenv("BOR_HISTORY_MAX_TURNS", raising=False)
|
||||
monkeypatch.delenv("BOR_HISTORY_MAX_CHARS", raising=False)
|
||||
s = _settings()
|
||||
assert s.history_max_turns == 40
|
||||
assert s.history_max_chars == 24_000
|
||||
|
||||
|
||||
def test_history_budget_env_overrides(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""``BOR_HISTORY_MAX_TURNS`` / ``BOR_HISTORY_MAX_CHARS`` override the
|
||||
defaults; ``0`` on either is the no-history kill switch (the
|
||||
pre-phase-74 two-message requests)."""
|
||||
monkeypatch.setenv("BOR_HISTORY_MAX_TURNS", "12")
|
||||
monkeypatch.setenv("BOR_HISTORY_MAX_CHARS", "5000")
|
||||
s = _settings()
|
||||
assert s.history_max_turns == 12
|
||||
assert s.history_max_chars == 5000
|
||||
monkeypatch.setenv("BOR_HISTORY_MAX_TURNS", "0")
|
||||
assert _settings().history_max_turns == 0
|
||||
monkeypatch.setenv("BOR_HISTORY_MAX_CHARS", "0")
|
||||
assert _settings().history_max_chars == 0
|
||||
|
||||
|
||||
def test_history_max_turns_rejects_negative(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""``0`` is the no-history kill switch — a negative value is a typo,
|
||||
so the validator fails loudly at startup (the ``agent_max_rounds``
|
||||
pattern)."""
|
||||
monkeypatch.setenv("BOR_HISTORY_MAX_TURNS", "-1")
|
||||
with pytest.raises(ValidationError, match="history_max_turns"):
|
||||
_settings()
|
||||
|
||||
|
||||
def test_history_max_chars_rejects_negative(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""A negative char budget is a typo — fail loudly at startup."""
|
||||
monkeypatch.setenv("BOR_HISTORY_MAX_CHARS", "-1")
|
||||
with pytest.raises(ValidationError, match="history_max_chars"):
|
||||
_settings()
|
||||
|
||||
|
||||
def test_agent_max_rounds_default_and_env_override(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||
"""Phase 45: the per-tool budgets are gone — ``BOR_AGENT_MAX_ROUNDS``
|
||||
(default 10) is the single agent-loop knob; ``0`` is the no-tools
|
||||
|
||||
Reference in New Issue
Block a user