feat(rag): pass chat history with prior thinking to the LLM
Build and Push Containers / build-and-push-app (push) Successful in 1m39s
Build and Push Containers / build-and-push-db (push) Successful in 11s

Phase 74 (TODO.md L4): a follow-up question now reaches the model WITH
the conversation so far — every prior user/brain turn and the prior
thinking blocks on brain turns (preserve-thinking) — while
POST /api/chat stays stateless (A10): the client provides the history
in the request body and the server stores nothing new.

Server (task 01):
- ChatRequest.history: optional list[HistoryTurn] (who: user|brain,
  text, optional thinking) — absent/empty keeps the request
  byte-identical to pre-phase-74 (the two-message [system, user]
  request; the kill-switch semantics are pinned in the integration
  suite).
- app.rag.prompts.history_to_messages: pure mapper — walks the turns
  newest-first against the settings budgets (history_max_turns=40 /
  history_max_chars=24000, BOR_HISTORY_MAX_TURNS /
  BOR_HISTORY_MAX_CHARS); a capped turn is dropped WHOLE (never cut
  mid-answer); the kept window is returned oldest-first; brain turns
  carry their thinking as reasoning_content (A4) only when
  non-empty.
- Both branches feed it: the deflected path splices it between the
  system prompt and the current user message (the phase-71 recovery
  still rebuilds from messages[1:]), the grounded agent receives
  run_agent(..., history=hist); llm.py's message params widen to
  list[dict[str, Any]] (string-only messages stay byte-identical on
  the wire — the SDK passes message dicts through verbatim).
- The per-turn log line (PLAN §9) gains history_msgs=N after
  kb_chars=N.
- Pins: tests/unit/test_history.py (mapper: mapping, reasoning
  gating, both budgets, drop-whole, ordering, empty default),
  tests/unit/test_config.py (the two settings + env overrides),
  tests/unit/test_agent.py (the history splice + the default),
  tests/integration/test_chat_api.py (deflected AND grounded forward
  the history incl. reasoning_content, no-history byte-identity, 422
  pins, the log field).

Client (task 02):
- runTurn — the single funnel for fresh send / phase-49 retry /
  phase-53 stale-regen — sends history = the conversation record
  minus the current question, with thinking only on brain records
  that streamed one (undefined drops the key from the JSON, the
  record's convention); the question is never duplicated into the
  history.

Wire proof (task 03):
- The mock's echo my history marker (HISTORY_TRIGGER) answers with
  the deterministic history echo — history: N prior messages; last
  answer tail: <last 24 chars>; thinking: yes|no — checked BEFORE
  the DEFLECT_MODE branch (like TABLE_TRIGGER), so it fires on both
  turn branches whatever the gate says; the module docstring records
  the user/assistant-only history invariant that keeps every
  existing (tool-result-classified) marker flow unaffected.
- tests/e2e/test_llm_history.py (isolated): a grounded follow-up and
  a deflected follow-up both receive history: 2 prior messages +
  thinking: yes + the byte-exact tail of turn 1's answer (derived
  from the persisted bor.chat.v1 record — the same array the client
  maps into the body); a cold start receives history: 0 prior
  messages / last answer tail: none / thinking: no.
- Regressions green in isolation: chat_rag, chat_history (phase 50),
  agent_document_tools, harness_aligned_tools, stop_generation,
  retry_answer, response_to_docs.
This commit is contained in:
2026-09-05 16:04:40 -04:00
parent a16130c71d
commit 055c0b5d85
29 changed files with 1418 additions and 18 deletions
+59 -1
View File
@@ -54,10 +54,12 @@ not the wording, so that contract is unchanged.
from __future__ import annotations
from collections.abc import Sequence
from typing import Any
from app.config import get_settings
from app.config import Settings, get_settings
from app.models import Document
from app.rag.retriever import TRUNCATION_MARKER
from app.schemas import HistoryTurn
#: PLAN §6 verbatim (line wrapping included); ``{relevance}`` is filled by
#: :func:`_base`.
@@ -166,6 +168,62 @@ TOOLS_SECTION: str = (
)
def history_to_messages(
history: Sequence[HistoryTurn],
settings: Settings,
) -> list[dict[str, Any]]:
"""Client-provided chat history → model messages (phase 74, TODO L4).
The ``POST /api/chat`` ``history`` (the client's prior turns, oldest
first) becomes the message block that sits between the system prompt
and the current user message — so a follow-up question reaches the
model together with the exchange so far, on BOTH turn branches
(the deflected path and the grounded agent).
Trimming (owner-locked A3, 2026-09-08): the turns are walked
**newest-first** and kept while BOTH budgets hold — the turn count
stays ≤ ``settings.history_max_turns`` and the cumulative chars
(``len(text) + len(thinking or "")`` per turn) stay ≤
``settings.history_max_chars``. A turn that would overflow either
remaining budget is DROPPED WHOLE — never cut mid-answer — and the
walk stops there, so the kept history is always the contiguous
newest window (the oldest turns are the ones dropped; ``0`` on
either budget yields ``[]`` — the pre-phase-74 behavior). The kept
turns are returned in chronological (oldest → newest) order.
Mapping (owner-locked A4, 2026-09-08): ``who="user"`` →
``{"role": "user", "content": text}``; ``who="brain"`` →
``{"role": "assistant", "content": text}`` plus
``"reasoning_content": thinking`` ONLY when *thinking* is
non-empty — the preserve-thinking wire convention
:mod:`app.rag.llm` already reads on the response side
(``delta.reasoning_content``), which is what keeps the owner's
preserve-thinking models carrying the reasoning chain forward.
Pure and side-effect free (no I/O) — unit-testable in isolation.
"""
kept: list[HistoryTurn] = []
chars = 0
for turn in reversed(history):
if len(kept) >= settings.history_max_turns:
break
size = len(turn.text) + len(turn.thinking or "")
if chars + size > settings.history_max_chars:
break
kept.append(turn)
chars += size
messages: list[dict[str, Any]] = []
for turn in reversed(kept):
if turn.who == "user":
messages.append({"role": "user", "content": turn.text})
continue
message: dict[str, Any] = {"role": "assistant", "content": turn.text}
if turn.thinking:
message["reasoning_content"] = turn.thinking
messages.append(message)
return messages
def _base(relevance: str) -> str:
if relevance not in ("HIGH", "LOW"):
raise ValueError(f"relevance must be HIGH or LOW, got {relevance!r}")