feat(rag): pass chat history with prior thinking to the LLM
Build and Push Containers / build-and-push-app (push) Successful in 1m39s
Build and Push Containers / build-and-push-db (push) Successful in 11s

Phase 74 (TODO.md L4): a follow-up question now reaches the model WITH
the conversation so far — every prior user/brain turn and the prior
thinking blocks on brain turns (preserve-thinking) — while
POST /api/chat stays stateless (A10): the client provides the history
in the request body and the server stores nothing new.

Server (task 01):
- ChatRequest.history: optional list[HistoryTurn] (who: user|brain,
  text, optional thinking) — absent/empty keeps the request
  byte-identical to pre-phase-74 (the two-message [system, user]
  request; the kill-switch semantics are pinned in the integration
  suite).
- app.rag.prompts.history_to_messages: pure mapper — walks the turns
  newest-first against the settings budgets (history_max_turns=40 /
  history_max_chars=24000, BOR_HISTORY_MAX_TURNS /
  BOR_HISTORY_MAX_CHARS); a capped turn is dropped WHOLE (never cut
  mid-answer); the kept window is returned oldest-first; brain turns
  carry their thinking as reasoning_content (A4) only when
  non-empty.
- Both branches feed it: the deflected path splices it between the
  system prompt and the current user message (the phase-71 recovery
  still rebuilds from messages[1:]), the grounded agent receives
  run_agent(..., history=hist); llm.py's message params widen to
  list[dict[str, Any]] (string-only messages stay byte-identical on
  the wire — the SDK passes message dicts through verbatim).
- The per-turn log line (PLAN §9) gains history_msgs=N after
  kb_chars=N.
- Pins: tests/unit/test_history.py (mapper: mapping, reasoning
  gating, both budgets, drop-whole, ordering, empty default),
  tests/unit/test_config.py (the two settings + env overrides),
  tests/unit/test_agent.py (the history splice + the default),
  tests/integration/test_chat_api.py (deflected AND grounded forward
  the history incl. reasoning_content, no-history byte-identity, 422
  pins, the log field).

Client (task 02):
- runTurn — the single funnel for fresh send / phase-49 retry /
  phase-53 stale-regen — sends history = the conversation record
  minus the current question, with thinking only on brain records
  that streamed one (undefined drops the key from the JSON, the
  record's convention); the question is never duplicated into the
  history.

Wire proof (task 03):
- The mock's echo my history marker (HISTORY_TRIGGER) answers with
  the deterministic history echo — history: N prior messages; last
  answer tail: <last 24 chars>; thinking: yes|no — checked BEFORE
  the DEFLECT_MODE branch (like TABLE_TRIGGER), so it fires on both
  turn branches whatever the gate says; the module docstring records
  the user/assistant-only history invariant that keeps every
  existing (tool-result-classified) marker flow unaffected.
- tests/e2e/test_llm_history.py (isolated): a grounded follow-up and
  a deflected follow-up both receive history: 2 prior messages +
  thinking: yes + the byte-exact tail of turn 1's answer (derived
  from the persisted bor.chat.v1 record — the same array the client
  maps into the body); a cold start receives history: 0 prior
  messages / last answer tail: none / thinking: no.
- Regressions green in isolation: chat_rag, chat_history (phase 50),
  agent_document_tools, harness_aligned_tools, stop_generation,
  retry_answer, response_to_docs.
This commit is contained in:
2026-09-05 16:04:40 -04:00
parent a16130c71d
commit 055c0b5d85
29 changed files with 1418 additions and 18 deletions
@@ -0,0 +1,18 @@
All criteria verified — no defects found, no fixes needed. Final report:
**Phase 73 (hidden_tab_stream) — final verification pass: ALL GREEN**
- Verified C1 fix: `pagehide` partial is correlated via `leavePartialIndex`; done/fallback/stop settles *replace* it in place (identity-guarded `rememberBrainTurn` in-place mode) — `frontend/assets/app.js`
- Verified C2 fix: `visibilitychange` re-arms the still-armed 120s pre-token guard on return to visible (hidden time no longer counts)
- Verified phase-48 teardown untouched: Stop / tab close / navigation still abort (`turnAbort`, `cancelStream`)
- Commit `a16130c fix(chat): keep generating while the tab is hidden`: unsigned (`--no-gpg-sign` ✓), body carries the root-cause line ("none of C1-C3 — Chromium 151 real mode: hidden tab neither stops the stream nor fires pagehide; C1 latent via synthetic pagehide; C2 confirmed to fire while hidden")
**Test / lint / coverage (exact commands):**
- `uv run pytest tests/e2e/test_hidden_tab_stream.py -v --no-cov` → 4 passed (36.8s)
- Isolation regressions: `test_stop_generation` 3 passed · `test_chat_persistence` 4 passed · `test_chat_rag` 3 passed · `test_chat_history` 5 passed · `test_retry_answer` 4 passed
- `uv run pytest --cov=app --cov-report=term-missing` → 1519 passed, `app/` coverage **99%** (>90% ✓)
- `uv run ruff check . && uv run pyright` → clean (0 errors)
**Completion criteria:** root-cause line in commit ✓ · phase E2E green in isolation (exactly-one-brain-turn pin) ✓ · real departures unchanged (stop/persistence green) ✓ · full suite + >90% + ruff/pyright ✓ · atomic unsigned Conventional-Commits commit; task files 01–03 in `complete/` (00_phase.md move is the harness's on success) ✓
**Deviations:** none. **Next pending phase:** `74_llm_chat_history`.
@@ -0,0 +1,78 @@
........................................................................ [ 4%]
........................................................................ [ 9%]
........................................................................ [ 14%]
........................................................................ [ 18%]
........................................................................ [ 23%]
........................................................................ [ 28%]
........................................................................ [ 33%]
........................................................................ [ 37%]
........................................................................ [ 42%]
........................................................................ [ 47%]
........................................................................ [ 52%]
........................................................................ [ 56%]
........................................................................ [ 61%]
........................................................................ [ 66%]
........................................................................ [ 71%]
........................................................................ [ 75%]
........................................................................ [ 80%]
........................................................................ [ 85%]
........................................................................ [ 90%]
........................................................................ [ 94%]
........................................................................ [ 99%]
....... [100%]
=============================== warnings summary ===============================
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
Name Stmts Miss Cover
-----------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 22 0 100%
app/api/chat.py 176 0 100%
app/api/chats.py 110 0 100%
app/api/config.py 7 0 100%
app/api/doc_drafts.py 93 0 100%
app/api/docs.py 50 0 100%
app/api/git_sources.py 212 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 8 0 100%
app/api/sync.py 101 0 100%
app/config.py 126 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 20 0 100%
app/core/caching.py 108 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/logging.py 13 0 100%
app/db.py 21 0 100%
app/main.py 52 0 100%
app/models.py 86 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 222 0 100%
app/rag/archive_upload.py 128 0 100%
app/rag/chunker.py 206 4 98%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 180 3 98%
app/rag/llm.py 216 0 100%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 65 0 100%
app/rag/retriever.py 150 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 205 0 100%
-----------------------------------------------
TOTAL 2946 12 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
validation OK
@@ -0,0 +1,18 @@
Task 03 complete — all gates green, phase committed.
**Implemented**
- `tests/e2e/test_hidden_tab_stream.py` (4 tests): synthetic `pagehide`+hidden `visibilityState` mid-stream → answer completes exactly once, `bor.chat.v1` holds exactly ONE brain turn (partial read pinned pre-settle, full byte-exact post-settle), auto-saved `saved_chats` row carries the same single turn (admin context), `query_log==1` (no phase-48 cancel); reload restores one brain bubble; no-event baseline; fake-clock pre-token guard re-arm test.
- Both pins verified discriminating: with the C1 replace disabled the main test catches `['user','brain','brain']`; with the re-arm disabled the C2 test catches the "stuck" banner. Fixtures restored byte-identical afterwards.
**Results (exact commands)**
- `uv run pytest tests/e2e/test_hidden_tab_stream.py -v --no-cov` → 4 passed
- Regressions in isolation: `test_chat_rag.py` 3✓, `test_chat_persistence.py` 4✓, `test_chat_history.py` 5✓, `test_stop_generation.py` 3✓, `test_retry_answer.py` 4✓
- `uv run pytest --cov=app --cov-report=term-missing` → 1519 passed, app/ 99% (>90%)
- `uv run ruff check . && uv run pyright` → clean, 0 errors
- Committed `a16130c` `fix(chat): keep generating while the tab is hidden` (`--no-gpg-sign`), task-01 root-cause line in the body; recorded the harness's 01/02 move + task-01 evidence.
**Decisions**
- C2 test uses `pretend to think slowly` (pure 3s pre-token silence) + Playwright fake clock: the task's suggested `think out loud then hesitate` streams scratchpad first, clearing the guard before its 4s pause — it cannot exercise the re-arm.
- Did not move `todo/73_hidden_tab_stream/` (harness rule; it holds `00_phase.md` + this task file for the harness to move on success).
**Next pending task:** `.agents/phases/todo/74_llm_chat_history/01_server_history_ingestion.md`
@@ -0,0 +1,78 @@
........................................................................ [ 4%]
........................................................................ [ 9%]
........................................................................ [ 14%]
........................................................................ [ 18%]
........................................................................ [ 23%]
........................................................................ [ 28%]
........................................................................ [ 33%]
........................................................................ [ 37%]
........................................................................ [ 42%]
........................................................................ [ 47%]
........................................................................ [ 52%]
........................................................................ [ 56%]
........................................................................ [ 61%]
........................................................................ [ 66%]
........................................................................ [ 71%]
........................................................................ [ 75%]
........................................................................ [ 80%]
........................................................................ [ 85%]
........................................................................ [ 90%]
........................................................................ [ 94%]
........................................................................ [ 99%]
....... [100%]
=============================== warnings summary ===============================
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
Name Stmts Miss Cover
-----------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 22 0 100%
app/api/chat.py 176 0 100%
app/api/chats.py 110 0 100%
app/api/config.py 7 0 100%
app/api/doc_drafts.py 93 0 100%
app/api/docs.py 50 0 100%
app/api/git_sources.py 212 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 8 0 100%
app/api/sync.py 101 0 100%
app/config.py 126 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 20 0 100%
app/core/caching.py 108 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/logging.py 13 0 100%
app/db.py 21 0 100%
app/main.py 52 0 100%
app/models.py 86 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 222 0 100%
app/rag/archive_upload.py 128 0 100%
app/rag/chunker.py 206 4 98%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 180 3 98%
app/rag/llm.py 216 0 100%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 65 0 100%
app/rag/retriever.py 150 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 205 0 100%
-----------------------------------------------
TOTAL 2946 12 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
validation OK
@@ -0,0 +1,16 @@
All done — full gate green. Final report:
**Implemented (phase 74, task 01 — server-side history ingestion):**
- `app/config.py`: `history_max_turns=40` / `history_max_chars=24000` (`BOR_`-overridable, fail-loud validators on negatives, house pattern)
- `app/schemas.py`: `HistoryTurn` + `ChatRequest.history` (default `[]`, schema ceiling 100)
- `app/rag/prompts.py`: pure `history_to_messages()` — newest-first walk, both budgets, drop-whole (contiguous newest window), chronological output, `reasoning_content` only when thinking non-empty
- `app/api/chat.py`: `hist` mapped once per turn; deflected branch splices it into `messages`; grounded passes `history=hist` to `run_agent`; log line gains `history_msgs=N` after `kb_chars=N`
- `app/rag/agent.py`: `run_agent(..., history=())` spliced between system and current user; obsolete `cast`s dropped; tool rounds / recovery / retries unchanged
- `app/rag/llm.py`: `chat`/`chat_stream`/`chat_stream_retried` widened to `list[dict[str, Any]]` (no transport change)
- Tests: 16 new in `tests/unit/test_history.py`, 4 in `test_config.py`, 3 in `test_agent.py`, 5 in `tests/integration/test_chat_api.py` (deflected + grounded forwarding incl. `reasoning_content`, no-history byte-identical pin, 422s, caplog `history_msgs`)
**Results:** `uv run pytest` → 1547 passed (baseline 1519); `uv run pytest --cov=app` → TOTAL 99% (>90%, all touched files 100%); `uv run ruff check .` → clean; `uv run pyright` → 0 errors; `uv run .agents/validate.sh` → exit 0.
**Notes:** `.agents/PLAN.md` is absent from the repo — worked from the phase overview + task spec (A2/A3/A4 assumptions). Trim semantics implemented as contiguous newest window (walk stops at first non-fitting turn). No commit, per task (task 03 commits).
**Next pending task:** `.agents/phases/todo/74_llm_chat_history/02_client_sends_history.md`
@@ -0,0 +1,78 @@
........................................................................ [ 4%]
........................................................................ [ 9%]
........................................................................ [ 13%]
........................................................................ [ 18%]
........................................................................ [ 23%]
........................................................................ [ 27%]
........................................................................ [ 32%]
........................................................................ [ 37%]
........................................................................ [ 41%]
........................................................................ [ 46%]
........................................................................ [ 51%]
........................................................................ [ 55%]
........................................................................ [ 60%]
........................................................................ [ 65%]
........................................................................ [ 69%]
........................................................................ [ 74%]
........................................................................ [ 79%]
........................................................................ [ 83%]
........................................................................ [ 88%]
........................................................................ [ 93%]
........................................................................ [ 97%]
................................... [100%]
=============================== warnings summary ===============================
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
Name Stmts Miss Cover
-----------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 22 0 100%
app/api/chat.py 177 0 100%
app/api/chats.py 110 0 100%
app/api/config.py 7 0 100%
app/api/doc_drafts.py 93 0 100%
app/api/docs.py 50 0 100%
app/api/git_sources.py 212 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 8 0 100%
app/api/sync.py 101 0 100%
app/config.py 140 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 20 0 100%
app/core/caching.py 108 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/logging.py 13 0 100%
app/db.py 21 0 100%
app/main.py 52 0 100%
app/models.py 86 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 222 0 100%
app/rag/archive_upload.py 128 0 100%
app/rag/chunker.py 206 4 98%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 180 3 98%
app/rag/llm.py 216 0 100%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 88 0 100%
app/rag/retriever.py 150 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 210 0 100%
-----------------------------------------------
TOTAL 2989 12 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
validation OK
@@ -0,0 +1,19 @@
All gates green and the diff is exactly the specified change. Final report:
## Task 02 complete — client sends conversation history (with thinking)
**Implemented**
- `frontend/assets/app.js` `runTurn` (single funnel for fresh send / phase-49 retry / phase-53 stale-regen): builds `history = conversation.slice(0, -1).map(...)` carrying exactly `{who, text, thinking}` (`thinking` only on brain records that streamed one — `undefined` drops the key) and sends `body: JSON.stringify({ message: text, history })`.
- Verified the invariant: in all three paths the sent question is the LAST `conversation` entry at fetch time (fresh send pushes it; retry/stale-regen pop the old answer and keep the question last).
- No UI change; no other files touched by this task (task 01's server work remains uncommitted in-tree for task 03's phase commit).
**Verification**
- Wire check (throwaway Playwright script, run against real app + deterministic mock, then deleted): 10/10 PASS — turn 1 `history: []`; turn 2 history = first Q/A pair with 2.7k-char `thinking` on the brain entry and no extra keys; retry's history byte-identical to turn 2 (popped answer excluded, question not duplicated). Server logs show `history_msgs=0 → 2 → 2`.
- `uv run pytest` → 1547 passed
- `uv run pytest --cov=app --cov-report=term-missing` → 99% (>90%)
- `uv run ruff check . && uv run pyright` → clean (0 errors)
- E2E regressions in isolation: `test_chat_rag.py` 3 passed, `test_retry_answer.py` 4 passed, `test_thinking_display.py` 5 passed
**Decisions** — none; change matches the task spec verbatim. No commit (task 03 owns the phase commit).
**Next pending task:** `74_llm_chat_history/03_mock_marker_e2e.md` (mock `echo my history` marker + story E2E + commit).
@@ -0,0 +1,78 @@
........................................................................ [ 4%]
........................................................................ [ 9%]
........................................................................ [ 13%]
........................................................................ [ 18%]
........................................................................ [ 23%]
........................................................................ [ 27%]
........................................................................ [ 32%]
........................................................................ [ 37%]
........................................................................ [ 41%]
........................................................................ [ 46%]
........................................................................ [ 51%]
........................................................................ [ 55%]
........................................................................ [ 60%]
........................................................................ [ 65%]
........................................................................ [ 69%]
........................................................................ [ 74%]
........................................................................ [ 79%]
........................................................................ [ 83%]
........................................................................ [ 88%]
........................................................................ [ 93%]
........................................................................ [ 97%]
................................... [100%]
=============================== warnings summary ===============================
.venv/lib/python3.13/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain_of_reese/.venv/lib/python3.13/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.13.13-final-0 _______________
Name Stmts Miss Cover
-----------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 22 0 100%
app/api/chat.py 177 0 100%
app/api/chats.py 110 0 100%
app/api/config.py 7 0 100%
app/api/doc_drafts.py 93 0 100%
app/api/docs.py 50 0 100%
app/api/git_sources.py 212 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 8 0 100%
app/api/sync.py 101 0 100%
app/config.py 140 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 20 0 100%
app/core/caching.py 108 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/logging.py 13 0 100%
app/db.py 21 0 100%
app/main.py 52 0 100%
app/models.py 86 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 222 0 100%
app/rag/archive_upload.py 128 0 100%
app/rag/chunker.py 206 4 98%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 180 3 98%
app/rag/llm.py 216 0 100%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 88 0 100%
app/rag/retriever.py 150 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 210 0 100%
-----------------------------------------------
TOTAL 2989 12 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
validation OK
+32 -4
View File
@@ -115,6 +115,23 @@ recovery, via the holder; deflected: this turn's filters), 0 on clean
turns (the field is uniform, the phase-67 ``retries=N`` pattern); the
recovery does not bump ``retries=N`` (it is not a phase-67
endpoint-retry).
Chat history (phase 74, TODO L4, owner-locked A2/A3/A4 2026-09-08):
``POST /api/chat`` accepts an optional ``history`` — the client's prior
turns, oldest first (the ``bor.chat.v1`` record minus the current
question; the endpoint stays stateless per A10 — nothing is stored). It
is mapped ONCE per turn by :func:`app.rag.prompts.history_to_messages`
— trimmed newest-first against the settings budgets
(``history_max_turns`` / ``history_max_chars``; a capped-out turn is
dropped whole, never truncated) — and fed to the model on BOTH turn
branches: the deflected path splices it between the system prompt and
the current user message (the phase-71 recovery still rebuilds from
``messages[1:]`` — unchanged), and the grounded agent receives it as
``run_agent(..., history=hist)``. Prior brain turns' thinking travels
as ``reasoning_content`` on the assistant message (the preserve-
thinking wire convention, A4). The per-turn log line records
``history_msgs=N`` after ``kb_chars=N`` (0 when the request carries no
history — the two-message request stays byte-identical).
"""
from __future__ import annotations
@@ -150,7 +167,7 @@ from app.rag.llm import (
chat_stream_retried, # phase 67: the retry-before-first-piece primitive
)
from app.rag.overview import load_kb_overview
from app.rag.prompts import build_deflect_prompt, build_high_prompt
from app.rag.prompts import build_deflect_prompt, build_high_prompt, history_to_messages
from app.rag.retriever import RetrievedChunk, retrieve, select_documents, weak_hit_titles
from app.rag.scaffolding import ScaffoldingFilter # phase 71: the streaming filter
from app.rag.suggestions import derive_suggestions
@@ -306,6 +323,14 @@ async def chat(
retries_used = 0 # phase 67: LLM requests restarted this turn (log line)
try:
settings = get_settings()
# Phase 74 (TODO L4): the client's prior turns, mapped ONCE
# per turn — trimmed newest-first against the settings
# budgets, assistant turns carrying their prior thinking as
# ``reasoning_content`` (A4). BOTH branches below (deflected
# + grounded agent) reuse the same block; an absent/empty
# history yields ``[]`` (the byte-identical two-message
# request, A2).
hist = history_to_messages(request.history, settings)
# 1. Embed the question.
# Phase 67: a dead embeddings endpoint is retried before any
@@ -384,8 +409,9 @@ async def chat(
).model_dump()
)
return
messages = [
messages: list[dict[str, Any]] = [
{"role": "system", "content": plan.system_prompt},
*hist, # phase 74: the trimmed prior turns (empty by default)
{"role": "user", "content": request.message},
]
@@ -434,6 +460,7 @@ async def chat(
seed_docs=plan.docs,
settings=settings,
holder=holder,
history=hist, # phase 74: the same trimmed prior turns
)
thinking_chars = 0
content_chars = 0 # phase 71: the turn's visible (clean) content
@@ -636,8 +663,8 @@ async def chat(
logger.info(
"question=%r embed_ms=%d top_score=%.3f fts_hits=%d summary_hits=%d tuning=%d "
"kb_chars=%d threshold=%.2f deflected=%s sources=%r thinking_chars=%d "
"tool_calls=%d total_ms=%d retries=%d scaffold_stripped=%d",
"kb_chars=%d history_msgs=%d threshold=%.2f deflected=%s sources=%r "
"thinking_chars=%d tool_calls=%d total_ms=%d retries=%d scaffold_stripped=%d",
request.message,
embed_ms,
plan.top_score,
@@ -645,6 +672,7 @@ async def chat(
plan.summary_hits,
plan.tuning_count,
plan.kb_chars,
len(hist),
settings.relevance_threshold,
plan.deflected,
source_paths,
+35
View File
@@ -80,6 +80,21 @@ class Settings(BaseSettings):
#: Flat seconds to wait between attempts (phase 67,
#: ``BOR_LLM_RETRY_DELAY``); the TODO-locked 5 s, no backoff.
llm_retry_delay: float = 5.0
# --- Chat history (phase 74, TODO L4: prior turns + prior thinking) ---
#: Newest client-provided history turns kept per ``POST /api/chat``
#: (phase 74, ``BOR_HISTORY_MAX_TURNS``): the request's ``history``
#: (the client's prior turns, stateless per A10) is walked
#: newest-first and the walk stops once this many turns are kept —
#: the oldest turns are the ones dropped. ``0`` = no history (the
#: pre-phase-74 two-message requests — the kill switch).
history_max_turns: int = 40
#: Total char budget for the kept history (phase 74,
#: ``BOR_HISTORY_MAX_CHARS``) — ``len(text) + len(thinking or "")``
#: per turn, so prior thinking blocks count against the same budget
#: as the answer text. A turn that would overflow the remaining
#: budget is dropped WHOLE (never cut mid-answer) and the walk stops
#: there — the kept history is always a contiguous newest window.
history_max_chars: int = 24_000
# --- RAG tuning ---
embedding_dim: int = 768 # verified against aipi /v1 (embed model)
@@ -294,6 +309,26 @@ class Settings(BaseSettings):
raise ValueError("upload_max_mb must be > 0 (MiB)")
return v
@field_validator("history_max_turns")
@classmethod
def _history_max_turns_non_negative(cls, v: int) -> int:
"""``0`` is the no-history kill switch (pre-phase-74 two-message
requests) — a negative value is a typo (the ``agent_max_rounds``
pattern)."""
if v < 0:
raise ValueError("history_max_turns must be >= 0 (0 = no history)")
return v
@field_validator("history_max_chars")
@classmethod
def _history_max_chars_non_negative(cls, v: int) -> int:
"""``0`` is the no-history kill switch (pre-phase-74 two-message
requests) — a negative value is a typo (the ``agent_max_rounds``
pattern)."""
if v < 0:
raise ValueError("history_max_chars must be >= 0 (chars)")
return v
@field_validator("docs_branch", "docs_base_branch")
@classmethod
def _docs_branch_tokens(cls, v: str, info: ValidationInfo) -> str:
+18 -4
View File
@@ -176,7 +176,7 @@ import logging
import re
from collections.abc import AsyncIterator, Sequence
from dataclasses import dataclass, field
from typing import Any, cast
from typing import Any
from sqlalchemy import select
from sqlalchemy.orm import Session
@@ -821,6 +821,7 @@ async def run_agent(
seed_docs: Sequence[Document],
settings: Settings,
holder: AgentHolder,
history: Sequence[dict[str, Any]] = (),
) -> AsyncIterator[StreamPiece | ToolCallPiece | RetryPiece]:
"""Run the grounded-turn tool loop, yielding every stream piece.
@@ -830,6 +831,18 @@ async def run_agent(
SSE ``retry`` events. After the loop finishes, *holder* carries the
read documents and the executed tool-call count (re-lists included).
History (phase 74, TODO L4): *history* is the client's prior turns
already mapped to model messages by
:func:`app.rag.prompts.history_to_messages` (trimmed newest-first
against the settings budgets; assistant turns carry their prior
thinking as ``reasoning_content``). It is spliced between the system
prompt and the current user message —
``[system, *history, user]`` — and everything downstream (the tool
rounds, the phase-71 recovery rebuilding from ``messages[1:]``, the
retry restarts) already operates on that one ``messages`` list,
unchanged. ``()`` (the default) keeps the pre-phase-74 two-message
request byte-identical.
Retries (phase 67, owner-locked A2): every model request goes through
:func:`chat_stream_retried` — a failed round is retried **before** its
first piece (same messages, ``settings.llm_retries`` restarts, a flat
@@ -862,6 +875,7 @@ async def run_agent(
"""
messages: list[dict[str, Any]] = [
{"role": "system", "content": system_prompt},
*history, # phase 74: the client's prior turns (empty by default)
{"role": "user", "content": user_message},
]
# Phase 45: no per-tool budgets — the tools stay offered for the
@@ -892,7 +906,7 @@ async def run_agent(
# a quiet no-op.
stream = chat_stream_retried(
llm,
cast("list[dict[str, str]]", messages),
messages,
tools=tools,
retries=settings.llm_retries,
delay=settings.llm_retry_delay,
@@ -948,7 +962,7 @@ async def run_agent(
recovery_filter = ScaffoldingFilter()
recovered = chat_stream_retried(
llm,
cast("list[dict[str, str]]", messages_recovered),
messages_recovered,
tools=None,
retries=settings.llm_retries,
delay=settings.llm_retry_delay,
@@ -1030,7 +1044,7 @@ async def run_agent(
final_filter = ScaffoldingFilter()
final = chat_stream_retried(
llm,
cast("list[dict[str, str]]", messages),
messages,
tools=None,
retries=settings.llm_retries,
delay=settings.llm_retry_delay,
+11 -3
View File
@@ -298,7 +298,7 @@ class LLMClient:
return vec
async def chat(
self, messages: list[dict[str, str]], model: str | None = None
self, messages: list[dict[str, Any]], model: str | None = None
) -> str:
"""One-shot (non-streaming) completion (A5 extended, phase 30).
@@ -341,12 +341,20 @@ class LLMClient:
async def chat_stream(
self,
messages: list[dict[str, str]],
messages: list[dict[str, Any]],
tools: list[dict[str, Any]] | None = None,
scaffolding: ScaffoldingFilter | None = None,
) -> AsyncGenerator[StreamPiece | ToolCallPiece, None]:
"""Stream assistant pieces from the chat model (PLAN A5/A15, phase 17).
Messages are passed to the request body VERBATIM: string-only
``{role, content}`` dicts are byte-identical on the wire to the
pre-phase-74 requests, and an assistant message may additionally
carry ``reasoning_content`` (the client's prior thinking, phase
74 — the same wire field the model uses for its OWN reasoning on
the response side; the ``openai`` SDK passes message dicts
through untouched, so no transport change).
``stream=True`` against the OpenAI-compatible endpoint, yielding
typed :class:`StreamPiece` values. Wire convention (verified live
against aipi's ``turbo`` on 2026-08-23): the model's reasoning
@@ -488,7 +496,7 @@ class LLMClient:
async def chat_stream_retried(
llm: LLMClient,
messages: list[dict[str, str]],
messages: list[dict[str, Any]],
*,
tools: list[dict[str, Any]] | None = None,
retries: int = 0,
+59 -1
View File
@@ -54,10 +54,12 @@ not the wording, so that contract is unchanged.
from __future__ import annotations
from collections.abc import Sequence
from typing import Any
from app.config import get_settings
from app.config import Settings, get_settings
from app.models import Document
from app.rag.retriever import TRUNCATION_MARKER
from app.schemas import HistoryTurn
#: PLAN §6 verbatim (line wrapping included); ``{relevance}`` is filled by
#: :func:`_base`.
@@ -166,6 +168,62 @@ TOOLS_SECTION: str = (
)
def history_to_messages(
history: Sequence[HistoryTurn],
settings: Settings,
) -> list[dict[str, Any]]:
"""Client-provided chat history → model messages (phase 74, TODO L4).
The ``POST /api/chat`` ``history`` (the client's prior turns, oldest
first) becomes the message block that sits between the system prompt
and the current user message — so a follow-up question reaches the
model together with the exchange so far, on BOTH turn branches
(the deflected path and the grounded agent).
Trimming (owner-locked A3, 2026-09-08): the turns are walked
**newest-first** and kept while BOTH budgets hold — the turn count
stays ≤ ``settings.history_max_turns`` and the cumulative chars
(``len(text) + len(thinking or "")`` per turn) stay ≤
``settings.history_max_chars``. A turn that would overflow either
remaining budget is DROPPED WHOLE — never cut mid-answer — and the
walk stops there, so the kept history is always the contiguous
newest window (the oldest turns are the ones dropped; ``0`` on
either budget yields ``[]`` — the pre-phase-74 behavior). The kept
turns are returned in chronological (oldest → newest) order.
Mapping (owner-locked A4, 2026-09-08): ``who="user"`` →
``{"role": "user", "content": text}``; ``who="brain"`` →
``{"role": "assistant", "content": text}`` plus
``"reasoning_content": thinking`` ONLY when *thinking* is
non-empty — the preserve-thinking wire convention
:mod:`app.rag.llm` already reads on the response side
(``delta.reasoning_content``), which is what keeps the owner's
preserve-thinking models carrying the reasoning chain forward.
Pure and side-effect free (no I/O) — unit-testable in isolation.
"""
kept: list[HistoryTurn] = []
chars = 0
for turn in reversed(history):
if len(kept) >= settings.history_max_turns:
break
size = len(turn.text) + len(turn.thinking or "")
if chars + size > settings.history_max_chars:
break
kept.append(turn)
chars += size
messages: list[dict[str, Any]] = []
for turn in reversed(kept):
if turn.who == "user":
messages.append({"role": "user", "content": turn.text})
continue
message: dict[str, Any] = {"role": "assistant", "content": turn.text}
if turn.thinking:
message["reasoning_content"] = turn.thinking
messages.append(message)
return messages
def _base(relevance: str) -> str:
if relevance not in ("HIGH", "LOW"):
raise ValueError(f"relevance must be HIGH or LOW, got {relevance!r}")
+43 -2
View File
@@ -26,9 +26,50 @@ class SuggestionList(BaseModel):
suggestions: list[str]
class ChatRequest(BaseModel):
message: str = Field(min_length=1, max_length=4000)
class HistoryTurn(BaseModel):
"""One prior chat turn the client sends with ``POST /api/chat``
(phase 74, TODO L4).
The endpoint stays stateless (A10): the client's ``bor.chat.v1``
conversation record (minus the question about to be asked) is
provided in the request body as ``history`` so a follow-up question
reaches the model together with the exchange so far — and, for
preserve-thinking models, with the prior brain turns' thinking (the
record has carried the ``thinking`` key since phase 17).
``thinking`` travels to the model as ``reasoning_content`` on the
assistant message (the wire convention :mod:`app.rag.llm` already
documents for the response side) — only when non-empty (A4).
``text`` mirrors :attr:`ChatMessage.text`'s answer shape; the
thinking cap is looser (scratchpads run longer than answers). These
are boundary sanity caps only — the real trimming budget is the
settings pair ``history_max_turns`` / ``history_max_chars``
(``app.config``, A3: a capped-out turn is dropped whole, never
truncated).
"""
who: Literal["user", "brain"]
text: str = Field(min_length=1, max_length=4000)
thinking: str | None = Field(default=None, max_length=32000)
class ChatRequest(BaseModel):
"""``POST /api/chat`` body: the current question plus the optional
prior turns (phase 74 — the client-provided history, stateless per
A10).
``history`` is the client's earlier turns, oldest first (the
``bor.chat.v1`` record minus the current question); the mapper
(:func:`app.rag.prompts.history_to_messages`) trims it newest-first
against the settings budgets and maps it to model messages. The
schema-level ``max_length=100`` is a DoS sanity ceiling only — the
config budgets do the real trimming (A3). Absent or empty keeps the
request byte-identical to pre-phase-74: the model sees exactly the
two-message ``[system, user]`` request.
"""
message: str = Field(min_length=1, max_length=4000)
history: list[HistoryTurn] = Field(default_factory=list, max_length=100)
class LoginRequest(BaseModel):
"""``POST /api/login`` body (phase 16): the single admin's password.
+22 -1
View File
@@ -1972,10 +1972,31 @@ async function runTurn(text, { reask = false } = {}) {
setUiState(UI_STATE.error, "That's taking a long time — the answer may be stuck.");
});
// Phase 74 (TODO L4): the conversation so far travels WITH the
// question — the `conversation` record minus the current question.
// The invariant every caller holds at fetch time: a fresh send just
// pushed the question (save point 1 above); the phase-49 retry and
// the phase-53 stale-regen (both through retryLastTurn) popped the
// old answer and keep the question as the last entry — so slice(0,
// -1) is exactly the prior turns, oldest first, and the question is
// never duplicated into the history. `thinking` rides only brain
// records that actually streamed one (phase 17's optional key —
// `undefined` drops it from the JSON, the record's convention);
// user turns and old/restored records without thinking send none
// (the server maps those to plain assistant messages). No other
// record key (`sources`, `tools`, `deflected`, `suggestions`,
// `stopped`) travels in the body — the server schema (task 01)
// accepts exactly {who, text, thinking}; `tools` metadata is
// display-only and was never part of the LLM wire.
const history = conversation.slice(0, -1).map((m) => ({
who: m.who,
text: m.text,
thinking: m.who === "brain" ? m.thinking || undefined : undefined,
}));
res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ message: text }),
body: JSON.stringify({ message: text, history }),
signal: turnAbort.signal, // phase 48: the Stop button aborts the fetch
});
if (!res.ok || !res.body) {
+86
View File
@@ -228,6 +228,30 @@ Implements just enough of the aipi surface:
``SUMMARY_MODE``), so a marker question always gets the table
answer; the E2E asks it against an on-topic fixture (HIGH gate) and
asserts non-deflection.
- user message containing ``echo my history``
(``HISTORY_TRIGGER``, phase 74, TODO L4 — chat history with prior
thinking reaches the LLM) -> the deterministic HISTORY ECHO, derived
statelessly from the request messages and byte-stable:
``history: N prior messages; last answer tail: <tail>; thinking:
yes|no (Deterministic mock answer for E2E.)`` where N = the count
of non-``system`` messages before the LAST ``user`` message
(everything the client sent as prior turns — the current question
itself is excluded), <tail> = the LAST 24 chars of the most recent
prior ``assistant`` message's content (``none`` when there is no
prior assistant message), and thinking is ``yes`` iff that prior
``assistant`` message carries a non-empty ``reasoning_content``
field (the client's phase-74 history mapping of the brain record's
``thinking`` — A4). Checked BEFORE the ``DEFLECT_MODE`` branch
(like ``TABLE_TRIGGER`` — the marker lives in the user message, a
deflection prompt never carries it), so a marker question always
gets the echo whatever the honesty gate says; the story E2E
(``tests/e2e/test_llm_history.py``) asserts the wire contents
byte-exactly against the conversation record the client persisted.
Invariant the marker relies on: the client history contains ONLY
``user``/``assistant`` messages — never ``tool``-role ones (the
client never sends tool calls/results) — so every existing marker
flow (which classifies statelessly from TOOL results and the LAST
user message) is unaffected by the now-always-present history.
Failure injection (phase 67, LLM retry, TODO.md L3) — deterministic
dead-endpoint behavior for the retry E2E suite (``tests/e2e/
@@ -415,6 +439,15 @@ TABLE_TRIGGER = "show me a table"
#: E2E asserts the rendered table shape, the escaped ``<img onerror>``
#: line (the XSS payload must survive the mock byte-for-byte), and the
#: wide table's ``scrollWidth > clientWidth`` inside the 46rem column.
#: Phase 74 (chat history, TODO L4): a user message containing this
#: substring (case-insensitive) gets the deterministic HISTORY ECHO
#: (``_history_echo`` below — see the module docstring): the prior-turn
#: count, the last 24 chars of the most recent prior answer, and
#: whether that prior answer carried ``reasoning_content``. Verified
#: 2026-09-08: no existing E2E question or fixture file contains the
#: phrase, so every other suite is unaffected.
HISTORY_TRIGGER = "echo my history"
TABLE_ANSWER = (
"Here's the shape, in a table:\n"
"\n"
@@ -1011,6 +1044,50 @@ def first_kb_bullet(system: str) -> str | None:
return None
def _history_echo(body: dict[str, Any]) -> str:
"""The phase-74 history echo (byte-stable, stateless over messages).
``history: N prior messages`` — N = the count of non-``system``
messages before the LAST ``user`` message (the client's phase-74
``history`` block: the prior turns only, the current question
itself excluded). ``last answer tail: <tail>`` — the LAST 24 chars
of the most recent prior ``assistant`` message's content, or
``none`` when there is no prior assistant message (the cold-start
pin: no phantom history). ``thinking: yes|no`` — ``yes`` iff that
prior assistant message carries a non-empty ``reasoning_content``
field (A4: the client's prior thinking, mapped by
``app.rag.prompts.history_to_messages``), ``no`` otherwise.
The invariant (see the module docstring): the client history is
``user``/``assistant``-only, so the last ``user`` message is always
the current question and every earlier non-system message is a
client-provided prior turn.
"""
msgs = _messages(body)
last_user = max(
(i for i, m in enumerate(msgs) if m.get("role") == "user"),
default=-1,
)
prior = [
m
for i, m in enumerate(msgs)
if i < last_user and m.get("role") != "system"
]
tail = "none"
thinking = "no"
for m in reversed(prior):
if m.get("role") == "assistant":
tail = str(m.get("content") or "")[-24:]
thinking = "yes" if str(m.get("reasoning_content") or "") else "no"
break
return (
f"history: {len(prior)} prior messages; "
f"last answer tail: {tail}; "
f"thinking: {thinking} "
"(Deterministic mock answer for E2E.)"
)
def compose_answer(body: dict[str, Any]) -> str:
system = _system(body)
user = _user(body)
@@ -1053,6 +1130,15 @@ def compose_answer(body: dict[str, Any]) -> str:
# against an on-topic fixture, where the gate is HIGH, and
# asserts non-deflection as part of the table test.
answer = TABLE_ANSWER
elif HISTORY_TRIGGER in user.lower():
# Phase 74 (TODO L4, chat history): the deterministic history
# echo — proves on the wire that the client's prior turns (and
# the prior thinking, as ``reasoning_content`` on the assistant
# messages) reached the model. Checked BEFORE the DEFLECT_MODE
# branch, like TABLE_TRIGGER: the marker lives in the user
# message, a deflection prompt never carries it, so a marker
# question always gets the echo whatever the gate says.
answer = _history_echo(body)
elif "DEFLECT_MODE" in system:
answer = (
"Ah — I haven't done anything like that, so I don't want to make stuff up! "
+226
View File
@@ -0,0 +1,226 @@
"""Phase 74 E2E (Playwright): prior turns + prior thinking reach the LLM.
TODO.md L4 (owner 2026-09-05): "Chat history isn't being passed to the
LLM. When the LLM responds and you ask a follow-up question the
previous question/answer isn't passed to the model. Since my models
support preserve thinking, make sure to pass previous thinking blocks
as well."
Run in isolation (DB must be up: ``podman compose up -d db``):
uv run pytest tests/e2e/test_llm_history.py -v --no-cov
The mock's ``echo my history`` marker (``HISTORY_TRIGGER``) answers
with a deterministic echo of the history block the model received —
``history: N prior messages; last answer tail: <last 24 chars of the
prior answer>; thinking: yes|no`` — so every assertion below is a
byte-exact pin on the wire contents. The prior answer's tail is
derived from the conversation record the client persisted
(localStorage ``bor.chat.v1``) — the SAME array task 02 maps into the
request body's ``history``, so what the record shows IS what the model
received (``thinking`` travels as ``reasoning_content`` on the
assistant message — A4).
The marker is checked BEFORE the mock's ``DEFLECT_MODE`` branch, so
the echo fires on BOTH turn branches — the branch under test is
discriminated separately (the grounded source chip / the
``is-deflected`` bubble class). The echo answers carry no tool
markup, so no marker tool flow is re-triggered by the now-always-
present (user/assistant-only) history.
The file name deliberately differs from phase 50's
``test_chat_history.py`` (save & view chat history — a different
story).
"""
from __future__ import annotations
import asyncio
import json
from pathlib import Path
from threading import Thread
from typing import Any
from playwright.sync_api import Page, expect
from sqlalchemy import text
from app.config import Settings
from app.db import SessionLocal
from app.rag.importer import ImportSummary, import_sources
from app.rag.llm import LLMClient
REPO = Path(__file__).resolve().parents[2]
FIXTURES = REPO / "tests" / "fixtures" / "docs"
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
STORAGE_KEY = "bor.chat.v1"
#: Turn 1 (both follow-up stories): on-topic (HIGH gate -> grounded)
#: and carries the phase-17 thinking trigger, so the brain record
#: streams a deterministic scratchpad into its ``thinking`` key.
T1 = "think out loud — how is my Kubernetes cluster set up?"
#: Turn 2, grounded story: on-topic + the phase-74 history echo marker.
T2_GROUNDED = "echo my history about my kubernetes cluster"
#: Turn 2, deflected story: OFF-topic (LOW gate -> deflected branch) +
#: the marker — ASSUMPTION A3: BOTH branches carry the history, and
#: the marker fires before the DEFLECT_MODE branch, so this is the
#: deflected path under test.
T2_DEFLECTED = "echo my history — how do I bake sourdough bread?"
async def _import_fixtures(mock_port: int) -> ImportSummary:
kwargs: dict[str, Any] = {"_env_file": None, "llm_base_url": f"http://127.0.0.1:{mock_port}/v1"}
settings = Settings(**kwargs) # pyright: ignore[reportCallIssue]
return await import_sources([FIXTURES], LLMClient(settings))
def _run_in_thread(coro: Any) -> Any:
"""Run a coroutine on a worker thread.
Playwright's sync API keeps an asyncio loop running on the test
thread, so ``asyncio.run`` cannot be called directly from a test
body.
"""
box: dict[str, Any] = {}
def runner() -> None:
try:
box["value"] = asyncio.run(coro)
except BaseException as e: # noqa: BLE001 — re-raised on the test thread
box["error"] = e
t = Thread(target=runner)
t.start()
t.join()
if "error" in box:
raise box["error"]
return box["value"]
def _reset_db(mock_port: int, seed: bool) -> ImportSummary | None:
"""Truncate the KB (+ query log + steering notes — deterministic
mock answers), then optionally re-import fixtures."""
with SessionLocal() as db:
db.execute(text("TRUNCATE chunks, documents, query_log, steering_notes"))
db.commit()
if not seed:
return None
return _run_in_thread(_import_fixtures(mock_port))
def _ask(page: Page, question: str) -> None:
"""Send one turn and wait until the answer has fully landed (the
``done`` event restored the Send button)."""
page.fill("#message-input", question)
page.click("#send-btn")
expect(page.locator(".msg.user .bubble").last).to_contain_text(question)
expect(page.locator(".msg.brain .bubble").last).to_contain_text(
MOCK_ANSWER_MARKER, timeout=60_000
)
expect(page.locator("#send-label")).to_have_text("Send")
def _record(page: Page) -> dict[str, Any]:
"""The persisted ``bor.chat.v1`` record (the same array task 02
maps into the request body's ``history``)."""
raw = page.evaluate(f"localStorage.getItem({STORAGE_KEY!r})")
return json.loads(raw) if raw else {"messages": []}
def _wait_record(page: Page, n_messages: int) -> dict[str, Any]:
"""Wait until the persisted record carries ``n_messages`` turns
(the ``done`` event's save point has landed in localStorage)."""
page.wait_for_function(
"""([key, n]) => {
const raw = localStorage.getItem(key);
const rec = raw ? JSON.parse(raw) : null;
return !!rec && rec.messages.length >= n;
}""",
arg=[STORAGE_KEY, n_messages],
timeout=15_000,
)
return _record(page)
def test_followup_receives_history_and_thinking(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 13 # A9 formats (phase 47 added quadlet+j2)
page.set_default_timeout(30_000)
# Cold start: no restored conversation — every prior turn the model
# sees on turn 2 is the one this test just sent.
page.add_init_script("localStorage.clear()")
page.goto(app_url)
# Turn 1 — grounded + the thinking trigger: the brain record must
# carry the streamed scratchpad in its ``thinking`` key.
_ask(page, T1)
brain1 = _wait_record(page, 2)["messages"][1]
assert brain1["who"] == "brain"
assert brain1["thinking"], "turn 1 must have streamed thinking into the record"
assert MOCK_ANSWER_MARKER in brain1["text"]
answer_tail = brain1["text"][-24:]
# Turn 2 — grounded + the echo marker: the model receives
# [system, user(T1), assistant(A1, reasoning_content), user(T2)]
# and the echo proves it byte-exactly.
_ask(page, T2_GROUNDED)
bubble = page.locator(".msg.brain .bubble").last
expect(bubble).to_contain_text("history: 2 prior messages", timeout=30_000)
expect(bubble).to_contain_text(f"last answer tail: {answer_tail}")
expect(bubble).to_contain_text("thinking: yes")
# Grounded proof — the echo fires in BOTH branches, so the branch
# is discriminated by the kubernetes.md source chip (the deflected
# turn carries no cited sources). Scoped to the LAST brain message:
# turn 1 cited kubernetes.md too.
expect(
page.locator(".msg.brain").last.locator(".source-chip", has_text="kubernetes.md")
).to_have_count(1)
def test_first_question_has_no_history(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 13 # A9 formats (phase 47 added quadlet+j2)
page.set_default_timeout(30_000)
page.add_init_script("localStorage.clear()")
page.goto(app_url)
# Cold start: the request body's history is empty — no phantom
# prior turns, no phantom thinking.
_ask(page, T2_GROUNDED)
bubble = page.locator(".msg.brain .bubble").last
expect(bubble).to_contain_text("history: 0 prior messages", timeout=30_000)
expect(bubble).to_contain_text("last answer tail: none")
expect(bubble).to_contain_text("thinking: no")
# Grounded: the echo question is on-topic (the chip proves the
# HIGH gate, not a deflection).
expect(
page.locator(".msg.brain").last.locator(".source-chip", has_text="kubernetes.md")
).to_have_count(1)
def test_deflected_followup_receives_history(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 13 # A9 formats (phase 47 added quadlet+j2)
page.set_default_timeout(30_000)
page.add_init_script("localStorage.clear()")
page.goto(app_url)
_ask(page, T1)
brain1 = _wait_record(page, 2)["messages"][1]
assert brain1["thinking"], "turn 1 must have streamed thinking into the record"
answer_tail = brain1["text"][-24:]
# Turn 2 — OFF-topic (LOW gate -> deflected branch) + the marker:
# the echo still arrives with the SAME history block (A3: both
# branches carry it — the marker is checked before the
# DEFLECT_MODE branch, so this test proves the deflected path).
_ask(page, T2_DEFLECTED)
bubble = page.locator(".msg.brain.is-deflected .bubble").last
bubble.wait_for(state="visible", timeout=30_000)
expect(bubble).to_contain_text("history: 2 prior messages", timeout=30_000)
expect(bubble).to_contain_text(f"last answer tail: {answer_tail}")
expect(bubble).to_contain_text("thinking: yes")
+170 -2
View File
@@ -89,7 +89,10 @@ class FakeRagLLM:
#: recovered answer endpoint for the pre-first-piece retry rule.
self.stream_fail_count = stream_fail_count
self.question_embeds: list[str] = []
self.seen_messages: list[list[dict[str, str]]] = []
#: Phase 74: assistant history messages may carry
#: ``reasoning_content`` — the dict values stay strings, but the
#: key set is wider than the pre-phase ``{role, content}`` shape.
self.seen_messages: list[list[dict[str, Any]]] = []
#: Every request's ``tools`` value (phase 37) — ``None`` is the
#: pre-phase request shape (the key is absent from the payload).
self.seen_tools: list[list[dict[str, Any]] | None] = []
@@ -138,7 +141,7 @@ class FakeRagLLM:
async def chat_stream(
self,
messages: list[dict[str, str]],
messages: list[dict[str, Any]],
tools: list[dict[str, Any]] | None = None,
scaffolding: ScaffoldingFilter | None = None,
):
@@ -1226,3 +1229,168 @@ def test_deflected_mixed_scaffolding_and_content_needs_no_recovery(
assert len(flaky.seen_messages) == 1 # the clean content stands — no recovery
lines = [r.getMessage() for r in caplog.records if "question=" in r.getMessage()]
assert lines and f"scaffold_stripped={len(span)}" in lines[-1]
# ---------- phase 74: client-provided history (with prior thinking) ----------
#: The client's prior turns (oldest first — the ``bor.chat.v1`` record
#: minus the current question): two user turns, two brain turns, the
#: FIRST brain turn carrying a prior thinking block (A4) and the second
#: not (the ``reasoning_content`` gate has both shapes on one request).
HISTORY: list[dict[str, Any]] = [
{"who": "user", "text": "What port does Tailscale run on?"},
{
"who": "brain",
"text": "Tailscale runs on 41641/udp.",
"thinking": "The Tailscale wire protocol uses 41641/udp.",
},
{"who": "user", "text": "And the subnet router?"},
{"who": "brain", "text": "The subnet router shares the same port."},
]
#: What :func:`app.rag.prompts.history_to_messages` must produce for
#: :data:`HISTORY` — chronological, ``reasoning_content`` ONLY on the
#: turn that had thinking.
HISTORY_MESSAGES: list[dict[str, Any]] = [
{"role": "user", "content": "What port does Tailscale run on?"},
{
"role": "assistant",
"content": "Tailscale runs on 41641/udp.",
"reasoning_content": "The Tailscale wire protocol uses 41641/udp.",
},
{"role": "user", "content": "And the subnet router?"},
{"role": "assistant", "content": "The subnet router shares the same port."},
]
def _stream_chat_with_history(
client: TestClient, message: str, history: list[dict[str, Any]]
) -> list[dict[str, Any]]:
"""Phase 74 variant of :func:`_stream_chat`: sends ``history`` (the
client's prior turns, oldest first) in the request body."""
with client.stream(
"POST", "/api/chat", json={"message": message, "history": history}
) as r:
assert r.status_code == 200
assert r.headers["content-type"].startswith("text/event-stream")
buf = ""
frames: list[dict[str, Any]] = []
for part in r.iter_text():
buf += part
while "\n\n" in buf:
frame, buf = buf.split("\n\n", 1)
frame = frame.strip()
if frame.startswith("data:"):
frames.append(json.loads(frame.removeprefix("data:").strip()))
assert buf.strip() == "", "stream must end on a frame boundary"
return frames
def test_deflected_turn_forwards_history_with_prior_thinking(
client,
db,
seeded_kb: FakeRagLLM,
caplog: pytest.LogCaptureFixture,
) -> None:
"""A DEFLECTED turn sends the prior turns — chronological, with the
prior brain turn's thinking as ``reasoning_content`` — between the
LOW system prompt and the current question (A2/A3/A4); the per-turn
log line carries ``history_msgs=4``."""
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb
try:
caplog.set_level(logging.INFO, logger="app.chat")
frames = _stream_chat_with_history(client, OFF_TOPIC, HISTORY)
finally:
fastapi_app.dependency_overrides.clear()
assert frames[-1]["type"] == "done"
assert frames[-1]["deflected"] is True
assert len(seeded_kb.seen_messages) == 1
(messages,) = seeded_kb.seen_messages
assert messages[0]["role"] == "system"
assert "DEFLECT_MODE" in messages[0]["content"] # the LOW prompt
assert messages[1:-1] == HISTORY_MESSAGES # the prior turns, chronological
assert messages[-1] == {"role": "user", "content": OFF_TOPIC}
lines = [r.getMessage() for r in caplog.records if "question=" in r.getMessage()]
assert lines and "history_msgs=4" in lines[-1]
def test_grounded_turn_forwards_history_through_the_agent(
client,
db,
seeded_kb: FakeRagLLM,
caplog: pytest.LogCaptureFixture,
) -> None:
"""The GROUNDED agent branch receives the same block: its first
request is ``[HIGH system, *history, current question]`` (the tool
rounds then append to that same list); ``history_msgs=4``."""
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb
try:
caplog.set_level(logging.INFO, logger="app.chat")
frames = _stream_chat_with_history(client, QUESTION, HISTORY)
finally:
fastapi_app.dependency_overrides.clear()
assert frames[-1]["type"] == "done"
assert frames[-1]["deflected"] is False
assert len(seeded_kb.seen_messages) == 1 # the canned answer ends the loop
(messages,) = seeded_kb.seen_messages
assert messages[0]["role"] == "system"
assert "<tools>" in messages[0]["content"] # the HIGH prompt
assert messages[1:-1] == HISTORY_MESSAGES
assert messages[-1] == {"role": "user", "content": QUESTION}
lines = [r.getMessage() for r in caplog.records if "question=" in r.getMessage()]
assert lines and "history_msgs=4" in lines[-1]
def test_request_without_history_sends_exactly_system_and_user(
client,
db,
seeded_kb: FakeRagLLM,
caplog: pytest.LogCaptureFixture,
) -> None:
"""Byte-identical pin (A2): a request WITHOUT ``history`` sends
exactly the two-message ``[system, user]`` request on BOTH branches
(deflected + grounded), and the per-turn log line carries
``history_msgs=0``."""
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb
try:
caplog.set_level(logging.INFO, logger="app.chat")
_stream_chat(client, OFF_TOPIC) # deflected branch
_stream_chat(client, QUESTION) # grounded branch
finally:
fastapi_app.dependency_overrides.clear()
assert len(seeded_kb.seen_messages) == 2
for messages in seeded_kb.seen_messages:
assert [m["role"] for m in messages] == ["system", "user"]
assert seeded_kb.seen_messages[0][1] == {"role": "user", "content": OFF_TOPIC}
assert seeded_kb.seen_messages[1][1] == {"role": "user", "content": QUESTION}
lines = [r.getMessage() for r in caplog.records if "question=" in r.getMessage()]
assert len(lines) == 2
assert all("history_msgs=0" in line for line in lines)
def test_history_rejects_unknown_who(client, db) -> None:
"""Schema pin: ``who`` is a ``Literal["user", "brain"]`` — anything
else is a 422 at the boundary (the same trust model as the saved-
chat ``ChatMessage``)."""
r = client.post(
"/api/chat",
json={"message": "hi", "history": [{"who": "alien", "text": "x"}]},
)
assert r.status_code == 422
def test_history_rejects_more_than_100_entries(client, db) -> None:
"""Schema pin: the DoS sanity ceiling is 100 turns — 101 is a 422
(the config budgets do the real trimming; this only keeps a
pathological body from wasting the mapper's work)."""
r = client.post(
"/api/chat",
json={
"message": "hi",
"history": [{"who": "user", "text": f"q{i}"} for i in range(101)],
},
)
assert r.status_code == 422
+77 -1
View File
@@ -30,7 +30,7 @@ import asyncio
import json
import logging
import uuid
from collections.abc import AsyncGenerator, AsyncIterator
from collections.abc import AsyncGenerator, AsyncIterator, Sequence
from copy import deepcopy
from typing import TYPE_CHECKING, Any, cast
@@ -115,7 +115,11 @@ async def _run(
holder: AgentHolder,
settings: Settings,
seed_docs: list[Document] | None = None,
history: Sequence[dict[str, Any]] = (),
) -> list[StreamPiece | ToolCallPiece | RetryPiece]:
"""Consume one ``run_agent`` turn; *history* (phase 74) is the
client's prior turns spliced between system and user (default
``()`` — the pre-phase-74 two-message request)."""
out: list[StreamPiece | ToolCallPiece | RetryPiece] = []
async for piece in run_agent(
cast("LLMClient", llm),
@@ -125,6 +129,7 @@ async def _run(
seed_docs=seed_docs or [],
settings=settings,
holder=holder,
history=history,
):
out.append(piece)
return out
@@ -458,6 +463,77 @@ def test_content_and_tool_call_in_one_stream_keeps_both(
assert llm.requests[1][0][3]["content"] == "0 documents:\n"
# ---------- phase 74: client history between system and user ----------
def test_run_agent_default_history_keeps_two_message_request() -> None:
"""No *history* (the default ``()``) → the model sees exactly the
pre-phase-74 two-message request ``[system, user]`` — byte-identical
behavior (owner-locked A2)."""
llm = ScriptedLLM([StreamPiece("content", "the answer")])
asyncio.run(_run(llm, AgentHolder(), _settings()))
(messages, _tools) = llm.requests[0]
assert messages == [
{"role": "system", "content": "SYSTEM_PROMPT"},
{"role": "user", "content": "QUESTION"},
]
def test_run_agent_places_history_between_system_and_user() -> None:
"""A non-empty *history* (the client's prior turns, already mapped by
``history_to_messages``) is spliced between the system prompt and the
CURRENT user message — oldest-first, with the assistant turn's prior
thinking riding on ``reasoning_content`` (A4). The current question
stays LAST."""
history = [
{"role": "user", "content": "old question"},
{
"role": "assistant",
"content": "old answer",
"reasoning_content": "old thinking",
},
]
llm = ScriptedLLM([StreamPiece("content", "the answer")])
pieces = asyncio.run(
_run(llm, AgentHolder(), _settings(), history=history)
)
assert [p for p in pieces if isinstance(p, StreamPiece)] == [
StreamPiece("content", "the answer")
]
(messages, _tools) = llm.requests[0]
assert messages == [
{"role": "system", "content": "SYSTEM_PROMPT"},
{"role": "user", "content": "old question"},
{
"role": "assistant",
"content": "old answer",
"reasoning_content": "old thinking",
},
{"role": "user", "content": "QUESTION"},
]
def test_run_agent_history_survives_a_tool_round(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The tool rounds append assistant/tool messages to the SAME
``messages`` list — the prior history stays in place between the
system prompt and the current question on the SECOND request too."""
monkeypatch.setattr(agent, "list_catalog", lambda db: [])
llm = ScriptedLLM(
[ToolCallPiece(id="call_1", name="ls", arguments={})],
[StreamPiece("content", "the answer")],
)
history = [{"role": "assistant", "content": "old answer"}]
asyncio.run(_run(llm, AgentHolder(), _settings(), history=history))
_first, second = llm.requests
assert second[0][:3] == [
{"role": "system", "content": "SYSTEM_PROMPT"},
{"role": "assistant", "content": "old answer"},
{"role": "user", "content": "QUESTION"},
]
# ---------- ls: full catalog + scoping ----------
+42
View File
@@ -147,6 +147,48 @@ def test_llm_retry_delay_rejects_negative(monkeypatch: pytest.MonkeyPatch) -> No
_settings()
def test_history_budget_defaults(monkeypatch: pytest.MonkeyPatch) -> None:
"""Phase 74 (TODO L4): the client-provided history is trimmed
newest-first against the newest 40 turns within a total of
24 000 chars (text + prior thinking, owner-locked A3)."""
monkeypatch.delenv("BOR_HISTORY_MAX_TURNS", raising=False)
monkeypatch.delenv("BOR_HISTORY_MAX_CHARS", raising=False)
s = _settings()
assert s.history_max_turns == 40
assert s.history_max_chars == 24_000
def test_history_budget_env_overrides(monkeypatch: pytest.MonkeyPatch) -> None:
"""``BOR_HISTORY_MAX_TURNS`` / ``BOR_HISTORY_MAX_CHARS`` override the
defaults; ``0`` on either is the no-history kill switch (the
pre-phase-74 two-message requests)."""
monkeypatch.setenv("BOR_HISTORY_MAX_TURNS", "12")
monkeypatch.setenv("BOR_HISTORY_MAX_CHARS", "5000")
s = _settings()
assert s.history_max_turns == 12
assert s.history_max_chars == 5000
monkeypatch.setenv("BOR_HISTORY_MAX_TURNS", "0")
assert _settings().history_max_turns == 0
monkeypatch.setenv("BOR_HISTORY_MAX_CHARS", "0")
assert _settings().history_max_chars == 0
def test_history_max_turns_rejects_negative(monkeypatch: pytest.MonkeyPatch) -> None:
"""``0`` is the no-history kill switch — a negative value is a typo,
so the validator fails loudly at startup (the ``agent_max_rounds``
pattern)."""
monkeypatch.setenv("BOR_HISTORY_MAX_TURNS", "-1")
with pytest.raises(ValidationError, match="history_max_turns"):
_settings()
def test_history_max_chars_rejects_negative(monkeypatch: pytest.MonkeyPatch) -> None:
"""A negative char budget is a typo — fail loudly at startup."""
monkeypatch.setenv("BOR_HISTORY_MAX_CHARS", "-1")
with pytest.raises(ValidationError, match="history_max_chars"):
_settings()
def test_agent_max_rounds_default_and_env_override(monkeypatch: pytest.MonkeyPatch) -> None:
"""Phase 45: the per-tool budgets are gone — ``BOR_AGENT_MAX_ROUNDS``
(default 10) is the single agent-loop knob; ``0`` is the no-tools
+214
View File
@@ -0,0 +1,214 @@
"""Unit: the client-history → model-messages mapper (phase 74, TODO L4).
``app.rag.prompts.history_to_messages`` is pure (no I/O) — every branch
is pinned here: the user/brain role mapping, the ``reasoning_content``
gating (prior thinking travels ONLY when non-empty — the preserve-
thinking wire convention, A4), the turn-count budget (newest kept,
oldest dropped), the char budget (``text`` + ``thinking`` accounted,
drop-WHOLE semantics — never cut mid-answer, A3), the budgets working
together, and the chronological (oldest → newest) order of the result.
"""
from __future__ import annotations
from typing import Any, Literal
from app.config import Settings
from app.rag.prompts import history_to_messages
from app.schemas import HistoryTurn
def _settings(**kwargs: Any) -> Settings:
kwargs.setdefault("_env_file", None)
return Settings(**kwargs) # pyright: ignore[reportCallIssue] (kwarg exists at runtime)
def _turn(
who: Literal["user", "brain"], text: str, thinking: str | None = None
) -> HistoryTurn:
return HistoryTurn(who=who, text=text, thinking=thinking)
# ---------- mapping ----------
def test_empty_history_yields_no_messages() -> None:
"""Absent client history (the pre-phase-74 request shape) → ``[]`` —
the caller then builds the byte-identical two-message request."""
assert history_to_messages([], _settings()) == []
def test_user_turn_maps_to_user_role() -> None:
got = history_to_messages([_turn("user", "What port does Tailscale use?")], _settings())
assert got == [{"role": "user", "content": "What port does Tailscale use?"}]
def test_brain_turn_without_thinking_maps_to_assistant_role() -> None:
"""No ``thinking`` key → a plain assistant message: NO
``reasoning_content`` key at all (A4 gating, ``None`` case)."""
got = history_to_messages([_turn("brain", "Tailscale runs on 41641/udp.")], _settings())
assert got == [{"role": "assistant", "content": "Tailscale runs on 41641/udp."}]
assert "reasoning_content" not in got[0]
def test_brain_turn_with_thinking_carries_reasoning_content() -> None:
"""A prior thinking block travels as ``reasoning_content`` on the
assistant message (A4 — the preserve-thinking wire convention the
response side already reads)."""
thinking = "Tailscale's wire protocol port is 41641/udp."
got = history_to_messages(
[_turn("brain", "Tailscale runs on 41641/udp.", thinking=thinking)],
_settings(),
)
assert got == [
{
"role": "assistant",
"content": "Tailscale runs on 41641/udp.",
"reasoning_content": thinking,
}
]
def test_brain_turn_with_empty_thinking_omits_reasoning_content() -> None:
"""``thinking=""`` is "empty" for the A4 gate — no
``reasoning_content`` key (an empty scratchpad carries nothing)."""
got = history_to_messages(
[_turn("brain", "Same answer.", thinking="")], _settings()
)
assert got == [{"role": "assistant", "content": "Same answer."}]
assert "reasoning_content" not in got[0]
def test_result_is_chronological_oldest_to_newest() -> None:
"""The input is oldest-first; the output must be too — the newest
turn ends up LAST, directly ahead of the current user message the
caller appends."""
turns = [
_turn("user", "q1"),
_turn("brain", "a1", thinking="t1"),
_turn("user", "q2"),
_turn("brain", "a2"),
_turn("user", "q3"),
]
got = history_to_messages(turns, _settings())
assert [m["role"] for m in got] == ["user", "assistant", "user", "assistant", "user"]
assert [m["content"] for m in got] == ["q1", "a1", "q2", "a2", "q3"]
assert got[1]["reasoning_content"] == "t1"
assert "reasoning_content" not in got[3]
# ---------- turn-count budget ----------
def test_turn_cap_keeps_newest_and_drops_oldest() -> None:
"""The newest ``history_max_turns`` turns are kept; the OLDEST are
the ones dropped (newest-first walk, stop at the count cap)."""
turns = [_turn("user", f"q{i}") for i in range(1, 6)] # q1 … q5, oldest first
got = history_to_messages(turns, _settings(history_max_turns=3, history_max_chars=10_000))
assert [m["content"] for m in got] == ["q3", "q4", "q5"]
def test_default_turn_cap_is_40() -> None:
"""45 turns under the DEFAULT caps (40 turns / 24 000 chars, short
texts so the char budget never binds) keep the newest 40."""
turns = [_turn("user", f"question number {i}") for i in range(1, 46)]
got = history_to_messages(turns, _settings())
assert len(got) == 40
assert got[0]["content"] == "question number 6" # the five oldest are gone
assert got[-1]["content"] == "question number 45"
# ---------- char budget ----------
def test_char_budget_counts_text_plus_thinking() -> None:
"""The per-turn size is ``len(text) + len(thinking or "")`` — prior
thinking blocks count against the same budget as the answer text."""
# Newest-first sizes: 5 + 100 (20+80) + 10; budget 110 keeps the
# newest two (105) and drops the oldest (115 > 110).
turns = [
_turn("user", "a" * 10), # oldest — dropped whole
_turn("brain", "b" * 20, thinking="c" * 80),
_turn("user", "d" * 5), # newest
]
got = history_to_messages(turns, _settings(history_max_turns=40, history_max_chars=110))
assert len(got) == 2
assert got[0]["content"] == "b" * 20
assert got[0]["reasoning_content"] == "c" * 80
assert got[1]["content"] == "d" * 5
def test_char_budget_exact_fit_is_kept() -> None:
"""Cumulative chars EQUAL to the cap fit (≤, not <) — the exact-fit
turn is kept, and the older turn that would push past is dropped."""
turns = [
_turn("user", "a" * 10), # oldest — 100+10=110 > 100, dropped
_turn("brain", "b" * 100), # newest — exactly the 100-char cap, kept
]
got = history_to_messages(turns, _settings(history_max_turns=40, history_max_chars=100))
assert [m["content"] for m in got] == ["b" * 100]
def test_overflowing_turn_is_dropped_whole_never_truncated() -> None:
"""A turn that would overflow the remaining budget is DROPPED WHOLE
(A3) — its text appears nowhere in the result, not even partially,
and the walk stops there (the kept history stays a contiguous
newest window)."""
big = "x" * 120 # alone it would overflow the 100-char budget
turns = [
_turn("user", "old question"),
_turn("brain", "old answer"),
_turn("brain", big), # newest — does not fit at all
]
got = history_to_messages(turns, _settings(history_max_turns=40, history_max_chars=100))
assert got == [] # the newest does not fit → nothing is kept
assert not any("x" in m["content"] for m in got)
def test_overflowing_middle_turn_stops_the_walk() -> None:
"""Newest-first: the newest fits, the NEXT (middle) turn would
overflow → it is dropped whole AND the walk stops — the oldest turn
is not sneaked in across the gap (no discontinuous history)."""
turns = [
_turn("user", "a" * 5), # oldest — never even considered
_turn("user", "b" * 51), # middle — 60+51=111 > 100, dropped whole
_turn("user", "c" * 60), # newest — fits (60 ≤ 100)
]
got = history_to_messages(turns, _settings(history_max_turns=40, history_max_chars=100))
assert [m["content"] for m in got] == ["c" * 60]
def test_zero_char_budget_yields_no_history() -> None:
"""``history_max_chars=0`` is a budget that fits nothing — the
kill-switch shape (no history, pre-phase-74 two-message request)."""
turns = [_turn("user", "q1"), _turn("brain", "a1")]
assert history_to_messages(turns, _settings(history_max_chars=0)) == []
def test_zero_turn_budget_yields_no_history() -> None:
"""``history_max_turns=0`` keeps no turns even though chars are free."""
turns = [_turn("user", "q1"), _turn("brain", "a1")]
assert history_to_messages(turns, _settings(history_max_turns=0)) == []
# ---------- budgets together ----------
def test_turn_cap_wins_when_chars_remain() -> None:
"""Both budgets in play: plenty of chars, a small turn cap — the
count cap stops the walk first (newest 3 of 5 kept)."""
turns = [_turn("user", f"q{i}") for i in range(1, 6)]
got = history_to_messages(turns, _settings(history_max_turns=3, history_max_chars=10_000))
assert len(got) == 3
assert [m["content"] for m in got] == ["q3", "q4", "q5"]
def test_char_cap_wins_when_turns_remain() -> None:
"""Symmetrically: plenty of turn budget, a tight char cap — the char
budget stops the walk (2 of 3 turns kept)."""
turns = [
_turn("user", "a" * 50), # oldest — dropped (50+60=110 > 100)
_turn("user", "b" * 60),
_turn("user", "c" * 40), # newest
]
got = history_to_messages(turns, _settings(history_max_turns=40, history_max_chars=100))
assert [m["content"] for m in got] == ["b" * 60, "c" * 40]