feat(chat): stream model thinking over SSE and show it in a collapsible block

This commit is contained in:
2026-08-24 09:52:27 -04:00
parent cbc263a4b2
commit b16deb2b1d
18 changed files with 1045 additions and 63 deletions
+34 -6
View File
@@ -5,7 +5,8 @@
> Anchors table are settled — do not re-litigate them in a phase. > Anchors table are settled — do not re-litigate them in a phase.
> **Revisions (2026-08-21, owner permission):** A7/A8/A9 revised (multi-format > **Revisions (2026-08-21, owner permission):** A7/A8/A9 revised (multi-format
> ingestion, hybrid FTS+vector retrieval, re-tuned honesty gate); dark tech > ingestion, hybrid FTS+vector retrieval, re-tuned honesty gate); dark tech
> theme (Phase 08); clickable document viewer (Phase 10). See roadmap §12. > theme (Phase 08); clickable document viewer (Phase 10); thinking display
> (Phase 17, owner permission 2026-08-23). See roadmap §12.
--- ---
@@ -134,14 +135,28 @@ All endpoints stateless (A10). Errors: standard JSON `{detail: str}`.
### SSE contract (`POST /api/chat`) ### SSE contract (`POST /api/chat`)
``` ```
data: {"type":"thinking","text":"…"}\n\n
data: {"type":"thinking","text":"…"}\n\n
data: {"type":"delta","text":"Hey! "}\n\n data: {"type":"delta","text":"Hey! "}\n\n
data: {"type":"delta","text":"Good "}\n\n data: {"type":"delta","text":"Good "}\n\n
... ...
data: {"type":"done","deflected":false,"sources":[{"source":"Homelab","path":"kubernetes.md","title":"Kubernetes Homelab Cluster"}],"suggestions":[]}\n\n data: {"type":"done","deflected":false,"sources":[{"source":"Homelab","path":"kubernetes.md","title":"Kubernetes Homelab Cluster"}],"suggestions":[]}\n\n
``` ```
Client rules: render deltas as they arrive; on `done` append source chips / Client rules: render deltas as they arrive; render `thinking` text in a
suggestion chips and clear the busy state; on HTTP/stream error show the collapsible block above the answer; auto-collapse on the first `delta`;
error banner + retry (never a stuck button). tolerate interleaved `thinking` events (append — never reopen once the
answer started); the `done` shape is unchanged (thinking never travels on
`done`); on `done` append source chips / suggestion chips and clear the
busy state; on HTTP/stream error show the error banner + retry (never a
stuck button).
> **SSE revision (phase 17, owner permission 2026-08-23):** the contract
> gains one event type — `{"type":"thinking","text":"…"}` — carrying the
> model's reasoning ahead of the `delta` events (the `turbo` model emits
> `delta.reasoning_content` chunks before the first content chunk, verified
> live 2026-08-23; `BOR_STREAM_THINKING=0` suppresses the frames
> server-side). `delta` and `done` shapes are unchanged — a recorded
> extension of A15, not a silent deviation.
--- ---
@@ -286,6 +301,7 @@ Rules:
|-------|----| |-------|----|
| **Idle** | Send button enabled, label "Send". | | **Idle** | Send button enabled, label "Send". |
| **Thinking (pre-token)** | 3-dot typing bubble + button disabled with spinner, label "Thinking…". | | **Thinking (pre-token)** | 3-dot typing bubble + button disabled with spinner, label "Thinking…". |
| **Thinking (model reasoning)** | Collapsible `.thinking` block streams open (replaces the typing dots as the live indicator), auto-collapses on the first answer token, toggleable afterwards, persisted with the message (phase 14); 120s guard clears on the first `thinking` *or* `delta` event. |
| **Streaming** | Deltas append live into the brain bubble; button stays busy. | | **Streaming** | Deltas append live into the brain bubble; button stays busy. |
| **Done (answer)** | Source chips under the bubble (mono, path-based); button re-enabled. | | **Done (answer)** | Source chips under the bubble (mono, path-based); button re-enabled. |
| **Done (deflected)** | Amber-bordered bubble + "Maybe try:" suggestion chips. | | **Done (deflected)** | Amber-bordered bubble + "Maybe try:" suggestion chips. |
@@ -293,6 +309,9 @@ Rules:
| **KB offline** | Amber banner at top of chat ("start Postgres…"); chat disabled with explanation. | | **KB offline** | Amber banner at top of chat ("start Postgres…"); chat disabled with explanation. |
| **Guard** | 120s client-side timeout → error state (a button can never sit "stuck" forever). | | **Guard** | 120s client-side timeout → error state (a button can never sit "stuck" forever). |
> The **Thinking (model reasoning)** row is a phase-17 addition (owner
> permission 2026-08-23) — see the §4 SSE revision.
### 7.5 Component inventory (ids used by tests) ### 7.5 Component inventory (ids used by tests)
`#messages` (stream), `#empty-state`, `#suggestions`, `.suggestion-chip`, `#messages` (stream), `#empty-state`, `#suggestions`, `.suggestion-chip`,
`#composer`, `#message-input`, `#send-btn` / `#send-label`, `#typing-indicator`, `#composer`, `#message-input`, `#send-btn` / `#send-label`, `#typing-indicator`,
@@ -301,7 +320,9 @@ Rules:
`#stat-last`, `#docs-table`, `#docs-tbody`, `#sources-empty`; viewer `#stat-last`, `#docs-table`, `#docs-tbody`, `#sources-empty`; viewer
(Phase 10): `/document.html`, `#doc-title`, `#doc-meta`, `#doc-content`, (Phase 10): `/document.html`, `#doc-title`, `#doc-meta`, `#doc-content`,
`.doc-raw`, `.format-badge`, `#doc-not-found`, `.doc-link` (Sources table `.doc-raw`, `.format-badge`, `#doc-not-found`, `.doc-link` (Sources table
path links). path links); thinking (phase 17, owner permission 2026-08-23):
`.thinking`, `.thinking-text` (collapsible thinking block; plain
`<summary>`, no id).
--- ---
@@ -322,7 +343,10 @@ path links).
- **App logs:** single-line `timestamp LEVEL logger :: message` on stdout; - **App logs:** single-line `timestamp LEVEL logger :: message` on stdout;
uvicorn access logs on. INFO by default (`BOR_LOG_LEVEL`). uvicorn access logs on. INFO by default (`BOR_LOG_LEVEL`).
- **Per-chat-turn log line (required):** - **Per-chat-turn log line (required):**
`question=… embed_ms=… top_score=… fts_hits=… threshold=… deflected=… sources=… total_ms=…` `question=… embed_ms=… top_score=… fts_hits=… tuning=N threshold=… deflected=… sources=… thinking_chars=… total_ms=…`
(`thinking_chars=` counts the turn's reasoning chars — phase 17, owner
permission 2026-08-23 — and is counted even when `BOR_STREAM_THINKING=0`
suppresses the frames.)
- **Importer logs:** per-file `added|updated|unchanged|pruned` + summary - **Importer logs:** per-file `added|updated|unchanged|pruned` + summary
(counts, embedding batches, total time). (counts, embedding batches, total time).
- **`query_log` table:** durable record of every question (score, deflection, - **`query_log` table:** durable record of every question (score, deflection,
@@ -385,6 +409,10 @@ ranks live hybrid results for a question (retrieval tuning).
| 08 | `08_story_dark_tech_theme.md` | `dark-tech-theme.md` | `tests/e2e/test_dark_tech_theme.py` | | 08 | `08_story_dark_tech_theme.md` | `dark-tech-theme.md` | `tests/e2e/test_dark_tech_theme.py` |
| 09 | `09_story_retrieval_quality.md` | `retrieval-quality.md` | `tests/e2e/test_retrieval_quality.py` | | 09 | `09_story_retrieval_quality.md` | `retrieval-quality.md` | `tests/e2e/test_retrieval_quality.py` |
| 10 | `10_story_document_viewer.md` | `document-viewer.md` | `tests/e2e/test_document_viewer.py` | | 10 | `10_story_document_viewer.md` | `document-viewer.md` | `tests/e2e/test_document_viewer.py` |
| 17 | `17_thinking_display.md` | `thinking-display.md` | `tests/e2e/test_thinking_display.py` |
> Row 17 (thinking display) added 2026-08-23 with owner permission — the
> A15 SSE extension recorded in §4.
Completion = unit+integration green, coverage >90%, story E2E green in Completion = unit+integration green, coverage >90%, story E2E green in
isolation, UI verification passed, **one `--no-gpg-sign` commit**. isolation, UI verification passed, **one `--no-gpg-sign` commit**.
+1
View File
@@ -16,6 +16,7 @@ BOR_LLM_API_KEY= # falls back to $AIPI_KEY, then "not-needed"
BOR_LLM_CHAT_MODEL=turbo BOR_LLM_CHAT_MODEL=turbo
BOR_LLM_EMBED_MODEL=embed BOR_LLM_EMBED_MODEL=embed
BOR_EMBEDDING_DIM=768 # verified 2026-08 via scripts/llm_probe.py BOR_EMBEDDING_DIM=768 # verified 2026-08 via scripts/llm_probe.py
BOR_STREAM_THINKING=1 # stream the model's thinking as `thinking` SSE events (0 to suppress)
# --- RAG tuning --- # --- RAG tuning ---
BOR_TOP_N_DOCS=2 BOR_TOP_N_DOCS=2
+15
View File
@@ -87,6 +87,21 @@ uv run uvicorn app.main:app --reload
anonymous visitors see a sign-in gate instead (the catalog is what the anonymous visitors see a sign-in gate instead (the catalog is what the
login locks; the document viewer itself stays open to everyone). login locks; the document viewer itself stays open to everyone).
## Thinking
The self-hosted `turbo` model reasons before it answers. That reasoning is
streamed with the turn as `thinking` SSE events and shown in a
**collapsible "Thinking" block** above the answer bubble: it opens and
fills in live while the model thinks, tucks itself away the moment the
first answer token lands, and stays click-toggleable afterwards. Thinking
persists with the message, so a reloaded conversation restores the block
(collapsed) alongside the answer. How much the model thinks — or whether
it thinks at all — is the model's call: turns without reasoning render
exactly as before.
To hide it, set `BOR_STREAM_THINKING=0` — the `thinking` events stop
(the per-turn log line still counts `thinking_chars`).
## Admin & sign-in ## Admin & sign-in
Brain of Reese has exactly **one account: the admin (you)**. Signing in Brain of Reese has exactly **one account: the admin (you)**. Signing in
+34 -5
View File
@@ -8,6 +8,13 @@ persona prompt (PLAN §6) → ``turbo`` streamed as ``delta`` events → final
Mid-stream failures become a structured ``error`` event; a pre-stream DB Mid-stream failures become a structured ``error`` event; a pre-stream DB
outage is a plain 503 JSON. outage is a plain 503 JSON.
Thinking (phase 17, PLAN §4 extension, owner permission 2026-08-23): the
model's reasoning arrives ahead of the answer and is streamed as
``thinking`` events before the ``delta`` events of the same turn. Each
turn's thinking is counted in the per-turn log line
(``thinking_chars=N``); ``BOR_STREAM_THINKING=0`` suppresses the
``thinking`` frames (the pieces are still counted).
Honesty gate (A8, revised 2026-08-21): LOW — deflection — only when the Honesty gate (A8, revised 2026-08-21): LOW — deflection — only when the
best cosine is strictly below ``BOR_RELEVANCE_THRESHOLD`` **and** no best cosine is strictly below ``BOR_RELEVANCE_THRESHOLD`` **and** no
candidate chunk FTS-matches the question (``fts_hits == 0``). A candidate chunk FTS-matches the question (``fts_hits == 0``). A
@@ -39,11 +46,22 @@ from app.api.steering import load_steering_notes
from app.config import Settings, get_settings from app.config import Settings, get_settings
from app.db import db_available, get_db from app.db import db_available, get_db
from app.models import Document, QueryLog from app.models import Document, QueryLog
from app.rag.llm import EmbeddingError, LLMClient, LLMError from app.rag.llm import (
EmbeddingError,
LLMClient,
LLMError,
StreamPiece, # noqa: F401 (phase 17 typing: chat_stream yields StreamPiece)
)
from app.rag.prompts import build_deflect_prompt, build_high_prompt from app.rag.prompts import build_deflect_prompt, build_high_prompt
from app.rag.retriever import RetrievedChunk, retrieve, select_documents, weak_hit_titles from app.rag.retriever import RetrievedChunk, retrieve, select_documents, weak_hit_titles
from app.rag.suggestions import derive_suggestions from app.rag.suggestions import derive_suggestions
from app.schemas import ChatDoneEvent, ChatErrorEvent, ChatRequest, SourceRef from app.schemas import (
ChatDoneEvent,
ChatErrorEvent,
ChatRequest,
ChatThinkingEvent,
SourceRef,
)
logger = logging.getLogger("app.chat") logger = logging.getLogger("app.chat")
router = APIRouter(tags=["chat"]) router = APIRouter(tags=["chat"])
@@ -202,9 +220,19 @@ async def chat(
] ]
# 3. Stream the answer (grounded, or an honest deflection). # 3. Stream the answer (grounded, or an honest deflection).
# Phase 17: thinking pieces stream as ``thinking`` events
# ahead of the ``delta`` events (PLAN §4 extension); the
# kill-switch (``BOR_STREAM_THINKING=0``) suppresses the
# frames, not the counting.
thinking_chars = 0
try: try:
async for piece in llm.chat_stream(messages): async for piece in llm.chat_stream(messages): # StreamPiece (phase 17)
yield sse_event({"type": "delta", "text": piece}) if piece.kind == "thinking":
thinking_chars += len(piece.text)
if settings.stream_thinking:
yield sse_event(ChatThinkingEvent(text=piece.text).model_dump())
else:
yield sse_event({"type": "delta", "text": piece.text})
except LLMError as e: except LLMError as e:
logger.error( logger.error(
"chat: LLM stream failed question=%r total_ms=%d — %s", "chat: LLM stream failed question=%r total_ms=%d — %s",
@@ -239,7 +267,7 @@ async def chat(
logger.info( logger.info(
"question=%r embed_ms=%d top_score=%.3f fts_hits=%d tuning=%d threshold=%.2f " "question=%r embed_ms=%d top_score=%.3f fts_hits=%d tuning=%d threshold=%.2f "
"deflected=%s sources=%r total_ms=%d", "deflected=%s sources=%r thinking_chars=%d total_ms=%d",
request.message, request.message,
embed_ms, embed_ms,
plan.top_score, plan.top_score,
@@ -248,6 +276,7 @@ async def chat(
settings.relevance_threshold, settings.relevance_threshold,
plan.deflected, plan.deflected,
source_paths, source_paths,
thinking_chars,
total_ms, total_ms,
) )
yield sse_event( yield sse_event(
+5
View File
@@ -41,6 +41,11 @@ class Settings(BaseSettings):
llm_api_key: str = "" llm_api_key: str = ""
llm_chat_model: str = "turbo" llm_chat_model: str = "turbo"
llm_embed_model: str = "embed" llm_embed_model: str = "embed"
#: Operator kill-switch for the ``thinking`` SSE events (phase 17,
#: ``BOR_STREAM_THINKING``; ``0``/``false`` → off). When off, thinking
#: pieces are still counted for the per-turn log line but never
#: emitted — the answer stream itself is unchanged.
stream_thinking: bool = True
# --- RAG tuning --- # --- RAG tuning ---
embedding_dim: int = 768 # verified against aipi /v1 (embed model) embedding_dim: int = 768 # verified against aipi /v1 (embed model)
+54 -14
View File
@@ -1,7 +1,12 @@
"""Async OpenAI-compatible client for the self-hosted aipi endpoint (PLAN A5). """Async OpenAI-compatible client for the self-hosted aipi endpoint (PLAN A5).
Provides the embeddings surface (importer, retrieval) and chat streaming Provides the embeddings surface (importer, retrieval) and chat streaming
(PLAN A15) for the RAG pipeline. (PLAN A15) for the RAG pipeline. Chat streaming yields typed
:class:`StreamPiece` values (phase 17): aipi's ``turbo`` model streams
its reasoning as ``delta.reasoning_content`` chunks (deepseek/litellm
wire convention, verified live 2026-08-23) **before** the answer's
``delta.content`` chunks, and reasoning counts against ``max_tokens``
(an answer can in principle be empty).
Fail-loud rule (PLAN A6): the ``chunks.embedding`` column is fixed at 768 Fail-loud rule (PLAN A6): the ``chunks.embedding`` column is fixed at 768
dimensions when the table is created, so a model that returns a different dimensions when the table is created, so a model that returns a different
@@ -12,7 +17,8 @@ from __future__ import annotations
import logging import logging
from collections.abc import AsyncIterator from collections.abc import AsyncIterator
from typing import cast from dataclasses import dataclass
from typing import Literal, cast
from openai import AsyncOpenAI from openai import AsyncOpenAI
from openai.types.chat import ChatCompletionMessageParam from openai.types.chat import ChatCompletionMessageParam
@@ -34,6 +40,19 @@ class LLMError(RuntimeError):
"""The chat-completions endpoint failed (network, HTTP, or mid-stream).""" """The chat-completions endpoint failed (network, HTTP, or mid-stream)."""
@dataclass(frozen=True)
class StreamPiece:
"""One piece of a streamed chat turn (phase 17, PLAN §4 extension).
``kind`` is ``"content"`` for answer text (an SSE ``delta`` frame)
or ``"thinking"`` for the model's reasoning (an SSE ``thinking``
frame). Frozen: pieces are immutable wire values, not accumulators.
"""
kind: Literal["content", "thinking"]
text: str
# aipi's local embedding model rejects requests over ~1024 input tokens # aipi's local embedding model rejects requests over ~1024 input tokens
# ("input is too large to process"). Batch by estimated tokens, with a # ("input is too large to process"). Batch by estimated tokens, with a
# safety margin under that cap — code-dense text can tokenize at ~3 # safety margin under that cap — code-dense text can tokenize at ~3
@@ -168,17 +187,30 @@ class LLMClient:
(vec,) = await self.embed([text]) (vec,) = await self.embed([text])
return vec return vec
async def chat_stream(self, messages: list[dict[str, str]]) -> AsyncIterator[str]: async def chat_stream(
"""Stream assistant text deltas from the chat model (PLAN A5/A15). self, messages: list[dict[str, str]]
) -> AsyncIterator[StreamPiece]:
"""Stream assistant pieces from the chat model (PLAN A5/A15, phase 17).
``stream=True`` against the OpenAI-compatible endpoint; yields only ``stream=True`` against the OpenAI-compatible endpoint, yielding
non-empty ``delta.content`` pieces. Any failure (network, HTTP, typed :class:`StreamPiece` values. Wire convention (verified live
malformed stream) surfaces as :class:`LLMError` so the API layer can against aipi's ``turbo`` on 2026-08-23): the model's reasoning
turn it into an SSE ``error`` event instead of a hung request. arrives as ``delta.reasoning_content`` chunks (deepseek/litellm
convention) **before** the first ``delta.content`` chunk, so in
practice thinking pieces precede content pieces. The ``openai``
SDK keeps unknown delta fields in ``model_extra``, so ``getattr``
is the right accessor — no raw-HTTP parsing is needed. A chunk
carrying both fields yields the thinking piece **first**.
Answers are allowed up to ``BOR_MAX_OUTPUT_TOKENS`` (default 32 768) Reasoning counts against ``max_tokens``: an answer can in principle
output tokens — the old hard 700-token cap cut long answers off be empty (thinking with no content) — the UI handles that.
mid-sentence (owner report 2026-08-22). Answers are allowed up to ``BOR_MAX_OUTPUT_TOKENS`` (default
32 768) output tokens — the old hard 700-token cap cut long
answers off mid-sentence (owner report 2026-08-22).
Any failure (network, HTTP, malformed stream) surfaces as
:class:`LLMError` so the API layer can turn it into an SSE
``error`` event instead of a hung request.
""" """
try: try:
# ``{role, content}`` dicts are exactly what the message params # ``{role, content}`` dicts are exactly what the message params
@@ -193,9 +225,17 @@ class LLMClient:
async for chunk in stream: async for chunk in stream:
if not chunk.choices: if not chunk.choices:
continue continue
piece = chunk.choices[0].delta.content delta = chunk.choices[0].delta
if piece: reasoning = getattr(delta, "reasoning_content", None)
yield piece if not reasoning:
# Future-proofing: the same wire convention under a
# shorter field name.
reasoning = getattr(delta, "reasoning", None)
if reasoning:
yield StreamPiece("thinking", reasoning)
content = delta.content
if content:
yield StreamPiece("content", content)
except LLMError: except LLMError:
raise raise
except Exception as e: # noqa: BLE001 — wrap transport-level failures except Exception as e: # noqa: BLE001 — wrap transport-level failures
+15
View File
@@ -45,6 +45,21 @@ class SourceRef(BaseModel):
title: str title: str
class ChatThinkingEvent(BaseModel):
"""SSE thinking event: one chunk of the model's reasoning (phase 17).
PLAN §4 extension (A15, owner permission 2026-08-23): frames of the
shape ``{type: "thinking", text: str}`` stream ahead of the
``delta`` frames in practice (the model reasons before it answers). The
client renders them in a collapsible "Thinking" block; the ``done``
event shape is unchanged and thinking text never travels on it.
Sibling of :class:`ChatErrorEvent`.
"""
type: str = "thinking"
text: str
class ChatDoneEvent(BaseModel): class ChatDoneEvent(BaseModel):
"""Final SSE event of a chat turn: metadata for the finished answer.""" """Final SSE event of a chat turn: metadata for the finished answer."""
+111 -16
View File
@@ -14,16 +14,28 @@
* • thinking — pre-token: typing dots + disabled "Thinking…" button; * • thinking — pre-token: typing dots + disabled "Thinking…" button;
* after 10s the indicator's aria-label shows elapsed * after 10s the indicator's aria-label shows elapsed
* seconds so screen-reader users are never left guessing. * seconds so screen-reader users are never left guessing.
* Phase 17: while the model streams reasoning (`thinking`
* SSE events), the live collapsible Thinking block IS the
* visible feedback (it replaces the typing dots; the UI
* state stays "thinking" — button still disabled,
* "Thinking…") and the 120s guard clears on the first
* thinking *or* delta event.
* • streaming — the first delta removes the dots and appends live into * • streaming — the first delta removes the dots and appends live into
* the answer bubble; the button stays busy until `done`. * the answer bubble (auto-collapsing the Thinking block,
* phase 17); the button stays busy until `done`.
* • error — red banner (role="alert") with an actionable retry hint; * • error — red banner (role="alert") with an actionable retry hint;
* the 120s guard (TURN_TIMEOUT_MS) catches hung pre-token * the 120s guard (TURN_TIMEOUT_MS) catches hung pre-token
* streams, so the button can never sit zombified. * streams and the sawDone guard (phase 17) catches a
* stream that dies after frames but before `done`, so the
* button can never sit zombified.
* *
* Conversation persistence (phase 14) makes the chat a durable LOCAL * Conversation persistence (phase 14) makes the chat a durable LOCAL
* session: the message list (raw text + turn metadata) lives in * session: the message list (raw text + turn metadata) lives in
* localStorage under the versioned key `bor.chat.v1` and is re-rendered on * localStorage under the versioned key `bor.chat.v1` and is re-rendered on
* load — refresh, tab close, and a trip to Sources never lose it. A10 is * load — refresh, tab close, and a trip to Sources never lose it. Phase
* 17: a brain record may carry an optional `thinking` field — the
* collapsed Thinking block is restored with it; records without it (old
* sessions) restore exactly as before, so no version bump. A10 is
* untouched: the API stays stateless, nothing is stored server-side. * untouched: the API stays stateless, nothing is stored server-side.
* "New chat" (#new-chat-btn) clears the key + the list back to the empty * "New chat" (#new-chat-btn) clears the key + the list back to the empty
* state. * state.
@@ -77,6 +89,10 @@ const SEND_STATUS = Object.freeze({
const TYPING_LABEL = "Brain of Reese is thinking"; const TYPING_LABEL = "Brain of Reese is thinking";
const ERROR_HINT = "Try again — if this persists, check the LLM is reachable."; const ERROR_HINT = "Try again — if this persists, check the LLM is reachable.";
/* A turn with no answer content (an empty stream, or reasoning that
exhausted max_tokens — phase 17) still renders a bubble, and this exact
text is what gets persisted: what the user saw is what is stored. */
const EMPTY_ANSWER_FALLBACK = "Hmm — that came back empty. Ask me again?";
/* Calm, don't remove: smooth scrolling is the one motion JS controls. */ /* Calm, don't remove: smooth scrolling is the one motion JS controls. */
const reducedMotion = const reducedMotion =
@@ -347,6 +363,33 @@ function removeTyping() {
document.querySelector("#typing-indicator")?.remove(); document.querySelector("#typing-indicator")?.remove();
} }
/* ---------- thinking block (phase 17) ----------
* The model's reasoning streams into a collapsible <details> block ABOVE
* the answer bubble: created OPEN on the first `thinking` event,
* auto-collapsed when the first answer token lands, and user-toggleable
* afterwards (native <details>/<summary> — a real focusable control).
* ensureThinkingBlock is idempotent (returns the existing block if any);
* closeThinkingBlock never reopens a block once the answer has started,
* so a late/interleaved `thinking` event only appends to the closed text. */
function ensureThinkingBlock(wrap) {
let block = wrap.querySelector(".thinking");
if (!block) {
block = document.createElement("details");
block.className = "thinking";
block.open = true;
block.innerHTML =
`<summary>Thinking</summary><div class="thinking-text"></div>`;
const body = wrap.querySelector(".msg-body");
body.insertBefore(block, body.querySelector(".bubble"));
}
return block;
}
function closeThinkingBlock(wrap) {
const block = wrap?.querySelector?.(".thinking");
if (block) block.open = false; // idempotent; no-op without a block
}
/* ---------- suggestions (shared chip component, phase 05) ---------- /* ---------- suggestions (shared chip component, phase 05) ----------
* *
* One component, two homes: the onboarding row in the empty state and the * One component, two homes: the onboarding row in the empty state and the
@@ -555,7 +598,8 @@ function appendMaybeTry(wrap, suggestions) {
* lives in localStorage under a versioned key; a format bump = clean start: * lives in localStorage under a versioned key; a format bump = clean start:
* *
* bor.chat.v1 → { v: 1, messages: [{ who: "user"|"brain", text, * bor.chat.v1 → { v: 1, messages: [{ who: "user"|"brain", text,
* sources?, deflected?, suggestions? }] } * sources?, deflected?, suggestions?,
* thinking? }] }
* *
* Only RAW TEXT is stored — restore re-renders it through the escape-first * Only RAW TEXT is stored — restore re-renders it through the escape-first
* markdown renderer, so no HTML is ever persisted. Save points: the user * markdown renderer, so no HTML is ever persisted. Save points: the user
@@ -631,6 +675,12 @@ function renderStoredMessage(m) {
return; return;
} }
const wrap = addMessage("brain", renderMarkdown(m.text), "auto"); const wrap = addMessage("brain", renderMarkdown(m.text), "auto");
if (m.thinking) {
// Phase 17: restore the thinking block COLLAPSED above the bubble.
const block = ensureThinkingBlock(wrap);
block.open = false;
block.querySelector(".thinking-text").innerHTML = renderMarkdown(m.thinking);
}
if (m.deflected) { if (m.deflected) {
wrap.classList.add("is-deflected"); wrap.classList.add("is-deflected");
appendMaybeTry(wrap, m.suggestions); appendMaybeTry(wrap, m.suggestions);
@@ -647,8 +697,11 @@ function restoreConversation() {
for (const m of conversation) renderStoredMessage(m); for (const m of conversation) renderStoredMessage(m);
} }
/* Brain message save point (on `done`): raw accumulated text only. An /* Brain message save point (on `done`): raw accumulated text + metadata.
empty stream keeps the "…" placeholder that was actually rendered. */ Phase 17: meta.thinking is optional — `undefined` drops the key from
the JSON, so turns without thinking persist exactly as before. An empty
answer keeps the fallback/"…" text that was actually rendered — what
the user saw is what is stored. */
function rememberBrainTurn(rawText, meta) { function rememberBrainTurn(rawText, meta) {
conversation.push({ who: "brain", text: rawText || "…", ...meta }); conversation.push({ who: "brain", text: rawText || "…", ...meta });
saveConversation(); saveConversation();
@@ -754,11 +807,16 @@ async function handleSend(e) {
let acc = ""; let acc = "";
let res = null; let res = null;
let aborted = false; // the 120s guard already took the turn to error let aborted = false; // the 120s guard already took the turn to error
// Phase 17 (thinking display): turn-local reasoning state.
let thinkingAcc = ""; // accumulated thinking text (persisted with the turn)
let sawThinking = false; // did any `thinking` frame arrive this turn?
let sawDone = false; // did the stream end with a `done` event?
try { try {
// thinking = pre-token: dots + busy button. The guard is armed so a // thinking = pre-token: dots + busy button. The guard is armed so a
// hung stream can never leave the button zombified; it clears on the // hung stream can never leave the button zombified; it clears on the
// first delta (entering "streaming") and on every terminal transition. // first thinking OR delta event (phase 17) and on every terminal
// transition.
setUiState(UI_STATE.thinking); setUiState(UI_STATE.thinking);
armTurnTimeout(() => { armTurnTimeout(() => {
aborted = true; aborted = true;
@@ -781,16 +839,34 @@ async function handleSend(e) {
} }
await readSSE(res, (ev) => { await readSSE(res, (ev) => {
if (aborted) return; if (aborted) return;
if (ev.type === "delta") { if (ev.type === "thinking") {
acc += ev.text || ""; // Phase 17: model reasoning — stream it live into the collapsible
if (!wrap) { // Thinking block. No setUiState here: the UI state stays
// First token: dots out, live bubble in; the button stays busy. // "thinking" (button still disabled with "Thinking…", #send-status
setUiState(UI_STATE.streaming); // unchanged) — the live block simply replaces the typing dots as
wrap = addMessage("brain", ""); // the visible feedback.
thinkingAcc += ev.text || "";
sawThinking = true;
clearTurnTimeout(); // the stream is alive — as the first delta says
if (!wrap) wrap = addMessage("brain", "");
removeTyping(); // the live block replaces the dots as feedback
const block = ensureThinkingBlock(wrap);
const textEl = block.querySelector(".thinking-text");
textEl.innerHTML = renderMarkdown(thinkingAcc); // escape-first, XSS-safe
if (block.open) {
textEl.scrollTop = textEl.scrollHeight; // pin the stream to the bottom
wrap.scrollIntoView({ behavior: SCROLL, block: "end" });
} }
} else if (ev.type === "delta") {
acc += ev.text || "";
if (uiState === UI_STATE.thinking) setUiState(UI_STATE.streaming);
if (!wrap) wrap = addMessage("brain", ""); // first token: live bubble in
closeThinkingBlock(wrap); // auto-collapse; idempotent, never reopens
wrap.querySelector(".bubble").innerHTML = renderMarkdown(acc); wrap.querySelector(".bubble").innerHTML = renderMarkdown(acc);
wrap.scrollIntoView({ behavior: SCROLL, block: "end" }); wrap.scrollIntoView({ behavior: SCROLL, block: "end" });
} else if (ev.type === "done") { } else if (ev.type === "done") {
sawDone = true;
closeThinkingBlock(wrap); // the turn is over: settle the block closed
if (!wrap) { if (!wrap) {
setUiState(UI_STATE.streaming); setUiState(UI_STATE.streaming);
wrap = addMessage("brain", "…"); wrap = addMessage("brain", "…");
@@ -801,9 +877,18 @@ async function handleSend(e) {
} }
appendSources(wrap, ev.sources); appendSources(wrap, ev.sources);
appendTuneButton(wrap); // every completed brain bubble is tunable appendTuneButton(wrap); // every completed brain bubble is tunable
// Thinking-without-answer (reasoning can exhaust max_tokens): the
// bubble gets the empty-answer fallback — what the user saw is
// what gets persisted.
const finalText = acc || (sawThinking ? EMPTY_ANSWER_FALLBACK : "");
if (!acc && sawThinking) {
wrap.querySelector(".bubble").innerHTML = renderMarkdown(finalText);
}
// Persistence save point 2: the answer lands only when the turn is // Persistence save point 2: the answer lands only when the turn is
// complete (raw text + the done metadata). // complete (raw text + the done metadata; phase 17: + optional
rememberBrainTurn(acc, { // thinking — `undefined` drops the key from the JSON).
rememberBrainTurn(finalText || acc, {
thinking: thinkingAcc || undefined,
deflected: !!ev.deflected, deflected: !!ev.deflected,
sources: ev.sources, sources: ev.sources,
suggestions: ev.suggestions, suggestions: ev.suggestions,
@@ -812,8 +897,18 @@ async function handleSend(e) {
throw new Error(ev.detail || "Something went wrong on my side."); throw new Error(ev.detail || "Something went wrong on my side.");
} }
}); });
// Stream-drop guard (phase 17): frames arrived but no `done` event —
// the connection died mid-turn. Say so; never settle silently into
// idle with a half bubble. The zero-frame case falls through to the
// existing empty-answer fallback below.
if (!sawDone && !aborted && (acc || thinkingAcc)) {
setUiState(
UI_STATE.error,
"The stream ended before my answer finished — try again?"
);
}
if (!aborted && !wrap) { if (!aborted && !wrap) {
const fallback = "Hmm — that came back empty. Ask me again?"; const fallback = EMPTY_ANSWER_FALLBACK;
const fwrap = addMessage("brain", fallback); const fwrap = addMessage("brain", fallback);
appendTuneButton(fwrap); appendTuneButton(fwrap);
rememberBrainTurn(fallback, {}); // persist what the user actually saw rememberBrainTurn(fallback, {}); // persist what the user actually saw
+54
View File
@@ -401,6 +401,58 @@ body::after {
border-color: var(--accent-line); border-color: var(--accent-line);
} }
/* Collapsible "Thinking" block (phase 17): the model's reasoning streams
open ABOVE the answer bubble, auto-collapses when the answer starts, and
stays user-toggleable (native <details>/<summary> — a real focusable
control: >=44px target, :focus-visible via the global rule). Summary
text is brand-ink on surface ≈8.7:1; the scratchpad body is ink-soft on
surface ≈6.9:1 — both AA. */
details.thinking {
background: var(--surface);
border: 1px solid var(--line);
border-left: 3px solid var(--brand-soft);
border-radius: var(--radius-sm);
margin: 0 0 0.5rem;
overflow: hidden;
}
details.thinking summary {
display: flex;
align-items: center;
gap: 0.5rem;
padding: 0.5rem 0.75rem;
min-height: 44px;
color: var(--brand-ink); /* 8.7:1 on --surface */
font-size: 0.9rem;
cursor: pointer;
list-style: none;
}
details.thinking summary::-webkit-details-marker { display: none; }
/* CSS chevron: ▸ rotates 90° when open (transition stills under
prefers-reduced-motion — see the reduced-motion block below). */
details.thinking summary::before {
content: "▸";
display: inline-block;
transition: transform 0.15s ease;
}
details.thinking[open] summary::before { transform: rotate(90deg); }
details.thinking summary:focus-visible {
outline: 3px solid var(--brand);
outline-offset: 2px;
}
/* The scratchpad is a scrollable, compact area (max-height keeps long
reasoning from pushing the answer off-screen while open). */
details.thinking .thinking-text {
padding: 0 0.75rem 0.75rem;
color: var(--ink-soft); /* 6.9:1 on --surface */
font-size: 0.875rem;
line-height: 1.55;
max-height: 320px;
overflow-y: auto;
}
/* The scratchpad is compact: tighten the renderer's paragraph/list margins. */
details.thinking .thinking-text p,
details.thinking .thinking-text ul { margin: 0 0 0.5rem; }
.msg-meta { .msg-meta {
font-size: 0.75rem; font-size: 0.75rem;
color: var(--ink-soft); color: var(--ink-soft);
@@ -631,6 +683,8 @@ body::after {
} }
@media (prefers-reduced-motion: reduce) { @media (prefers-reduced-motion: reduce) {
.typing span { animation: none; opacity: 0.7; } .typing span { animation: none; opacity: 0.7; }
/* Phase 17 thinking block: the chevron stills (no rotation motion). */
details.thinking summary::before { transition: none; }
} }
/* ---------- Empty state & suggestions ---------- */ /* ---------- Empty state & suggestions ---------- */
+65 -7
View File
@@ -15,6 +15,9 @@ Implements just enough of the aipi surface:
- otherwise -> upbeat answer quoting the provided document context - otherwise -> upbeat answer quoting the provided document context
- user message containing ``pretend to think slowly`` -> 3s warm-up delay - user message containing ``pretend to think slowly`` -> 3s warm-up delay
(used by the loading-feedback story). (used by the loading-feedback story).
- user message containing ``think out loud`` -> the answer is preceded by
~800 chars of deterministic ``reasoning_content`` chunks (the
thinking-display story, phase 17).
- system prompt containing ``<tuning>`` (phase 15, steering notes) -> - system prompt containing ``<tuning>`` (phase 15, steering notes) ->
the composed answer ends with `` (tuning: <first note line>)`` — the composed answer ends with `` (tuning: <first note line>)`` —
makes prompt injection observable in the UI deterministically. makes prompt injection observable in the UI deterministically.
@@ -79,6 +82,13 @@ LONG_ANSWER_TRIGGER = "write a long answer"
LONG_ANSWER_LINES = 40 LONG_ANSWER_LINES = 40
LONG_ANSWER_END = "LONG-ANSWER-END" LONG_ANSWER_END = "LONG-ANSWER-END"
#: Phase 17 (thinking-display story): a user message containing this
#: substring (case-insensitive) is answered with a deterministic
#: ``reasoning_content`` stream ahead of the content — same convention as
#: the other user-message triggers above. Existing E2E questions do not
#: contain the substring, so every other suite is unaffected.
THINKING_TRIGGER = "think out loud"
def long_answer() -> str: def long_answer() -> str:
"""~900-word deterministic walkthrough (phase 11): numbered steps plus """~900-word deterministic walkthrough (phase 11): numbered steps plus
@@ -142,6 +152,32 @@ def compose_answer(body: dict[str, Any]) -> str:
return answer return answer
def compose_thinking(body: dict[str, Any]) -> str:
"""Deterministic reasoning scratchpad (thinking-display story, phase 17).
A fixed 4-line "Step 1… Step 4" template quoting the first ~60 chars
of the user question: unique per question, byte-stable across runs,
~700–900 chars total (≈ 60–75 frames at the mock's 12-char/0.02s
pacing). The ``Step 2: Check my notes`` line fragment is what the E2E
assertions key off.
"""
q = _user(body).strip()[:60]
return (
f"Step 1: Read the question carefully — “{q}” — and figure out what kind of "
"answer it wants (a how-to, a lookup, or a design decision) before touching "
"the docs, so I don't over- or under-answer.\n"
"Step 2: Check my notes for the closest match. The homelab kubernetes file "
"is the obvious candidate, but I should also consider whether a deployments "
"note covers the same ground better.\n"
"Step 3: Re-read the relevant sections top to bottom so every specific — "
"hosts, versions, ports, schedules — is exact as written rather than "
"remembered, and note which document each fact comes from.\n"
"Step 4: Draft the answer around those specifics, keep it tight with short "
"paragraphs and bullets where it helps, cite the documents by path, and "
"double-check that nothing is invented."
)
@app.post("/__shutdown__") @app.post("/__shutdown__")
def shutdown() -> dict[str, Any]: def shutdown() -> dict[str, Any]:
"""Test hook (loading-feedback story): terminate this mock process to """Test hook (loading-feedback story): terminate this mock process to
@@ -186,11 +222,31 @@ def embeddings(body: dict[str, Any]) -> dict[str, Any]:
} }
def _sse_stream(answer: str, delay: float) -> Any: def _sse_stream(answer: str, delay: float, thinking: str = "") -> Any:
"""SSE frames for one chat completion (phase 17: + reasoning).
When ``thinking`` is non-empty its 12-char slices go out FIRST as
``delta.reasoning_content`` frames — same 0.02s cadence and envelope
as the content frames, the aipi wire convention (reasoning before
content). Without ``thinking`` the output is byte-identical to the
content-only stream, so the other story suites are unaffected.
"""
model = "turbo" model = "turbo"
chunk_id = f"chatcmpl-{uuid.uuid4()}" chunk_id = f"chatcmpl-{uuid.uuid4()}"
if delay: if delay:
time.sleep(delay) time.sleep(delay)
for piece in re.findall(r".{1,12}", thinking, re.S):
payload = {
"id": chunk_id,
"object": "chat.completion.chunk",
"created": int(time.time()),
"model": model,
"choices": [
{"index": 0, "delta": {"reasoning_content": piece}, "finish_reason": None}
],
}
yield f"data: {json_dumps(payload)}\n\n"
time.sleep(0.02)
for piece in re.findall(r".{1,12}", answer, re.S): for piece in re.findall(r".{1,12}", answer, re.S):
payload = { payload = {
"id": chunk_id, "id": chunk_id,
@@ -239,25 +295,27 @@ def _apply_max_tokens(answer: str, max_tokens: Any) -> str:
def chat_completions(body: dict[str, Any]) -> Any: def chat_completions(body: dict[str, Any]) -> Any:
answer = _apply_max_tokens(compose_answer(body), body.get("max_tokens")) answer = _apply_max_tokens(compose_answer(body), body.get("max_tokens"))
delay = 3.0 if "pretend to think slowly" in _user(body) else 0.0 delay = 3.0 if "pretend to think slowly" in _user(body) else 0.0
thinking = compose_thinking(body) if THINKING_TRIGGER in _user(body).lower() else ""
if not body.get("stream"): if not body.get("stream"):
message: dict[str, Any] = {"role": "assistant", "content": answer}
if thinking:
# Harmless future-proofing: the app only uses streaming, but a
# non-streaming client that reads the field gets the reasoning.
message["reasoning_content"] = thinking
return { return {
"id": f"chatcmpl-{uuid.uuid4()}", "id": f"chatcmpl-{uuid.uuid4()}",
"object": "chat.completion", "object": "chat.completion",
"created": int(time.time()), "created": int(time.time()),
"model": body.get("model", "turbo"), "model": body.get("model", "turbo"),
"choices": [ "choices": [
{ {"index": 0, "message": message, "finish_reason": "stop"}
"index": 0,
"message": {"role": "assistant", "content": answer},
"finish_reason": "stop",
}
], ],
"usage": {"prompt_tokens": 100, "completion_tokens": 50, "total_tokens": 150}, "usage": {"prompt_tokens": 100, "completion_tokens": 50, "total_tokens": 150},
} }
return StreamingResponse( return StreamingResponse(
_sse_stream(answer, delay), _sse_stream(answer, delay, thinking=thinking),
media_type="text/event-stream", media_type="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"}, headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
) )
+281
View File
@@ -0,0 +1,281 @@
"""Phase 17 E2E (Playwright, mock-only): the model's "thinking" display.
Story: ``.agent/user_stories/thinking-display.md``
Run in isolation (DB must be up: ``podman compose up -d db``):
uv run pytest tests/e2e/test_thinking_display.py -v --no-cov
MOCK-ONLY suite: ``E2E_REAL_LLM=1`` is not supported here. The real
``turbo`` thinks on *every* turn, which would break the no-thinking
regression test (scenario 4) — the deterministic mock's ``think out loud``
trigger (mock_llm.py) keeps all five scenarios reproducible.
Test → story mapping (Playwright Mapping Rule):
1. ``test_thinking_block_streams_open_then_collapses``
2. ``test_thinking_toggle_after_done``
3. ``test_thinking_restored_after_reload``
4. ``test_no_thinking_block_without_trigger``
5. ``test_thinking_with_deflection``
Determinism note: the mock paces every SSE frame at 0.02s and the thinking
text is ~700–900 chars (≈ 60–75 frames ≈ 1.2–1.5s) before the first
content frame, so "attach → assert open" runs well inside the open window
on headless Chromium; all other assertions are made after the send button
re-enables (fully settled state).
"""
from __future__ import annotations
import asyncio
import json
import re
from collections.abc import Iterator
from pathlib import Path
from threading import Thread
from typing import Any
import pytest
from playwright.sync_api import Page, expect
from sqlalchemy import text
from app.config import Settings
from app.db import SessionLocal
from app.rag.importer import ImportSummary, import_sources
from app.rag.llm import LLMClient
REPO = Path(__file__).resolve().parents[2]
FIXTURES = REPO / "tests" / "fixtures" / "docs"
THINK_QUESTION = "think out loud — how is my kubernetes cluster set up?"
PLAIN_QUESTION = "How is my Kubernetes cluster set up?"
THINK_DEFLECT_QUESTION = "think out loud — tell me about quantum wormhole cooling"
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
DEFLECT_PHRASE = r"haven't done anything like that"
#: Line fragment the mock's deterministic scratchpad must carry — the
#: suite keys off it (mock_llm.compose_thinking).
THINKING_FRAGMENT = "Step 2: Check my notes"
STORAGE_KEY = "bor.chat.v1"
#: Phase-10 viewer URL + phase-13 back=/ (byte-identical to the chip the
#: persistence suite pins — grounded-turn sources are unchanged by 17).
CHIP_HREF = "/document.html?source=docs&path=homelab%2Fkubernetes.md&back=%2F"
async def _import_fixtures(mock_port: int) -> ImportSummary:
kwargs: dict[str, Any] = {"_env_file": None, "llm_base_url": f"http://127.0.0.1:{mock_port}/v1"}
settings = Settings(**kwargs) # pyright: ignore[reportCallIssue]
return await import_sources([FIXTURES], LLMClient(settings))
def _run_in_thread(coro: Any) -> Any:
"""Run a coroutine on a worker thread.
Playwright's sync API keeps an asyncio loop running on the test thread,
so ``asyncio.run`` cannot be called directly from a test body.
"""
box: dict[str, Any] = {}
def runner() -> None:
try:
box["value"] = asyncio.run(coro)
except BaseException as e: # noqa: BLE001 — re-raised on the test thread
box["error"] = e
t = Thread(target=runner)
t.start()
t.join()
if "error" in box:
raise box["error"]
return box["value"]
def _reset_db(mock_port: int, seed: bool) -> ImportSummary | None:
"""Truncate the KB (and query log), then optionally re-import fixtures."""
with SessionLocal() as db:
db.execute(text("TRUNCATE chunks, documents, query_log"))
db.commit()
if not seed:
return None
return _run_in_thread(_import_fixtures(mock_port))
@pytest.fixture()
def seeded_kb(mock_llm: int, db_ready: None) -> Iterator[None]:
"""A fresh KB seeded from ``tests/fixtures/docs`` (8 docs, A9 formats),
truncated again on teardown. ``db_ready`` (conftest) skips with clear
instructions when Postgres is down."""
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8
yield
_reset_db(mock_llm, seed=False)
def send_and_wait(page: Page, question: str) -> None:
"""Type into #message-input, submit via #composer, then wait until the
last brain message settles (send button re-enabled, label "Send")."""
page.fill("#message-input", question)
page.evaluate("() => document.querySelector('#composer').requestSubmit()")
expect(page.locator(".msg.user .bubble").last).to_contain_text(question)
# The mock streams at 0.02s/chunk, so thinking + answer land in a few
# seconds — 30s is generous on headless Chromium.
expect(page.locator(".msg.brain .bubble").last).not_to_have_text("", timeout=30_000)
expect(page.locator("#send-btn")).to_be_enabled(timeout=30_000)
expect(page.locator("#send-label")).to_have_text("Send")
# ---------------------------------------------------------------------------
# 1. Streaming: the block attaches OPEN at the first thinking event, then
# auto-collapses when the first answer token lands
# ---------------------------------------------------------------------------
def test_thinking_block_streams_open_then_collapses(
page: Page, app_url: str, seeded_kb: None
) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
page.fill("#message-input", THINK_QUESTION)
page.evaluate("() => document.querySelector('#composer').requestSubmit()")
expect(page.locator(".msg.user .bubble").last).to_contain_text(THINK_QUESTION)
# The block attaches at the FIRST thinking event — before any answer
# token — and is created OPEN.
details = page.locator(".msg.brain").last.locator("details.thinking")
details.wait_for(state="attached", timeout=10_000)
# The ~800-char thinking stream (≈1.3s) keeps the block open right
# after attach — assert while it is still streaming.
expect(details).to_have_attribute("open", "")
expect(details.locator(".thinking-text")).not_to_have_text("")
# First answer token: the block auto-collapses and stays closed.
bubble = page.locator(".msg.brain .bubble").last
expect(bubble).not_to_have_text("", timeout=30_000)
expect(details).not_to_have_attribute("open")
# Settled: full scratchpad, grounded mock answer, source chip(s),
# and the re-enabled send button.
expect(details.locator(".thinking-text")).to_contain_text(THINKING_FRAGMENT)
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER)
chip = page.locator(".msg.brain .source-chip", has_text="kubernetes.md")
expect(chip.first).to_be_visible()
expect(chip.first).to_have_attribute("href", CHIP_HREF)
expect(page.locator("#send-btn")).to_be_enabled()
expect(page.locator("#send-label")).to_have_text("Send")
# ---------------------------------------------------------------------------
# 2. Toggle: after settle the block is closed; the summary re-opens it
# (a real keyboard-focusable control) and closes it again
# ---------------------------------------------------------------------------
def test_thinking_toggle_after_done(page: Page, app_url: str, seeded_kb: None) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
send_and_wait(page, THINK_QUESTION)
last = page.locator(".msg.brain").last
details = last.locator("details.thinking")
expect(details).to_have_count(1)
expect(details).not_to_have_attribute("open") # auto-collapsed at first token
# The summary is a real, keyboard-focusable control.
details.locator("summary").focus()
assert page.evaluate("() => document.activeElement.tagName") == "SUMMARY"
# Open: the full scratchpad is visible.
details.locator("summary").click()
expect(details).to_have_attribute("open", "")
text_el = details.locator(".thinking-text")
expect(text_el).to_be_visible()
expect(text_el).to_contain_text(THINKING_FRAGMENT)
expect(text_el).to_contain_text("nothing is invented")
# Closed again — user control in both directions.
details.locator("summary").click()
expect(details).not_to_have_attribute("open")
expect(text_el).not_to_be_visible()
# ---------------------------------------------------------------------------
# 3. Persistence: the thinking block (and its text) survives a reload,
# restored COLLAPSED — phase-14 restore path + phase-17 field
# ---------------------------------------------------------------------------
def test_thinking_restored_after_reload(page: Page, app_url: str, seeded_kb: None) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
send_and_wait(page, THINK_QUESTION)
# The live block is collapsed; capture what it shows and what the
# turn persisted (raw text, same as what was rendered).
details = page.locator(".msg.brain").last.locator("details.thinking")
expect(details).not_to_have_attribute("open")
captured = details.locator(".thinking-text").text_content()
assert captured
# The persisted raw text is the same scratchpad (renderMarkdown turns
# the line breaks into <br>, which textContent drops — compare without
# whitespace).
raw = json.loads(
page.evaluate(f"() => localStorage.getItem('{STORAGE_KEY}')")
)["messages"][1]["thinking"]
assert re.sub(r"\s+", "", raw) == re.sub(r"\s+", "", captured)
page.reload()
expect(page.locator("#empty-state")).to_be_hidden()
restored = page.locator(".msg.brain").last.locator("details.thinking")
expect(restored).to_have_count(1)
expect(restored).not_to_have_attribute("open") # restored COLLAPSED
expect(restored.locator(".thinking-text")).to_have_text(captured)
# Answer bubble + source chip are intact (phase-14 restore path).
expect(page.locator(".msg.brain .bubble").last).to_contain_text(MOCK_ANSWER_MARKER)
chip = page.locator(".msg.brain .source-chip", has_text="kubernetes.md")
expect(chip.first).to_have_attribute("href", CHIP_HREF)
# ---------------------------------------------------------------------------
# 4. No thinking, no block: a model/turn that emits no reasoning renders
# exactly as before (no layout regression)
# ---------------------------------------------------------------------------
def test_no_thinking_block_without_trigger(page: Page, app_url: str, seeded_kb: None) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
send_and_wait(page, PLAIN_QUESTION)
# No trigger → no thinking events → no block anywhere on the page.
expect(page.locator("details.thinking")).to_have_count(0)
# The turn itself is complete and grounded, exactly as before phase 17.
chip = page.locator(".msg.brain .source-chip", has_text="kubernetes.md")
expect(chip.first).to_be_visible()
expect(chip.first).to_have_attribute("href", CHIP_HREF)
# ---------------------------------------------------------------------------
# 5. Coexistence: the honesty gate (deflection) and the thinking block
# on the same turn
# ---------------------------------------------------------------------------
def test_thinking_with_deflection(page: Page, app_url: str, seeded_kb: None) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
send_and_wait(page, THINK_DEFLECT_QUESTION)
last = page.locator(".msg.brain").last
# The honesty gate fired: amber deflected bubble + "Maybe try" chips.
expect(last).to_have_class(re.compile(r"is-deflected"))
expect(last.locator(".bubble")).to_contain_text(
re.compile(DEFLECT_PHRASE, re.IGNORECASE)
)
chips = last.locator(".maybe-try .suggestion-chip")
expect(chips.first).to_be_visible()
assert chips.count() >= 2
# And the thinking block came along, closed, with its scratchpad.
details = last.locator("details.thinking")
expect(details).to_have_count(1)
expect(details).not_to_have_attribute("open")
expect(details.locator(".thinking-text")).to_contain_text(THINKING_FRAGMENT)
+80 -3
View File
@@ -28,7 +28,7 @@ from app.config import Settings, get_settings
from app.main import app as fastapi_app from app.main import app as fastapi_app
from app.models import Chunk, QueryLog from app.models import Chunk, QueryLog
from app.rag.importer import import_sources from app.rag.importer import import_sources
from app.rag.llm import EmbeddingError, LLMError from app.rag.llm import EmbeddingError, LLMError, StreamPiece
FIXTURES = Path(__file__).resolve().parents[1] / "fixtures" / "docs" FIXTURES = Path(__file__).resolve().parents[1] / "fixtures" / "docs"
QUESTION = "How is my Kubernetes cluster set up?" QUESTION = "How is my Kubernetes cluster set up?"
@@ -53,6 +53,7 @@ class FakeRagLLM:
def __init__( def __init__(
self, self,
answer: str = "Hey — you've got this! Talos, Cilium, three nodes. 🧠", answer: str = "Hey — you've got this! Talos, Cilium, three nodes. 🧠",
thinking: str = "",
embed_error: Exception | None = None, embed_error: Exception | None = None,
stream_error: Exception | None = None, stream_error: Exception | None = None,
fail_mid_stream: bool = False, fail_mid_stream: bool = False,
@@ -60,6 +61,7 @@ class FakeRagLLM:
self.settings = Settings(_env_file=None) # pyright: ignore[reportCallIssue] self.settings = Settings(_env_file=None) # pyright: ignore[reportCallIssue]
self.embed_batches = 0 self.embed_batches = 0
self.answer = answer self.answer = answer
self.thinking = thinking
self.embed_error = embed_error self.embed_error = embed_error
self.stream_error = stream_error self.stream_error = stream_error
self.fail_mid_stream = fail_mid_stream self.fail_mid_stream = fail_mid_stream
@@ -77,14 +79,20 @@ class FakeRagLLM:
return _token_vec(text) return _token_vec(text)
async def chat_stream(self, messages: list[dict[str, str]]): async def chat_stream(self, messages: list[dict[str, str]]):
"""Typed stream (phase 17): ``thinking`` slices (same 12-char
cadence as content) **before** the content pieces. With the
default ``thinking=""`` this yields content-only pieces — today's
behavior, new yield type."""
self.seen_messages.append(messages) self.seen_messages.append(messages)
if self.stream_error is not None: if self.stream_error is not None:
raise self.stream_error raise self.stream_error
if self.fail_mid_stream: if self.fail_mid_stream:
yield "partial " yield StreamPiece("content", "partial ")
raise LLMError("mid-stream dropout") raise LLMError("mid-stream dropout")
for i in range(0, len(self.thinking), 12):
yield StreamPiece("thinking", self.thinking[i : i + 12])
for i in range(0, len(self.answer), 12): for i in range(0, len(self.answer), 12):
yield self.answer[i : i + 12] yield StreamPiece("content", self.answer[i : i + 12])
@pytest.fixture() @pytest.fixture()
@@ -150,6 +158,75 @@ def test_chat_streams_deltas_then_done_with_sources(client, db, seeded_kb: FakeR
assert "HONESTY GATE" in system["content"] assert "HONESTY GATE" in system["content"]
def test_chat_streams_thinking_before_deltas(client, db, seeded_kb: FakeRagLLM) -> None:
"""Phase 17: ``thinking`` frames precede every ``delta`` frame and
reassemble to the model's reasoning; the ``done`` contract is
unchanged."""
thinker = FakeRagLLM(
thinking=(
"Step 1: parse the question. Step 2: check the kubernetes doc. "
"Step 3: name Talos, Cilium, three nodes. Step 4: answer."
)
)
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: thinker
try:
_, _, frames = _stream_chat(client, QUESTION)
finally:
fastapi_app.dependency_overrides.clear()
thinking = [f for f in frames if f.get("type") == "thinking"]
deltas = [f for f in frames if f.get("type") == "delta"]
assert len(thinking) >= 1 # genuinely streamed
assert len(deltas) >= 2
# Every thinking frame precedes every delta frame.
ordered = [f["type"] for f in frames if f["type"] in ("thinking", "delta")]
assert ordered == ["thinking"] * len(thinking) + ["delta"] * len(deltas)
assert all(set(f.keys()) == {"type", "text"} for f in thinking)
assert "".join(f["text"] for f in thinking) == thinker.thinking
assert "".join(d["text"] for d in deltas) == thinker.answer
# Done still last; sources unchanged by the thinking extension.
done = frames[-1]
assert done["type"] == "done"
assert done["deflected"] is False
assert done["suggestions"] == []
assert done["sources"][0]["path"] == "homelab/kubernetes.md"
assert done["sources"][0]["source"] == "docs"
assert not any(f.get("type") == "error" for f in frames)
def test_chat_thinking_suppressed_when_disabled(
client, db, monkeypatch: pytest.MonkeyPatch
) -> None:
"""Phase 17 kill-switch: ``BOR_STREAM_THINKING=0`` drops every
``thinking`` frame; the delta stream is byte-identical to the
thinking-free case."""
thinker = FakeRagLLM(thinking="hidden reasoning that must never reach the wire")
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: thinker
# Same honesty gate the conftest/module already use (mock-calibrated
# 0.30 from the environment) — only the kill-switch changes.
live = get_settings()
monkeypatch.setattr(
chat_api,
"get_settings",
lambda: Settings(
_env_file=None, # pyright: ignore[reportCallIssue]
relevance_threshold=live.relevance_threshold,
stream_thinking=False,
),
)
try:
_, _, frames = _stream_chat(client, QUESTION)
finally:
fastapi_app.dependency_overrides.clear()
assert not any(f.get("type") == "thinking" for f in frames)
deltas = [f for f in frames if f.get("type") == "delta"]
assert "".join(d["text"] for d in deltas) == thinker.answer
assert frames[-1]["type"] == "done"
assert not any(f.get("type") == "error" for f in frames)
def test_chat_writes_query_log_row(client, db, seeded_kb: FakeRagLLM) -> None: def test_chat_writes_query_log_row(client, db, seeded_kb: FakeRagLLM) -> None:
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb
try: try:
+2 -1
View File
@@ -20,6 +20,7 @@ from app.api import chat as chat_api
from app.config import Settings from app.config import Settings
from app.main import app as fastapi_app from app.main import app as fastapi_app
from app.models import Document, QueryLog from app.models import Document, QueryLog
from app.rag.llm import StreamPiece
from app.rag.retriever import RetrievedChunk, weak_hit_titles from app.rag.retriever import RetrievedChunk, weak_hit_titles
from app.rag.suggestions import MAX_SUGGESTIONS, derive_suggestions from app.rag.suggestions import MAX_SUGGESTIONS, derive_suggestions
@@ -276,7 +277,7 @@ class _CannedLLM:
async def chat_stream(self, messages: list[dict[str, str]]): async def chat_stream(self, messages: list[dict[str, str]]):
self.seen.append(messages) self.seen.append(messages)
for i in range(0, len(self.answer), 12): for i in range(0, len(self.answer), 12):
yield self.answer[i : i + 12] yield StreamPiece("content", self.answer[i : i + 12])
class _FakeSteeringResult: class _FakeSteeringResult:
+70 -3
View File
@@ -105,11 +105,15 @@ def test_save_points_user_on_send_and_brain_on_done() -> None:
assert user_push < js.find('fetch("/api/chat"'), ( assert user_push < js.find('fetch("/api/chat"'), (
"the user message must be saved before the turn starts" "the user message must be saved before the turn starts"
) )
# Brain save point is wired into the done handler with full metadata. # Brain save point is wired into the done handler with full metadata
# (phase 17: the persisted text is finalText — the empty-answer
# fallback substitution — and the optional thinking field rides along
# in the same meta object).
done_idx = js.find('ev.type === "done"') done_idx = js.find('ev.type === "done"')
assert done_idx != -1 assert done_idx != -1
done_block = js[done_idx : done_idx + 900] done_block = js[done_idx : done_idx + 1300]
assert "rememberBrainTurn(acc" in done_block assert "rememberBrainTurn(finalText || acc" in done_block
assert "thinking: thinkingAcc || undefined" in done_block
assert "deflected: !!ev.deflected" in done_block assert "deflected: !!ev.deflected" in done_block
assert "sources: ev.sources" in done_block assert "sources: ev.sources" in done_block
assert "suggestions: ev.suggestions" in done_block assert "suggestions: ev.suggestions" in done_block
@@ -181,3 +185,66 @@ def test_new_chat_button_style_contract() -> None:
assert mobile, "mobile media query missing" assert mobile, "mobile media query missing"
assert ".new-chat-label { display: none; }" in mobile.group(1) assert ".new-chat-label { display: none; }" in mobile.group(1)
assert ".new-chat-btn svg { display: block; }" in mobile.group(1) assert ".new-chat-btn svg { display: block; }" in mobile.group(1)
def test_brain_turn_persists_optional_thinking_field() -> None:
"""Phase 17: the done save point carries `thinking: thinkingAcc ||
undefined` — `undefined` drops the key from the JSON, so turns without
thinking persist byte-identical to before (no version bump). A
thinking-without-answer turn (reasoning exhausts max_tokens) renders
+ persists the shared empty-answer fallback: what the user saw is what
is stored."""
js = _js()
done_idx = js.find('ev.type === "done"')
error_idx = js.find('ev.type === "error"')
assert -1 < done_idx < error_idx, "done branch missing from the turn handler"
branch = js[done_idx:error_idx]
assert "thinking: thinkingAcc || undefined" in branch
assert (
'const finalText = acc || (sawThinking ? EMPTY_ANSWER_FALLBACK : "")'
in branch
)
assert "renderMarkdown(finalText)" in branch, (
"the substituted fallback must render into the bubble"
)
def test_restore_renders_collapsed_thinking_block() -> None:
"""Phase 17: a stored brain message carrying `thinking` re-renders the
block COLLAPSED above its bubble (escape-first markdown, as everywhere
else in the persistence contract); messages without the field render
exactly as before — no block."""
js = _js()
fn_start = js.find("function renderStoredMessage")
assert fn_start != -1
body = js[fn_start : js.find("\n}\n", fn_start)]
assert "if (m.thinking)" in body
assert "ensureThinkingBlock(wrap)" in body
assert "block.open = false" in body, "restored blocks must be collapsed"
assert "renderMarkdown(m.thinking)" in body
def test_thinking_block_css_uses_phase08_tokens() -> None:
"""Phase 17 styling (Phase-08 tokens, WCAG AA): the block frame, the
≥44px summary control (brand-ink ≈8.7:1 on surface) and the scrollable
scratchpad (ink-soft ≈6.9:1 on surface, 320px cap)."""
css = _css()
block = re.search(r"details\.thinking \{([\s\S]*?)\n\}", css)
assert block, "styles.css must style details.thinking"
body = block.group(1)
assert "var(--surface)" in body
assert "var(--line)" in body
assert "var(--brand-soft)" in body
assert "var(--radius-sm)" in body
summary = re.search(r"details\.thinking summary \{([\s\S]*?)\n\}", css)
assert summary, "the summary must be a styled focusable control"
sbody = summary.group(1)
assert "min-height: 44px" in sbody
assert "var(--brand-ink)" in sbody
assert "cursor: pointer" in sbody
text = re.search(r"details\.thinking \.thinking-text \{([\s\S]*?)\n\}", css)
assert text, "the .thinking-text scroll area must be styled"
tbody = text.group(1)
assert "var(--ink-soft)" in tbody
assert "max-height: 320px" in tbody
assert "overflow-y: auto" in tbody
+15
View File
@@ -36,6 +36,8 @@ def test_defaults_match_locked_decisions(monkeypatch: pytest.MonkeyPatch) -> Non
assert s.top_n_docs >= 1 assert s.top_n_docs >= 1
# Owner instruction 2026-08-22: answers may run up to 32 768 tokens. # Owner instruction 2026-08-22: answers may run up to 32 768 tokens.
assert s.max_output_tokens == 32_768 assert s.max_output_tokens == 32_768
# Phase 17: the model's thinking streams by default (kill-switch off).
assert s.stream_thinking is True
assert len(s.suggestions) >= 3 assert len(s.suggestions) >= 3
# A9 (revised): the import scope covers the seven A9 formats. # A9 (revised): the import scope covers the seven A9 formats.
assert s.import_extension_set == { assert s.import_extension_set == {
@@ -57,6 +59,19 @@ def test_max_output_tokens_env_override(monkeypatch) -> None:
assert s.max_output_tokens == 1234 assert s.max_output_tokens == 1234
def test_stream_thinking_default_true_and_env_parse(monkeypatch: pytest.MonkeyPatch) -> None:
"""Phase 17 kill-switch (``BOR_STREAM_THINKING``): on by default,
``0``/``false`` turn the ``thinking`` SSE frames off."""
assert _settings().stream_thinking is True
assert _settings(stream_thinking=False).stream_thinking is False
monkeypatch.setenv("BOR_STREAM_THINKING", "0")
assert _settings().stream_thinking is False
monkeypatch.setenv("BOR_STREAM_THINKING", "false")
assert _settings().stream_thinking is False
monkeypatch.setenv("BOR_STREAM_THINKING", "1")
assert _settings().stream_thinking is True
def test_import_extensions_env_override_is_a_csv_list(monkeypatch) -> None: def test_import_extensions_env_override_is_a_csv_list(monkeypatch) -> None:
monkeypatch.setenv("BOR_IMPORT_EXTENSIONS", "md,yml") monkeypatch.setenv("BOR_IMPORT_EXTENSIONS", "md,yml")
s = _settings() s = _settings()
+98
View File
@@ -91,3 +91,101 @@ def test_busy_button_style_tokens() -> None:
assert re.search(r"\.spinner \{[^}]*width: 16px", css) assert re.search(r"\.spinner \{[^}]*width: 16px", css)
assert "Thinking…" in js assert "Thinking…" in js
assert 'sendLabel.textContent' in js assert 'sendLabel.textContent' in js
# ---------- thinking display (phase 17) ----------
def test_thinking_event_is_a_first_class_turn_branch() -> None:
"""Phase 17: `thinking` SSE frames stream live into the collapsible
Thinking block — the typing dots make way, the 120s pre-token guard
clears (the stream is alive), and the text renders through the
escape-first markdown renderer (XSS-safe). While open, the stream is
pinned to the bottom of the block."""
js = _js()
thinking_idx = js.find('ev.type === "thinking"')
delta_idx = js.find('ev.type === "delta"')
assert -1 < thinking_idx < delta_idx, "the turn handler must branch on thinking frames"
branch = js[thinking_idx:delta_idx]
assert "thinkingAcc += ev.text" in branch
assert "sawThinking = true" in branch
assert "clearTurnTimeout()" in branch, "first thinking frame clears the 120s guard"
assert "removeTyping()" in branch, "the live block replaces the typing dots"
assert "ensureThinkingBlock(wrap)" in branch
assert "renderMarkdown(thinkingAcc)" in branch, "escape-first renderer (XSS-safe)"
assert "textEl.scrollTop = textEl.scrollHeight" in branch, "bottom-pinned while open"
def test_thinking_block_helpers_are_idempotent() -> None:
"""ensureThinkingBlock returns the existing `.thinking` details or
creates it OPEN above the .bubble; closeThinkingBlock is a no-op
without a block and never reopens one once the answer started."""
js = _js()
fn = js.find("function ensureThinkingBlock")
assert fn != -1, "ensureThinkingBlock must exist (near addTyping/removeTyping)"
body = js[fn : js.find("\n}\n", fn)]
assert "block.open = true" in body, "created open — the stream is the show"
assert "insertBefore" in body
assert 'querySelector(".bubble")' in body, "the block sits ABOVE the bubble"
fn2 = js.find("function closeThinkingBlock")
assert fn2 != -1, "closeThinkingBlock must exist"
body2 = js[fn2 : js.find("\n}\n", fn2)]
assert "block.open = false" in body2
def test_delta_branch_collapses_block_and_transitions_to_streaming() -> None:
"""The first answer delta transitions thinking → streaming (even when
thinking created the wrap first) and auto-collapses the block —
idempotent, and it never reopens once the answer started."""
js = _js()
delta_idx = js.find('ev.type === "delta"')
done_idx = js.find('ev.type === "done"')
assert -1 < delta_idx < done_idx
branch = js[delta_idx:done_idx]
assert "uiState === UI_STATE.thinking" in branch
assert "setUiState(UI_STATE.streaming)" in branch
assert "closeThinkingBlock(wrap)" in branch
def test_done_branch_sets_sawdone_and_closes_block() -> None:
"""On `done` the turn marks itself complete (sawDone — the stream-drop
guard keys off it) and settles the thinking block closed."""
js = _js()
done_idx = js.find('ev.type === "done"')
error_idx = js.find('ev.type === "error"')
assert -1 < done_idx < error_idx
branch = js[done_idx:error_idx]
assert "sawDone = true" in branch
assert "closeThinkingBlock(wrap)" in branch
def test_stream_drop_guard_reports_severed_stream() -> None:
"""A stream that delivered frames but no `done` event ends in the error
state (never a silent idle with a half bubble); the zero-frame case
falls through to the existing empty-answer fallback. The guard runs
after readSSE, before that fallback."""
js = _js()
assert "let sawDone = false" in js
assert re.search(r"if \(!sawDone && !aborted && \(acc \|\| thinkingAcc\)\)", js), (
"sawDone stream-drop guard missing after readSSE"
)
assert "The stream ended before my answer finished" in js
sse_idx = js.find("await readSSE(res,")
guard_idx = js.find("!sawDone && !aborted")
fallback_idx = js.find("!aborted && !wrap")
assert -1 < sse_idx < guard_idx < fallback_idx, (
"guard must sit between readSSE and the zero-frame fallback"
)
def test_thinking_chevron_stills_under_reduced_motion() -> None:
"""Phase 17: the only motion in the thinking block (the summary
chevron rotation) is disabled under prefers-reduced-motion."""
css = _css()
blocks = re.findall(
r"@media \(prefers-reduced-motion: reduce\) \{([\s\S]*?)\n\}", css
)
assert any(
"details.thinking summary::before" in b and "transition: none" in b
for b in blocks
), "chevron transition must still under reduced motion"
+95 -7
View File
@@ -16,7 +16,13 @@ from typing import Any
import pytest import pytest
from app.config import Settings from app.config import Settings
from app.rag.llm import EmbeddingDimensionError, EmbeddingError, LLMClient, LLMError from app.rag.llm import (
EmbeddingDimensionError,
EmbeddingError,
LLMClient,
LLMError,
StreamPiece,
)
def _settings(**kwargs: Any) -> Settings: def _settings(**kwargs: Any) -> Settings:
@@ -236,11 +242,21 @@ def test_single_oversized_text_fails_actionably() -> None:
# ---------- chat streaming (phase 03) ---------- # ---------- chat streaming (phase 03) ----------
def _chunk(content: str | None = "text", empty: bool = False): def _chunk(
"""One fake ChatCompletionChunk (``choices[].delta.content`` shape).""" content: str | None = "text", empty: bool = False, reasoning: str | None = None
):
"""One fake ChatCompletionChunk (``choices[].delta`` shape).
``reasoning_content`` is present on the delta only when *reasoning*
is not None — mirroring the real wire, where the field exists only
when the model sends it.
"""
if empty: if empty:
return SimpleNamespace(choices=[]) return SimpleNamespace(choices=[])
return SimpleNamespace(choices=[SimpleNamespace(delta=SimpleNamespace(content=content))]) delta: SimpleNamespace = SimpleNamespace(content=content)
if reasoning is not None:
delta.reasoning_content = reasoning
return SimpleNamespace(choices=[SimpleNamespace(delta=delta)])
class _FakeChatStream: class _FakeChatStream:
@@ -284,7 +300,7 @@ def _make_stream_client(
return llm, completions return llm, completions
async def _collect(llm: LLMClient, messages: list[dict[str, str]]) -> list[str]: async def _collect(llm: LLMClient, messages: list[dict[str, str]]) -> list[StreamPiece]:
return [p async for p in llm.chat_stream(messages)] return [p async for p in llm.chat_stream(messages)]
@@ -293,7 +309,13 @@ def test_chat_stream_yields_deltas_in_order() -> None:
[_chunk("Hey "), _chunk("you've "), _chunk("got this! 🧠")] [_chunk("Hey "), _chunk("you've "), _chunk("got this! 🧠")]
) )
pieces = asyncio.run(_collect(llm, [{"role": "user", "content": "q"}])) pieces = asyncio.run(_collect(llm, [{"role": "user", "content": "q"}]))
assert pieces == ["Hey ", "you've ", "got this! 🧠"] # Content-only chunks yield content pieces in wire order.
assert [(p.kind, p.text) for p in pieces] == [
("content", "Hey "),
("content", "you've "),
("content", "got this! 🧠"),
]
assert all(isinstance(p, StreamPiece) for p in pieces)
def test_chat_stream_uses_locked_generation_params() -> None: def test_chat_stream_uses_locked_generation_params() -> None:
@@ -322,7 +344,73 @@ def test_chat_stream_max_tokens_comes_from_settings() -> None:
def test_chat_stream_skips_empty_deltas_and_choiceless_chunks() -> None: def test_chat_stream_skips_empty_deltas_and_choiceless_chunks() -> None:
llm, _ = _make_stream_client([_chunk("a"), _chunk(empty=True), _chunk(None), _chunk("b")]) llm, _ = _make_stream_client([_chunk("a"), _chunk(empty=True), _chunk(None), _chunk("b")])
assert asyncio.run(_collect(llm, [{"role": "user", "content": "q"}])) == ["a", "b"] pieces = asyncio.run(_collect(llm, [{"role": "user", "content": "q"}]))
assert [(p.kind, p.text) for p in pieces] == [("content", "a"), ("content", "b")]
def test_chat_stream_maps_reasoning_content_to_thinking_pieces() -> None:
"""The verified aipi wire field (``delta.reasoning_content``) maps to
``thinking`` pieces; content chunks are untouched by the presence of
reasoning elsewhere in the stream."""
llm, _ = _make_stream_client(
[
_chunk("", reasoning="Step 1: parse the question."),
_chunk("", reasoning="Step 2: cite the doc."),
_chunk("Talos."),
]
)
pieces = asyncio.run(_collect(llm, [{"role": "user", "content": "q"}]))
assert [(p.kind, p.text) for p in pieces] == [
("thinking", "Step 1: parse the question."),
("thinking", "Step 2: cite the doc."),
("content", "Talos."),
]
def test_chat_stream_falls_back_to_reasoning_field() -> None:
"""Future-proofing: a bare ``delta.reasoning`` field (no
``reasoning_content``) is picked up by the fallback getattr."""
chunk = SimpleNamespace(
choices=[
SimpleNamespace(delta=SimpleNamespace(content="ans", reasoning="why not"))
]
)
llm, _ = _make_stream_client([chunk])
pieces = asyncio.run(_collect(llm, [{"role": "user", "content": "q"}]))
assert [(p.kind, p.text) for p in pieces] == [
("thinking", "why not"),
("content", "ans"),
]
def test_chat_stream_thinking_yields_before_content_in_chunk() -> None:
"""One chunk carrying both fields yields the thinking piece first."""
llm, _ = _make_stream_client([_chunk("answer", reasoning="hmm")])
pieces = asyncio.run(_collect(llm, [{"role": "user", "content": "q"}]))
assert [(p.kind, p.text) for p in pieces] == [
("thinking", "hmm"),
("content", "answer"),
]
def test_chat_stream_interleaved_thinking_and_content_order_preserved() -> None:
"""The piece sequence must match the chunk sequence exactly — a late
or interleaved thinking chunk is emitted at its wire position."""
llm, _ = _make_stream_client(
[
_chunk("", reasoning="t1"),
_chunk("c1"),
_chunk("", reasoning="t2"),
_chunk("c2"),
]
)
pieces = asyncio.run(_collect(llm, [{"role": "user", "content": "q"}]))
assert [(p.kind, p.text) for p in pieces] == [
("thinking", "t1"),
("content", "c1"),
("thinking", "t2"),
("content", "c2"),
]
def test_chat_stream_wraps_failures_as_llm_error() -> None: def test_chat_stream_wraps_failures_as_llm_error() -> None:
+16 -1
View File
@@ -4,7 +4,7 @@ from __future__ import annotations
import json import json
from app.api.chat import sse_event from app.api.chat import sse_event
from app.schemas import ChatErrorEvent from app.schemas import ChatErrorEvent, ChatThinkingEvent
def _payload(frame: str) -> dict: def _payload(frame: str) -> dict:
@@ -61,3 +61,18 @@ def test_error_event_shape_is_type_and_detail_only() -> None:
dumped = ChatErrorEvent(detail="The chat model dropped the connection").model_dump() dumped = ChatErrorEvent(detail="The chat model dropped the connection").model_dump()
assert set(dumped.keys()) == {"type", "detail"} assert set(dumped.keys()) == {"type", "detail"}
assert dumped["type"] == "error" # default — call sites never spell it out assert dumped["type"] == "error" # default — call sites never spell it out
def test_thinking_frame_serializes_exactly() -> None:
"""Phase 17 (PLAN §4 extension): the ``thinking`` frame is exactly
``{type: "thinking", text: str}`` — the sibling shape of ``delta``
the client's readSSE handler will branch on."""
frame = sse_event(ChatThinkingEvent(text="Step 1: check the docs…").model_dump())
assert frame == 'data: {"type": "thinking", "text": "Step 1: check the docs…"}\n\n'
assert _payload(frame) == {"type": "thinking", "text": "Step 1: check the docs…"}
def test_thinking_event_shape_is_type_and_text_only() -> None:
dumped = ChatThinkingEvent(text="hmm").model_dump()
assert set(dumped.keys()) == {"type", "text"}
assert dumped["type"] == "thinking" # default — call sites never spell it out