Files
brain-of-reese/tests/e2e/test_nav_switch_keeps_stream.py
T
ducoterra ffa919b8bf fix(chat): keep in-flight answers alive across in-app view switches
Root cause (owner repro, verified in a real browser 2026-09-06): the
five navbar views (Chat, RAG, Sources, Tuning, History) were separate
HTML documents, so a navbar click was a REAL cross-document navigation
— the chat page unloaded, the in-flight SSE fetch was aborted, and the
phase-48 teardown (app/api/chat.py `finally`, "chat: turn cancelled")
stopped the model. Observed: send question -> click RAG mid-stream ->
click Chat -> the answer never finished: no `query_log` row, and on
return a dangling question with no brain record (the pre-token pagehide
partial persist skips because `acc` is empty).

Phase-48 LOCKED-DECISION REFINEMENT (owner-confirmed 2026-09-06,
flagged per AGENTS.md rule 3, not silently deviated): "real navigation
cancels the fetch" now means LEAVING THE APP — tab close,
external/other-document navigation, the Stop button. In-app navbar
switches are client-side view switches and no longer cancel.

Fix — Option A (SPA shell), chosen over B (Service Worker owns the
stream) and C (server-side turn registry + resume):
- frontend/index.html is the shell: ONE `<main id="main">` holds the
  five `<section class="view">` blocks; hidden views carry BOTH
  `hidden` and `inert` (WCAG — no focus/keyboard traversal). The
  shared header, the single `doc-modal-*` skeleton, and the
  `#app-version` footer each exist exactly once; the per-view copies
  from the four folded pages are dropped.
- New frontend/assets/router.js (vanilla module — no framework, no
  bundler, No-CDN rule intact): lazy-imports a view module on FIRST
  show only (mount-once, hide-forever — the chat view's in-flight SSE
  reader persists across switches; that persistence IS the fix);
  intercepts same-shell navbar links with preventDefault +
  history.pushState (never a document load); handles popstate; single
  writer of `.nav-link` active state (is-active + aria-current),
  document.title, and the per-view meta description (values carried
  over from the old pages' heads, brand-resolved at write time).
- Each folded page's JS becomes `export async function mount(root)` —
  root-scoped queries; `initSharedHeader()` dropped (the header boots
  once in the shell via the chat module; the admin flag comes from the
  same cached `fetchIsAdmin()` promise — zero extra requests).
- app/main.py: a small list-driven route factory serves the shell for
  /tuning.html, /sources.html, /git-sources.html, /history.html —
  registered AFTER the API routers and BEFORE the static catch-all
  (routes-first). The phase-33 caching middleware applies no-cache +
  `?v=` rewriting unchanged; app/core/caching.py needed NO change
  (the view paths did not change — pinned by the integration tests).
- The four old view .html files are DELETED (one source of truth);
  deep links to the old URLs keep working (the router picks the view
  from the pathname); `/?chat=<id>` is unaffected; the Containerfile
  bundles router.js (inlining the lazy view modules) and drops the
  folded page files.
- app/schemas.py: HistoryTurn.text cap 4000 -> 32000 — the shell
  keeps long saved answers in the chat, and the old cap (stricter than
  the 24_000-char total history budget) 422-rejected any second turn
  in such a chat (found by the phase-42 E2E suite on the shell).

Boundaries: login.html, shared.html, doc-edit.html, document.html
REMAIN separate documents (flow pages, not navbar tabs); a mid-stream
navigation to doc-edit/document.html still cancels per phase 48
(follow-up candidate, out of scope). The SSE API is unchanged. Real
departures still cancel the turn — phase 48 intact (pinned by
tests/e2e/test_stop_generation.py, unchanged, and by the new suite's
real-departure control).

Tests:
- Phase-20 suite REWRITTEN to the new semantics
  (tests/e2e/test_sources_midstream_bug.py): a navbar switch no longer
  cancels — the stream survives the switch and the FULL answer
  settles; the pagehide partial persist REMAINS for real departures
  (the partial's exact shape — first streamed chunk prefix, no done
  metadata — is still pinned there).
- NEW story suite tests/e2e/test_nav_switch_keeps_stream.py (mock
  LLM): the owner repro (send -> RAG mid-stream -> Chat: window
  sentinel survives = same document, FULL answer, exactly one brain
  turn in bor.chat.v1, exactly one settled query_log row, auto-saved
  row matches) + the same mid-stream switch against the other three
  views + the real-departure-still-cancels control + the no-switch
  baseline.
- tests/unit/test_frontend_router.py: source-level pins of the router
  invariants (click interceptor targets ONLY same-shell view paths,
  pushState-only switches, mount-once guard, hidden+inert pair,
  single-writer active state/title); shell-route integration tests
  (each folded path serves the shell with no-cache + `?v=` body; a
  non-view path still 404s); the file-reading unit pins re-pointed at
  the shell (the four view files are gone — the shell is the source
  of truth).

Verification (this commit): full suite green — 1565 unit+integration
tests, app/ coverage 99% (>90% floor); ruff + pyright clean; the
phase's E2E suites green in isolation (house protocol, AGENTS.md rule
9). Owner repro verified in a real browser against the real LLM
(dev server :8010, headful Chromium): "tell me about everquest" ->
RAG mid-stream -> Chat — the answer completed with one brain bubble
and no error banner, `query_log` gained exactly one settled row
(deflected=True: the dev KB holds no EverQuest docs — the settle, not
the topic, is the proof), zero "chat: turn cancelled" lines for that
turn; the control (real navigation to /shared.html mid-stream) still
cancelled (no settled row, the cancel line logged, the partial
persisted on return). Screenshots: .agents/screenshots/76_manual_*.

Phase 76 (76_spa_nav_shell) complete — moved to
.agents/phases/complete/.
2026-09-06 06:31:31 -04:00

526 lines
22 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Phase 76 E2E (Playwright): in-app view switches never halt a
generating answer — the owner repro, pinned against the deterministic
mock LLM.
Source: owner repro, verified in a real browser 2026-09-06 — send a
question → click **RAG** in the navbar mid-stream → click **Chat** →
the answer never finished (every navbar view was a separate document,
so the navbar click was a REAL cross-document navigation: the chat
page unloaded, the in-flight fetch was aborted, and the phase-48
teardown cancelled the turn — no ``query_log`` row, a dangling
question on return). Phase 76 folded the five navbar views into ONE
HTML shell: a navbar click is a CLIENT-SIDE view switch
(``history.pushState`` + show/hide), so the in-flight SSE reader in
the hidden chat view keeps streaming and the answer COMPLETES when
the user returns.
Run in isolation (DB must be up: ``podman compose up -d db``):
uv run pytest tests/e2e/test_nav_switch_keeps_stream.py -v --no-cov
Timing is deterministic by construction: the mock's ``write a long
answer`` trigger streams a ~5400-char answer at 12 chars / 0.02 s
(~8–9 s of content), so the mid-stream switch window is wide.
The "same document" proof (the canonical pattern from the phase
overview): a ``window`` sentinel set before the nav click is still
readable after the switch — a real document load would wipe ``window``
globals. The ``performance.getEntriesByType("navigation")`` length is
deliberately NOT used: a real load resets that counter to 1 in the
fresh document, so it cannot distinguish pushState from a reload.
Test → story mapping (Playwright Mapping Rule):
1. ``test_rag_switch_mid_stream_completes`` — THE OWNER REPRO: send →
RAG mid-stream → Chat; the FULL answer completes, ``bor.chat.v1``
holds EXACTLY ONE brain turn, the server SETTLED the turn (one
``query_log`` row — no phase-48 ``turn cancelled``), and the
auto-saved row (admin, ``persistConversation``) carries the same
single full turn.
2. ``test_every_nav_view_keeps_stream`` — the same mid-stream switch
against the other three views (Sources/git-sources, Tuning,
History): one send, one switch, one return, full answer + settled
``query_log`` row each time.
3. ``test_real_departure_still_cancels`` — the phase-48 CONTROL (the
locked contract survives the phase): a genuine cross-document
departure (``/shared.html`` — a stable document for a signed-in
session; ``/login.html`` is deliberately avoided because it
auto-redirects a signed-in admin straight back into the shell)
still aborts the fetch, leaves NO ``query_log`` row, and the
page-20/73 leave-save lands the partial in the EXACT shape pinned
by ``tests/e2e/test_sources_midstream_bug.py::test_partial_answer_
survives_real_departure_midstream`` (mirrored, not re-invented).
The overlap with that suite is on purpose — different stories:
phase 20 pins the partial shape, this phase pins that the
navbar-switch path no longer cancels while real departures still
do.
4. ``test_baseline_no_switch_still_completes`` — the long question
with NO navigation completes identically (guards against the
shell fold changing the ordinary path).
"""
from __future__ import annotations
import asyncio
import json
import re
import time
from pathlib import Path
from threading import Thread
from typing import Any
import httpx
from playwright.sync_api import Locator, Page, expect
from sqlalchemy import text
from app.config import Settings
from app.db import SessionLocal
from app.rag.importer import ImportSummary, import_sources
from app.rag.llm import LLMClient
from e2e.auth_helpers import login
from e2e.mock_llm import LONG_ANSWER_END, LONG_ANSWER_LINES, long_answer
REPO = Path(__file__).resolve().parents[2]
FIXTURES = REPO / "tests" / "fixtures" / "docs"
#: The phase-11 on-topic long-answer phrasing (house pattern,
#: ``test_hidden_tab_stream.py`` / ``test_stop_generation.py``): the
#: honesty gate is HIGH and the ~900-word answer streams for ~8–9 s
#: (12 chars / 0.02 s) — the guaranteed mid-stream window.
LONG_QUESTION = "How is my Kubernetes cluster set up? write a long answer"
#: The mock's byte-stable long answer — the EXACT string the stream
#: delivers, so "the full answer" is an exact comparison, not a
#: contains check.
LONG_ANSWER = long_answer()
#: The mock's first 12-char content slice (the same cut ``_sse_stream``
#: makes) — the real-departure partial must START with it (raw text,
#: pre-render); the rendered first line keeps the list-item form (the
#: markdown renderer converts the "1. " marker into a list item,
#: pinned by test_long_answers).
FIRST_CHUNK_RAW = re.findall(r".{1,12}", LONG_ANSWER, re.S)[0]
FIRST_LINE_DOM = "Step 1: configure node-1"
STORAGE_KEY = "bor.chat.v1"
#: The typing indicator is itself a .msg.brain — exclude its bubble.
ANSWER = ".msg.brain .bubble:not(.typing)"
#: The other three navbar views (test 2): the nav link, the view's
#: URL (pushState target), and an admin-visible content marker inside
#: the view (proof the view actually showed — the RAG view gets the
#: same treatment with ``#docs-tbody tr`` in test 1).
OTHER_VIEWS: tuple[tuple[str, str, str], ...] = (
("#nav-git-sources", "/git-sources.html", "#git-sources-content"),
("#nav-tuning", "/tuning.html", "#tune-save"),
("#nav-history", "/history.html", "#history-table-wrap"),
)
# ---------------------------------------------------------------------------
# KB seeding (house pattern: TRUNCATE-then-import)
# ---------------------------------------------------------------------------
async def _import_fixtures(mock_port: int) -> ImportSummary:
kwargs: dict[str, Any] = {"_env_file": None, "llm_base_url": f"http://127.0.0.1:{mock_port}/v1"}
settings = Settings(**kwargs) # pyright: ignore[reportCallIssue]
return await import_sources([FIXTURES], LLMClient(settings))
def _run_in_thread(coro: Any) -> Any:
"""Run a coroutine on a worker thread.
Playwright's sync API keeps an asyncio loop running on the test thread,
so ``asyncio.run`` cannot be called directly from a test body.
"""
box: dict[str, Any] = {}
def runner() -> None:
try:
box["value"] = asyncio.run(coro)
except BaseException as e: # noqa: BLE001 — re-raised on the test thread
box["error"] = e
t = Thread(target=runner)
t.start()
t.join()
if "error" in box:
raise box["error"]
return box["value"]
def _reset_db(mock_port: int) -> ImportSummary:
"""House reset + the prompt-shaping tables: steering notes and the
KB overview would otherwise append deterministic suffixes to every
mock answer and break the exact-text assertions."""
with SessionLocal() as db:
db.execute(
text("TRUNCATE chunks, documents, query_log, steering_notes, kb_overview")
)
db.commit()
return _run_in_thread(_import_fixtures(mock_port))
def _query_log_count() -> int:
"""The settled-row count over the whole (truncated) log.
The query log finalizes a row ONLY when the LLM finished AND the
persistence succeeded (phase 48); a cancelled turn — a real
departure mid-stream, or one before the first token — leaves no
settled row, so the count IS the settled-row signal (0 = cancelled,
1 = settled; the house pattern from test_hidden_tab_stream.py).
"""
with SessionLocal() as db:
return db.execute(text("SELECT count(*) FROM query_log")).scalar_one()
# ---------------------------------------------------------------------------
# Shared flows
# ---------------------------------------------------------------------------
def _stored_parsed(page: Page) -> dict[str, Any]:
raw = page.evaluate(f"() => localStorage.getItem('{STORAGE_KEY}')")
assert raw is not None, "the conversation key must exist in localStorage"
return json.loads(raw)
def _no_error_banner(page: Page) -> None:
"""The never-stale contract, shell-scoped: the hidden views ship
their own role=alert surfaces (sync/upload banners, …) that are
inert while their view is hidden — so the pin is that NO alert is
VISIBLE, whatever the document carries hidden (the phase-20
rewrite's shell form)."""
expect(page.locator('[role="alert"]:visible')).to_have_count(0)
def _ask_long(page: Page) -> None:
page.fill("#message-input", LONG_QUESTION)
page.click("#send-btn")
expect(page.locator(".msg.user .bubble").last).to_contain_text(LONG_QUESTION)
def _wait_streaming(page: Page, answer: Locator) -> Locator:
"""Wait until answer text is visibly streaming (a few delta frames
rendered — the mid-stream moment, well inside the ~8–9 s stream)."""
answer.wait_for(state="visible", timeout=30_000)
partial = ""
deadline = time.monotonic() + 15
while time.monotonic() < deadline:
partial = answer.inner_text()
if len(partial.split()) >= 8:
break
time.sleep(0.05)
assert len(partial.split()) >= 8, "no answer deltas before the view switch"
# In flight at the switch: the button IS the enabled Stop control.
expect(page.locator("#send-label")).to_have_text("Stop")
expect(page.locator("#send-btn")).to_have_class(re.compile(r"is-stop"))
return answer
def _wait_done(page: Page, answer: Locator) -> str:
"""Wait for the ``done`` settle: the Send button is back and the
bubble carries the unique final line — no error banner on the way."""
expect(page.locator("#send-label")).to_have_text("Send", timeout=30_000)
expect(page.locator("#send-btn")).not_to_have_class(re.compile(r"is-stop"))
expect(answer).to_contain_text(LONG_ANSWER_END, timeout=30_000)
_no_error_banner(page)
return answer.inner_text()
def _assert_full_answer(text: str) -> None:
"""The bubble carries the FULL mock answer — every one of the 40
numbered steps plus the unique final line (a truncated stream
would be missing its tail)."""
for i in range(1, LONG_ANSWER_LINES + 1):
assert f"Step {i}: configure node-{i}" in text, f"step {i} missing from the answer"
assert LONG_ANSWER_END in text
def _assert_one_brain_turn(page: Page) -> dict[str, Any]:
"""``bor.chat.v1`` holds EXACTLY ONE brain turn for the question,
and its text is the COMPLETE mock answer byte-for-byte (the
``done`` settle's record — the settle, not a partial)."""
stored = _stored_parsed(page)
assert stored["v"] == 1
msgs = stored["messages"]
assert [m["who"] for m in msgs] == ["user", "brain"], (
"exactly ONE brain turn for the question: "
f"{[m['who'] for m in msgs]}"
)
assert msgs[0]["text"] == LONG_QUESTION
brain = msgs[1]
assert brain["text"] == LONG_ANSWER, "the record's text is the FULL answer"
assert brain.get("deflected") is False, "the done metadata rides the record"
return brain
def _admin_cookies(page: Page) -> dict[str, str]:
"""The signed session cookies the browser holds after a form login —
used to call the admin API with plain httpx (the test's API side
sees exactly what the signed-in browser sees)."""
return {c["name"]: c["value"] for c in page.context.cookies() if "name" in c and "value" in c}
def _delete_rows_by_title(app_url: str, cookies: dict[str, str], title: str) -> None:
"""Best-effort cleanup of the auto-saved row (a 404 is fine)."""
r = httpx.get(f"{app_url}/api/chats", timeout=10, cookies=cookies)
if r.status_code != 200:
return
for c in r.json()["chats"]:
if c["title"] == title:
httpx.delete(f"{app_url}/api/chats/{c['id']}", timeout=10, cookies=cookies)
def _wait_row_full(app_url: str, cookies: dict[str, str], title: str) -> dict[str, Any]:
"""Poll the auto-saved row until it carries the full answer as a
single brain turn (the ``done`` settle's fire-and-forget
``persistConversation`` PUT is the last writer)."""
deadline = time.monotonic() + 15
last: list[dict[str, Any]] = []
while time.monotonic() < deadline:
r = httpx.get(f"{app_url}/api/chats", timeout=10, cookies=cookies)
rows = (
[c for c in r.json()["chats"] if c["title"] == title]
if r.status_code == 200
else []
)
for c in rows:
row = httpx.get(f"{app_url}/api/chats/{c['id']}", timeout=10, cookies=cookies).json()
brains = [m for m in row["messages"] if m["who"] == "brain"]
if len(brains) == 1 and brains[0]["text"] == LONG_ANSWER:
return row
last = [row]
time.sleep(0.2)
raise AssertionError(
"the auto-saved row never held the full answer as exactly one brain turn; last: "
f"{last!r}"
)
# ---------------------------------------------------------------------------
# 1. THE OWNER REPRO: send → RAG mid-stream → Chat — the FULL answer
# completes, one brain turn, one settled query_log row, and the
# auto-saved row carries the same single full turn
# ---------------------------------------------------------------------------
def test_rag_switch_mid_stream_completes(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm)
assert summary is not None and summary.added == 13 # A9 formats (phase 47 added quadlet+j2)
page.set_default_timeout(30_000)
# Admin login: the admin-only #nav-sources link is revealed, and
# the auto-save row (the persistConversation path) is reachable,
# so the saved-chat side gets pinned too.
login(page, app_url, next="/")
expect(page.locator("#nav-sources")).to_be_visible()
cookies = _admin_cookies(page)
_delete_rows_by_title(app_url, cookies, LONG_QUESTION) # stale rows from crashed runs
_ask_long(page)
answer = _wait_streaming(page, page.locator(ANSWER))
# THE SWITCH (mid-stream): the window sentinel set BEFORE the click
# is still readable AFTER it — the canonical same-document proof
# (a real navigation would have wiped window globals).
page.evaluate("() => { window.__shell_boot = 'phase76'; }")
page.click("#nav-sources")
expect(page).to_have_url(app_url + "/sources.html")
assert page.evaluate("() => window.__shell_boot") == "phase76", (
"a real navigation would have wiped the window sentinel — "
"the switch must be same-document"
)
# The RAG view actually showed (the fixture docs' rows are listed)
# and the chat view is hidden (the stream fills it in the
# background — that persistence IS the fix).
expect(page.locator("#docs-tbody tr").first).to_be_visible(timeout=15_000)
expect(page.locator("#view-chat")).to_be_hidden()
# Stay on the RAG view while the stream keeps running (the switch
# is ~t+2 s; the full answer needs ~8–9 s).
page.wait_for_timeout(2000)
# Back to the chat (the header link — a router-intercepted
# switch, still same-document).
page.click('a.nav-link[href="/"]')
expect(page).to_have_url(app_url + "/")
assert page.evaluate("() => window.__shell_boot") == "phase76"
# The answer COMPLETED — the bubble carries the FULL mock answer
# (every step line + the unique final line), no error banner.
done_text = _wait_done(page, answer)
_assert_full_answer(done_text)
# Storage: EXACTLY ONE brain turn — the FULL answer with done
# metadata (the settle, not a partial).
page.wait_for_timeout(500)
_assert_one_brain_turn(page)
# Server side: the turn SETTLED — exactly one query_log row, so no
# phase-48 "turn cancelled" teardown fired for an in-app switch.
assert _query_log_count() == 1, "a completed turn must finalize its query_log row"
# Auto-save (admin): the row carries the same single full brain
# turn (the shared record shape).
try:
row = _wait_row_full(app_url, cookies, LONG_QUESTION)
brains = [m for m in row["messages"] if m["who"] == "brain"]
assert len(brains) == 1, "the saved row holds exactly one brain turn"
assert brains[0]["text"] == LONG_ANSWER
finally:
_delete_rows_by_title(app_url, cookies, LONG_QUESTION)
# ---------------------------------------------------------------------------
# 2. The same mid-stream switch against the other three views — one
# send, one switch, one return, full answer + settled row each time
# ---------------------------------------------------------------------------
def test_every_nav_view_keeps_stream(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm)
assert summary is not None and summary.added == 13
page.set_default_timeout(30_000)
# Admin: every navbar link (incl. the three below) is revealed by
# the whoami gate.
login(page, app_url, next="/")
for i, (nav_sel, view_path, marker) in enumerate(OTHER_VIEWS, start=1):
expect(page.locator(nav_sel)).to_be_visible()
_ask_long(page)
# The CURRENT turn's bubble (the conversation accumulates one
# full turn per iteration — the latest pair is the pin).
answer = _wait_streaming(page, page.locator(ANSWER).last)
# Mid-stream switch to this view — same-document (sentinel).
page.evaluate("() => { window.__shell_boot = 'phase76'; }")
page.click(nav_sel)
expect(page).to_have_url(app_url + view_path)
assert page.evaluate("() => window.__shell_boot") == "phase76", (
f"a real navigation to {view_path} would have wiped the sentinel"
)
# The view actually showed (its admin content is up) and the
# chat view is hidden (the stream fills it in the background).
expect(page.locator(marker)).to_be_visible(timeout=15_000)
expect(page.locator("#view-chat")).to_be_hidden()
# Let the stream run while this view is up, then return to the
# chat (still same-document).
page.wait_for_timeout(2000)
page.click('a.nav-link[href="/"]')
expect(page).to_have_url(app_url + "/")
assert page.evaluate("() => window.__shell_boot") == "phase76"
# The answer COMPLETED — FULL mock answer, no error banner.
done_text = _wait_done(page, answer)
_assert_full_answer(done_text)
# Every turn SETTLED: exactly one query_log row per completed
# turn so far (a cancelled turn would leave no row).
page.wait_for_timeout(500)
assert _query_log_count() == i, (
f"turn {i} must finalize exactly one settled query_log row"
)
# Storage: the latest pair is the question + ONE brain turn
# carrying the FULL answer (each turn appended, none
# cancelled, none truncated).
msgs = _stored_parsed(page)["messages"]
assert msgs[-2] == {"who": "user", "text": LONG_QUESTION}
assert msgs[-1]["who"] == "brain"
assert msgs[-1]["text"] == LONG_ANSWER
# ---------------------------------------------------------------------------
# 3. The phase-48 CONTROL: a REAL cross-document departure still
# cancels the fetch (the locked contract survives the phase) — and
# the page-20/73 partial persist lands in the exact phase-20 shape
# ---------------------------------------------------------------------------
def test_real_departure_still_cancels(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm)
assert summary is not None and summary.added == 13
page.set_default_timeout(30_000)
login(page, app_url, next="/")
_ask_long(page)
_wait_streaming(page, page.locator(ANSWER))
# THE DEPARTURE: a REAL cross-document navigation (NOT a navbar
# link — those are view switches now). The fetch is aborted by the
# unload, which is the point (phase 48). /shared.html is a plain
# document with a stable state for a signed-in session — unlike
# /login.html, which auto-redirects a signed-in admin straight
# back into the shell.
page.goto(app_url + "/shared.html")
expect(page).to_have_url(app_url + "/shared.html")
expect(page.locator("#shared-title")).to_be_visible(timeout=15_000)
# The turn was CANCELLED — the phase-48 query_log row only lands
# when the LLM finished AND the persistence succeeded, so a
# cancelled mid-stream turn must leave NO settled row.
assert _query_log_count() == 0, (
"a cancelled mid-stream turn must not finalize a query_log row"
)
# Return to the chat — the page-20/73 leave-save is intact: the
# question AND the already-streamed partial are both rendered.
page.goto(app_url + "/")
expect(page.locator("#empty-state")).to_be_hidden()
expect(page.locator(".msg.user .bubble")).to_have_count(1)
expect(page.locator(".msg.user .bubble").first).to_contain_text(LONG_QUESTION)
restored = page.locator(".msg.brain .bubble")
expect(restored).to_have_count(1)
expect(restored.first).to_contain_text(FIRST_LINE_DOM)
_no_error_banner(page)
# The EXACT phase-20 partial shape (mirrored from
# test_sources_midstream_bug.py::test_partial_answer_survives_real_
# departure_midstream — do not invent a new shape): exactly one
# brain turn, raw text STARTING with the first streamed chunk,
# SHORTER than the full answer, and NO done metadata (no
# sources/deflected/suggestions/thinking — the turn never
# settled when it was written).
msgs = _stored_parsed(page)["messages"]
assert [m["who"] for m in msgs] == ["user", "brain"]
assert msgs[0]["text"] == LONG_QUESTION
brain = msgs[1]
assert brain["text"].startswith(FIRST_CHUNK_RAW)
assert len(brain["text"]) < len(LONG_ANSWER), "the stored answer must be partial"
assert brain["text"] != LONG_ANSWER
assert "sources" not in brain
assert "deflected" not in brain
assert "suggestions" not in brain
assert "thinking" not in brain
# ---------------------------------------------------------------------------
# 4. Baseline: the long question with NO navigation completes
# identically (guards against the shell fold changing the ordinary
# path)
# ---------------------------------------------------------------------------
def test_baseline_no_switch_still_completes(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
"""The normal path, untouched by the shell: the long answer
completes identically without any view switch (guards against an
over-eager router/view change altering the ordinary settle)."""
summary = _reset_db(mock_llm)
assert summary is not None and summary.added == 13
page.set_default_timeout(30_000)
page.goto(app_url)
_ask_long(page)
answer = _wait_streaming(page, page.locator(ANSWER))
done_text = _wait_done(page, answer)
_assert_full_answer(done_text)
_assert_one_brain_turn(page)
assert _query_log_count() == 1