Files
brain-of-reese/tests/e2e/test_chat_persistence.py
T
ducoterra ffa919b8bf fix(chat): keep in-flight answers alive across in-app view switches
Root cause (owner repro, verified in a real browser 2026-09-06): the
five navbar views (Chat, RAG, Sources, Tuning, History) were separate
HTML documents, so a navbar click was a REAL cross-document navigation
— the chat page unloaded, the in-flight SSE fetch was aborted, and the
phase-48 teardown (app/api/chat.py `finally`, "chat: turn cancelled")
stopped the model. Observed: send question -> click RAG mid-stream ->
click Chat -> the answer never finished: no `query_log` row, and on
return a dangling question with no brain record (the pre-token pagehide
partial persist skips because `acc` is empty).

Phase-48 LOCKED-DECISION REFINEMENT (owner-confirmed 2026-09-06,
flagged per AGENTS.md rule 3, not silently deviated): "real navigation
cancels the fetch" now means LEAVING THE APP — tab close,
external/other-document navigation, the Stop button. In-app navbar
switches are client-side view switches and no longer cancel.

Fix — Option A (SPA shell), chosen over B (Service Worker owns the
stream) and C (server-side turn registry + resume):
- frontend/index.html is the shell: ONE `<main id="main">` holds the
  five `<section class="view">` blocks; hidden views carry BOTH
  `hidden` and `inert` (WCAG — no focus/keyboard traversal). The
  shared header, the single `doc-modal-*` skeleton, and the
  `#app-version` footer each exist exactly once; the per-view copies
  from the four folded pages are dropped.
- New frontend/assets/router.js (vanilla module — no framework, no
  bundler, No-CDN rule intact): lazy-imports a view module on FIRST
  show only (mount-once, hide-forever — the chat view's in-flight SSE
  reader persists across switches; that persistence IS the fix);
  intercepts same-shell navbar links with preventDefault +
  history.pushState (never a document load); handles popstate; single
  writer of `.nav-link` active state (is-active + aria-current),
  document.title, and the per-view meta description (values carried
  over from the old pages' heads, brand-resolved at write time).
- Each folded page's JS becomes `export async function mount(root)` —
  root-scoped queries; `initSharedHeader()` dropped (the header boots
  once in the shell via the chat module; the admin flag comes from the
  same cached `fetchIsAdmin()` promise — zero extra requests).
- app/main.py: a small list-driven route factory serves the shell for
  /tuning.html, /sources.html, /git-sources.html, /history.html —
  registered AFTER the API routers and BEFORE the static catch-all
  (routes-first). The phase-33 caching middleware applies no-cache +
  `?v=` rewriting unchanged; app/core/caching.py needed NO change
  (the view paths did not change — pinned by the integration tests).
- The four old view .html files are DELETED (one source of truth);
  deep links to the old URLs keep working (the router picks the view
  from the pathname); `/?chat=<id>` is unaffected; the Containerfile
  bundles router.js (inlining the lazy view modules) and drops the
  folded page files.
- app/schemas.py: HistoryTurn.text cap 4000 -> 32000 — the shell
  keeps long saved answers in the chat, and the old cap (stricter than
  the 24_000-char total history budget) 422-rejected any second turn
  in such a chat (found by the phase-42 E2E suite on the shell).

Boundaries: login.html, shared.html, doc-edit.html, document.html
REMAIN separate documents (flow pages, not navbar tabs); a mid-stream
navigation to doc-edit/document.html still cancels per phase 48
(follow-up candidate, out of scope). The SSE API is unchanged. Real
departures still cancel the turn — phase 48 intact (pinned by
tests/e2e/test_stop_generation.py, unchanged, and by the new suite's
real-departure control).

Tests:
- Phase-20 suite REWRITTEN to the new semantics
  (tests/e2e/test_sources_midstream_bug.py): a navbar switch no longer
  cancels — the stream survives the switch and the FULL answer
  settles; the pagehide partial persist REMAINS for real departures
  (the partial's exact shape — first streamed chunk prefix, no done
  metadata — is still pinned there).
- NEW story suite tests/e2e/test_nav_switch_keeps_stream.py (mock
  LLM): the owner repro (send -> RAG mid-stream -> Chat: window
  sentinel survives = same document, FULL answer, exactly one brain
  turn in bor.chat.v1, exactly one settled query_log row, auto-saved
  row matches) + the same mid-stream switch against the other three
  views + the real-departure-still-cancels control + the no-switch
  baseline.
- tests/unit/test_frontend_router.py: source-level pins of the router
  invariants (click interceptor targets ONLY same-shell view paths,
  pushState-only switches, mount-once guard, hidden+inert pair,
  single-writer active state/title); shell-route integration tests
  (each folded path serves the shell with no-cache + `?v=` body; a
  non-view path still 404s); the file-reading unit pins re-pointed at
  the shell (the four view files are gone — the shell is the source
  of truth).

Verification (this commit): full suite green — 1565 unit+integration
tests, app/ coverage 99% (>90% floor); ruff + pyright clean; the
phase's E2E suites green in isolation (house protocol, AGENTS.md rule
9). Owner repro verified in a real browser against the real LLM
(dev server :8010, headful Chromium): "tell me about everquest" ->
RAG mid-stream -> Chat — the answer completed with one brain bubble
and no error banner, `query_log` gained exactly one settled row
(deflected=True: the dev KB holds no EverQuest docs — the settle, not
the topic, is the proof), zero "chat: turn cancelled" lines for that
turn; the control (real navigation to /shared.html mid-stream) still
cancelled (no settled row, the cancel line logged, the partial
persisted on return). Screenshots: .agents/screenshots/76_manual_*.

Phase 76 (76_spa_nav_shell) complete — moved to
.agents/phases/complete/.
2026-09-06 06:31:31 -04:00

302 lines
12 KiB
Python

"""Phase 14 E2E (Playwright): the chat conversation survives a refresh.
Story: ``.agents/user_stories/chat-persistence.md``
Run in isolation (DB must be up: ``podman compose up -d db``):
uv run pytest tests/e2e/test_chat_persistence.py -v --no-cov
The conversation is a durable LOCAL session (localStorage key
``bor.chat.v1`` — A10 keeps the API stateless). Each test gets a fresh
browser context (the shared conftest's ``page`` fixture calls
``browser.new_page``), so localStorage is clean by construction: the
fresh-context tests start with the empty state exactly as before phase 14.
Test → story mapping (Playwright Mapping Rule):
1. ``test_conversation_survives_reload``
2. ``test_deflected_turn_restores_styling``
3. ``test_new_chat_clears_conversation``
4. ``test_persists_across_page_navigation``
"""
from __future__ import annotations
import asyncio
import json
import re
from pathlib import Path
from threading import Thread
from typing import Any
from playwright.sync_api import Page, expect
from sqlalchemy import text
from app.config import Settings
from app.db import SessionLocal
from app.rag.importer import ImportSummary, import_sources
from app.rag.llm import LLMClient
from e2e.auth_helpers import login
REPO = Path(__file__).resolve().parents[2]
FIXTURES = REPO / "tests" / "fixtures" / "docs"
QUESTION = "How is my Kubernetes cluster set up?"
OFF_TOPIC = "How do I bake sourdough bread?"
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
DEFLECT_PHRASE = r"haven't done anything like that"
STORAGE_KEY = "bor.chat.v1"
#: Phase 10 viewer URL + phase 13 back=/ (the restored chip must be
#: byte-identical to the live-rendered one).
CHIP_HREF = "/document.html?source=docs&path=homelab%2Fkubernetes.md&back=%2F"
async def _import_fixtures(mock_port: int) -> ImportSummary:
kwargs: dict[str, Any] = {"_env_file": None, "llm_base_url": f"http://127.0.0.1:{mock_port}/v1"}
settings = Settings(**kwargs) # pyright: ignore[reportCallIssue]
return await import_sources([FIXTURES], LLMClient(settings))
def _run_in_thread(coro: Any) -> Any:
"""Run a coroutine on a worker thread.
Playwright's sync API keeps an asyncio loop running on the test thread,
so ``asyncio.run`` cannot be called directly from a test body.
"""
box: dict[str, Any] = {}
def runner() -> None:
try:
box["value"] = asyncio.run(coro)
except BaseException as e: # noqa: BLE001 — re-raised on the test thread
box["error"] = e
t = Thread(target=runner)
t.start()
t.join()
if "error" in box:
raise box["error"]
return box["value"]
def _reset_db(mock_port: int, seed: bool) -> ImportSummary | None:
"""Truncate the KB (and query log), then optionally re-import fixtures."""
with SessionLocal() as db:
db.execute(text("TRUNCATE chunks, documents, query_log"))
db.commit()
if not seed:
return None
return _run_in_thread(_import_fixtures(mock_port))
def _stored(page: Page) -> str | None:
"""Raw localStorage payload for the chat (None when the key is absent)."""
return page.evaluate(f"() => localStorage.getItem('{STORAGE_KEY}')")
def _stored_parsed(page: Page) -> dict[str, Any]:
raw = _stored(page)
assert raw is not None, "the conversation key must exist in localStorage"
return json.loads(raw)
def _ask(page: Page, question: str) -> None:
"""Send one turn and wait until the grounded answer has fully landed."""
page.fill("#message-input", question)
page.click("#send-btn")
expect(page.locator(".msg.user .bubble").last).to_contain_text(question)
expect(page.locator(".msg.brain .bubble").last).to_contain_text(
MOCK_ANSWER_MARKER, timeout=30_000
)
expect(page.locator("#send-btn")).to_be_enabled()
expect(page.locator("#send-label")).to_have_text("Send")
def _ask_deflected(page: Page, question: str) -> None:
"""Send an off-topic turn and wait until the deflected answer landed."""
page.fill("#message-input", question)
page.click("#send-btn")
bubble = page.locator(".msg.brain.is-deflected .bubble").first
bubble.wait_for(state="visible", timeout=30_000)
expect(bubble).to_contain_text(re.compile(DEFLECT_PHRASE, re.IGNORECASE), timeout=30_000)
expect(page.locator("#send-btn")).to_be_enabled()
expect(page.locator("#send-label")).to_have_text("Send")
# ---------------------------------------------------------------------------
# 1. Refresh: the whole conversation comes back exactly as left
# ---------------------------------------------------------------------------
def test_conversation_survives_reload(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 13 # A9 formats (phase 47 added quadlet+j2)
page.set_default_timeout(30_000)
page.goto(app_url)
_ask(page, QUESTION)
# The turn is persisted: versioned payload, RAW text (no HTML), and the
# brain message carries the done metadata.
stored = _stored_parsed(page)
assert stored["v"] == 1
assert [m["who"] for m in stored["messages"]] == ["user", "brain"]
assert stored["messages"][0]["text"] == QUESTION
brain = stored["messages"][1]
assert MOCK_ANSWER_MARKER in brain["text"]
assert "<" not in brain["text"], "persisted brain text must be raw, not rendered HTML"
assert brain["deflected"] is False
assert any(s["path"] == "homelab/kubernetes.md" for s in brain["sources"])
# Refresh — the same context keeps its localStorage.
page.reload()
expect(page.locator("#empty-state")).to_be_hidden()
# Both bubbles restored: text + the source chip with the exact viewer URL.
expect(page.locator(".msg.user .bubble")).to_have_count(1)
expect(page.locator(".msg.user .bubble")).to_contain_text(QUESTION)
bubble = page.locator(".msg.brain .bubble")
expect(bubble).to_have_count(1)
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER)
chip = page.locator(".msg.brain .source-chip", has_text="kubernetes.md")
expect(chip).to_have_count(1)
expect(chip.first).to_have_attribute("href", CHIP_HREF)
expect(chip.first).not_to_have_attribute("target") # phase 26: modal, not a new tab
# The restore is read-only: storage still holds the same two messages.
assert [m["who"] for m in _stored_parsed(page)["messages"]] == ["user", "brain"]
# And the restored chat is live: a follow-up turn extends it.
_ask(page, "What about the nodes?")
expect(page.locator(".msg.user .bubble")).to_have_count(2)
assert len(_stored_parsed(page)["messages"]) == 4
# ---------------------------------------------------------------------------
# 2. Deflected turn: amber styling + "Maybe try" chips survive a refresh
# ---------------------------------------------------------------------------
def test_deflected_turn_restores_styling(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
_reset_db(mock_llm, seed=True)
page.set_default_timeout(30_000)
page.goto(app_url)
_ask_deflected(page, OFF_TOPIC)
# The deflected metadata (suggestions) is persisted with the answer.
stored = _stored_parsed(page)
brain = stored["messages"][-1]
assert brain["who"] == "brain"
assert brain["deflected"] is True
assert len(brain["suggestions"]) >= 2
page.reload()
# Amber deflected bubble + "Maybe try" chips come back, styled.
expect(page.locator("#empty-state")).to_be_hidden()
expect(page.locator(".msg.user .bubble")).to_contain_text(OFF_TOPIC)
restored = page.locator(".msg.brain.is-deflected .bubble")
expect(restored).to_have_count(1)
expect(restored.first).to_contain_text(re.compile(DEFLECT_PHRASE, re.IGNORECASE))
style = restored.first.evaluate("el => getComputedStyle(el)")
assert style["backgroundColor"] == "rgb(43, 33, 16)" # --accent-bg (dark theme)
assert style["borderTopColor"] == "rgb(245, 158, 11)" # --accent-line
# The chips are restored from the stored suggestions — same texts, order,
# and still one-tap-submittable (the shared chip component).
chips = page.locator(".msg.brain.is-deflected .maybe-try .suggestion-chip")
expect(chips.first).to_be_visible()
restored_texts = [chips.nth(i).inner_text() for i in range(chips.count())]
assert restored_texts == [s.strip() for s in brain["suggestions"] if s.strip()]
chips.first.click()
expect(page.locator("#message-input")).to_have_value("")
expect(page.locator(".msg.user .bubble")).to_have_count(2, timeout=30_000)
expect(page.locator(".msg.brain .bubble").nth(1)).to_contain_text(
MOCK_ANSWER_MARKER, timeout=30_000
)
# ---------------------------------------------------------------------------
# 3. "New chat": clear the conversation, back to the empty state
# ---------------------------------------------------------------------------
def test_new_chat_clears_conversation(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
_reset_db(mock_llm, seed=True)
page.set_default_timeout(30_000)
page.goto(app_url)
_ask(page, QUESTION)
expect(page.locator(".msg")).to_have_count(2)
assert _stored(page) is not None
# The reset control: ghost pill in the chat header, ≥44px, accessible name.
btn = page.locator("#new-chat-btn")
expect(btn).to_be_visible()
expect(btn).to_have_attribute("type", "button")
expect(btn).to_have_attribute("aria-label", "New chat")
box = btn.bounding_box()
assert box is not None and box["height"] >= 44
btn.click()
# Conversation gone, empty state + suggestions back, storage key cleared.
expect(page.locator(".msg")).to_have_count(0)
expect(page.locator("#empty-state")).to_be_visible()
expect(page.locator("#suggestions .suggestion-chip").first).to_be_visible(timeout=15_000)
assert _stored(page) is None, "New chat must clear the localStorage key"
# Confirmation via the existing live region (#send-status, aria-live=polite).
expect(page.locator("#send-status")).to_contain_text("New chat started")
# And it is a clean slate: a fresh turn starts a fresh conversation.
_ask(page, QUESTION)
stored = _stored_parsed(page)
assert [m["who"] for m in stored["messages"]] == ["user", "brain"]
assert stored["messages"][0]["text"] == QUESTION
# ---------------------------------------------------------------------------
# 4. Navigation: a trip to Sources and back keeps the conversation
# ---------------------------------------------------------------------------
def test_persists_across_page_navigation(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
_reset_db(mock_llm, seed=True)
page.set_default_timeout(30_000)
page.goto(app_url)
_ask(page, QUESTION)
_ask_deflected(page, OFF_TOPIC) # mixed conversation: grounded + deflected
# A trip to Sources (phase 16: the catalog is admin-only — the trip
# starts with a real form login). Phase 76 (task 02): the shell
# carries the chat view (with its New chat button) in the DOM on
# EVERY view — hidden + inert — so the button EXISTS here but must
# be HIDDEN (the view-scoped absence pattern; it left the shared
# bar at owner request, 2026-08-28 — pinned in
# tests/e2e/test_shared_header.py).
login(page, app_url, next="/sources.html")
expect(page.locator("#docs-tbody tr").first).to_be_visible(timeout=15_000)
expect(page.locator("#new-chat-btn")).to_be_hidden()
# Back to the chat: the conversation is exactly as left — both turns,
# the source chip, and the amber deflected bubble with its chips.
page.goto(app_url + "/")
expect(page.locator("#empty-state")).to_be_hidden()
expect(page.locator(".msg.user .bubble")).to_have_count(2)
expect(page.locator(".msg.user .bubble").first).to_contain_text(QUESTION)
expect(page.locator(".msg.user .bubble").nth(1)).to_contain_text(OFF_TOPIC)
expect(page.locator(".msg.brain .bubble").first).to_contain_text(MOCK_ANSWER_MARKER)
expect(page.locator(".msg.brain .source-chip", has_text="kubernetes.md")).to_have_count(1)
deflected = page.locator(".msg.brain.is-deflected .bubble")
expect(deflected).to_have_count(1)
expect(deflected.first).to_contain_text(re.compile(DEFLECT_PHRASE, re.IGNORECASE))
maybe_chip = page.locator(".msg.brain.is-deflected .maybe-try .suggestion-chip")
expect(maybe_chip.first).to_be_visible()
# The New chat control is back on the chat page.
expect(page.locator("#new-chat-btn")).to_be_visible()