Files
brain-of-reese/tests/e2e/test_thinking_scroll.py
T
ducoterra bc70ce36e0 feat(chat): render markdown tables in answers, viewer, and thinking
GFM pipe tables in the shared renderer (TODO.md L6): a table-protection
pass in frontend/assets/markdown.js (fences -> tables -> escape order)
pulls each header+separator+body block out as a placeholder, renders
cells escape-first with the same inline transforms, and reinserts a
semantic <table class="md-table"> inside a horizontal-overflow
.md-table-wrap — so a pipe table in a chat answer, the document
viewer/modal, and the thinking block all render the same semantic
table. Fences win over tables; lone pipes stay text.

- styles.css: .md-table palette rules (PLAN §7.2 tokens, no motion);
  min-width: max-content so a WIDE table keeps its natural width and
  the wrapper is the real scroller (width:100% alone wrapped the wide
  table's cells — proven by the new E2E).
- mock_llm.py: TABLE_TRIGGER ("show me a table") -> byte-stable
  TABLE_ANSWER (3-column table, <img onerror> XSS probe line, wide
  5-column table), checked before DEFLECT_MODE like SUMMARY_MODE.
- tests/fixtures/docs/homelab/tables.md: 3x3 pipe table + pipe-heavy
  fenced block (viewer/fence subject); the shared fixture set grows
  8 -> 9 docs, so every suite pinning the count (added/formats/
  stat-docs/EXPECTED_ROWS) is updated accordingly.
- tests/e2e/test_markdown_tables.py (new, story suite): chat table
  shape + non-deflection, wide-table wrapper scroll (no page
  overflow), XSS probe inert, viewer modal table, fence-not-a-table,
  lone pipe stays text.
- tests/e2e/test_agent_document_tools.py: fix a pre-existing flake —
  the "Calling tool…" label window is ~0.4 s at the mock's 0.1 s
  tool-frame pacing, and a polling expect could stride over it
  (failed 3 of 5 runs on the committed baseline). The pre-submit
  MutationObserver record is the deterministic source of truth; the
  racy to_have_text gate is gone.

uv run pytest: 738 passed, app/ coverage 99% (TOTAL unchanged);
ruff + pyright clean; story E2E 6/6 in isolation; regression E2E
suites (chat_rag, document_viewer, document_summaries, smoke) green.
2026-08-28 03:35:50 -04:00

704 lines
31 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Phase 43 E2E (Playwright, mock-only): the Thinking window scrolls again.
Story: ``.agent/user_stories/thinking-scroll-back.md``
Run in isolation (DB must be up: ``podman compose up -d db``):
uv run pytest tests/e2e/test_thinking_scroll.py -v --no-cov
MOCK-ONLY suite: ``E2E_REAL_LLM=1`` is not supported here — the real
``turbo`` decides its own reasoning length and pacing, and this story's
contract (320px clip, follow-while-pinned at the window bottom, pause on
scroll-up, resume on return) needs the deterministic mock's long, paced
scratchpad (phase-21 length: ``compose_thinking`` ≈ 2 700 chars ≈ 4.5s of
12-char/0.02s frames, overflowing the window by ~2x).
This suite REPLACES the deleted phase-21 ``test_thinking_no_scroll.py``
(owner direction 2026-08-27, ``TODO.md`` L7, roadmap A2 — the 2026-08-24
no-scroll choice is reversed): ``details.thinking .thinking-text`` is
user-scrollable again (``overflow-y: auto``, 320px clip kept), and the
phase-17 bottom-pin is GATED in app.js (``THINKING_NEAR_BOTTOM_PX = 32``)
— the window follows the live tail only while the user is pinned near
its bottom; scrolling up pauses the follow, returning resumes it.
Determinism notes:
* Tests 1–4 key off the mock's ``think out loud then hesitate`` trigger
(phase 20): the thinking stream, then a deterministic 4s pre-content
pause with the block still OPEN and no further pin frames — a frozen
live tail (test 1) plus generous, stable mid-stream windows (tests
2–4 sample while the block is provably open and the bubble empty).
* Per-chunk pin: the gate re-runs on EVERY thinking frame; the user is
never re-pinned once scrolled up (distance > 32px only grows as the
content grows), and setting ``scrollTop`` back to the bottom re-arms
it on the very next chunk (a 12-char frame renders ≤ one line ≈ 22px,
always inside the 32px band — so "at the bottom" survives the whole
stream, which is exactly what test 2 samples).
* User-scroll gestures: wheel scrolling is 1:1 pixel-deterministic on
headless Chromium 151, and keyboard scrolling is exercised after
making the scratchpad focusable (``tabindex="0"`` + a real click — the
plain div is not keyboard-focusable by design; the attribute is test
scaffolding for the gesture, not app state). The mouse-drag gesture is
performed literally; headless Chromium's OVERLAY scrollbars are not
grabbable by synthetic mouse events (verified against a bare
overflow-auto probe page), so the drag assertion is that the window
stays where the user put it (no snap back to the live tail) — wheel
and keyboard are the proven movers.
* Mid-stream scroll-up/down (tests 3–4) is programmatic
(``el.scrollTop = …``), the phase-18 convention for simulating the
user's scroll-away: the pin is checked per frame in JS, so the value
set is exactly what the gate sees.
Test → story mapping (Playwright Mapping Rule):
1. ``test_thinking_window_user_scrollable`` — frozen live tail (4s
hesitation): wheel up, ``Home``, mouse-drag up — ``scrollTop``
moves; the window shows earlier content.
2. ``test_thinking_window_follows_while_pinned`` — live stream: at the
bottom, after the 2nd-to-last and the last chunk the window is
pinned to the tail (±1px); the last chunk's text renders inside the
visible rectangle.
3. ``test_thinking_window_stops_on_scroll_up`` — mid-stream: scrolled
up ~half the window, ``scrollTop`` is stable (±1px) over the next
≥5 chunks — no re-pin.
4. ``test_thinking_window_resumes_on_return`` — from the paused state,
``scrollTop`` back to the bottom: on the next chunk the window is
re-pinned to the tail (±1px).
5. ``test_thinking_window_css_contract`` — computed ``overflow-y`` is
``auto``, ``max-height`` is 320px, and the clip is real.
6. ``test_answer_bubble_still_scrollable`` — regression (phase 11):
long answer, the page scrolls, the bubble's overflow is untouched.
7. ``test_restored_collapsed_thinking_unaffected`` — regression
(phase 17): a settled thinking turn reloads collapsed with full text.
"""
from __future__ import annotations
import asyncio
import json
import re
from collections.abc import Iterator
from pathlib import Path
from threading import Thread
from typing import Any
import pytest
from playwright.sync_api import Page, expect
from sqlalchemy import text
from app.config import Settings
from app.db import SessionLocal
from app.rag.importer import ImportSummary, import_sources
from app.rag.llm import LLMClient
from e2e.mock_llm import compose_thinking
REPO = Path(__file__).resolve().parents[2]
FIXTURES = REPO / "tests" / "fixtures" / "docs"
#: Phase-17 trigger question (grounded turn, thinking + short answer).
THINK_QUESTION = "think out loud — how is my kubernetes cluster set up?"
#: Phase-20 hesitation trigger: the long phase-17/21 thinking stream, then
#: a deterministic 4s pause before the first content frame — a frozen
#: live tail with the block still open (the user-scroll test window).
HESITATE_QUESTION = "think out loud then hesitate — how is my kubernetes cluster set up?"
#: Phase-11 long-answer trigger (regression test 6).
LONG_QUESTION = "How is my Kubernetes cluster set up? write a long answer"
LONG_ANSWER_END = "LONG-ANSWER-END"
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
#: Line fragment the mock's deterministic scratchpad must carry (phase 17
#: convention, mock_llm.compose_thinking).
THINKING_FRAGMENT = "Step 2: Check my notes"
STORAGE_KEY = "bor.chat.v1"
SELECTOR = ".msg.brain details.thinking .thinking-text"
#: The story's "within 1px" tolerance for tail-pin claims.
TAIL_TOL = 1
async def _import_fixtures(mock_port: int) -> ImportSummary:
kwargs: dict[str, Any] = {"_env_file": None, "llm_base_url": f"http://127.0.0.1:{mock_port}/v1"}
settings = Settings(**kwargs) # pyright: ignore[reportCallIssue]
return await import_sources([FIXTURES], LLMClient(settings))
def _run_in_thread(coro: Any) -> Any:
"""Run a coroutine on a worker thread (Playwright owns the test loop)."""
box: dict[str, Any] = {}
def runner() -> None:
try:
box["value"] = asyncio.run(coro)
except BaseException as e: # noqa: BLE001 — re-raised on the test thread
box["error"] = e
t = Thread(target=runner)
t.start()
t.join()
if "error" in box:
raise box["error"]
return box["value"]
def _reset_db(mock_port: int, seed: bool) -> ImportSummary | None:
"""Truncate the KB (and query log), then optionally re-import fixtures."""
with SessionLocal() as db:
db.execute(text("TRUNCATE chunks, documents, query_log"))
db.commit()
if not seed:
return None
return _run_in_thread(_import_fixtures(mock_port))
@pytest.fixture()
def seeded_kb(mock_llm: int, db_ready: None) -> Iterator[None]:
"""A fresh KB seeded from ``tests/fixtures/docs`` (9 docs since phase
44, A9 formats), truncated again on teardown (same fixture shape as
the phase-17/21 suites)."""
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 9
yield
_reset_db(mock_llm, seed=False)
def send_and_wait(page: Page, question: str) -> None:
"""Type into #message-input, submit via #composer, then wait until the
last brain message settles (send button re-enabled, label "Send")."""
page.fill("#message-input", question)
page.evaluate("() => document.querySelector('#composer').requestSubmit()")
expect(page.locator(".msg.user .bubble").last).to_contain_text(question)
# The mock streams at 0.02s/chunk, so thinking + answer land in a few
# seconds — 30s is generous on headless Chromium.
expect(page.locator(".msg.brain .bubble").last).not_to_have_text("", timeout=30_000)
expect(page.locator("#send-btn")).to_be_enabled(timeout=30_000)
expect(page.locator("#send-label")).to_have_text("Send")
def submit(page: Page, question: str) -> None:
"""Submit a question and confirm the user bubble landed (the caller
then drives the mid-stream assertions itself)."""
page.fill("#message-input", question)
page.evaluate("() => document.querySelector('#composer').requestSubmit()")
expect(page.locator(".msg.user .bubble").last).to_contain_text(question)
def _scroll_sample(page: Page) -> dict[str, float]:
"""scrollTop / scrollHeight / clientHeight of the Thinking window."""
return page.evaluate(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
return {{ top: el.scrollTop, height: el.scrollHeight, client: el.clientHeight }}; }}"""
)
def _at_tail(sample: dict[str, float]) -> bool:
"""True when the window is pinned to the live tail: the bottom of the
content is visible (``scrollTop`` clamped at ``scrollHeight -
clientHeight``, within ``TAIL_TOL`` — the gated pin's effect)."""
return abs(sample["top"] - (sample["height"] - sample["client"])) <= TAIL_TOL
def _wait_text_stable(page: Page, timeout_ms: int = 30_000) -> None:
"""Wait until the scratchpad text stops growing for 300ms.
The mock paces frames at 0.02s, so a 300ms still length means the
thinking stream has ended — with the hesitation trigger, the 4s
pre-content pause (block still open, no further pin frames) is then
running."""
page.wait_for_function(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
if (!el) return false;
const len = el.innerText.length;
const now = performance.now();
if (!window.__thinkProbe) window.__thinkProbe = {{ len, at: now }};
const p = window.__thinkProbe;
if (len !== p.len) {{ p.len = len; p.at = now; return false; }}
return now - p.at >= 300; }}""",
timeout=timeout_ms,
)
def _wait_text_contains(page: Page, marker: str, timeout_ms: int = 30_000) -> None:
"""Wait until the rendered scratchpad (whitespace-insensitive) contains
``marker`` — a deterministic probe for a given point in the stream."""
page.wait_for_function(
f"""(tail) => {{ const el = document.querySelector('{SELECTOR}');
return !!el && el.innerText.replace(/\\s+/g, '').includes(tail); }}""",
arg=marker,
timeout=timeout_ms,
)
def _full_text_len() -> int:
"""Whitespace-stripped length of the mock's FULL scratchpad for the
hesitation question (what the rendered innerText will reach at the
end of the thinking stream)."""
expected = compose_thinking(
{"messages": [{"role": "user", "content": HESITATE_QUESTION}]}
)
return len(re.sub(r"\s+", "", expected))
def _rendered_len(page: Page) -> int:
"""Whitespace-stripped length of the rendered scratchpad (same
measure as ``_full_text_len`` — innerText adds nothing: the
scratchpad renders as one <p> with <br> line breaks, no escaping).
"""
return page.evaluate(
f"""() => document.querySelector('{SELECTOR}')
.innerText.replace(/\\s+/g, '').length"""
)
def _first_text_node(page: Page) -> dict[str, Any]:
"""Rect of the scratchpad's FIRST text node vs the window's box — the
"earlier content is visible" probe (the mock's text starts with
``Step 1: Read the question carefully``)."""
return page.evaluate(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
const box = el.getBoundingClientRect();
const walker = document.createTreeWalker(el, NodeFilter.SHOW_TEXT);
const first = walker.nextNode();
const range = document.createRange();
range.selectNodeContents(first);
const r = range.getBoundingClientRect();
return {{ nodeText: first ? first.textContent.slice(0, 8) : "",
nodeTop: r.top, boxTop: box.top, boxBottom: box.bottom }}; }}"""
)
# ---------------------------------------------------------------------------
# 1. User scroll restored: wheel, keyboard (Home), and mouse-drag on the
# frozen live tail (4s hesitation) move / hold the window; earlier
# content becomes visible
# ---------------------------------------------------------------------------
def test_thinking_window_user_scrollable(
page: Page, app_url: str, seeded_kb: None
) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
submit(page, HESITATE_QUESTION)
expect(page.locator(".msg.user .bubble").last).to_contain_text(HESITATE_QUESTION)
details = page.locator(".msg.brain").last.locator("details.thinking")
details.wait_for(state="attached", timeout=10_000)
# Premise: the long scratchpad overflows the 320px clip (the story's
# "window" is only real once it clips).
page.wait_for_function(
f"() => {{ const el = document.querySelector('{SELECTOR}');"
" return !!el && el.scrollHeight > el.clientHeight; }",
timeout=30_000,
)
# The thinking stream has ENDED (4s hesitation pause running): no more
# pin frames, the block is still open, and the tail is frozen — any
# user scroll persists and is observable here.
_wait_text_stable(page)
expect(details).to_have_attribute("open", "")
expect(page.locator(".msg.brain .bubble").last).to_have_text("")
# Precondition: the gated pin left the window at the live tail.
before = _scroll_sample(page)
assert before["height"] > before["client"], "the window must overflow"
assert _at_tail(before), "the pin must have left the window at the tail"
box = details.locator(".thinking-text").bounding_box()
assert box
cx, cy = box["x"] + box["width"] / 2, box["y"] + box["height"] / 2
page.mouse.move(cx, cy)
# Wheel back (up) — the window (not the page) must scroll; headless
# Chromium applies the delta 1:1 (verified: 150 -> exactly 150px).
WHEEL_UP = 160
page.mouse.wheel(0, -WHEEL_UP)
page.wait_for_timeout(200)
wheeled = _scroll_sample(page)
moved = min(WHEEL_UP, before["top"]) # clamped at the window's top
assert wheeled["top"] <= before["top"] - moved + TAIL_TOL, (
f"wheel up did not scroll the window: {before['top']} -> {wheeled['top']}"
)
assert wheeled["top"] < before["top"], "the window must have scrolled up"
# Keyboard: the plain div is not keyboard-focusable by design — make
# it so (test scaffolding), give it REAL focus with a click, and press
# Home: the window must jump to its top (proven 1:1 on headless
# Chromium 151: click + Home -> 0, End -> maxScroll).
details.locator(".thinking-text").evaluate(
"el => { el.setAttribute('tabindex', '0'); }"
)
page.mouse.click(cx, cy)
page.keyboard.press("Home")
page.wait_for_timeout(200)
homed = _scroll_sample(page)
assert homed["top"] <= TAIL_TOL, f"Home must land the window at its top: {homed}"
# Earlier content is visible: the scratchpad's FIRST line ("Step 1…")
# now sits at the top of the visible window (it was off-screen while
# pinned to the tail).
first = _first_text_node(page)
assert first["nodeText"].startswith("Step 1"), "the first line must be Step 1"
assert first["boxTop"] - 1 <= first["nodeTop"] < first["boxTop"] + 60, (
f"the earlier content must be visible at the window top: {first}"
)
# Mouse drag up: the literal gesture over the window. Headless
# Chromium's overlay scrollbars are not grabbable by synthetic mouse
# events (bare overflow-auto probe: wheel + click/keys move the
# window; a thumb/track drag never does — no classic scrollbar is
# rendered, and `::-webkit-scrollbar` does not force one), so the
# contract proven here is that the drag leaves the window EXACTLY
# where the user put it — no snap back to the live tail (the
# ungated-pin regression would re-yank it). Wheel + keyboard above
# are the proven movers.
page.mouse.move(cx, cy)
page.mouse.down()
page.mouse.move(cx, cy - 80, steps=5)
page.mouse.up()
page.wait_for_timeout(200)
after = _scroll_sample(page)
assert abs(after["top"] - homed["top"]) <= TAIL_TOL, (
f"a mouse drag must not move the window from the user's position: "
f"{homed['top']} -> {after['top']}"
)
assert not _at_tail(after), "the window must stay scrolled up (no re-pin)"
# Still inside the pure-thinking window.
expect(details).to_have_attribute("open", "")
expect(page.locator(".msg.brain .bubble").last).to_have_text("")
# ---------------------------------------------------------------------------
# 2. Follow while pinned: at the window's bottom, the gated pin keeps the
# window glued to the live tail — sampled at the 2nd-to-last and the
# last chunk — and the last chunk's text renders inside the window
# ---------------------------------------------------------------------------
def test_thinking_window_follows_while_pinned(
page: Page, app_url: str, seeded_kb: None
) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
submit(page, HESITATE_QUESTION)
expect(page.locator(".msg.user .bubble").last).to_contain_text(HESITATE_QUESTION)
details = page.locator(".msg.brain").last.locator("details.thinking")
details.wait_for(state="attached", timeout=10_000)
# The exact deterministic scratchpad the mock will stream, sliced the
# same way the mock's _sse_stream does (12-char chunks). The user
# never scrolls: at the bottom, every 12-char frame renders at most
# one line (≈22px — single <p>, <br> line breaks), always inside the
# 32px band, so the gate passes on EVERY frame and the pin holds.
expected = compose_thinking(
{"messages": [{"role": "user", "content": HESITATE_QUESTION}]}
)
pieces = re.findall(r".{1,12}", expected, re.S)
ws = re.sub(r"\s+", "", "".join(pieces))
#: 12 rendered chars ending at the 2nd-to-last chunk.
marker_second_last = re.sub(r"\s+", "", "".join(pieces[:-1]))[-12:]
#: 12 rendered chars at the very end (the last chunk).
marker_last = ws[-12:]
# During the stream: once the 2nd-to-last chunk has landed, the window
# is pinned to the tail (the invariant holds at EVERY chunk while the
# user is at the bottom).
_wait_text_contains(page, marker_second_last)
assert _at_tail(_scroll_sample(page)), "not at the tail after chunk N-1"
# After the last chunk: the 4s hesitation pause holds this state with
# the block still open — sample the tail pin, then the geometry.
_wait_text_contains(page, marker_last)
_wait_text_stable(page)
expect(details).to_have_attribute("open", "")
expect(page.locator(".msg.brain .bubble").last).to_have_text("")
sample = _scroll_sample(page)
assert _at_tail(sample), f"not at the tail after the last chunk: {sample}"
# Geometry: the last chunk's text (the final text node of the
# scratchpad) renders INSIDE the visible rectangle, and a hit-test at
# the box's bottom lands inside .thinking-text.
geo = page.evaluate(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
const box = el.getBoundingClientRect();
const walker = document.createTreeWalker(el, NodeFilter.SHOW_TEXT);
let last = null;
while (walker.nextNode()) last = walker.currentNode;
const range = document.createRange();
range.selectNodeContents(last);
const r = range.getBoundingClientRect();
const hit = document.elementFromPoint(box.left + 10, box.bottom - 5);
return {{
nodeVisible: r.bottom > box.top && r.top < box.bottom,
nodeBottomInBox: r.bottom <= box.bottom + 1,
hitInside: hit ? el.contains(hit) : false,
}}; }}"""
)
assert geo["nodeVisible"], "the last chunk's text is outside the window"
assert geo["nodeBottomInBox"], "the last chunk's text is clipped off the bottom"
assert geo["hitInside"], "a hit-test at the box bottom missed .thinking-text"
# ---------------------------------------------------------------------------
# 3. Paused on scroll-up: mid-stream, scrolled up ~half the window, the
# gated pin never re-pins — scrollTop is stable across ≥5 chunks
# ---------------------------------------------------------------------------
def test_thinking_window_stops_on_scroll_up(
page: Page, app_url: str, seeded_kb: None
) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
submit(page, HESITATE_QUESTION)
expect(page.locator(".msg.user .bubble").last).to_contain_text(HESITATE_QUESTION)
details = page.locator(".msg.brain").last.locator("details.thinking")
details.wait_for(state="attached", timeout=10_000)
# Mid-stream: the window has REAL headroom (≥100px of hidden content
# — the first overflowing frame, a few px, is not the story's
# "window") and the (untouched, pinned) window is at the tail — the
# user now scrolls up ~half the window.
page.wait_for_function(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
return !!el && el.scrollHeight - el.clientHeight >= 100; }}""",
timeout=30_000,
)
sample = _scroll_sample(page)
assert _at_tail(sample), "the pin must hold at the tail before the scroll-up"
total = _full_text_len()
grown = _rendered_len(page)
assert total - grown >= 84, "the stream must have ≥5 chunks left mid-stream"
top0 = page.evaluate(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
el.scrollTop = Math.round((el.scrollHeight - el.clientHeight) / 2);
return el.scrollTop; }}"""
)
assert top0 >= 40, "half the window must be a real scroll-up"
len0 = _rendered_len(page)
# Over the NEXT ≥5 CHUNKS (84 chars = 7 × 12) the window must not
# move: scrolled-up readers are never re-pinned (the distance from
# the bottom only grows as content grows).
page.wait_for_function(
f"""(minLen) => {{ const el = document.querySelector('{SELECTOR}');
return !!el && el.innerText.length >= minLen; }}""",
arg=len0 + 84,
timeout=15_000,
)
after = _scroll_sample(page)
assert abs(after["top"] - top0) <= TAIL_TOL, (
f"the window must hold still over ≥5 chunks (no re-pin): "
f"{top0} -> {after['top']}"
)
# Still inside the pure-thinking window (the stability claim is only
# meaningful while the block is open and no content frame has landed).
expect(details).to_have_attribute("open", "")
expect(page.locator(".msg.brain .bubble").last).to_have_text("")
# ---------------------------------------------------------------------------
# 4. Resumes on return: from the paused state, scrollTop back to the
# bottom — on the NEXT chunk the window is re-pinned to the tail
# ---------------------------------------------------------------------------
def test_thinking_window_resumes_on_return(
page: Page, app_url: str, seeded_kb: None
) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
submit(page, HESITATE_QUESTION)
expect(page.locator(".msg.user .bubble").last).to_contain_text(HESITATE_QUESTION)
details = page.locator(".msg.brain").last.locator("details.thinking")
details.wait_for(state="attached", timeout=10_000)
# Mid-stream: real window headroom (≥100px) + pinned tail (test-3
# premises).
page.wait_for_function(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
return !!el && el.scrollHeight - el.clientHeight >= 100; }}""",
timeout=30_000,
)
sample = _scroll_sample(page)
assert _at_tail(sample), "the pin must hold at the tail before the scroll-up"
total = _full_text_len()
grown = _rendered_len(page)
assert total - grown >= 96, "the stream must have enough chunks left to pause + resume"
# Scroll up ~half the window and confirm the PAUSED state (≥5 chunks,
# stable within 1px — the same invariant as test 3).
top0 = page.evaluate(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
el.scrollTop = Math.round((el.scrollHeight - el.clientHeight) / 2);
return el.scrollTop; }}"""
)
assert top0 >= 40, "half the window must be a real scroll-up"
len0 = page.evaluate(f"() => document.querySelector('{SELECTOR}').innerText.length")
page.wait_for_function(
f"""(minLen) => {{ const el = document.querySelector('{SELECTOR}');
return !!el && el.innerText.length >= minLen; }}""",
arg=len0 + 84,
timeout=15_000,
)
assert abs(_scroll_sample(page)["top"] - top0) <= TAIL_TOL, (
"precondition: the window must be paused (test-3 invariant)"
)
len1 = _rendered_len(page)
# Return to the bottom: the gate re-arms on the very next chunk (the
# distance from the bottom is 0, and a 12-char frame renders ≤ one
# line — always inside the 32px band), so the pin fires immediately.
page.evaluate(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
el.scrollTop = el.scrollHeight; }}"""
)
page.wait_for_function(
f"""(minLen) => {{ const el = document.querySelector('{SELECTOR}');
return !!el && el.innerText.length > minLen; }}""",
arg=len1,
timeout=15_000,
)
after = _scroll_sample(page)
assert _at_tail(after), f"the window must be re-pinned on the next chunk: {after}"
assert after["top"] > top0, "the window must have returned toward the tail"
# Still inside the pure-thinking window.
expect(details).to_have_attribute("open", "")
expect(page.locator(".msg.brain .bubble").last).to_have_text("")
# ---------------------------------------------------------------------------
# 5. CSS contract: overflow-y auto (user scroll restored), 320px clip
# kept, and the clip is real
# ---------------------------------------------------------------------------
def test_thinking_window_css_contract(
page: Page, app_url: str, seeded_kb: None
) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
send_and_wait(page, THINK_QUESTION)
details = page.locator(".msg.brain").last.locator("details.thinking")
expect(details).not_to_have_attribute("open") # auto-collapsed
details.locator("summary").click() # open for measurement
expect(details).to_have_attribute("open", "")
style = page.evaluate(
f"""() => {{ const el = document.querySelector('{SELECTOR}');
const cs = getComputedStyle(el);
return {{ overflowY: cs.overflowY, maxHeight: cs.maxHeight,
scrollHeight: el.scrollHeight, clientHeight: el.clientHeight }}; }}"""
)
assert style["overflowY"] == "auto", "the window must be user-scrollable again"
assert style["maxHeight"] == "320px", "the 320px clip must stay"
# The clip is real, not cosmetic: the long scratchpad overflows it.
assert style["scrollHeight"] > style["clientHeight"]
# ---------------------------------------------------------------------------
# 6. Regression (phase 11): the answer bubble is untouched — a long answer
# still grows the page, and normal (user) scrolling of the answer works
# ---------------------------------------------------------------------------
def test_answer_bubble_still_scrollable(
page: Page, app_url: str, seeded_kb: None
) -> None:
page.set_default_timeout(45_000)
page.goto(app_url)
page.fill("#message-input", LONG_QUESTION)
page.click("#send-btn")
expect(page.locator(".msg.user .bubble").last).to_contain_text(LONG_QUESTION)
# The ~900-word answer streams to completion (phase-11 contract).
bubble = page.locator(".msg.brain .bubble").last
expect(bubble).to_contain_text(LONG_ANSWER_END, timeout=60_000)
expect(page.locator("#send-btn")).to_be_enabled()
# The answer bubble keeps its existing overflow (this phase only
# touched .thinking-text) — it is NOT a clipped scroll window.
overflow_y = page.evaluate(
"() => { const els = document.querySelectorAll('.msg.brain .bubble');"
" return getComputedStyle(els[els.length - 1]).overflowY; }"
)
assert overflow_y != "hidden", "the answer bubble must keep its scroll behavior"
# A long answer grows the PAGE — and the page still scrolls normally
# (phase 42: no autoscroll — the page sits wherever the user left it,
# so move it to the bottom first (phase-18 convention for simulating
# the user's scroll); the user's wheel then moves it up and down).
state0 = page.evaluate(
"() => { window.scrollTo(0, document.documentElement.scrollHeight);"
" return { y: window.scrollY, sh: document.documentElement.scrollHeight,"
" ch: window.innerHeight }; }"
)
assert state0["sh"] > state0["ch"], "the long answer must make the page scrollable"
assert state0["y"] > 100, "the page must have scrolled to its bottom"
box = bubble.bounding_box()
assert box
viewport = page.viewport_size
assert viewport # the conftest `page` fixture is fixed at 1280x800
# A point in the visible lower part of the answer area (the page is
# at the bottom, so the bubble's lower edge is in the viewport).
mx = box["x"] + box["width"] / 2
my = max(50.0, min(box["y"] + box["height"] - 60.0, viewport["height"] - 100))
page.mouse.move(mx, my)
# Headless Chromium applies wheel scrolling through an async momentum
# pipeline — let each gesture settle before reading the position.
page.mouse.wheel(0, -400) # wheel up: away from the newest content
page.wait_for_timeout(500)
y_up = page.evaluate("() => window.scrollY")
page.mouse.wheel(0, 400) # wheel back down
page.wait_for_timeout(500)
y_down = page.evaluate("() => window.scrollY")
assert y_up < state0["y"] - 100, (
f"the page must scroll up on wheel: y {state0['y']} -> {y_up}"
)
assert y_down > y_up, f"the page must scroll back down on wheel: {y_up} -> {y_down}"
# ---------------------------------------------------------------------------
# 7. Regression (phase 17): a stored thinking turn restores a COLLAPSED
# block with its full text (the restore path is untouched by this
# phase — with the pin gated on `block.open`, a restored closed block
# is never auto-pinned)
# ---------------------------------------------------------------------------
def test_restored_collapsed_thinking_unaffected(
page: Page, app_url: str, seeded_kb: None
) -> None:
page.set_default_timeout(30_000)
page.goto(app_url)
send_and_wait(page, THINK_QUESTION)
details = page.locator(".msg.brain").last.locator("details.thinking")
expect(details).not_to_have_attribute("open") # auto-collapsed
captured = details.locator(".thinking-text").text_content()
assert captured
page.reload()
expect(page.locator("#empty-state")).to_be_hidden()
restored = page.locator(".msg.brain").last.locator("details.thinking")
expect(restored).to_have_count(1)
expect(restored).not_to_have_attribute("open") # restored COLLAPSED
expect(restored.locator(".thinking-text")).to_have_text(captured)
# Opening the restored block still shows the full scratchpad, and the
# answer + persistence are intact.
restored.locator("summary").click()
expect(restored).to_have_attribute("open", "")
expect(restored.locator(".thinking-text")).to_contain_text(THINKING_FRAGMENT)
expect(page.locator(".msg.brain .bubble").last).to_contain_text(MOCK_ANSWER_MARKER)
raw = json.loads(page.evaluate(f"() => localStorage.getItem('{STORAGE_KEY}')"))[
"messages"
][1]["thinking"]
assert re.sub(r"\s+", "", raw) == re.sub(r"\s+", "", captured)