Files
brain-of-reese/tests/e2e/test_tool_call_compact.py
T
ducoterra c851d1a1c5
Build and Push Containers / build-and-push-app (push) Successful in 2m2s
Build and Push Containers / build-and-push-db (push) Successful in 12s
phase: 117_tool_call_compact
All verification complete — every gate green. Final report:

**Phase 117 final verification pass — all criteria verified** (tasks 01–04 were already in `complete/`; I re-verified the full gate myself, no code changes needed)
- **Verified:** `appendToolLine` wraps lines in `details.tool-calls-disclosure` (open live, `closeToolCalls` at delta/done/stop + restore-closed, count via textContent); `addToolLines` parity (closed); `.tool-call` deboxed / `code` dechipped; new disclosure/summary CSS; only the one pinned CSS assert updated
- `uv run pytest` → **2466 passed**
- `uv run pytest --cov=app --cov-report=term-missing` → **99%** on `app/` (>90% floor)
- E2E in isolation: `test_tool_call_compact.py` **4 passed** (fold/expand/deboxed-flow/restored-folded); `test_agent_document_tools.py` **4 passed**; `test_big_read_progress.py` **4 passed** (live disclosure open + `.tool-elapsed` visible); `test_thinking_display.py` **5 passed**; `test_smoke.py` **3 passed**
- `uv run ruff check . && uv run pyright` → clean (pyright 0 errors, 0 warnings)
- Completion criteria: pins 1–4 ✅ (E2E above); live "calling tool" state byte-identical ✅ (`test_frontend_tool_states.py` 10 passed incl. `test_calling_tool_label_strings`); scope ✅ — `git diff` limited to the 3 frontend assets + 2 unit tests + E2E story + phase files, **0 changes in `app/`**
- Notable: `TODO.md` carries a **pre-existing uncommitted owner edit** (new retrieval-context idea) already in the tree before this pass — left untouched, will ride along in the harness commit; pre-existing committed quirk: the brain-wave commit's styles.css comment mislabels itself "Phase 117" (cosmetic, out of this diff)
- No commit made (harness commits + moves the phase); **next pending phase: none** — `todo/` holds only phase 117
2026-09-15 18:26:38 -04:00

425 lines
18 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Phase 117 E2E (Playwright, mock-only): compact, well-wrapped tool-call lines.
Source: owner visual-glitch report (2026-09-15, mobile viewport,
``https://brain.reeseapps.com``) — "how much space the tool calls take
up, and the tool call text is spit and wrapped poorly": one completed
answer stacked 6+ full-width bordered tool-call cards (one per ``tool``
SSE frame), each a complete card (accent left border + surface
background + radius + a mono path chip inside), with the label flex
item breaking mid-word (``Rea``/``ding``) and the path wrapping to
three lines indented to the right of that broken label.
Phase 117's fix (frontend-only, D1–D6 in ``00_phase.md``): the
per-call lines ride in ONE native ``<details>`` disclosure — a compact
"Tool calls (N)" summary, open while the turn is live and FOLDED when
the answer begins (the first ``delta``), on ``done``, on stop/abort,
and on restore — and the lines themselves are deboxed inline-flow
text (no card, no flex: the label + ``<code>`` run as one continuous
sentence and the path wraps to the left edge).
Run in isolation (DB must be up: ``podman compose up -d db``):
uv run pytest tests/e2e/test_tool_call_compact.py -v --no-cov
MOCK-ONLY suite: the marker question drives the SAME deterministic
3-call flow ``test_agent_document_tools.py`` pins — ``ls`` →
``ls(scoped)`` → ``read`` (``tests/e2e/mock_llm.py``: the user message
contains ``use your tools`` and the system prompt carries the
``<tools>`` section of the HIGH prompt) — over the byte-identical
two-document KB fixture (mirrored from that suite so the flow is
deterministic). No slow proxy: the fold is about a COMPLETED turn (it
folds on the first delta), so the fast mock is the deterministic
driver.
Completion marker: the mock's single-read flow answers
``Read <source/path>. <first 80 chars of the read document>`` — the
deterministic quote below (``MOCK_ANSWER_MARKER``), NOT the plain
grounded answer's "Deterministic mock answer for E2E" tail (the tool
flow never reaches that branch). Its presence in the bubble proves the
first delta has landed — the turn is COMPLETE and the disclosure has
already folded (phase 117, D3: the fold rides the first delta).
Test → phase mapping (Playwright Mapping Rule):
1. ``test_tool_calls_fold_to_one_line_after_turn`` — pin 1: a
completed 3-call turn renders ONE visible "Tool calls (3)" summary;
the disclosure is CLOSED (no ``open`` attribute); the three
``.tool-call`` lines are present in the DOM but folded (the first
line is not visible); the answer is present.
2. ``test_summary_click_expands_the_calls`` — pin 2: tapping the
native ``<summary>`` (a real focusable toggle, AA) opens the
disclosure; the three lines become visible, in order, with the
phase-37/94 pinned text.
3. ``test_expanded_line_is_deboxed_inline_flow`` — pin 3: the
expanded "Reading" line's computed ``display`` is NOT ``flex``
(the label + path are one inline run, not two flex items — the
debox), and the label runs directly into the path in the same run.
4. ``test_restored_turn_renders_folded`` — pin 4: a same-context
RELOAD (the phase-14/50 persisted conversation restores) renders
the turn FOLDED — closed disclosure, "Tool calls (3)" summary,
three lines in the DOM, first line hidden — no auto-expand on load.
"""
from __future__ import annotations
import hashlib
from collections.abc import Callable
from datetime import UTC, datetime
from playwright.sync_api import Page, expect
from sqlalchemy import text
from sqlalchemy.orm import Session
from app.db import SessionLocal
from app.models import Chunk, Document, GitSource
from e2e.auth_helpers import login
from tests.e2e.mock_llm import embed_text
# --------------------------------------------------------------------------
# KB fixture — byte-identical to the phase-37 suite (the mirror)
# --------------------------------------------------------------------------
SEED_SOURCE = "Homelab"
SEED_PATH = "aws-route53.md"
SEED_SP = f"{SEED_SOURCE}/{SEED_PATH}"
READ_SOURCE = "Deployments"
READ_PATH = "example-record-file.json"
READ_SP = f"{READ_SOURCE}/{READ_PATH}"
#: The retrievable document: references the JSON file "for the exact
#: JSON shape of reeselink.json" but never includes it. The repeated
#: record-file lines carry the marker question's key tokens (aws,
#: route53, hosted, zone, reeselink, json, exact, shape) — verified
#: ≈0.69 cosine against the mock's embeddings (E2E threshold 0.30)
#: plus FTS hits, so the turn is solidly grounded (HIGH prompt →
#: ``<tools>``).
ROUTE53_CONTENT = (
"# AWS Route 53 Notes\n\n"
"## Record file\n\n"
+ (
"The aws route53 hosted zone for reeselink keeps every record in "
"reseelink.json — the exact JSON shape of reeselink.json is "
"documented in example-record-file.json.\n"
)
* 10
+ "\n## Sync job\n\n"
"A cron job pushes reeselink.json to the aws route53 hosted zone "
"every fifteen minutes; the diff is applied through the route53 api.\n"
)
#: The referenced document: the exact JSON shape, seeded WITHOUT chunks
#: (indexed + catalogued + readable, but never a retrieval candidate).
#: Its ``(source, path)`` sorts FIRST in the catalog — the line the
#: mock's drill + read land on.
RECORD_FILE_CONTENT = (
'{ "version": 3, "comment": "ReeseLink hosted zone records — the exact '
'JSON shape of reeselink.json",\n'
' "hosted_zone_id": "Z0RESEELINK01",\n'
' "record_sets": [\n'
' { "name": "www.reeselink.example", "type": "A", "ttl": 300,\n'
' "resource_records": [ { "value": "10.0.0.20" } ] },\n'
' { "name": "api.reeselink.example", "type": "CNAME", "ttl": 300,\n'
' "resource_records": [ { "value": "www.reeselink.example" } ] }\n'
" ]\n"
"}\n"
)
assert "\n" not in RECORD_FILE_CONTENT[:63] # the quote's content part stays one line
#: The SAME marker that drives the mock's 3-tool flow (mirrored from
#: ``test_agent_document_tools.py`` — deterministic ``ls`` →
#: ``ls(scoped)`` → ``read``).
MARKER_QUESTION = (
"Use your tools: what is the exact JSON shape of reeselink.json "
"for my aws route53 hosted zone?"
)
#: The mock's deterministic single-read answer —
#: "Read <source/path>. <first 80 chars of the read document>"
#: (the phase-37 shape). This suite's completion marker: its presence
#: in the bubble proves the FIRST DELTA has landed — the turn is
#: COMPLETE and the disclosure has already folded (phase 117, D3).
MOCK_ANSWER_MARKER = f"Read {READ_SP}."
#: The phase-117 selectors (scoped to the conversation column).
SUMMARY = "#messages .msg.brain .tool-calls-summary"
DISCLOSURE = "#messages .msg.brain .tool-calls-disclosure"
LINES = "#messages .msg.brain .tool-call"
# --------------------------------------------------------------------------
# DB seeding (TRUNCATE-then-seed, cf. test_agent_document_tools.py)
# --------------------------------------------------------------------------
def _seed(db: Session) -> None:
"""The phase-37 two-document pair, byte-identical (see the module
docstring), plus the phase-94 registry rows: both sources
registered, ``Deployments`` FIRST — the mock's drill (first
source of the top-level listing) lands on the JSON file
deterministically, independent of the operator's
``BOR_GIT_SOURCES`` (a non-empty table ignores the env fallback)."""
# COMMIT between the inserts (not flush): ``added_at`` is
# ``server_default now()`` — the transaction timestamp — and the
# tie-break is the random uuid ``id``, so one-transaction rows order
# nondeterministically.
db.add(GitSource(url=READ_SOURCE, kind="local"))
db.commit()
db.add(GitSource(url=SEED_SOURCE, kind="local"))
md = Document(
source=SEED_SOURCE,
path=SEED_PATH,
full_path=f"/tmp/{SEED_PATH}",
title="AWS Route 53 Notes",
content=ROUTE53_CONTENT,
content_hash=hashlib.sha256(ROUTE53_CONTENT.encode()).hexdigest(),
indexed_at=datetime.now(UTC),
# Phase 106 (D5): explicit dates — byte-stable prompts/quotes
# (the mock's first-80-chars read quote carries the date line).
created_at=datetime(2024, 6, 15, tzinfo=UTC),
)
db.add(md)
db.flush()
# One chunk carrying the mock's own embedding → genuine token
# overlap between the marker question and this document.
db.add(
Chunk(
document_id=md.id,
position=0,
content=ROUTE53_CONTENT,
embedding=embed_text(ROUTE53_CONTENT),
)
)
# The referenced JSON: indexed, catalogued, readable — but NO
# chunks, so retrieval never puts it in context.
db.add(
Document(
source=READ_SOURCE,
path=READ_PATH,
full_path=f"/tmp/{READ_PATH}",
title="Example Record File",
content=RECORD_FILE_CONTENT,
content_hash=hashlib.sha256(RECORD_FILE_CONTENT.encode()).hexdigest(),
indexed_at=datetime.now(UTC),
created_at=datetime(2024, 6, 15, tzinfo=UTC),
)
)
def _reset_db(seed: Callable[[Session], None] | None = None) -> None:
"""Truncate the KB (plus the prompt-shaping tables and the
phase-55 auto-save rows), then re-seed — the E2E isolation
pattern. ``steering_notes`` / ``kb_overview`` are truncated too, so
the HIGH prompt is exactly ``<relevance>`` + ``<documents>`` +
``<tools>`` regardless of leftovers — byte-stable prompts,
byte-stable answers."""
with SessionLocal() as db:
db.execute(
text(
"TRUNCATE chunks, documents, query_log, "
"steering_notes, kb_overview, saved_chats, git_sources"
)
)
db.commit()
if seed is not None:
seed(db)
db.commit()
# --------------------------------------------------------------------------
# Page helpers
# --------------------------------------------------------------------------
def _disclosure_open(page: Page) -> bool:
"""The disclosure's open state: a closed ``<details>`` has NO
``open`` attribute (``get_attribute`` → ``None``); an open one
carries it (any value — the native boolean attribute is empty)."""
return page.get_attribute(DISCLOSURE, "open") is not None
def _submit_tools_turn(page: Page, app_url: str) -> None:
"""Log in as admin, submit the marker question, and wait for the
answer bubble — i.e. the turn is COMPLETE (the disclosure has
already folded on the first delta).
Auth: the phase-37 agent-tools E2E's pattern (the mock flow + the
house pattern for this flow win)."""
login(page, app_url, next="/")
page.fill("#message-input", MARKER_QUESTION)
page.click("#send-btn")
# The user bubble lands synchronously with the submit handler.
expect(page.locator(".msg.user .bubble").last).to_contain_text(MARKER_QUESTION)
# The answer's deterministic quote — its presence proves the first
# delta has landed (the fold rides that frame, phase 117 D3)…
expect(
page.locator(".msg.brain .bubble").last
).to_contain_text(MOCK_ANSWER_MARKER, timeout=30_000)
# …and the turn settled: button recovered (phase-48 — the in-flight
# button is the enabled Stop control, so the label assertion carries
# the settle wait).
expect(page.locator("#send-btn")).to_be_enabled(timeout=30_000)
expect(page.locator("#send-label")).to_have_text("Send", timeout=30_000)
def _expand_disclosure(page: Page) -> None:
"""Tap the summary — the native ``<summary>`` is a real focusable
toggle (AA), so a plain click is the user contract."""
page.locator(SUMMARY).click()
# --------------------------------------------------------------------------
# 1. A completed 3-call turn folds to one "Tool calls (3)" line
# --------------------------------------------------------------------------
def test_tool_calls_fold_to_one_line_after_turn(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
"""Pin 1 (the space fix): after the deterministic ``ls`` →
``ls(scoped)`` → ``read`` turn COMPLETES, the record is ONE visible
"Tool calls (3)" summary — the disclosure is CLOSED (no ``open``
attribute: the first answer delta folded it, D3), the three
``.tool-call`` lines are still in the DOM (the permanent record)
but hidden (the first line is not visible), and the answer bubble
is present."""
page.set_default_timeout(30_000)
_reset_db(_seed)
_submit_tools_turn(page, app_url)
# One compact summary line, visible, with the live count.
summary = page.locator(SUMMARY)
expect(summary).to_be_visible()
expect(summary).to_contain_text("Tool calls (3)")
# Folded at rest: a closed <details> carries no `open` attribute.
assert page.get_attribute(DISCLOSURE, "open") is None
# The lines are present in the DOM (the permanent record) but
# folded: the FIRST line is not rendered.
expect(page.locator(LINES)).to_have_count(3)
expect(page.locator(LINES).first).not_to_be_visible()
# The answer is present (the mock's deterministic single-read quote).
expect(page.locator(".msg.brain .bubble").last).to_contain_text(
MOCK_ANSWER_MARKER
)
# --------------------------------------------------------------------------
# 2. Tapping the summary expands the calls
# --------------------------------------------------------------------------
def test_summary_click_expands_the_calls(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
"""Pin 2 (the expand contract): clicking the native ``<summary>``
opens the disclosure (the ``open`` attribute appears) and all
three ``.tool-call`` lines become visible, in order, with the
phase-37/94 pinned text (the drill line is "Listing documents in
<source>" — the top level lists sources only)."""
page.set_default_timeout(30_000)
_reset_db(_seed)
_submit_tools_turn(page, app_url)
assert not _disclosure_open(page) # folded after the turn (pin 1)
_expand_disclosure(page)
# The native boolean attribute is now present (any value — it is
# empty in the DOM; "true" is the attribute-present shorthand).
assert page.get_attribute(DISCLOSURE, "open") is not None
# All three lines, in order, visible with their pinned text.
lines = page.locator(LINES)
expect(lines).to_have_count(3)
expect(lines.first).to_be_visible()
expect(lines.nth(0)).to_be_visible()
expect(lines.nth(0)).to_contain_text("Listing documents")
expect(lines.nth(1)).to_be_visible()
expect(lines.nth(1)).to_contain_text("Listing documents in")
expect(lines.nth(1)).to_contain_text(READ_SOURCE)
expect(lines.nth(2)).to_be_visible()
expect(lines.nth(2)).to_contain_text("Reading ")
expect(lines.nth(2)).to_contain_text(READ_SP)
# --------------------------------------------------------------------------
# 3. An expanded line is deboxed inline-flow text (not two flex items)
# --------------------------------------------------------------------------
def test_expanded_line_is_deboxed_inline_flow(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
"""Pin 3 (the wrapping fix): on the expanded "Reading" line the
computed ``display`` is NOT ``flex`` — the pre-phase-117 card had
``display: flex; align-items: baseline``, which made the label text
node and the ``<code>`` path two separate flex items (the label
broke mid-word, the path wrapped indented). Deboxed (D4), the label
+ inline ``<code>`` flow as one continuous run: the path wraps to
the left edge like a normal sentence."""
page.set_default_timeout(30_000)
_reset_db(_seed)
_submit_tools_turn(page, app_url)
_expand_disclosure(page)
# The visible "Reading" line (the third of the three).
line = page.locator(LINES).nth(2)
expect(line).to_be_visible()
expect(line).to_contain_text("Reading ")
# The debox: NOT a flex item pair — one inline run. (As a flex item
# of the .tool-calls column its used display blockifies to "block";
# the pre-phase-117 card computed "flex" — the discriminator.)
display = page.evaluate(
"""() => {
const lines = document.querySelectorAll('#messages .msg.brain .tool-call');
const el = Array.from(lines).find(
(l) => l.textContent.includes('Reading '));
return el ? getComputedStyle(el).display : null;
}"""
)
assert display is not None and display != "flex", (
"the expanded .tool-call line must be deboxed (label + path one "
f"inline run, not two flex items) — computed display: {display!r}"
)
# --------------------------------------------------------------------------
# 4. A reload restores the turn FOLDED (no auto-expand on load)
# --------------------------------------------------------------------------
def test_restored_turn_renders_folded(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
"""Pin 4 (the restore path): after the turn completes (the
phase-50/55 auto-save persisted it — the local ``bor.chat.v1``
record is written at the save points, so the reload sees the full
conversation), a same-context RELOAD re-renders the persisted
record through the restore path (``renderStoredMessage`` →
``appendToolLine`` → ``closeToolCalls``): the disclosure is
present, CLOSED (no ``open`` attribute — no auto-expand on load),
the summary reads "Tool calls (3)", the three lines are in the DOM
and the first is hidden (folded on load)."""
page.set_default_timeout(30_000)
_reset_db(_seed)
_submit_tools_turn(page, app_url)
page.reload()
expect(page.locator("#empty-state")).to_be_hidden(timeout=30_000)
# The restored brain message's answer (the restore is synchronous at
# boot — the persisted record is byte-stable for this fixture).
expect(
page.locator(".msg.brain .bubble").last
).to_contain_text(MOCK_ANSWER_MARKER, timeout=30_000)
# Folded on load: one compact summary line, disclosure closed.
expect(page.locator(SUMMARY)).to_be_visible()
expect(page.locator(SUMMARY)).to_contain_text("Tool calls (3)")
assert page.get_attribute(DISCLOSURE, "open") is None
# The three lines are in the DOM (the permanent record) but folded:
# the first line is not rendered — no auto-expand.
expect(page.locator(LINES)).to_have_count(3)
expect(page.locator(LINES).first).not_to_be_visible()