Files
brain-of-reese/tests/e2e/test_ls_tree_drilldown.py
T
ducoterra 35d65d2f25
Build and Push Containers / build-and-push-app (push) Successful in 2m10s
Build and Push Containers / build-and-push-db (push) Successful in 13s
feat(rag): summarize single-document folders (MIN_DOCS_PER_FOLDER 2 → 1)
Relax the phase-94 folder-summary scope rule from ≥ 2 documents to
≥ 1: a folder (or source root) is a candidate while ANY document
lives under it, so single-file folders and single-file source roots
get their own lite-written description. A row is now pruned only
when its folder loses its last document (vanishes from the
catalogue).

The constant is the single source of truth, so the flip propagates
to the generator's candidate set, the prune pass, the
missing_folder_summaries gap probe (the next sync self-heals the
new gaps), and the KB-tree summary_pending markers (1-doc folders /
sources now read "Summary pending" until their row lands).

Docstrings/comments across app/, scripts/import_docs.py, and the
E2E fixtures updated to the ≥ 1 wording. Unit + integration tests
updated to the new semantics (the pruned-below-minimum scenario is
now a folder losing its LAST doc; single-doc folders are pinned as
candidates/pending). Full suite: 2314 passed, app coverage 99%;
ruff + pyright clean; folder-summary E2E stories pass in isolation
(ls_tree_drilldown, sync_summary_visibility, kb_tree,
kb_tree_nav, document_dates, oneshot_llm_retry).
2026-09-14 08:57:57 -04:00

778 lines
32 KiB
Python

"""Phase 94 task 04 E2E (Playwright, mock-only): the ``ls`` drill-down tree.
The dedicated story suite for ``94_ls_tree_drilldown`` (owner TODO.md L4):
the LLM drills ``ls()`` → ``ls(source)`` → ``ls(source/folder)`` →
``read(source/file)`` through the real UI, the sync-time folder summaries
(must exist after a changed sync under the deterministic mock's
``FOLDER_SUMMARY_MODE`` branch) show up in the ``ls`` output, and the
50-line cap holds on a wide folder.
Run in isolation (DB must be up: ``podman compose up -d db``):
uv run pytest tests/e2e/test_ls_tree_drilldown.py -v --no-cov
MOCK-ONLY suite: ``E2E_REAL_LLM=1`` is not supported — the real ``turbo``
does whatever it does with the tools, while this story's gate is the
deterministic SCRIPTED drill-down flow in ``tests/e2e/mock_llm.py``
(``DRILL_TRIGGER``: user message contains ``drill down the tree`` —
the question carries its own tool call after the colon,
``drill down the tree: ls [target]`` / ``drill down the tree: read
source/path`` — **and** the system prompt carries the ``<tools>``
section of the HIGH prompt). The mock echoes the received tool result
into its final answer (the house scripted-turn way of asserting on tool
results — the mock is the only E2E lens on the LLM's context), so every
tree-level assertion below lands on the rendered answer; the DOM
assertions cover the tool lines + the answer.
KB fixture — a host temp dir tree (``tmp_path_factory``; the app runs on
the same host) with TWO registered local sources (the
``test_local_directory_sources.py`` registration + real-Sync pattern —
registration through the authenticated API, the real in-process
``POST /api/sync`` pipeline; no git anywhere):
* ``alpha/`` — ``root-note.md`` at the source root, ``one/`` (2 docs),
``two/`` (2 docs — the read target lives here) and ``wide/`` (51 tiny
files — the 50-line cap subject);
* ``beta/`` — ``gamma/`` (2 docs).
Every fixture doc carries the words ``drill down the tree`` in its body,
so every scripted question (which contains the trigger phrase) FTS-matches
at least one chunk — the honesty gate is HIGH for all turns regardless
of the mock's cosine distribution, and the ``<tools>`` section is present
(the flow's precondition).
The mock's canned ``FOLDER_SUMMARY_MODE`` branch (phase 94 task 01)
stores, per existing folder (the ≥ 1-doc rule — this fixture holds
no single-doc folders), the deterministic one-liner
``Fixture folder summary for <source>[/<folder>].`` — the ``synced_kb``
module fixture pins those exact rows in ``folder_summaries`` after the
sync, and the drill answers assert on them in the ``ls`` output.
Test → story mapping (Playwright Mapping Rule):
1. ``test_drill_down_sources_folders_files_and_read`` — the four
scripted turns: ``ls()`` (the ``🔎 Listing documents`` line + the
top-level shape — one ``— N documents`` line per source + the
canned source-root summaries), ``ls(alpha)`` (folder lines with
their summaries + the root file line), ``ls(alpha/two)`` (the exact
``source: X | path: Y | title: Z`` file lines), and the grounded
``read(alpha/two/two-a.md)`` (the ``📄 Reading …`` line + the answer
citing the document, the phase-37 assertion pattern). The read
target is DELIBERATELY a top-2 retrieval document for its question
(the question names the file's path, so the file self-matches the
hybrid gate deterministically): the agent's phase-72 dedupe returns
``ALREADY_IN_CONTEXT`` (a refusal — counts in nothing), and the
mock answers FROM THE ``<documents>`` PROMPT with the same citation
shape (``Already in context: Read <sp>. <first 80 chars>``) — the
document text reached the model either way, and the prefix pins
that the dedupe notice itself reached it.
2. ``test_wide_folder_holds_the_fifty_line_cap`` — ``ls(alpha/wide)``
on the 51-file folder: the mock's echo carries exactly 50 file lines
+ the ``…and 1 more documents in this folder — use grep (pattern)…``
note; the 51st file never reaches the model.
3. ``test_not_a_folder_teaching_and_scripted_recovery`` — the mock calls
``ls(alpha/nope)``; the NOT-A-FOLDER teaching line (the argument
echoed, the parent's subfolders listed) is visible to the model —
the mock's scripted recovery branch keys on receiving it — and the
scripted ``ls(alpha)`` recovery lands (the answer is the parent's
listing; the loop ends in one refusal + one correction, not at the
round cap).
"""
from __future__ import annotations
import json
import os
import re
import subprocess
import sys
import time
from collections.abc import Iterator
from pathlib import Path
from typing import Any
import httpx
import pytest
from playwright.sync_api import Locator, Page, expect
from sqlalchemy import select, text
from app.config import Settings as _Settings
from app.db import SessionLocal
from app.models import FolderSummary, QueryLog
from e2e.auth_helpers import login
from e2e.conftest import (
ADMIN_PASSWORD,
SESSION_SECRET,
USE_REAL_LLM,
_wait_http,
)
REPO = Path(__file__).resolve().parents[2]
# Phase 79 (task 04, full inventory): the conftest session app owns its
# port in a combined run — this module app binds its own port instead
# (a same-port second uvicorn dies on bind and would drive the wrong
# server). Env-overridable.
APP_PORT = int(os.environ.get("E2E_APP_PORT_LSTREE", "8136"))
APP_URL = f"http://127.0.0.1:{APP_PORT}"
# --------------------------------------------------------------------------
# Fixture documents (deterministic, token-controlled)
# --------------------------------------------------------------------------
ALPHA = "alpha"
BETA = "beta"
ROOT_NOTE = "root-note.md"
TWO_A = "two/two-a.md"
TWO_B = "two/two-b.md"
READ_SP = f"{ALPHA}/{TWO_A}"
ALPHA_COUNT = 56 # 1 root note + 2 one/ + 2 two/ + 51 wide/
BETA_COUNT = 2
TOTAL_DOCS = ALPHA_COUNT + BETA_COUNT
WIDE_COUNT = 51
LS_MAX_FILE_LINES = 50 # app.rag.agent.LS_MAX_FILE_LINES — the cap under test
#: Every fixture body carries ``drill down the tree`` (the trigger
#: phrase's words): every scripted question FTS-matches at least one
#: chunk → HIGH gate → the ``<tools>`` section the flow keys on.
DRILL_LEAD = "The drill down the tree fixture note"
def _md(title: str, body: str) -> str:
return f"# {title}\n\n{body}\n"
#: The read target — its FIRST line is ≥ 80 chars, so the mock's
#: first-80-chars quote (the phase-37 single-read shape) is newline-free
#: and the rendered-text assertion matches it verbatim. Pinned by the
#: assert below.
TWO_A_TITLE = (
"Alpha Two A — the drill-down read target for the alpha two folder "
"listing turn in the brain of reese fixture"
)
TWO_A_CONTENT = _md(
TWO_A_TITLE,
f"{DRILL_LEAD} for alpha two: this document covers topic A of the "
"alpha source tree; it is the file the scripted read turn opens "
"from the alpha/two listing.",
)
assert "\n" not in TWO_A_CONTENT[:80] # the quote must stay one line
#: The sync-time folder summaries the mock's canned ``FOLDER_SUMMARY_MODE``
#: branch stores (task 01's byte-stable template), in
#: ``(source, folder_path)`` order: one row per existing folder (the
#: ≥ 1-doc recursive-subtree rule — this fixture holds no single-doc
#: folders) — the ``""`` rows are the source roots.
EXPECTED_SUMMARIES: list[tuple[str, str, str]] = [
(ALPHA, "", f"Fixture folder summary for {ALPHA}."),
(ALPHA, "one", f"Fixture folder summary for {ALPHA}/one."),
(ALPHA, "two", f"Fixture folder summary for {ALPHA}/two."),
(ALPHA, "wide", f"Fixture folder summary for {ALPHA}/wide."),
(BETA, "", f"Fixture folder summary for {BETA}."),
(BETA, "gamma", f"Fixture folder summary for {BETA}/gamma."),
]
assert [
(source, folder) for source, folder, _s in EXPECTED_SUMMARIES
] == sorted((source, folder) for source, folder, _s in EXPECTED_SUMMARIES)
# --- the pinned tree levels (app.rag.agent's phase-94 templates) -------
#: ``ls()`` — the top level: sources in registry order (alpha registered
#: first), each with its recursive count + stored source-root summary.
TOP_HEADER = "2 sources:"
TOP_LINES = [
f"{ALPHA} — {ALPHA_COUNT} documents",
EXPECTED_SUMMARIES[0][2],
f"{BETA} — {BETA_COUNT} documents",
EXPECTED_SUMMARIES[4][2],
]
#: ``ls(alpha)`` — the source root: the subfolder lines (path order)
#: with their stored summaries, then the root's own file line (the
#: canonical ``read``/``grep`` identity format, unchanged).
SOURCE_HEADER = f"{ALPHA} — 1 documents, 3 folders:"
SOURCE_LINES = [
f"one/ — 2 documents: {EXPECTED_SUMMARIES[1][2]}",
f"two/ — 2 documents: {EXPECTED_SUMMARIES[2][2]}",
f"wide/ — {WIDE_COUNT} documents: {EXPECTED_SUMMARIES[3][2]}",
f"source: {ALPHA} | path: {ROOT_NOTE} | title: Alpha Root Note",
]
#: ``ls(alpha/two)`` — a leaf folder: the file lines in EXACTLY the
#: existing ``source: X | path: Y | title: Z`` format, path order.
FOLDER_HEADER = f"{ALPHA}/two — 2 documents, 0 folders:"
FOLDER_LINES = [
f"source: {ALPHA} | path: {TWO_A} | title: {TWO_A_TITLE}",
f"source: {ALPHA} | path: {TWO_B} | title: Alpha Two B",
]
#: ``ls(alpha/wide)`` — the 50-line cap: 51 files → 50 lines (path
#: order: wide-01 … wide-50) + one deterministic grep-pointer note;
#: wide-51 never reaches the model.
WIDE_HEADER = f"{ALPHA}/wide — {WIDE_COUNT} documents, 0 folders:"
WIDE_FIRST = f"source: {ALPHA} | path: wide/wide-01.md | title: Wide 01"
WIDE_LAST = f"source: {ALPHA} | path: wide/wide-50.md | title: Wide 50"
WIDE_NOTE = (
"…and 1 more documents in this folder — use grep "
"(pattern) to find a specific one."
)
# --- the scripted turns (the mock's ``DRILL_TRIGGER`` questions) -------
TOP_QUESTION = "Drill down the tree: ls — what sources are indexed?"
SOURCE_QUESTION = f"Drill down the tree: ls {ALPHA} — what's in source {ALPHA}?"
FOLDER_QUESTION = f"Drill down the tree: ls {ALPHA}/two — list that folder"
READ_QUESTION = f"Drill down the tree: read {READ_SP} — read the file"
WIDE_QUESTION = f"Drill down the tree: ls {ALPHA}/wide — how many files does this folder hold?"
NOPE_QUESTION = f"Drill down the tree: ls {ALPHA}/nope — is there such a folder?"
#: The read target is a top-2 retrieval document for its question (the
#: question names the path — the file self-matches the hybrid gate
#: deterministically), so the read gets the phase-72 ALREADY_IN_CONTEXT
#: dedupe and the mock answers from the ``<documents>`` prompt with the
#: same citation shape, prefixed (the suite pins the dedupe path).
READ_ANSWER_PREFIX = f"Already in context: Read {READ_SP}."
READ_ANSWER_QUOTE = TWO_A_CONTENT[:80]
# --------------------------------------------------------------------------
# Fixtures
# --------------------------------------------------------------------------
@pytest.fixture(scope="module")
def drill_dirs(tmp_path_factory: pytest.TempPathFactory) -> tuple[Path, Path]:
"""The two-source temp tree (see the module docstring): the app
server runs on the same host, so the paths are visible to it. The
directory NAMES are the source names (``kind=local`` → the
directory's basename, phase 38)."""
root = tmp_path_factory.mktemp("bor_ls_tree")
alpha = root / ALPHA
beta = root / BETA
(alpha / "one").mkdir(parents=True)
(alpha / "two").mkdir(parents=True)
(alpha / "wide").mkdir(parents=True)
(beta / "gamma").mkdir(parents=True)
(alpha / ROOT_NOTE).write_text(
_md(
"Alpha Root Note",
f"{DRILL_LEAD} at the alpha source root: this file sits "
"directly under the alpha source, not in any folder.",
),
encoding="utf-8",
)
(alpha / "one" / "one-a.md").write_text(
_md(
"Alpha One A",
f"{DRILL_LEAD} for alpha one: this document covers topic A "
"of the alpha source tree.",
),
encoding="utf-8",
)
(alpha / "one" / "one-b.md").write_text(
_md(
"Alpha One B",
f"{DRILL_LEAD} for alpha one: this document covers topic B "
"of the alpha source tree.",
),
encoding="utf-8",
)
(alpha / TWO_A).write_text(TWO_A_CONTENT, encoding="utf-8")
(alpha / TWO_B).write_text(
_md(
"Alpha Two B",
f"{DRILL_LEAD} for alpha two: this document covers topic B "
"of the alpha source tree.",
),
encoding="utf-8",
)
for nn in range(1, WIDE_COUNT + 1):
(alpha / "wide" / f"wide-{nn:02d}.md").write_text(
_md(
f"Wide {nn:02d}",
f"One of {WIDE_COUNT} tiny files in the alpha wide "
f"folder: {DRILL_LEAD.lower()} line {nn:02d}.",
),
encoding="utf-8",
)
(beta / "gamma" / "gamma-a.md").write_text(
_md(
"Beta Gamma A",
f"{DRILL_LEAD} for beta gamma: this document covers topic A "
"of the beta source tree.",
),
encoding="utf-8",
)
(beta / "gamma" / "gamma-b.md").write_text(
_md(
"Beta Gamma B",
f"{DRILL_LEAD} for beta gamma: this document covers topic B "
"of the beta source tree.",
),
encoding="utf-8",
)
assert (alpha / TWO_A).is_file() and (beta / "gamma" / "gamma-b.md").is_file()
return alpha, beta
@pytest.fixture(scope="module")
def app_server(mock_llm: int, drill_dirs: tuple[Path, Path]) -> Iterator[str]:
"""The real app under test — per-module app (the conftest pattern,
cf. ``test_local_directory_sources.py``): NO ``BOR_GIT_SOURCES``
(the env fallback is git-only — the sources here are DB-registered
local directories), the mock LLM, the mock-calibrated threshold,
and the leak-guarded code defaults. The session app is never
started in this isolated run, so no port clash."""
env = dict(os.environ)
env.pop("DEBUGPY", None)
env["BOR_ENVIRONMENT"] = "e2e"
env["BOR_STATIC_DIR"] = str(REPO / "frontend")
env["BOR_LLM_BASE_URL"] = (
"https://aipi.reeseapps.com/v1"
if USE_REAL_LLM
else f"http://127.0.0.1:{mock_llm}/v1"
)
# Mock-calibrated threshold (conftest pattern): every scripted
# question FTS-matches the fixture docs (the ``drill down the tree``
# words), so the gate is HIGH either way.
env["BOR_RELEVANCE_THRESHOLD"] = "0.30"
# Phase 67: instant retry waits + the code-default budget (the
# conftest leak-guard pattern).
env["BOR_LLM_RETRY_DELAY"] = "0"
env["BOR_LLM_RETRIES"] = str(_Settings.model_fields["llm_retries"].default)
env.setdefault(
"BOR_DATABASE_URL",
"postgresql+psycopg://reese:reese@localhost:5432/brain_of_reese",
)
# Phase 16: admin auth must be set or create_app() refuses to boot.
env["BOR_ADMIN_PASSWORD"] = ADMIN_PASSWORD
env["BOR_SESSION_SECRET"] = SESSION_SECRET
# The repo's .env file carries the owner's BOR_GIT_SOURCES (the app
# reads it from cwd) — override it with an EMPTY value (the env var
# beats the .env file): the registry must hold EXACTLY the two
# local directories this suite registers (a leftover env git list
# would pollute the top-level ``ls`` the whole story asserts on).
env["BOR_GIT_SOURCES"] = ""
# Leak guards (conftest pattern): an operator's local (gitignored)
# .env cannot leak corpus-specific settings into the app under test.
env["BOR_DOCS_REPO"] = ""
env["BOR_SUGGESTIONS"] = json.dumps(
_Settings.model_fields["suggestions"].default
)
env["BOR_INPUT_PLACEHOLDER"] = _Settings.model_fields["input_placeholder"].default
env["BOR_FOOTER_TEXT"] = _Settings.model_fields["footer_text"].default
proc = subprocess.Popen(
[sys.executable, "-m", "uvicorn", "app.main:app",
"--host", "127.0.0.1", "--port", str(APP_PORT), "--log-level", "warning"],
cwd=REPO,
env=env,
)
try:
_wait_http(f"{APP_URL}/api/health")
yield APP_URL
finally:
proc.terminate()
try:
proc.wait(timeout=10)
except subprocess.TimeoutExpired:
proc.kill()
@pytest.fixture(scope="module")
def app_url(app_server: str) -> str:
return app_server
def _truncate_all() -> None:
"""Fresh registry + KB (the E2E isolation pattern): the E2E suites
share one Postgres, so a leftover git_sources row would pollute the
top-level ``ls`` and a leftover document would show up in the
folder listings the drill answers assert on byte-exactly."""
with SessionLocal() as db:
db.execute(
text(
"TRUNCATE chunks, documents, query_log, steering_notes, "
"kb_overview, git_sources, folder_summaries"
)
)
db.commit() # without the commit the TRUNCATE rolls back (the house pattern)
db.commit()
def _wait_sync_done_http(client: httpx.Client, timeout_s: float = 180.0) -> dict[str, Any]:
"""Poll the (cookie-authenticated) status endpoint until the run
reaches a terminal state (the test_local_directory_sources pattern,
over plain httpx — this fixture has no browser page yet)."""
deadline = time.monotonic() + timeout_s
body: dict[str, Any] = {}
while time.monotonic() < deadline:
r = client.get("/api/sync/status")
assert r.status_code == 200, r.text
body = r.json()
if body["state"] in ("success", "failed"):
return body
time.sleep(0.5)
raise AssertionError(f"sync did not reach a terminal state: {body}")
@pytest.fixture(scope="module")
def synced_kb(app_server: str, drill_dirs: tuple[Path, Path]) -> None:
"""The story's precondition: the folder-structured KB synced under
the deterministic mock.
Registers the two temp directories through the authenticated API
(the ``test_local_directory_sources.py`` pattern — ``alpha`` FIRST,
committed separately, so the registry order — ``(added_at, id)`` —
lists alpha before beta, the top-level ``ls`` order the suite
asserts), runs the REAL in-process sync (``POST /api/sync`` —
walk → chunk → embed → overview → folder summaries → version bump),
and pins the stored folder summaries: the mock's canned
``FOLDER_SUMMARY_MODE`` branch (task 01) makes the sync store one
deterministic row per existing folder (the ≥ 1-doc rule — this
fixture holds no single-doc folders) — the drill turns' answers
assert on that exact text.
"""
alpha, beta = drill_dirs
_truncate_all()
with httpx.Client(base_url=app_server, timeout=30.0) as client:
r = client.post("/api/login", json={"password": ADMIN_PASSWORD})
assert r.status_code == 204, r.text
r = client.post(
"/api/git-sources", json={"kind": "local", "path": str(alpha)}
)
assert r.status_code == 201, r.text
time.sleep(0.05) # distinct added_at: alpha before beta (registry order)
r = client.post(
"/api/git-sources", json={"kind": "local", "path": str(beta)}
)
assert r.status_code == 201, r.text
r = client.post("/api/sync")
assert r.status_code == 202, r.text
body = _wait_sync_done_http(client)
assert body["state"] == "success", body
detail = body["detail"]
assert detail["added"] == TOTAL_DOCS, detail
assert detail["pruned"] == 0, detail
assert detail["overview"] is True, detail
# The change-gated folder summaries (phase 94 task 02) landed: one
# row per existing folder (the ≥ 1-doc rule — this fixture holds
# no single-doc folders), the mock's byte-stable text (the drill
# answers quote exactly these lines).
with SessionLocal() as db:
rows = db.execute(
select(
FolderSummary.source, FolderSummary.folder_path,
FolderSummary.summary,
).order_by(FolderSummary.source, FolderSummary.folder_path)
).all()
assert [(s, f, t) for s, f, t in rows] == EXPECTED_SUMMARIES, rows
@pytest.fixture(autouse=True)
def _clean(db_ready: None) -> Iterator[None]:
"""Per-test query_log isolation (the KB itself is module-scoped —
the drill turns never change it, so the folder summaries and the
registry persist across the tests of this module)."""
with SessionLocal() as db:
db.execute(text("TRUNCATE query_log"))
db.commit()
yield
with SessionLocal() as db:
db.execute(text("TRUNCATE query_log"))
db.commit()
# --------------------------------------------------------------------------
# Page helpers (the test_agent_document_tools house pattern)
# --------------------------------------------------------------------------
#: Captures the raw SSE ``data:`` payloads of the /api/chat stream
#: (a response clone read in the background) — wire-level assertions
#: for the ``tool`` frames, independent of the UI rendering.
SSE_HOOK = """
() => {
if (window.__sseInstalled) return;
window.__sseInstalled = true;
window.__sseFrames = [];
const origFetch = window.fetch;
window.fetch = async function (...args) {
const res = await origFetch.apply(this, args);
try {
const url = typeof args[0] === 'string' ? args[0] : args[0].url;
if (url.includes('/api/chat')) {
res.clone().text().then((bodyText) => {
for (const block of bodyText.split('\\n\\n')) {
const line = block.trim();
if (line.startsWith('data: ')) {
window.__sseFrames.push(line.slice(6));
}
}
});
}
} catch (e) { /* non-clonable responses: ignored */ }
return res;
};
}
"""
def _install_page_hooks(page: Page) -> None:
page.evaluate(SSE_HOOK)
def _frames(page: Page) -> list[dict]:
"""The SSE frames captured since the last submit (``_submit``
clears the buffer), once the hook's background read settles."""
deadline = time.monotonic() + 30.0
while True:
raw = page.evaluate("() => window.__sseFrames || []")
parsed = [json.loads(line) for line in raw if line]
if any(f.get("type") == "done" for f in parsed):
return parsed
if time.monotonic() > deadline:
raise AssertionError(
f"SSE hook captured no `done` frame (frames so far: "
f"{len(parsed)}) — hook install failed?"
)
time.sleep(0.05)
def _tool_frames(frames: list[dict]) -> list[dict]:
return [f for f in frames if f.get("type") == "tool"]
def _submit(page: Page, question: str) -> None:
page.evaluate("window.__sseFrames = []")
page.fill("#message-input", question)
page.click("#send-btn")
# The user bubble lands synchronously with the submit handler.
expect(page.locator(".msg.user .bubble").last).to_contain_text(question)
def _wait_settled(page: Page) -> None:
"""The turn is complete: answer text in the bubble, button recovered
(the phase-48 settle wait, the test_agent_document_tools helper)."""
expect(page.locator(".msg.brain .bubble").last).not_to_have_text("", timeout=30_000)
expect(page.locator("#send-btn")).to_be_enabled(timeout=30_000)
expect(page.locator("#send-label")).to_have_text("Send", timeout=30_000)
def _last_brain(page: Page) -> Locator:
return page.locator(".msg.brain").last
def _assert_turn(
page: Page,
expected_tools: list[dict[str, Any]],
expected_answer_lines: list[str],
) -> None:
"""One scripted drill turn, fully asserted: the wire carries exactly
the expected ``tool`` frames (ahead of the first ``delta``), the
bubble carries the expected answer lines (the mock's echo of the
tool results the model received), and the turn was grounded (the
``done`` frame is not deflected)."""
frames = _frames(page)
assert _tool_frames(frames) == expected_tools, _tool_frames(frames)
if expected_tools:
first_delta = next(
i for i, f in enumerate(frames) if f.get("type") == "delta"
)
assert all(
i < first_delta
for i, f in enumerate(frames)
if f.get("type") == "tool"
)
done = next(f for f in frames if f.get("type") == "done")
assert done["deflected"] is False, done
bubble = _last_brain(page).locator(".bubble")
for line in expected_answer_lines:
expect(bubble).to_contain_text(line)
def _query_log_rows() -> list[QueryLog]:
with SessionLocal() as db:
return list(db.scalars(select(QueryLog)).all())
# --------------------------------------------------------------------------
# 1. The scripted drill: sources → folders (summaries) → files → read
# --------------------------------------------------------------------------
def test_drill_down_sources_folders_files_and_read(
page: Page, app_url: str, synced_kb: None, db_ready: None
) -> None:
page.set_default_timeout(30_000)
login(page, app_url, next="/")
_install_page_hooks(page)
# --- turn 1: ls() — the top level (sources + source-root summaries) -
_submit(page, TOP_QUESTION)
_wait_settled(page)
lines = _last_brain(page).locator(".tool-call")
expect(lines).to_have_count(1)
# The no-arg ls line — NOT a scoped "Listing documents in …" one
# (regex match: string expectations normalize whitespace, so the
# scope check must be a byte-exact pattern).
expect(lines.nth(0)).to_contain_text("Listing documents")
expect(lines.nth(0)).not_to_have_text(re.compile(r"Listing documents in"))
_assert_turn(
page,
[{"type": "tool", "name": "ls", "argument": None}],
[TOP_HEADER, *TOP_LINES],
)
# --- turn 2: ls(alpha) — the source root (subfolders + root files) --
_submit(page, SOURCE_QUESTION)
_wait_settled(page)
lines = _last_brain(page).locator(".tool-call")
expect(lines).to_have_count(1)
expect(lines.nth(0)).to_contain_text(f"Listing documents in {ALPHA}")
_assert_turn(
page,
[{"type": "tool", "name": "ls", "argument": ALPHA}],
[SOURCE_HEADER, *SOURCE_LINES],
)
# --- turn 3: ls(alpha/two) — the folder level (the file lines) ------
_submit(page, FOLDER_QUESTION)
_wait_settled(page)
lines = _last_brain(page).locator(".tool-call")
expect(lines).to_have_count(1)
expect(lines.nth(0)).to_contain_text(f"Listing documents in {ALPHA}/two")
_assert_turn(
page,
[{"type": "tool", "name": "ls", "argument": f"{ALPHA}/two"}],
[FOLDER_HEADER, *FOLDER_LINES],
)
# --- turn 4: read(alpha/two/two-a.md) — the grounded read ------------
_submit(page, READ_QUESTION)
_wait_settled(page)
lines = _last_brain(page).locator(".tool-call")
expect(lines).to_have_count(1)
# The phase-37 "Reading <source/path>" line (the combined identity).
expect(lines.nth(0)).to_contain_text(f"Reading {READ_SP}")
frames = _frames(page)
assert _tool_frames(frames) == [
{"type": "tool", "name": "read", "argument": READ_SP}
], _tool_frames(frames)
bubble = _last_brain(page).locator(".bubble")
# The dedupe notice reached the model (the prefix pins it — the read
# was refused as ALREADY_IN_CONTEXT because the target is a top-2
# retrieval document)…
expect(bubble).to_contain_text(READ_ANSWER_PREFIX)
# …and the answer still cites the document: the mock quotes the
# FIRST 80 chars of the target's text from the ``<documents>``
# prompt (the refusal's instruction — answer from that text; the
# quote is newline-free, pinned above).
expect(bubble).to_contain_text(READ_ANSWER_QUOTE)
done = next(f for f in frames if f.get("type") == "done")
assert done["deflected"] is False, done
# The read document is in the turn's sources (retrieval + agent-read,
# deduped — the grounded-turn record).
assert any(
s["path"] == TWO_A and s["source"] == ALPHA for s in done["sources"]
), done["sources"]
# Durable records: all four turns grounded, in order, the read turn
# logging the read document.
rows = _query_log_rows()
assert [r.question for r in rows] == [
TOP_QUESTION, SOURCE_QUESTION, FOLDER_QUESTION, READ_QUESTION
]
assert all(r.deflected is False for r in rows)
assert READ_SP in rows[3].sources, rows[3].sources
# --------------------------------------------------------------------------
# 2. The 50-line cap: a 51-file folder costs the model 50 lines + the note
# --------------------------------------------------------------------------
def test_wide_folder_holds_the_fifty_line_cap(
page: Page, app_url: str, synced_kb: None, db_ready: None
) -> None:
page.set_default_timeout(30_000)
login(page, app_url, next="/")
_install_page_hooks(page)
_submit(page, WIDE_QUESTION)
_wait_settled(page)
lines = _last_brain(page).locator(".tool-call")
expect(lines).to_have_count(1)
expect(lines.nth(0)).to_contain_text(f"Listing documents in {ALPHA}/wide")
frames = _frames(page)
assert _tool_frames(frames) == [
{"type": "tool", "name": "ls", "argument": f"{ALPHA}/wide"}
], _tool_frames(frames)
done = next(f for f in frames if f.get("type") == "done")
assert done["deflected"] is False, done
# The mock echoed the listing VERBATIM: the header carries the
# PRE-cap count (51 — the cap hides lines, not the truth), the
# file lines stop at 50, and the one deterministic grep-pointer note
# folds the 51st file away.
bubble = _last_brain(page).locator(".bubble")
text = bubble.text_content() or ""
assert WIDE_HEADER in text, text
assert WIDE_FIRST in text, text
assert WIDE_LAST in text, text
assert WIDE_NOTE in text, text
assert f"wide/wide-{WIDE_COUNT}.md" not in text, text # the 51st file: gone
# Exactly LS_MAX_FILE_LINES file lines reached the model.
assert text.count(f"source: {ALPHA} | path: wide/") == LS_MAX_FILE_LINES, text
row = _query_log_rows()
assert len(row) == 1
assert row[0].deflected is False
# --------------------------------------------------------------------------
# 3. The NOT-A-FOLDER teaching is visible to the model; the scripted
# recovery (ls of the parent level) works
# --------------------------------------------------------------------------
def test_not_a_folder_teaching_and_scripted_recovery(
page: Page, app_url: str, synced_kb: None, db_ready: None
) -> None:
page.set_default_timeout(30_000)
login(page, app_url, next="/")
_install_page_hooks(page)
_submit(page, NOPE_QUESTION)
_wait_settled(page)
# Two tool lines: the scripted misuse, then the scripted recovery —
# the mock's recovery branch fires ONLY when it RECEIVES the
# NOT-A-FOLDER teaching line (``'alpha/nope' is not a folder —
# alpha has: one/ two/ wide/`` — the argument echoed, the parent's
# subfolders listed): the teaching being visible to the model is
# exactly what the second call proves (the phase-72
# self-correction contract, now carrying the tree's teaching).
lines = _last_brain(page).locator(".tool-call")
expect(lines).to_have_count(2)
expect(lines.nth(0)).to_contain_text(f"Listing documents in {ALPHA}/nope")
expect(lines.nth(1)).to_contain_text(f"Listing documents in {ALPHA}")
frames = _frames(page)
assert _tool_frames(frames) == [
{"type": "tool", "name": "ls", "argument": f"{ALPHA}/nope"},
{"type": "tool", "name": "ls", "argument": ALPHA},
], _tool_frames(frames)
done = next(f for f in frames if f.get("type") == "done")
assert done["deflected"] is False, done
# The loop ended in ONE refusal + ONE correction: the final answer is
# the recovery's PARENT listing (not the round-cap, not an echo of
# the refusal) — the model self-corrected and got the tree level.
bubble = _last_brain(page).locator(".bubble")
for line in [SOURCE_HEADER, *SOURCE_LINES]:
expect(bubble).to_contain_text(line)
row = _query_log_rows()
assert len(row) == 1
assert row[0].deflected is False