feat(chat): render markdown tables in answers, viewer, and thinking

GFM pipe tables in the shared renderer (TODO.md L6): a table-protection
pass in frontend/assets/markdown.js (fences -> tables -> escape order)
pulls each header+separator+body block out as a placeholder, renders
cells escape-first with the same inline transforms, and reinserts a
semantic <table class="md-table"> inside a horizontal-overflow
.md-table-wrap — so a pipe table in a chat answer, the document
viewer/modal, and the thinking block all render the same semantic
table. Fences win over tables; lone pipes stay text.

- styles.css: .md-table palette rules (PLAN §7.2 tokens, no motion);
  min-width: max-content so a WIDE table keeps its natural width and
  the wrapper is the real scroller (width:100% alone wrapped the wide
  table's cells — proven by the new E2E).
- mock_llm.py: TABLE_TRIGGER ("show me a table") -> byte-stable
  TABLE_ANSWER (3-column table, <img onerror> XSS probe line, wide
  5-column table), checked before DEFLECT_MODE like SUMMARY_MODE.
- tests/fixtures/docs/homelab/tables.md: 3x3 pipe table + pipe-heavy
  fenced block (viewer/fence subject); the shared fixture set grows
  8 -> 9 docs, so every suite pinning the count (added/formats/
  stat-docs/EXPECTED_ROWS) is updated accordingly.
- tests/e2e/test_markdown_tables.py (new, story suite): chat table
  shape + non-deflection, wide-table wrapper scroll (no page
  overflow), XSS probe inert, viewer modal table, fence-not-a-table,
  lone pipe stays text.
- tests/e2e/test_agent_document_tools.py: fix a pre-existing flake —
  the "Calling tool…" label window is ~0.4 s at the mock's 0.1 s
  tool-frame pacing, and a polling expect could stride over it
  (failed 3 of 5 runs on the committed baseline). The pre-submit
  MutationObserver record is the deterministic source of truth; the
  racy to_have_text gate is gone.

uv run pytest: 738 passed, app/ coverage 99% (TOTAL unchanged);
ruff + pyright clean; story E2E 6/6 in isolation; regression E2E
suites (chat_rag, document_viewer, document_summaries, smoke) green.
This commit is contained in:
2026-08-28 03:35:50 -04:00
parent 27b7cb96d5
commit bc70ce36e0
30 changed files with 967 additions and 73 deletions
+56
View File
@@ -65,6 +65,17 @@ Implements just enough of the aipi surface:
section, or with the tool conversation not yet started and no tools
offered — e.g. budgets 0/0) behave exactly as today. ``E2E_REAL_LLM=1``
ignores the mock entirely (the real model does what it does).
- user message containing ``show me a table`` (phase 44, markdown
tables, TODO.md L6) -> the fixed table answer (``TABLE_ANSWER``):
a 3-column service table, an ``<img onerror>`` XSS probe line, and
a deliberately wide 5-column table — byte-stable, so the story E2E
can assert the rendered ``<table class="md-table">`` shape, the
escaped XSS line, and the wrapper's horizontal scroll inside the
46rem column. Checked BEFORE the ``DEFLECT_MODE`` branch (a
deflection prompt never carries the marker, same reasoning as
``SUMMARY_MODE``), so a marker question always gets the table
answer; the E2E asks it against an on-topic fixture (HIGH gate) and
asserts non-deflection.
``max_tokens`` is honored deterministically (token ≈ whitespace word),
like a real endpoint: an answer longer than the cap is truncated. This
@@ -163,6 +174,37 @@ _DOCUMENTS_BLOCK_RE = re.compile(r"<documents>.*?</documents>", re.S)
#: contain the phrase, so every other suite is unaffected.
TOOLS_TRIGGER = "use your tools"
#: Phase 44 (markdown-tables story, TODO.md L6): a user message
#: containing this substring (case-insensitive) gets the fixed table
#: answer (``TABLE_ANSWER`` below) — a 3-column table, an XSS probe
#: line, and a deliberately wide table (see the module docstring).
#: Existing E2E questions do not contain the phrase, so every other
#: suite is unaffected.
TABLE_TRIGGER = "show me a table"
#: The fixed table answer (phase 44) — byte-stable on purpose: the story
#: E2E asserts the rendered table shape, the escaped ``<img onerror>``
#: line (the XSS payload must survive the mock byte-for-byte), and the
#: wide table's ``scrollWidth > clientWidth`` inside the 46rem column.
TABLE_ANSWER = (
"Here's the shape, in a table:\n"
"\n"
"| Service | Port | Host |\n"
"|---|---|---|\n"
"| Caddy | 80 | homelab-gw |\n"
"| GitLab | 8929 | homelab-git |\n"
"| ntfy | 2087 | homelab-ntfy |\n"
"\n"
"<img src=x onerror=alert(1)>\n"
"\n"
"And the wide one:\n"
"\n"
"| A very long column header to force overflow | Second column with "
"some padding text | Third column | Fourth | Fifth |\n"
"|---|---|---|---|---|\n"
"| value-one | value-two | value-three | value-four | value-five |"
)
#: The agent's ``read_document`` tool-result prefix (app.rag.agent
#: ``_execute_tool``): ``"Document <source/path>:\n<content>"``.
@@ -298,6 +340,20 @@ def compose_answer(body: dict[str, Any]) -> str:
answer = "Knowledge base outline:\n- " + " ".join(
TOKEN_RE.findall(user.lower())[:8]
)
elif TABLE_TRIGGER in user.lower():
# Markdown tables (phase 44, TODO.md L6): the story E2E's
# deterministic table answer — a 3-column table, the
# <img onerror> XSS probe line (it must survive the mock
# byte-for-byte so the E2E can prove the renderer neutralizes
# it), and a wide 5-column table (guarantees scrollWidth >
# clientWidth inside the 46rem column). Byte-stable. Checked
# BEFORE the DEFLECT_MODE branch: a deflection prompt never
# carries the marker (it lives in the user message, same
# reasoning as SUMMARY_MODE), so a marker question always gets
# the table answer, whatever the gate says; the E2E asks it
# against an on-topic fixture, where the gate is HIGH, and
# asserts non-deflection as part of the table test.
answer = TABLE_ANSWER
elif "DEFLECT_MODE" in system:
answer = (
"Ah — I haven't done anything like that, so I don't want to make stuff up! "
+2 -2
View File
@@ -208,10 +208,10 @@ def test_admin_login_unlocks_sources_and_tuning(
login(page, app_url)
expect(page).to_have_url(app_url + "/sources.html")
expect(page.locator("#sources-gate")).to_be_hidden()
expect(page.locator("#stat-docs")).to_have_text("8")
expect(page.locator("#stat-docs")).to_have_text("9") # phase 44: +tables.md
expect(page.locator("#stat-chunks")).not_to_have_text("–")
expect(page.locator("#docs-table")).to_be_visible()
expect(page.locator("#docs-tbody tr")).to_have_count(8)
expect(page.locator("#docs-tbody tr")).to_have_count(9)
# Chat: the tuning UI is back — header toggle with count badge,
# Sign out instead of Sign in, Tune under the answer.
+8 -6
View File
@@ -327,14 +327,16 @@ def test_marker_question_lists_reads_and_quotes(
_install_page_hooks(page)
_submit(page, MARKER_QUESTION)
# While a tool runs the button carries the "calling tool" label: the
# first `tool` frame sets it and it holds until the FIRST answer
# delta (the agent loop completes before the answer stream) — so the
# poll issued right after the click must catch it inside that window.
expect(page.locator("#send-label")).to_have_text("Calling tool…", timeout=20_000)
# The "calling tool" label window is transient: the first `tool`
# frame sets it and it holds until the FIRST answer delta (the agent
# loop completes before the answer stream) — ~0.4 s at the mock's
# 0.1 s tool-frame pacing. A polling expect can stride straight over
# that window (observed flake, fixed in phase 44 task 03), so the
# pre-submit MutationObserver record below is the deterministic
# source of truth for the label transition.
_wait_settled(page)
# The label transition is also recorded deterministically (no race):
# The label transition, recorded deterministically (no race):
# Thinking… → Calling tool… → … → Send.
labels = page.evaluate("() => window.__labels")
assert "Calling tool…" in labels, labels
+1 -1
View File
@@ -128,7 +128,7 @@ def test_conversation_survives_reload(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8 # A9 formats
assert summary is not None and summary.added == 9 # A9 formats (phase 44 added tables.md)
page.set_default_timeout(30_000)
page.goto(app_url)
_ask(page, QUESTION)
+1 -1
View File
@@ -76,7 +76,7 @@ def test_on_topic_question_streams_grounded_answer(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8 # A9 formats
assert summary is not None and summary.added == 9 # A9 formats (phase 44 added tables.md)
page.set_default_timeout(30_000)
page.goto(app_url)
+1 -1
View File
@@ -114,7 +114,7 @@ def _seed_kb(mock_port: int) -> ImportSummary:
db.execute(text("TRUNCATE chunks, documents, query_log"))
db.commit()
summary = _run_in_thread(_import_fixtures(mock_port))
assert summary is not None and summary.added == 8 # A9 formats
assert summary is not None and summary.added == 9 # A9 formats (phase 44 added tables.md)
return summary
+1 -1
View File
@@ -7,7 +7,7 @@ Run in isolation (DB must be up: ``podman compose up -d db``):
The fixture KB is a story-dedicated directory
(``tests/fixtures/summary_kb/`` — the shared ``tests/fixtures/docs/``
stays at its 8 pinned files) with two documents:
stays at its 9 pinned files) with two documents:
* ``quadlet/qwen-llamacpp.yaml`` — a non-markdown A9 doc. At import the
mock ``lite`` model (``SUMMARY_MODE`` marker, ``tests/e2e/mock_llm.py``)
+1 -1
View File
@@ -336,7 +336,7 @@ def test_edit_note_steers_answer(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8 # A9 formats
assert summary is not None and summary.added == 9 # A9 formats (phase 44 added tables.md)
page.set_default_timeout(30_000)
_open_tuning(page, app_url)
+1 -1
View File
@@ -84,7 +84,7 @@ def test_off_topic_question_deflects_honestly(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8 # A9 formats
assert summary is not None and summary.added == 9 # A9 formats (phase 44 added tables.md)
page.set_default_timeout(30_000)
page.goto(app_url)
expect(page.locator("#kb-banner")).to_be_hidden()
+7 -5
View File
@@ -40,6 +40,7 @@ EXPECTED_ROWS = (
"homelab/networking/static-dns.json",
"homelab/scripts/uptime_probe.py",
"homelab/ssh/ssh_aliases.txt",
"homelab/tables.md", # phase 44: the markdown-tables fixture
)
@@ -85,13 +86,14 @@ def test_sources_page_lists_indexed_docs(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
# Eight A9-format files are imported; .hidden/junk.md is out of scope
# (A9 revised — hidden path components are never walked).
assert summary is not None and summary.added == 8
assert summary.formats == {"md": 4, "yaml": 1, "json": 1, "py": 1, "txt": 1}
# Nine A9-format files are imported (phase 44 added homelab/tables.md);
# .hidden/junk.md is out of scope (A9 revised — hidden path components
# are never walked).
assert summary is not None and summary.added == 9
assert summary.formats == {"md": 5, "yaml": 1, "json": 1, "py": 1, "txt": 1}
login(page, app_url) # phase 16: the catalog is admin-only
expect(page.locator("#stat-docs")).to_have_text("8")
expect(page.locator("#stat-docs")).to_have_text("9")
# Phase 30: non-markdown fixtures each gained one ``is_summary`` chunk,
# so the Sources total is content chunks + summary chunks.
expect(page.locator("#stat-chunks")).to_have_text(
+2 -2
View File
@@ -170,7 +170,7 @@ def test_on_topic_answer_echoes_kb_overview(
the mock's echo of the ``<knowledge_base>`` section's first bullet —
only possible if the section reached the LLM prompt."""
summary = _reset_db(mock_llm, seed=True, overview=OVERVIEW)
assert summary is not None and summary.added == 8 # A9 formats
assert summary is not None and summary.added == 9 # phase 44 added tables.md
assert _overview_row() is not None # the row the turn must inject
bubble = _ask(page, app_url, QUESTION, KB_ECHO)
@@ -250,7 +250,7 @@ def test_mock_generated_outline_is_stored_and_echoed(
stores the byte-stable 8-token digest of the generator's document
list, and a chat turn echoes its first bullet."""
summary = _reset_db(mock_llm, seed=True, overview=None)
assert summary is not None and summary.added == 8
assert summary is not None and summary.added == 9 # phase 44 added tables.md
assert _overview_row() is None # direct import never regenerates
kwargs: dict[str, Any] = {
+1 -1
View File
@@ -176,7 +176,7 @@ def test_typing_indicator_during_slow_think(
"""AC1/AC5: the 3s mock warm-up must show the typing indicator for
>=2s before any text appears, then it is gone once the answer lands."""
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8 # A9 formats
assert summary is not None and summary.added == 9 # A9 formats (phase 44 added tables.md)
page.set_default_timeout(30_000)
page.goto(app_url)
+350
View File
@@ -0,0 +1,350 @@
"""Phase 44 E2E (Playwright): GFM pipe tables in the shared renderer.
Story: ``.agent/user_stories/markdown-tables.md``
Run in isolation (DB must be up: ``podman compose up -d db``):
uv run pytest tests/e2e/test_markdown_tables.py -v --no-cov
Seeding reuses the real importer against ``tests/fixtures/docs/`` with
the deterministic mock embeddings (same pattern as ``test_chat_rag.py``).
The mock's ``TABLE_TRIGGER`` (``show me a table``, phase 44 task 02)
returns the byte-stable table answer: a 3-column service table, an
``<img onerror>`` XSS probe line, and a deliberately wide 5-column
table. The phase-44 fixture ``homelab/tables.md`` (a 3×3 pipe table
plus a pipe-heavy fenced block) is the viewer/fence subject — the
document viewer is database-only, so the imported row is enough.
Test → story mapping (Playwright Mapping Rule):
1. ``test_chat_table_renders`` — the brain bubble carries
``<div class="md-table-wrap"><table class="md-table">`` with a
``<thead>`` of three ``<th scope="col">`` (Service/Port/Host), the
expected body cells, no raw ``|---|`` separator text, and the turn is
NOT deflected (the honesty-gate interplay is part of the contract).
2. ``test_wide_table_scrolls`` — the wide table's wrapper has
``scrollWidth > clientWidth`` and horizontal scroll moves it; the
page itself has no horizontal overflow (the 46rem column holds).
3. ``test_table_xss_safe`` — the ``<img onerror>`` line renders as
visible, escaped text: zero injected ``<img>`` nodes, no dialog.
4. ``test_viewer_table_renders`` — the fixture's pipe table opens from
the Sources table (admin) in the modal and renders the same
``<table class="md-table">`` (shared renderer, story AC6).
5. ``test_fence_not_a_table`` — the fixture's pipe-heavy fenced block
renders ``<pre><code>``; the only ``<table>`` in the document is the
real pipe table (fences win, story AC3).
6. ``test_plain_pipe_stays_text`` — a grounded prose answer with a lone
``|`` (the mock echoes the question) renders as text, no
``<table>`` (story AC4).
"""
from __future__ import annotations
import asyncio
from pathlib import Path
from threading import Thread
from typing import Any
from playwright.sync_api import Page, expect
from sqlalchemy import text
from app.config import Settings
from app.db import SessionLocal
from app.rag.importer import ImportSummary, import_sources
from app.rag.llm import LLMClient
from e2e.auth_helpers import login
REPO = Path(__file__).resolve().parents[2]
FIXTURES = REPO / "tests" / "fixtures" / "docs"
#: Carries the mock's ``TABLE_TRIGGER`` ("show me a table") and is
#: on-topic (the fixture set answers it — FTS-OR grounds it, so the
#: turn is HIGH and the suite can assert non-deflection).
QUESTION = "Show me a table of my homelab services?"
#: Grounded kubernetes question with a single ``|`` in the prose — the
#: mock's default branch echoes the question (first 80 chars), so the
#: lone pipe lands in the rendered answer.
PLAIN_QUESTION = "How is my Kubernetes cluster set up? A lone | in prose stays text."
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
TABLES_PATH = "homelab/tables.md"
WIDE_HEADER = "A very long column header to force overflow"
XSS_LINE = "<img src=x onerror=alert(1)>"
EXPECTED_HEADER = ["Service", "Port", "Host"]
EXPECTED_ROWS = [
["Caddy", "80", "homelab-gw"],
["GitLab", "8929", "homelab-git"],
["ntfy", "2087", "homelab-ntfy"],
]
async def _import_fixtures(mock_port: int) -> ImportSummary:
kwargs: dict[str, Any] = {"_env_file": None, "llm_base_url": f"http://127.0.0.1:{mock_port}/v1"}
settings = Settings(**kwargs) # pyright: ignore[reportCallIssue]
return await import_sources([FIXTURES], LLMClient(settings))
def _run_in_thread(coro: Any) -> Any:
"""Run a coroutine on a worker thread.
Playwright's sync API keeps an asyncio loop running on the test
thread, so ``asyncio.run`` cannot be called directly from a test
body (the established house helper).
"""
box: dict[str, Any] = {}
def runner() -> None:
try:
box["value"] = asyncio.run(coro)
except BaseException as e: # noqa: BLE001 — re-raised on the test thread
box["error"] = e
t = Thread(target=runner)
t.start()
t.join()
if "error" in box:
raise box["error"]
return box["value"]
def _reset_db(mock_port: int, seed: bool) -> ImportSummary | None:
"""Truncate the KB (+ the global prompt-state rows), then optionally
re-import the fixtures (9 docs since phase 44 added tables.md)."""
with SessionLocal() as db:
db.execute(
text("TRUNCATE chunks, documents, query_log, steering_notes, kb_overview")
)
db.commit()
if not seed:
return None
return _run_in_thread(_import_fixtures(mock_port))
def _ask_table_answer(page: Page, app_url: str) -> Any:
"""Drive the trigger question and return the brain bubble once the
whole byte-stable table answer has streamed in (the wide table's
last cell lands last)."""
page.goto(app_url)
page.fill("#message-input", QUESTION)
page.click("#send-btn")
bubble = page.locator(".msg.brain .bubble").first
bubble.wait_for(state="visible", timeout=30_000)
expect(bubble).to_contain_text("value-five", timeout=30_000)
# Non-deflection is part of the table contract (honesty gate interplay).
expect(page.locator(".msg.brain.is-deflected")).to_have_count(0)
return bubble
def _open_tables_doc_modal(page: Page, app_url: str) -> None:
"""Admin → Sources → the tables.md row → same-page document modal."""
login(page, app_url) # phase 16: the Sources catalog is admin-only
row = page.locator("#docs-tbody tr", has_text=TABLES_PATH)
expect(row).to_have_count(1)
row.locator("td:nth-child(2) a.doc-link").click()
expect(page.locator(".doc-modal")).to_be_visible()
expect(page.locator("#doc-modal-title")).to_have_text("Service Port Table")
# ---------------------------------------------------------------------------
# 1. Chat: the pipe table renders as a semantic table
# ---------------------------------------------------------------------------
def test_chat_table_renders(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 9 # phase 44: +tables.md
page.set_default_timeout(30_000)
bubble = _ask_table_answer(page, app_url)
# Both tables of the answer rendered: the 3-column service table and
# the wide one — each in its horizontal-overflow wrapper.
tables = bubble.locator("table.md-table")
expect(tables).to_have_count(2)
expect(bubble.locator(".md-table-wrap")).to_have_count(2)
# The 3×3 table: <thead> of three <th scope="col"> + the body cells
# (the whole answer has already streamed in — the DOM is settled).
first = tables.nth(0)
headers = first.locator("thead th[scope='col']")
expect(headers).to_have_count(3)
assert headers.all_inner_texts() == EXPECTED_HEADER
rows = first.locator("tbody tr")
expect(rows).to_have_count(3)
for i, cells in enumerate(EXPECTED_ROWS):
assert rows.nth(i).locator("td").all_inner_texts() == cells
# The raw markdown must not survive: no separator row, no raw header
# row as text anywhere in the bubble.
bubble_text = bubble.inner_text()
assert "|---|" not in bubble_text, "the |---| separator leaked into the bubble"
assert "| Service | Port | Host |" not in bubble_text, "the raw header row leaked"
# Grounded retrieval: the table fixture is the top source chip.
chip = page.locator(".msg.brain .source-chip", has_text="homelab/tables.md")
expect(chip).to_have_count(1)
# ---------------------------------------------------------------------------
# 2. Wide table: the wrapper scrolls, the page does not
# ---------------------------------------------------------------------------
def test_wide_table_scrolls(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
_reset_db(mock_llm, seed=True)
page.set_default_timeout(30_000)
bubble = _ask_table_answer(page, app_url)
# The wide table (5 columns, one deliberately long header) sits in
# ITS wrapper — the 3-column table's wrapper is not the scroller.
wrap = bubble.locator(".md-table-wrap", has=page.locator("th", has_text=WIDE_HEADER))
expect(wrap).to_have_count(1)
scroll_width, client_width = wrap.evaluate(
"el => [el.scrollWidth, el.clientWidth]"
)
assert scroll_width > client_width, (
f"the wide table must overflow its wrapper "
f"(scrollWidth {scroll_width} <= clientWidth {client_width})"
)
# Horizontal scrolling (scrollLeft) moves the wrapper's content.
before = wrap.evaluate("el => el.scrollLeft")
wrap.evaluate("el => { el.scrollLeft = 120; }")
after = wrap.evaluate("el => el.scrollLeft")
assert after > before, "the wrapper must scroll horizontally"
# The 46rem chat column must not break the page: no horizontal
# document overflow (PLAN §7.1).
page_scroll, page_client = page.evaluate(
"() => [document.documentElement.scrollWidth, document.documentElement.clientWidth]"
)
assert page_scroll <= page_client, (
f"the page overflowed horizontally ({page_scroll} > {page_client})"
)
# ---------------------------------------------------------------------------
# 3. XSS-safe: the <img onerror> probe renders inert text
# ---------------------------------------------------------------------------
def test_table_xss_safe(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
_reset_db(mock_llm, seed=True)
page.set_default_timeout(30_000)
dialogs: list[str] = []
def _catch(d) -> None: # a fired dialog == the probe executed
dialogs.append(d.message)
d.dismiss()
page.on("dialog", _catch)
_ask_table_answer(page, app_url)
state = page.evaluate(
"""() => {
const el = document.querySelector('.msg.brain .bubble');
return {
imgs: el.querySelectorAll('img').length,
onerror: el.querySelectorAll('[onerror]').length,
text: el.innerText,
html: el.innerHTML,
};
}"""
)
assert state["imgs"] == 0, "the XSS probe became a live <img> element"
assert state["onerror"] == 0, "an onerror attribute survived into the DOM"
# The escaped tag renders as VISIBLE text (the escape-first contract).
assert XSS_LINE in state["text"], "the probe line must be visible text"
assert "&lt;img src=x onerror=alert(1)&gt;" in state["html"]
assert dialogs == [], f"dialog fired — the probe executed: {dialogs}"
# ---------------------------------------------------------------------------
# 4. Shared renderer: the viewer/modal renders the fixture's table
# ---------------------------------------------------------------------------
def test_viewer_table_renders(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
_reset_db(mock_llm, seed=True)
page.set_default_timeout(30_000)
_open_tables_doc_modal(page, app_url)
# The same <table class="md-table"> shape the chat bubble gets — the
# shared renderer (story AC6) serves the viewer too.
table = page.locator("#doc-modal-content table.md-table")
expect(table).to_have_count(1)
headers = table.locator("thead th[scope='col']")
expect(headers).to_have_count(3)
assert headers.all_inner_texts() == EXPECTED_HEADER
rows = table.locator("tbody tr")
expect(rows).to_have_count(3)
for i, cells in enumerate(EXPECTED_ROWS):
assert rows.nth(i).locator("td").all_inner_texts() == cells
assert (
"|---|" not in page.locator("#doc-modal-content").inner_text()
), "the separator row leaked into the viewer"
# ---------------------------------------------------------------------------
# 5. Fences win: the pipe-heavy fenced block is code, never a table
# ---------------------------------------------------------------------------
def test_fence_not_a_table(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
_reset_db(mock_llm, seed=True)
page.set_default_timeout(30_000)
_open_tables_doc_modal(page, app_url)
# The fixture's ``` block (pipe table inside) renders as code —
# fence protection runs before the table pass (story AC3).
pre = page.locator("#doc-modal-content pre code")
expect(pre).to_have_count(1)
expect(pre).to_contain_text("caddy", timeout=30_000)
code_text = pre.inner_text()
assert "| Service | Port |" in code_text, "the fenced header line must stay raw"
assert "|----------|------|" in code_text, "the fenced separator must stay raw"
assert "| caddy | 80 |" in code_text
assert "| gitlab | 8929 |" in code_text
# Exactly ONE table in the whole document — the real pipe table. The
# fenced rows (lowercase "caddy"/"gitlab") must not become cells.
table = page.locator("#doc-modal-content table.md-table")
expect(table).to_have_count(1)
cells = table.locator("th, td").all_inner_texts()
assert "caddy" not in cells and "gitlab" not in cells, (
"the fenced pipe block was parsed as a table"
)
# ---------------------------------------------------------------------------
# 6. Non-tables stay put: a lone pipe in grounded prose renders as text
# ---------------------------------------------------------------------------
def test_plain_pipe_stays_text(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
_reset_db(mock_llm, seed=True)
page.set_default_timeout(30_000)
page.goto(app_url)
page.fill("#message-input", PLAIN_QUESTION)
page.click("#send-btn")
bubble = page.locator(".msg.brain .bubble").first
bubble.wait_for(state="visible", timeout=30_000)
expect(bubble).to_contain_text(MOCK_ANSWER_MARKER, timeout=30_000)
# Grounded (the kubernetes FTS hit), not deflected — this is the
# default-answer path, so the echoed question is what we assert on.
expect(page.locator(".msg.brain.is-deflected")).to_have_count(0)
# A single "|" in prose is not a table (no header + separator pair).
expect(bubble.locator("table")).to_have_count(0)
expect(bubble.locator(".md-table-wrap")).to_have_count(0)
assert "A lone | in prose stays text" in bubble.inner_text()
+4 -4
View File
@@ -123,11 +123,11 @@ def _reset_db(mock_port: int, seed: bool) -> ImportSummary | None:
@pytest.fixture()
def seeded_kb(mock_llm: int, db_ready: None) -> Iterator[None]:
"""A fresh KB seeded from ``tests/fixtures/docs`` (8 docs, A9 formats),
truncated again on teardown. ``db_ready`` (conftest) skips with clear
instructions when Postgres is down."""
"""A fresh KB seeded from ``tests/fixtures/docs`` (9 docs since phase
44, A9 formats), truncated again on teardown. ``db_ready`` (conftest)
skips with clear instructions when Postgres is down."""
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8
assert summary is not None and summary.added == 9
yield
_reset_db(mock_llm, seed=False)
+6 -5
View File
@@ -87,9 +87,10 @@ def test_multi_format_import_hidden_doc_excluded(
"""A9 (revised): all seven formats import; hidden (dot) paths never do."""
summary = _reset_db(mock_llm, seed=True)
assert summary is not None
# Eight A9-format fixture files; .hidden/junk.md must never be walked.
assert summary.added == 8
assert summary.formats == {"md": 4, "yaml": 1, "json": 1, "py": 1, "txt": 1}
# Nine A9-format fixture files (phase 44 added homelab/tables.md);
# .hidden/junk.md must never be walked.
assert summary.added == 9
assert summary.formats == {"md": 5, "yaml": 1, "json": 1, "py": 1, "txt": 1}
# Phase 16: the catalog is admin-only — perform the real form login,
# then call the API with the signed cookie the browser now holds.
@@ -102,7 +103,7 @@ def test_multi_format_import_hidden_doc_excluded(
r = httpx.get(f"{app_url}/api/docs", timeout=10, cookies=cookies)
assert r.status_code == 200
docs = r.json()["documents"]
assert len(docs) == 8
assert len(docs) == 9
assert all(".hidden" not in d["path"] for d in docs)
assert {d["path"] for d in docs} >= {
"homelab/container_gitlab/gitlab.md",
@@ -113,7 +114,7 @@ def test_multi_format_import_hidden_doc_excluded(
}
# The Sources page (we're already on it, signed in) reflects the set.
expect(page.locator("#stat-docs")).to_have_text("8")
expect(page.locator("#stat-docs")).to_have_text("9")
expect(page.locator("#docs-tbody tr", has_text=".hidden")).to_have_count(0)
+4 -4
View File
@@ -154,11 +154,11 @@ def _no_error_banner(page: Page) -> None:
@pytest.fixture()
def seeded_kb(mock_llm: int, db_ready: None) -> Iterator[None]:
"""A fresh KB seeded from ``tests/fixtures/docs`` (8 docs, A9 formats),
truncated again on teardown. ``db_ready`` (conftest) skips with clear
instructions when Postgres is down."""
"""A fresh KB seeded from ``tests/fixtures/docs`` (9 docs since phase
44, A9 formats), truncated again on teardown. ``db_ready`` (conftest)
skips with clear instructions when Postgres is down."""
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8
assert summary is not None and summary.added == 9
yield
_reset_db(mock_llm, seed=False)
+1 -1
View File
@@ -127,7 +127,7 @@ def test_tune_under_answer_persists_and_steers(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8 # A9 formats
assert summary is not None and summary.added == 9 # A9 formats (phase 44 added tables.md)
page.set_default_timeout(30_000)
page.goto(app_url)
login(page, app_url, next="/") # phase 16: tuning is admin-only
+1 -1
View File
@@ -63,7 +63,7 @@ def _seed_kb(mock_port: int) -> ImportSummary:
db.execute(text("TRUNCATE chunks, documents, query_log"))
db.commit()
summary = _run_in_thread(_import_fixtures(mock_port))
assert summary is not None and summary.added == 8 # A9 formats
assert summary is not None and summary.added == 9 # A9 formats (phase 44 added tables.md)
return summary
+4 -4
View File
@@ -99,11 +99,11 @@ def _reset_db(mock_port: int, seed: bool) -> ImportSummary | None:
@pytest.fixture()
def seeded_kb(mock_llm: int, db_ready: None) -> Iterator[None]:
"""A fresh KB seeded from ``tests/fixtures/docs`` (8 docs, A9 formats),
truncated again on teardown. ``db_ready`` (conftest) skips with clear
instructions when Postgres is down."""
"""A fresh KB seeded from ``tests/fixtures/docs`` (9 docs since phase
44, A9 formats), truncated again on teardown. ``db_ready`` (conftest)
skips with clear instructions when Postgres is down."""
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8
assert summary is not None and summary.added == 9
yield
_reset_db(mock_llm, seed=False)
+4 -4
View File
@@ -148,11 +148,11 @@ def _reset_db(mock_port: int, seed: bool) -> ImportSummary | None:
@pytest.fixture()
def seeded_kb(mock_llm: int, db_ready: None) -> Iterator[None]:
"""A fresh KB seeded from ``tests/fixtures/docs`` (8 docs, A9 formats),
truncated again on teardown (same fixture shape as the phase-17/21
suites)."""
"""A fresh KB seeded from ``tests/fixtures/docs`` (9 docs since phase
44, A9 formats), truncated again on teardown (same fixture shape as
the phase-17/21 suites)."""
summary = _reset_db(mock_llm, seed=True)
assert summary is not None and summary.added == 8
assert summary is not None and summary.added == 9
yield
_reset_db(mock_llm, seed=False)
+3 -3
View File
@@ -17,8 +17,8 @@ The oversized documents are seeded directly via SQLAlchemy (a
``documents`` row + 2–3 ``chunks`` rows whose embeddings are the mock's
own deterministic bag-of-words vectors, so the question's live mock
embedding genuinely overlaps — no fixture files added:
``tests/fixtures/docs/`` stays at its 8 files, other suites pin
``summary.added == 8``).
``tests/fixtures/docs/`` stays at its 9 files (phase 44), other suites
pin ``summary.added == 9``).
"""
from __future__ import annotations
@@ -301,7 +301,7 @@ def test_small_document_path_unchanged(
path, byte-identical to before — no marker, kubernetes.md cited."""
_reset_db(None)
summary = _run_in_thread(_import_fixtures(mock_llm))
assert summary.added == 8 # A9 formats (fixture set unchanged)
assert summary.added == 9 # A9 formats (phase 44 added tables.md)
bubble = _ask(page, app_url, SMALL_QUESTION)
expect(bubble).to_contain_text(SMALL_QUESTION, timeout=30_000)