diff --git a/.agent/phases/todo/31_kb_overview_prompt/01_migration_kb_overview.md b/.agent/phases/complete/31_kb_overview_prompt/01_migration_kb_overview.md similarity index 100% rename from .agent/phases/todo/31_kb_overview_prompt/01_migration_kb_overview.md rename to .agent/phases/complete/31_kb_overview_prompt/01_migration_kb_overview.md diff --git a/.agent/phases/todo/31_kb_overview_prompt/02_overview_generator.md b/.agent/phases/complete/31_kb_overview_prompt/02_overview_generator.md similarity index 100% rename from .agent/phases/todo/31_kb_overview_prompt/02_overview_generator.md rename to .agent/phases/complete/31_kb_overview_prompt/02_overview_generator.md diff --git a/.agent/phases/todo/31_kb_overview_prompt/03_prompt_injection.md b/.agent/phases/complete/31_kb_overview_prompt/03_prompt_injection.md similarity index 100% rename from .agent/phases/todo/31_kb_overview_prompt/03_prompt_injection.md rename to .agent/phases/complete/31_kb_overview_prompt/03_prompt_injection.md diff --git a/.agent/phases/todo/31_kb_overview_prompt/04_import_trigger.md b/.agent/phases/complete/31_kb_overview_prompt/04_import_trigger.md similarity index 100% rename from .agent/phases/todo/31_kb_overview_prompt/04_import_trigger.md rename to .agent/phases/complete/31_kb_overview_prompt/04_import_trigger.md diff --git a/.agent/user_stories/kb-overview-prompt.md b/.agent/user_stories/kb-overview-prompt.md new file mode 100644 index 0000000..da30b27 --- /dev/null +++ b/.agent/user_stories/kb-overview-prompt.md @@ -0,0 +1,80 @@ +# Story: KB Overview in the System Prompt (lite-generated knowledge-base outline) + +**Phase:** `31_kb_overview_prompt.md` · **E2E:** `tests/e2e/test_kb_overview.py` + +## Narrative + +As **any user of Reese**, I ask questions without knowing what the +knowledge base actually covers. Retrieval only reveals the answer +documents after it has run, and when it finds nothing the deflection has +no picture of what Reese DOES know. I want the system prompt to carry, +on every turn, a plain-text outline of the **basic categories** of +everything that has been read — so the agent knows roughly what its +knowledge base contains before retrieval returns documents. The outline +is **generated by the `lite` model** from the document catalogue +(source, path, title, first summary line per document) and **stored in +a single row**, so it is regenerated whenever an import changes the +knowledge base — never per chat turn. + +- **Given** a knowledge base that was imported (with at least one + document added or updated, or no outline stored yet) +- **When** I ask any question +- **Then** the system prompt of that turn — grounded (HIGH) or + deflected (LOW) — contains a `` section with the + stored outline (budgeted by `BOR_KB_OVERVIEW_MAX_CHARS`, overflow + marked with the shared `[…truncated…]` marker), and the per-turn log + line records `kb_chars=N`; when no outline is stored, both prompts + are byte-identical to the pre-phase text. + +## Acceptance criteria +1. Migration 0005: single-row `kb_overview` table (`id INTEGER PK + DEFAULT 1`, `content TEXT NOT NULL DEFAULT ''`, + `updated_at TIMESTAMPTZ`) + `KbOverview` model — no other schema + change (reversible, integration-tested up/down). +2. `app/rag/overview.py`: the `KB_OVERVIEW_MODE` prompt builder + (document list capped at `BOR_OVERVIEW_INPUT_MAX_CHARS`, overflow + cut at the cap with the shared truncation marker), `load_kb_overview` + (one indexed PK lookup, `""` when the row is missing/empty), and + `regenerate_overview` (best-effort upsert into the single row with a + fresh UTC `updated_at`, `overview: regenerated docs=… chars=…` log + line; a `lite` failure logs and leaves the previous outline intact; + zero documents → no model call, any existing row untouched). +3. Prompt injection (phase 15 convention): the `` + section is budgeted by `BOR_KB_OVERVIEW_MAX_CHARS` (default 4000) + with the shared `[…truncated…]` overflow, ordered + `` → `` → `` → mode body in + **both** the HIGH and LOW prompts; a missing/empty row keeps both + prompts **byte-identical** to the pre-phase text (unit-asserted). +4. Chat turn: `plan_turn` reads the stored row per turn (no per-turn + LLM call) and records `TurnPlan.kb_chars`; the per-turn log line + records `kb_chars=N` after `tuning=N` (PLAN §9 extension). +5. Import trigger: after a run that added/updated at least one document + (or no row exists yet), `import_docs` regenerates the outline with + the same `lite` model (`BOR_LLM_SUMMARY_MODEL` — A5 extended); + unchanged re-imports and `--limit` debug runs never burn a `lite` + call; the summary line ends `overview=updated|skipped|failed` and a + `lite` failure never changes the import's exit code. +6. `.env.example` + README document `BOR_KB_OVERVIEW_MAX_CHARS` / + `BOR_OVERVIEW_INPUT_MAX_CHARS` and the regeneration behavior; the + deterministic mock LLM answers `KB_OVERVIEW_MODE` with a byte-stable + outline (first 8 tokens of the document list) and echoes the + `` section's first bullet as `(kb: …)` — the + `(tuning: …)` steering-echo precedent. +7. Unit + integration green, `app/` coverage >90% (TOTAL ≥ pre-change), + story E2E green in isolation, `ruff` + `pyright` clean, one + `--no-gpg-sign` commit. + +## Playwright Mapping Rule +`tests/e2e/test_kb_overview.py` — one story, one file, run in +isolation. It seeds the `kb_overview` row directly in the DB (the +import-trigger path is integration-tested) with a recognizable +multi-bullet outline, imports the shared fixture KB, asks an on-topic +question and asserts the rendered brain answer ends with +`(kb: Kubernetes cluster and node maintenance notes)` (the mock's echo +of the injected section's first bullet — only possible if the +`` section reached the LLM prompt), asks an off-topic +question and asserts the deflected answer carries the same echo (the +section is in the LOW prompt too), deletes the row and asserts a fresh +answer has no `(kb: …)` suffix (absence end-to-end), and runs the real +`regenerate_overview` against the mock to assert the `KB_OVERVIEW_MODE` +branch's byte-stable outline is stored and echoed by a subsequent turn. diff --git a/.env.example b/.env.example index ee6239b..d5bd7f1 100644 --- a/.env.example +++ b/.env.example @@ -25,6 +25,8 @@ BOR_RELEVANCE_THRESHOLD=0.62 # answer when best cosine >= this OR an FTS hi BOR_MAX_OUTPUT_TOKENS=32768 # max answer length in tokens (answers must not be cut off) BOR_STEERING_MAX_CHARS=8000 # char budget for the (steering notes) prompt section BOR_SUMMARY_MAX_CHARS=12000 # cap on document content sent to the lite summary model (phase 30) +BOR_KB_OVERVIEW_MAX_CHARS=4000 # char budget for the prompt section (phase 31) +BOR_OVERVIEW_INPUT_MAX_CHARS=40000 # cap on the document list sent to the lite model for the KB outline (phase 31) BOR_CHUNK_TARGET_CHARS=2000 BOR_CHUNK_OVERLAP_CHARS=200 BOR_EMBED_BATCH_SIZE=16 diff --git a/README.md b/README.md index 8d8fc92..66bb07d 100644 --- a/README.md +++ b/README.md @@ -247,6 +247,14 @@ embedded. stdlib `ast`) and their title comes from the file stem. - Unchanged files are **not re-embedded** — only new/changed ones, so refreshes are cheap. +- After a run that **changed** the knowledge base (at least one document + added or updated), the import also regenerates the stored **KB + overview** — a plain-text outline of the KB's basic categories that + every chat turn injects into the system prompt as `` + (phase 31). The regeneration is best-effort and change-gated: unchanged + re-imports and `--limit` debug runs skip it (no `lite` call), and a + `lite`-model failure leaves the previous outline intact without failing + the import. The summary line ends `overview=updated|skipped|failed`. - To sanity-check the LLM backend (models + embedding dimension) after any aipi change: `uv run python -m scripts.llm_probe`. @@ -454,6 +462,8 @@ served locally (no CDN), `BOR_ENVIRONMENT=production`. | `BOR_SOURCES_DIR` | `~/bor-sources` | where the `BOR_GIT_SOURCES` repos are cloned/pulled (one subdirectory per repo) | | `BOR_STEERING_MAX_CHARS` | `8000` | char budget for the `` (steering notes) prompt section | | `BOR_SUMMARY_MAX_CHARS` | `12000` | cap on document content sent to the `lite` summary model at import (see *Document summaries*) | +| `BOR_KB_OVERVIEW_MAX_CHARS` | `4000` | char budget for the `` (KB overview) prompt section | +| `BOR_OVERVIEW_INPUT_MAX_CHARS` | `40000` | cap on the document list sent to the `lite` model when generating the KB overview | | `BOR_SUGGESTIONS` | built-in list | JSON list of onboarding chips | | `BOR_ADMIN_PASSWORD` | *(required)* | the single admin's password (plaintext, `.env`); app refuses to start when empty | | `BOR_SESSION_SECRET` | *(required)* | signing key for the `bor_session` cookie; `python -c 'import secrets;print(secrets.token_hex(32))'` | diff --git a/alembic/versions/0005_kb_overview.py b/alembic/versions/0005_kb_overview.py new file mode 100644 index 0000000..8f7518c --- /dev/null +++ b/alembic/versions/0005_kb_overview.py @@ -0,0 +1,48 @@ +"""kb_overview: single-row, lite-generated knowledge-base outline + +Revision ID: 0005 +Revises: 0004 +Create Date: 2026-08-25 + +Phase 31 (kb-overview-prompt story, A13 — one additive, reversible table, +no other schema change): + +* ``kb_overview`` — a **single-row** table (``id INTEGER PK DEFAULT 1``) + holding a plain-text outline of the knowledge base's basic categories, + generated by the aipi ``lite`` model (phase 30's ``LLMClient.chat``) + whenever an import changes the KB. Chat turns only *read* the row (one + indexed PK lookup) and inject it into the system prompt as the + ```` section, so the agent knows roughly what the KB + contains before retrieval returns documents. ``content`` defaults to the + empty string — no row / empty row means the prompts are byte-identical + to the pre-phase text (phase 15 convention); ``updated_at`` is stamped + server-side on every upsert. +""" +from __future__ import annotations + +import sqlalchemy as sa + +from alembic import op + +revision = "0005" +down_revision = "0004" +branch_labels = None +depends_on = None + + +def upgrade() -> None: + op.create_table( + "kb_overview", + sa.Column("id", sa.Integer(), primary_key=True, server_default=sa.text("1")), + sa.Column("content", sa.Text(), nullable=False, server_default=sa.text("''")), + sa.Column( + "updated_at", + sa.DateTime(timezone=True), + nullable=False, + server_default=sa.func.now(), + ), + ) + + +def downgrade() -> None: + op.drop_table("kb_overview") diff --git a/app/api/chat.py b/app/api/chat.py index 42eeb4e..bd4eeaf 100644 --- a/app/api/chat.py +++ b/app/api/chat.py @@ -36,6 +36,16 @@ assembly is untouched. ``TurnPlan.summary_hits`` counts the hit chunks with ``is_summary`` whose parent document landed in the selected top-N context, and the per-turn log line records ``summary_hits=N`` after ``fts_hits`` (PLAN §9 line extension). + +KB overview (phase 31): the lite-generated outline of the knowledge +base (single ``kb_overview`` row) is read per turn (one indexed PK +lookup — no LLM call) and injected into **both** prompts as the +```` section, ordered ```` → +```` → ```` → mode body. With an empty row the +prompts stay byte-identical to the pre-phase text (phase 15 +convention); ``TurnPlan.kb_chars`` records the length of the stored +outline (0 when absent) and the per-turn log line records +``kb_chars=N`` after ``tuning=N`` (PLAN §9 line extension). """ from __future__ import annotations @@ -60,6 +70,7 @@ from app.rag.llm import ( LLMError, StreamPiece, # noqa: F401 (phase 17 typing: chat_stream yields StreamPiece) ) +from app.rag.overview import load_kb_overview from app.rag.prompts import build_deflect_prompt, build_high_prompt from app.rag.retriever import RetrievedChunk, retrieve, select_documents, weak_hit_titles from app.rag.suggestions import derive_suggestions @@ -107,12 +118,17 @@ class TurnPlan: #: Hit chunks with ``is_summary`` whose parent document made it into #: *docs* (phase 30; per-turn log line ``summary_hits=N``). summary_hits: int = 0 + #: Length of the stored KB overview injected as the + #: ```` section (phase 31; per-turn log line + #: ``kb_chars=N``). 0 when no non-empty row exists. + kb_chars: int = 0 def plan_turn( chunks: Sequence[RetrievedChunk], settings: Settings, notes: Sequence[str] | None = None, + kb_overview: str | None = None, ) -> TurnPlan: """Apply the honesty gate (A8, revised) and assemble prompt + context. @@ -133,11 +149,20 @@ def plan_turn( when non-empty, both the HIGH and the LOW prompt carry the ```` section; with no notes the prompts are unchanged. + *kb_overview* is the stored KB outline (phase 31, one PK lookup per + turn): when non-empty, both prompts carry the ```` + section (between ```` and ````) and + ``kb_chars`` records the outline's length; with no outline the + prompts are byte-identical to the pre-phase text and ``kb_chars`` + is 0. + ``summary_hits`` (phase 30) counts the hit chunks with ``is_summary`` whose parent document is among the selected top-N documents — both the HIGH and the LOW branch record it. """ steering = list(notes or []) + kb_text = (kb_overview or "").strip() + kb_chars = len(kb_text) best_cosine = max((c.cosine for c in chunks), default=0.0) fts_hits = sum(1 for c in chunks if c.fts_hit) docs = select_documents(chunks, n=settings.top_n_docs) @@ -148,22 +173,24 @@ def plan_turn( best_cosine, fts_hits, False, - build_high_prompt(docs, notes=steering), + build_high_prompt(docs, notes=steering, kb_overview=kb_text), docs, [], len(steering), summary_hits, + kb_chars, ) titles = weak_hit_titles(chunks) return TurnPlan( best_cosine, fts_hits, True, - build_deflect_prompt(titles, notes=steering), + build_deflect_prompt(titles, notes=steering, kb_overview=kb_text), docs, derive_suggestions(titles, settings.suggestions), len(steering), summary_hits, + kb_chars, ) @@ -216,8 +243,11 @@ async def chat( settings = get_settings() try: steering_notes = load_steering_notes(db) + # KB overview (phase 31): one indexed PK lookup per turn — the + # outline is generated at import time, never per chat turn. + kb_overview = load_kb_overview(db) chunks = retrieve(db, request.message, question_vec) - plan = plan_turn(chunks, settings, notes=steering_notes) + plan = plan_turn(chunks, settings, notes=steering_notes, kb_overview=kb_overview) except Exception: # noqa: BLE001 — DB failure mid-turn logger.exception( "chat: retrieval failed question=%r total_ms=%d", @@ -284,13 +314,14 @@ async def chat( logger.info( "question=%r embed_ms=%d top_score=%.3f fts_hits=%d summary_hits=%d tuning=%d " - "threshold=%.2f deflected=%s sources=%r thinking_chars=%d total_ms=%d", + "kb_chars=%d threshold=%.2f deflected=%s sources=%r thinking_chars=%d total_ms=%d", request.message, embed_ms, plan.top_score, plan.fts_hits, plan.summary_hits, plan.tuning_count, + plan.kb_chars, settings.relevance_threshold, plan.deflected, source_paths, diff --git a/app/config.py b/app/config.py index e16c62b..73d8795 100644 --- a/app/config.py +++ b/app/config.py @@ -76,6 +76,17 @@ class Settings(BaseSettings): #: and the shared ``[…truncated…]`` marker is appended (see #: ``app.rag.summarizer``). summary_max_chars: int = 12_000 + #: Char budget for the ```` section of the system prompt + #: (phase 31: lite-generated KB overview, ``app.rag.overview``). The + #: newest-fitting prefix of the stored outline is kept and the overflow + #: is replaced by the shared ``[…truncated…]`` marker (phase 15 + #: convention — ``app.rag.prompts``). + kb_overview_max_chars: int = 4_000 + #: Cap on the document list (source/path/title/first summary line per + #: row) sent to the ``lite`` overview model in one call (phase 31, + #: ``app.rag.overview``). Overflow is cut at the cap and the shared + #: ``[…truncated…]`` marker is appended (summarizer convention). + overview_input_max_chars: int = 40_000 # --- Hybrid retrieval (A7, revised 2026-08-21) --- # cosine top-N ∪ Postgres FTS top-N, fused with Reciprocal Rank Fusion diff --git a/app/models.py b/app/models.py index 7b14601..1e27adc 100644 --- a/app/models.py +++ b/app/models.py @@ -10,6 +10,9 @@ Data model — see ``.agent/PLAN.md`` §Data Model: the deflection decision, and latency. * ``steering_notes`` — owner tuning notes injected into the system prompt of every chat turn (phase 15, ```` section). +* ``kb_overview`` — single-row lite-generated outline of the KB's basic + categories, injected as the ```` + section of every chat turn (phase 31). """ from __future__ import annotations @@ -109,3 +112,24 @@ class SteeringNote(Base): id: Mapped[uuid.UUID] = mapped_column(UUID(as_uuid=True), primary_key=True, default=uuid.uuid4) note: Mapped[str] = mapped_column(Text) # trimmed, 1–2000 chars (API-enforced) created_at: Mapped[datetime] = mapped_column(DateTime(timezone=True), server_default=func.now()) + + +class KbOverview(Base): + """Single-row, lite-generated outline of the knowledge base (phase 31). + + Exactly one row (``id = 1``, enforced by the migration 0005 server + defaults) holds a plain-text outline of the KB's basic categories, + generated by the aipi ``lite`` model whenever an import changes the KB. + Chat turns only read this row (one indexed PK lookup) and inject it into + the system prompt of every turn as the ```` section — an + empty row means the section is absent and the prompt stays + byte-identical to the pre-phase text (phase 15 convention). + """ + + __tablename__ = "kb_overview" + + id: Mapped[int] = mapped_column(Integer, primary_key=True, server_default="1") + content: Mapped[str] = mapped_column(Text, server_default="") # the outline text + updated_at: Mapped[datetime] = mapped_column( + DateTime(timezone=True), server_default=func.now() + ) diff --git a/app/rag/overview.py b/app/rag/overview.py new file mode 100644 index 0000000..cca8ba9 --- /dev/null +++ b/app/rag/overview.py @@ -0,0 +1,202 @@ +"""KB overview generator (phase 31, task 02). + +Builds the ``KB_OVERVIEW_MODE`` prompt from the document catalogue +(source, path, title, first summary line per document), calls the aipi +``lite`` model through the one-shot ``LLMClient.chat`` (phase 30, same +model as document summaries — no new model management), and stores the +plain-text outline in the single ``kb_overview`` row (migration 0005). + +Chat turns never generate an overview — they read the stored row +(:func:`load_kb_overview`, one indexed PK lookup) and inject it into the +system prompt of every turn as the ```` section +(phase 31, task 03). Generation happens at import time (phase 31, task +04) and is **best-effort and change-gated**: callers invoke it only when +an import added/updated at least one document (or no row exists yet), +and a ``lite`` failure logs and leaves the previous outline intact — an +old outline is better than none (phase locked decisions). + +The ``KB_OVERVIEW_MODE`` marker follows the ``SUMMARY_MODE`` / +``DEFLECT_MODE`` convention: the deterministic E2E mock LLM keys on it +in the system prompt (wired in phase 31, task 05). +""" +from __future__ import annotations + +import logging +from collections.abc import Sequence +from datetime import UTC, datetime +from typing import Protocol + +from sqlalchemy import select +from sqlalchemy.orm import Session + +from app.config import Settings, get_settings +from app.db import SessionLocal +from app.models import Document, KbOverview +from app.rag.llm import LLMError +from app.rag.retriever import TRUNCATION_MARKER + +logger = logging.getLogger("app.rag.overview") + +#: System-prompt marker for KB overview generation — the E2E mock LLM keys +#: on it (same convention as ``SUMMARY_MODE`` / ``DEFLECT_MODE``). +KB_OVERVIEW_MODE = "KB_OVERVIEW_MODE" + +#: Locked instruction for the ``lite`` model (phase 31): the overview is +#: the agent's *a priori* picture of the knowledge base, so it must be a +#: compact plain-text outline of basic categories and topics — strictly +#: grounded in the document list, nothing invented. +OVERVIEW_INSTRUCTION = ( + "From the document list below, write a compact plain-text outline of " + "the basic categories and topics this knowledge base covers. Group by " + "source where useful, use `-` bullet lines, at most ~1500 characters, " + "no markdown headings, and no topics not present in the list." +) + +#: Full system prompt: marker first (the mock's key), then the instruction. +SYSTEM_PROMPT = f"{KB_OVERVIEW_MODE}: {OVERVIEW_INSTRUCTION}" + + +class OverviewLLM(Protocol): + """The one-shot chat surface the overview generator needs. + + :class:`app.rag.llm.LLMClient` satisfies it; unit tests pass a + duck-typed fake (``chat`` + ``settings``) instead — same pattern as + the summarizer's ``SummaryLLM`` protocol. + """ + + settings: Settings + + async def chat( + self, messages: list[dict[str, str]], model: str | None = None + ) -> str: ... + + +def _first_summary_line(summary: str | None) -> str: + """First line of a stored summary, stripped; ``''`` when absent. + + The stored summary (phase 30) ends in the code-appended + ``Source: …`` pointer line, so its first line is the model-written + lead sentence — the best one-line picture of the document for the + outline. Blank/whitespace-only summaries yield ``''`` as well. + """ + if not summary: + return "" + for line in summary.splitlines(): + stripped = line.strip() + if stripped: + return stripped + return "" + + +def build_overview_prompt( + rows: Sequence[tuple[str, str, str, str | None]], max_chars: int | None = None +) -> tuple[str, str]: + """The ``(system, user)`` message pair for one overview call. + + Each row is ``(source, path, title, summary)``. + + * ``system`` — :data:`SYSTEM_PROMPT`: the ``KB_OVERVIEW_MODE`` marker + + the locked instruction. + * ``user`` — one line per document, + ``source — path — title — {first line of summary}``, joined with + newlines. The summary field is omitted when the document has no + summary (markdown docs and the fail-soft path — no dangling + dash). The list is capped at *max_chars* (default + ``BOR_OVERVIEW_INPUT_MAX_CHARS``): overflow is cut exactly at the + cap and the shared ``[…truncated…]`` marker is appended on its own + line, so the model never sees more than the cap and the cut is + visible (summarizer convention, phase 30). + """ + limit = max_chars if max_chars is not None else get_settings().overview_input_max_chars + lines: list[str] = [] + for source, path, title, summary in rows: + line = f"{source} — {path} — {title}" + first = _first_summary_line(summary) + if first: + line += f" — {first}" + lines.append(line) + user = "\n".join(lines) + if len(user) > limit: + user = user[:limit] + "\n" + TRUNCATION_MARKER + return SYSTEM_PROMPT, user + + +def load_kb_overview(db: Session) -> str: + """The stored KB outline for prompt injection (phase 31, task 03). + + One indexed PK lookup (the single row, ``id = 1``) returning + ``content`` trimmed. Returns ``""`` when the row is missing or + empty — the chat path then omits the ```` section + entirely, keeping the prompt byte-identical to the pre-phase text + (phase 15 convention). + """ + row = db.get(KbOverview, 1) + if row is None: + return "" + return (row.content or "").strip() + + +async def regenerate_overview( + llm: OverviewLLM, session: Session | None = None +) -> bool: + """Generate a fresh KB outline with the ``lite`` model and upsert it. + + Best-effort and change-gated by the caller (phase locked decisions): + the import script (phase 31, task 04) and the admin sync (phase 32) + invoke this only when an import added/updated at least one document + (or no row exists yet). *session* may be supplied (the import + script's own session, tests); a private one is opened and closed + otherwise — same convention as ``import_sources``. + + * Zero documents → the existing row (if any) is left untouched and + ``False`` is returned — no KB, no outline, no wasted model call. + * Happy path → the model's text is upserted into the single row + (``id = 1``, fresh UTC ``updated_at``), committed, and + ``overview: regenerated docs=… chars=…`` is logged; returns + ``True``. Idempotent by construction: one row, always ``id = 1``. + * :class:`LLMError` → ``overview: regeneration failed — …`` is + logged and ``False`` is returned with the previous row intact — + an old outline is better than none. + """ + owns_session = session is None + if session is None: + session = SessionLocal() + try: + result = session.execute( + select(Document.source, Document.path, Document.title, Document.summary) + .order_by(Document.source, Document.path) + ).all() + rows: list[tuple[str, str, str, str | None]] = [ + (source, path, title, summary) + for source, path, title, summary in result + ] + if not rows: + logger.info( + "overview: no documents in the knowledge base — " + "leaving any existing outline untouched" + ) + return False + + system, user = build_overview_prompt(rows) + try: + text = await llm.chat( + [{"role": "system", "content": system}, {"role": "user", "content": user}], + model=llm.settings.llm_summary_model, + ) + except LLMError as e: + logger.error("overview: regeneration failed — %s", e) + return False + + now = datetime.now(UTC) + row = session.get(KbOverview, 1) + if row is None: + session.add(KbOverview(id=1, content=text, updated_at=now)) + else: + row.content = text + row.updated_at = now + session.commit() + logger.info("overview: regenerated docs=%d chars=%d", len(rows), len(text)) + return True + finally: + if owns_session: + session.close() diff --git a/app/rag/prompts.py b/app/rag/prompts.py index b7efbed..4899c55 100644 --- a/app/rag/prompts.py +++ b/app/rag/prompts.py @@ -15,6 +15,14 @@ Steering (phase 15): when the owner has stored tuning notes, both modes carry a ```` section between ``…`` and the mode body. With zero notes the prompt is byte-identical to the pre-steering text. + +KB overview (phase 31): when the single ``kb_overview`` row holds a +lite-generated outline of the knowledge base, both modes carry a +```` section between ``…`` and +the ```` section (order: ```` → +```` → ```` → mode body) — the agent knows +roughly what the KB contains before retrieval. With an empty row the +prompt is byte-identical to the pre-phase text. """ from __future__ import annotations @@ -51,6 +59,13 @@ _STEERING_INTRO = ( "Where these instructions conflict with the defaults above, follow the owner:\n" ) +#: One-line intro of the ```` section (phase 31): the +#: lite-generated outline is the agent's a-priori picture of the KB. +_KB_INTRO = ( + "The basic categories of everything in this knowledge base " + "(generated at import time):\n" +) + def _base(relevance: str) -> str: if relevance not in ("HIGH", "LOW"): @@ -95,9 +110,57 @@ def build_steering_section(notes: Sequence[str], max_chars: int | None = None) - return "" -def build_high_prompt(documents: Sequence[Document], notes: Sequence[str] | None = None) -> str: - """Grounded turn: locked persona (+ steering) + full texts of the top - documents.""" +def build_kb_section(overview: str, max_chars: int | None = None) -> str: + """The ```` section of the system prompt (phase 31). + + * No outline (or only whitespace) → ``""`` — callers then build the + prompt exactly as before, so a no-overview prompt is byte-identical + to the pre-phase text (phase 15 convention). + * Otherwise: the intro line + the stored outline, capped at + *max_chars* (default ``BOR_KB_OVERVIEW_MAX_CHARS``). When the budget + cannot hold the whole outline, the longest-fitting prefix is kept + and the overflow is replaced by the shared ``[…truncated…]`` marker + on its own line — the exact :func:`build_steering_section` pattern, + including its pathological-budget handling (never exceed the cap; + bare marker when even one outline character does not fit). + """ + text = str(overview or "").strip() + if not text: + return "" + limit = max_chars if max_chars is not None else get_settings().kb_overview_max_chars + if limit <= 0: + return "" + + def render(cut: int) -> str: + lines = [text[:cut]] + if cut < len(text): + lines.append(TRUNCATION_MARKER) + return f"\n{_KB_INTRO}" + "\n".join(lines) + "\n" + + for cut in range(len(text), 0, -1): + rendered = render(cut) + if len(rendered) <= limit: + return rendered + # Pathological budget: not even one outline character fits. The + # section must still respect the cap — the bare marker when it fits, + # else none (steering precedent, phase 15). + if len(TRUNCATION_MARKER) <= limit: + return TRUNCATION_MARKER + return "" + + +def build_high_prompt( + documents: Sequence[Document], + notes: Sequence[str] | None = None, + kb_overview: str | None = None, +) -> str: + """Grounded turn: locked persona (+ steering, + KB overview) + full +texts of the top documents. + + Section order (phase 31): ```` → ```` → + ```` → ````; empty steering/overview omit their + section, keeping the prompt byte-identical to the pre-phase text. + """ blocks = [ f'\n' f"{doc.content}\n" @@ -107,21 +170,34 @@ def build_high_prompt(documents: Sequence[Document], notes: Sequence[str] | None body = "\n\n".join(blocks) if blocks else ( "(no documents matched — do not invent specifics)" ) - section = build_steering_section(notes or []) prompt = _base("HIGH") - if section: - prompt += "\n" + section + for part in (build_kb_section(kb_overview or ""), build_steering_section(notes or [])): + if part: + prompt += "\n" + part return prompt + "\n\n" + body + "\n" -def build_deflect_prompt(titles: Sequence[str], notes: Sequence[str] | None = None) -> str: - """Deflection turn: weak-hit titles only (no document content).""" +def build_deflect_prompt( + titles: Sequence[str], + notes: Sequence[str] | None = None, + kb_overview: str | None = None, +) -> str: + """Deflection turn: weak-hit titles only (no document content). + + Section order (phase 31): ```` → ```` → + ```` → ``DEFLECT_MODE`` body; empty steering/overview omit + their section, keeping the prompt byte-identical to the pre-phase text. + """ weak = "\n".join(f"- {t}" for t in titles) if titles else "(nothing close at all)" - section = build_steering_section(notes or []) - mid = f"\n{section}\n" if section else "\n" + mid = "\n".join( + part + for part in (build_kb_section(kb_overview or ""), build_steering_section(notes or [])) + if part + ) + gap = f"\n{mid}\n" if mid else "\n" return ( _base("LOW") - + mid + + gap + "DEFLECT_MODE: retrieval was weak — the titles below are the closest " "your notes come to the question. They are titles only; do not pretend " "they answer it. Use them to propose 2-3 alternative questions.\n" diff --git a/scripts/import_docs.py b/scripts/import_docs.py index ad6302f..762cdf1 100644 --- a/scripts/import_docs.py +++ b/scripts/import_docs.py @@ -27,6 +27,17 @@ caches) is skipped, along with non-content dirs (``.venv``, ``build``). Re-runs are cheap: files are diffed by sha256 and unchanged ones are not re-embedded; ``--prune`` also drops documents whose files no longer match the format filter. + +After a run that **changed** the knowledge base (at least one document +added or updated — or no outline stored yet), the single ``kb_overview`` +row is regenerated with the ``lite`` model (phase 31): the plain-text +outline of the KB's basic categories that every chat turn injects into +the system prompt as ````. The regeneration is +**best-effort and change-gated** — unchanged re-imports and ``--limit`` +debug runs never burn a ``lite`` call, and a ``lite`` failure only +reports ``overview=failed`` on the summary line: the import's exit code +is about files, and the previous outline stays (an old outline is better +than none). """ from __future__ import annotations @@ -40,8 +51,11 @@ from pathlib import Path from app.config import Settings, get_settings from app.core.debugging import configure_debugging from app.core.logging import configure_logging -from app.rag.importer import import_sources +from app.db import SessionLocal +from app.models import KbOverview +from app.rag.importer import ImportSummary, import_sources from app.rag.llm import LLMClient +from app.rag.overview import regenerate_overview from scripts.git_sync import GitSyncError, clone_or_pull logger = logging.getLogger("scripts.import_docs") @@ -116,6 +130,18 @@ def _resolve_sources(cli_sources: list[Path] | None, settings: Settings) -> list return [path.expanduser() for path in DEFAULT_SOURCES] +def _overview_row_exists() -> bool: + """Whether the single ``kb_overview`` row (id = 1) is present. + + One indexed PK lookup (phase 31, task 04): a missing outline after an + unchanged re-import — e.g. the first run after migration 0005 — still + gets a fresh outline, while a present one is left untouched until the + KB actually changes. + """ + with SessionLocal() as session: + return session.get(KbOverview, 1) is not None + + def main(argv: list[str] | None = None) -> int: args = build_parser().parse_args(argv) settings = get_settings() @@ -144,16 +170,49 @@ def main(argv: list[str] | None = None) -> int: return 1 llm = LLMClient() - summary = asyncio.run(import_sources(sources, llm, prune=args.prune, limit=args.limit)) + + async def _run() -> tuple[ImportSummary, str]: + """Import, then (change-gated) refresh the stored KB overview. + + One event loop, one ``LLMClient`` (phase 31, task 04): the + outline that every chat prompt injects as ```` is + regenerated only when this run added/updated at least one + document — or when no outline exists yet after a walk that + actually saw files (e.g. the first run after migration 0005). + ``--limit`` debug runs (an incomplete walk must not rewrite the + outline — mirrors the ``--prune``-with-``--limit`` guard) and + unchanged re-imports never burn a ``lite`` call, and a ``lite`` + failure only flips the status token (``failed``) — the import's + exit code is unchanged. + """ + summary = await import_sources(sources, llm, prune=args.prune, limit=args.limit) + if args.limit is not None: + logger.info("overview: skipped (--limit)") + return summary, "skipped" + if summary.added + summary.updated == 0: + if summary.files == 0: + logger.info("overview: skipped (nothing imported)") + return summary, "skipped" + if _overview_row_exists(): + logger.info("overview: skipped (KB unchanged)") + return summary, "skipped" + # No outline yet after an unchanged re-import (e.g. the first + # run after migration 0005) — fall through and generate one. + ok = await regenerate_overview(llm) + return summary, "updated" if ok else "failed" + + summary, overview_status = asyncio.run(_run()) print( f"import_docs: files={summary.files} added={summary.added} " f"updated={summary.updated} unchanged={summary.unchanged} " f"pruned={summary.pruned} errors={summary.errors} chunks={summary.chunks} " f"embed_batches={summary.embed_batches} summaries={summary.summaries} " - f"summary_errors={summary.summary_errors} formats={summary.format_counts()}" + f"summary_errors={summary.summary_errors} formats={summary.format_counts()} " + f"overview={overview_status}" ) # Non-zero if any file failed, so cron/CI notice — the rest of the KB - # was imported and the failed files are retried on the next run. + # was imported and the failed files are retried on the next run. The + # overview is best-effort: a failed outline never changes the exit code. return 1 if summary.errors else 0 diff --git a/tests/e2e/mock_llm.py b/tests/e2e/mock_llm.py index 3dddb5f..d73c333 100644 --- a/tests/e2e/mock_llm.py +++ b/tests/e2e/mock_llm.py @@ -15,6 +15,9 @@ Implements just enough of the aipi surface: tokens of the user message (the summarizer puts the capped document content there) — byte-stable for a given fixture (document summaries, phase 30) + - ``KB_OVERVIEW_MODE`` -> the deterministic outline: the first 8 tokens + of the user message (the generator puts the document list there) — + byte-stable for a given KB (KB overview, phase 31) - ``DEFLECT_MODE`` -> honest "I haven't done anything like that" answer - otherwise -> upbeat answer quoting the provided document context - user message containing ``pretend to think slowly`` -> 3s warm-up delay @@ -30,6 +33,9 @@ Implements just enough of the aipi surface: - system prompt containing ```` (phase 15, steering notes) -> the composed answer ends with `` (tuning: )`` — makes prompt injection observable in the UI deterministically. + - system prompt containing ```` (phase 31, KB overview) + -> the composed answer ends with `` (kb: )`` — + the same echo convention for the overview's prompt injection. - user message containing ``show the end of your notes`` (phase 24, whole-document context) -> the answer quotes the **last 160 chars of the document context** — a tail echo, byte-stable across runs, so a @@ -155,6 +161,30 @@ def first_tuning_note(system: str) -> str | None: return None +#: First ``-`` bullet line of a ```` section (phase 31). +_KB_BLOCK_RE = re.compile(r"\n(.*?)\n", re.S) +_KB_BULLET_RE = re.compile(r"^-(?:\s+(.*))?$") + + +def first_kb_bullet(system: str) -> str | None: + """The first outline bullet in the system prompt, or ``None``. + + The stored outline (phase 31) is ``-`` bullet lines (see + ``app.rag.overview.OVERVIEW_INSTRUCTION``); the mock echoes the first + one into its answer as `` (kb: )`` — the exact + :func:`first_tuning_note` convention, so prompt injection of the + ```` section is observable in the UI. + """ + block = _KB_BLOCK_RE.search(system) + if not block: + return None + for line in block.group(1).splitlines(): + m = _KB_BULLET_RE.match(line.strip()) + if m: + return (m.group(1) or "").strip() + return None + + def compose_answer(body: dict[str, Any]) -> str: system = _system(body) user = _user(body) @@ -172,6 +202,17 @@ def compose_answer(body: dict[str, Any]) -> str: f"This document covers " f"{' '.join(TOKEN_RE.findall(user.lower())[:24])}." ) + elif "KB_OVERVIEW_MODE" in system: + # KB overview (phase 31): the ``lite`` stand-in returns the + # deterministic outline — the first 8 tokens of the user message + # (the generator puts the document list there). Byte-stable for a + # given KB, so the stored row is a pure function of the fixture. + # Checked BEFORE the DEFLECT_MODE branch, like SUMMARY_MODE, so a + # prompt that ever carries both markers cannot shadow the + # overview call. + answer = "Knowledge base outline:\n- " + " ".join( + TOKEN_RE.findall(user.lower())[:8] + ) elif "DEFLECT_MODE" in system: answer = ( "Ah — I haven't done anything like that, so I don't want to make stuff up! " @@ -202,6 +243,13 @@ def compose_answer(body: dict[str, Any]) -> str: note = first_tuning_note(system) if note: answer = f"{answer} (tuning: {note})" + # KB overview (phase 31): when the system prompt carries + # , the answer ends with the first outline bullet — + # mirrors the steering echo exactly (appended after it, so the kb + # suffix is the last thing rendered). + bullet = first_kb_bullet(system) + if bullet: + answer = f"{answer} (kb: {bullet})" return answer diff --git a/tests/e2e/test_kb_overview.py b/tests/e2e/test_kb_overview.py new file mode 100644 index 0000000..37e5c0c --- /dev/null +++ b/tests/e2e/test_kb_overview.py @@ -0,0 +1,281 @@ +"""Phase 31 E2E (Playwright): the stored KB overview reaches every turn. + +Story: ``.agent/user_stories/kb-overview-prompt.md`` +Run in isolation (DB must be up: ``podman compose up -d db``): + + uv run pytest tests/e2e/test_kb_overview.py -v --no-cov + +The story: a single-row ``kb_overview`` table holds a lite-generated +plain-text outline of the KB's basic categories; every chat turn — HIGH +(grounded) and LOW (deflected) alike — injects it into the system prompt +as the ```` section, so the agent knows roughly what the +KB contains before retrieval returns. The mock LLM echoes the section's +first bullet into its answer (`` (kb: )`` — the +``(tuning: …)`` steering-echo precedent), which makes the prompt change +observable in the rendered UI deterministically. + +The row is seeded **directly in the DB** so the INJECTION path is what +is under test (the import-time trigger is integration-tested in +``tests/integration/test_import_docs_overview.py``). The outline is +multi-line so the echo must pick the FIRST bullet — not the intro line, +not a later line. The last test additionally runs the real +``regenerate_overview`` against the mock to cover the mock's +``KB_OVERVIEW_MODE`` generation branch end-to-end (stored outline is the +byte-stable 8-token digest of the generator's document list). + +Test → story mapping (Playwright Mapping Rule): +1. ``test_on_topic_answer_echoes_kb_overview`` +2. ``test_deflected_answer_still_echoes_kb_overview`` +3. ``test_absent_row_yields_no_kb_echo`` +4. ``test_mock_generated_outline_is_stored_and_echoed`` +""" +from __future__ import annotations + +import asyncio +import re +from pathlib import Path +from threading import Thread +from typing import Any + +from playwright.sync_api import Page, expect +from sqlalchemy import select, text + +from app.config import Settings +from app.db import SessionLocal +from app.models import Document, KbOverview, QueryLog +from app.rag.importer import ImportSummary, import_sources +from app.rag.llm import LLMClient +from app.rag.overview import build_overview_prompt, regenerate_overview +from tests.e2e.mock_llm import TOKEN_RE + +REPO = Path(__file__).resolve().parents[2] +FIXTURES = REPO / "tests" / "fixtures" / "docs" +QUESTION = "How is my Kubernetes cluster set up?" +OFF_TOPIC = "How do I bake sourdough bread?" +MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E" + +#: The stored outline: multi-line ``-`` bullets (the generator's locked +#: format). The echo must pick the FIRST bullet — the parsing (skip the +#: intro line, strip the dash) is part of the story. +OVERVIEW = ( + "- Kubernetes cluster and node maintenance notes\n" + "- Backup schedules and restore runbooks\n" + "- Networking: static DNS and kafkabridge" +) +FIRST_BULLET = "Kubernetes cluster and node maintenance notes" +KB_ECHO = f"(kb: {FIRST_BULLET})" + + +# --- Importer + thread helpers (test_document_summaries.py pattern) --- + + +async def _import_fixtures(mock_port: int) -> ImportSummary: + kwargs: dict[str, Any] = { + "_env_file": None, + "llm_base_url": f"http://127.0.0.1:{mock_port}/v1", + } + settings = Settings(**kwargs) # pyright: ignore[reportCallIssue] + return await import_sources([FIXTURES], LLMClient(settings)) + + +def _run_in_thread(coro: Any) -> Any: + """Run a coroutine on a worker thread. + + Playwright's sync API keeps an asyncio loop running on the test thread, + so ``asyncio.run`` cannot be called directly from a test body. + """ + box: dict[str, Any] = {} + + def runner() -> None: + try: + box["value"] = asyncio.run(coro) + except BaseException as e: # noqa: BLE001 — re-raised on the test thread + box["error"] = e + + t = Thread(target=runner) + t.start() + t.join() + if "error" in box: + raise box["error"] + return box["value"] + + +def _reset_db( + mock_port: int, seed: bool, overview: str | None = None +) -> ImportSummary | None: + """Truncate the KB (+ query log, steering, kb_overview), re-import the + fixtures, and optionally seed the single ``kb_overview`` row.""" + with SessionLocal() as db: + db.execute( + text("TRUNCATE chunks, documents, query_log, steering_notes, kb_overview") + ) + db.commit() + if overview is not None: + db.add(KbOverview(id=1, content=overview)) + db.commit() + if not seed: + return None + return _run_in_thread(_import_fixtures(mock_port)) + + +def _ask(page: Page, app_url: str, question: str, done_text: str) -> Any: + """Submit *question* and wait until the streamed answer has fully + landed. + + ``done_text`` is the deterministic last thing the answer contains for + its path: the ``(kb: …)`` echo (grounded AND deflected turns with a + stored row — it is appended after everything else) or the mock + marker (the no-row control, where the echo is absent by design). + """ + page.set_default_timeout(30_000) + page.goto(app_url) + page.fill("#message-input", question) + page.click("#send-btn") + bubble = page.locator(".msg.brain .bubble") + bubble.first.wait_for(state="visible", timeout=30_000) + expect(bubble.first).to_contain_text(done_text, timeout=30_000) + expect(page.locator("#send-btn")).to_be_enabled() + expect(page.locator("#send-label")).to_have_text("Send") + return bubble.first + + +def _last_query_log() -> QueryLog: + with SessionLocal() as db: + rows = db.scalars(select(QueryLog)).all() + assert len(rows) == 1, f"expected exactly one query_log row, got {len(rows)}" + return rows[0] + + +def _overview_row() -> KbOverview | None: + with SessionLocal() as db: + return db.get(KbOverview, 1) + + +def _doc_rows() -> list[tuple[str, str, str, str | None]]: + with SessionLocal() as db: + rows = db.execute( + select(Document.source, Document.path, Document.title, Document.summary) + .order_by(Document.source, Document.path) + ).all() + return [(source, path, title, summary) for source, path, title, summary in rows] + + +# --- 1. Injection: grounded turn ------------------------------------------ + + +def test_on_topic_answer_echoes_kb_overview( + page: Page, app_url: str, mock_llm: int, db_ready: None +) -> None: + """On-topic question with a stored row: the rendered answer ENDS with + the mock's echo of the ```` section's first bullet — + only possible if the section reached the LLM prompt.""" + summary = _reset_db(mock_llm, seed=True, overview=OVERVIEW) + assert summary is not None and summary.added == 8 # A9 formats + assert _overview_row() is not None # the row the turn must inject + + bubble = _ask(page, app_url, QUESTION, KB_ECHO) + + # The echo is the LAST thing in the bubble (the section was in the + # HIGH prompt) and names the FIRST outline bullet (not the intro, + # not a later line). + expect(bubble).to_have_text( + re.compile(rf"{re.escape(KB_ECHO)}\s*$"), timeout=30_000 + ) + # Grounded, not deflected — the overview does not change the gate. + expect(page.locator(".msg.brain.is-deflected")).to_have_count(0) + row = _last_query_log() + assert row.question == QUESTION + assert row.deflected is False + + +# --- 2. Injection: deflected turn ------------------------------------------ + + +def test_deflected_answer_still_echoes_kb_overview( + page: Page, app_url: str, mock_llm: int, db_ready: None +) -> None: + """Off-topic question (LOW path): the deflected answer STILL carries + the echo — the ```` section is in the LOW prompt too + (that is what lets the model offer real alternatives).""" + _reset_db(mock_llm, seed=True, overview=OVERVIEW) + + # The deflection answer carries no mock marker — the kb echo is its + # deterministic end, so it doubles as the completion signal. + bubble = _ask(page, app_url, OFF_TOPIC, KB_ECHO) + + expect(page.locator(".msg.brain.is-deflected")).to_have_count(1) + expect(bubble).to_have_text(re.compile(r"haven't done anything like that")) + expect(bubble).to_have_text( + re.compile(rf"{re.escape(KB_ECHO)}\s*$"), timeout=30_000 + ) + row = _last_query_log() + assert row.question == OFF_TOPIC + assert row.deflected is True + + +# --- 3. Absence: no row → no echo ------------------------------------------ + + +def test_absent_row_yields_no_kb_echo( + page: Page, app_url: str, mock_llm: int, db_ready: None +) -> None: + """Delete the row: a fresh question's answer carries NO ``(kb: …)`` + suffix — the section is absent and the prompt behaves exactly like + the pre-phase text (byte-identity is unit-asserted; this proves it + end-to-end through the rendered answer).""" + summary = _reset_db(mock_llm, seed=True, overview=OVERVIEW) + assert summary is not None + # Now delete the row (the absence control). + with SessionLocal() as db: + db.execute(text("TRUNCATE kb_overview")) + db.commit() + assert _overview_row() is None + + bubble = _ask(page, app_url, QUESTION, MOCK_ANSWER_MARKER) + + expect(bubble).to_contain_text(MOCK_ANSWER_MARKER) + expect(bubble).not_to_contain_text("(kb:") + row = _last_query_log() + assert row.deflected is False # grounded either way — only the echo is gone + + +# --- 4. Generation: KB_OVERVIEW_MODE branch through the real regenerator ---- + + +def test_mock_generated_outline_is_stored_and_echoed( + page: Page, app_url: str, mock_llm: int, db_ready: None +) -> None: + """The mock's ``KB_OVERVIEW_MODE`` branch is exercised end-to-end: + the real ``regenerate_overview`` (``LLMClient`` pointed at the mock) + stores the byte-stable 8-token digest of the generator's document + list, and a chat turn echoes its first bullet.""" + summary = _reset_db(mock_llm, seed=True, overview=None) + assert summary is not None and summary.added == 8 + assert _overview_row() is None # direct import never regenerates + + kwargs: dict[str, Any] = { + "_env_file": None, + "llm_base_url": f"http://127.0.0.1:{mock_llm}/v1", + } + settings = Settings(**kwargs) # pyright: ignore[reportCallIssue] + ok = _run_in_thread(regenerate_overview(LLMClient(settings))) + assert ok is True + + row = _overview_row() + assert row is not None and row.id == 1 + # Byte-stable expectation: the mock digests the generator's user + # message (the same document list, the same cap) to its first 8 tokens. + _, user = build_overview_prompt(_doc_rows()) + expected = "Knowledge base outline:\n- " + " ".join( + TOKEN_RE.findall(user.lower())[:8] + ) + assert row.content == expected + first_bullet = row.content.splitlines()[1].removeprefix("- ").strip() + assert first_bullet # the digest line is non-empty + + # A grounded chat turn echoes the STORED outline's first bullet. + bubble = _ask(page, app_url, QUESTION, f"(kb: {first_bullet})") + expect(bubble).to_have_text( + re.compile(rf"\(kb: {re.escape(first_bullet)}\)\s*$"), timeout=30_000 + ) + expect(page.locator(".msg.brain.is-deflected")).to_have_count(0) diff --git a/tests/integration/test_import_docs_overview.py b/tests/integration/test_import_docs_overview.py new file mode 100644 index 0000000..13735a4 --- /dev/null +++ b/tests/integration/test_import_docs_overview.py @@ -0,0 +1,222 @@ +"""Integration test: ``import_docs`` regenerates the KB overview (phase 31, task 04). + +Drives ``scripts.import_docs.main()`` end to end against the local compose +Postgres with a deterministic fake LLM (no live aipi, no git — explicit +``--source`` dirs and fresh settings, the phase 28 test's mocking style), +covering the change-gated overview trigger: + +- a KB-changing import → ``kb_overview`` row written, the ``lite`` ``chat`` + called exactly once, summary line ends ``overview=updated``; +- an unchanged re-import → ``chat`` **not** called again, ``overview=skipped``; +- a ``lite`` failure → exit code still ``0`` (the import itself was fine), + ``overview=failed``, the previous row untouched; +- a ``--limit`` debug run with changes → ``overview=skipped``; +- an empty source run with no row → no row created, ``overview=skipped``. +""" +from __future__ import annotations + +import logging +from collections.abc import Iterator +from pathlib import Path + +import pytest +from sqlalchemy import text +from sqlalchemy.orm import Session + +from app.config import Settings +from app.models import KbOverview +from app.rag.llm import LLMError +from scripts import import_docs +from tests.fakes import FakeEmbedder + + +class FailingChatEmbedder(FakeEmbedder): + """A ``lite`` model that always fails (drives the fail-soft path).""" + + async def chat(self, messages: list[dict[str, str]], model: str | None = None) -> str: + self.chat_calls.append(list(messages)) + raise LLMError("simulated lite-model failure (test sentinel)") + + +def _row(db: Session) -> KbOverview | None: + """The stored ``kb_overview`` row (freshly reloaded).""" + db.expire_all() + return db.get(KbOverview, 1) + + +def _run_main( + monkeypatch: pytest.MonkeyPatch, + llm: FakeEmbedder, + argv: list[str], + capsys: pytest.CaptureFixture[str], +) -> tuple[int, str]: + """Run ``import_docs.main`` with fresh settings, a fake LLM, and a + fail-loud git mock (``--source`` always wins, so git must stay idle).""" + monkeypatch.setattr( + import_docs, + "get_settings", + lambda: Settings(_env_file=None), # pyright: ignore[reportCallIssue] + ) + + def _no_git(url: str, dest: Path | str) -> Path: + raise AssertionError("git sync must not run with explicit --source") + + monkeypatch.setattr(import_docs, "clone_or_pull", _no_git) + monkeypatch.setattr(import_docs, "LLMClient", lambda: llm) + rc = import_docs.main(argv) + return rc, capsys.readouterr().out + + +@pytest.fixture() +def src(tmp_path: Path) -> Path: + """A source dir with two markdown docs (md → no summary chat calls).""" + root = tmp_path / "MyDocs" + root.mkdir() + (root / "alpha.md").write_text("# Alpha\n\nFirst document.\n", encoding="utf-8") + (root / "beta.md").write_text("# Beta\n\nSecond document.\n", encoding="utf-8") + return root + + +@pytest.fixture(autouse=True) +def _clean_kb(db: Session) -> Iterator[None]: + db.execute(text("TRUNCATE chunks, documents, kb_overview")) + db.commit() + yield + db.execute(text("TRUNCATE chunks, documents, kb_overview")) + db.commit() + + +def test_changed_import_writes_overview_row( + db: Session, + src: Path, + monkeypatch: pytest.MonkeyPatch, + capsys: pytest.CaptureFixture[str], +) -> None: + llm = FakeEmbedder() + records: list[logging.LogRecord] = [] + + class _Sink(logging.Handler): + def emit(self, record: logging.LogRecord) -> None: + records.append(record) + + overview_logger = logging.getLogger("app.rag.overview") + sink = _Sink() + overview_logger.addHandler(sink) + overview_logger.setLevel(logging.INFO) + try: + rc, out = _run_main(monkeypatch, llm, ["--source", str(src)], capsys) + finally: + overview_logger.removeHandler(sink) + + assert rc == 0 + assert "added=2" in out + assert out.rstrip().endswith("overview=updated") + # Exactly one lite call — the overview itself (markdown files never + # get a summary, so nothing else may touch ``chat``). + assert len(llm.chat_calls) == 1 + by_role = {m["role"]: m["content"] for m in llm.chat_calls[0]} + assert "KB_OVERVIEW_MODE" in by_role["system"] + # One line per doc: source — path — title (no summary for markdown). + assert "MyDocs — alpha.md — Alpha" in by_role["user"] + assert "MyDocs — beta.md — Beta" in by_role["user"] + # The model's outline lands in the single row. + row = _row(db) + assert row is not None + assert row.content == "Summary of MyDocs" + assert row.updated_at is not None + # The phase's required log line (PLAN §9). + assert any("overview: regenerated docs=2 chars=" in r.getMessage() for r in records) + + +def test_unchanged_reimport_does_not_call_lite( + db: Session, + src: Path, + monkeypatch: pytest.MonkeyPatch, + capsys: pytest.CaptureFixture[str], +) -> None: + llm = FakeEmbedder() + rc, out = _run_main(monkeypatch, llm, ["--source", str(src)], capsys) + assert rc == 0 + assert out.rstrip().endswith("overview=updated") + assert len(llm.chat_calls) == 1 + assert _row(db) is not None + + # Same hashes → no KB change → no lite call, previous outline kept. + rc, out = _run_main(monkeypatch, llm, ["--source", str(src)], capsys) + assert rc == 0 + assert "unchanged=2" in out + assert out.rstrip().endswith("overview=skipped") + assert len(llm.chat_calls) == 1 # no new lite call + row = _row(db) + assert row is not None and row.content == "Summary of MyDocs" + + +def test_lite_failure_is_fail_soft( + db: Session, + src: Path, + monkeypatch: pytest.MonkeyPatch, + capsys: pytest.CaptureFixture[str], +) -> None: + good = FakeEmbedder() + rc, out = _run_main(monkeypatch, good, ["--source", str(src)], capsys) + assert rc == 0 + assert out.rstrip().endswith("overview=updated") + previous = _row(db) + assert previous is not None + previous_content = previous.content + + # A KB-changing run whose ``lite`` model fails: the import still + # succeeds (exit 0) and the previous outline stays untouched. + (src / "alpha.md").write_text("# Alpha\n\nChanged content.\n", encoding="utf-8") + bad = FailingChatEmbedder() + rc, out = _run_main(monkeypatch, bad, ["--source", str(src)], capsys) + assert rc == 0 # a failed outline must not fail the import + assert "updated=1" in out + assert out.rstrip().endswith("overview=failed") + assert len(bad.chat_calls) == 1 # the (failed) attempt was made + row = _row(db) + assert row is not None + assert row.content == previous_content # previous row untouched + + +def test_limit_run_skips_overview( + db: Session, + src: Path, + monkeypatch: pytest.MonkeyPatch, + capsys: pytest.CaptureFixture[str], +) -> None: + llm = FakeEmbedder() + rc, out = _run_main(monkeypatch, llm, ["--source", str(src)], capsys) + assert rc == 0 + assert out.rstrip().endswith("overview=updated") + assert len(llm.chat_calls) == 1 + + # An incomplete walk must not rewrite the outline (mirrors the + # --prune-with---limit guard). + (src / "alpha.md").write_text("# Alpha\n\nChanged content.\n", encoding="utf-8") + rc, out = _run_main(monkeypatch, llm, ["--source", str(src), "--limit", "1"], capsys) + assert rc == 0 + assert "updated=1" in out + assert out.rstrip().endswith("overview=skipped") + assert len(llm.chat_calls) == 1 # --limit never burns a lite call + row = _row(db) + assert row is not None and row.content == "Summary of MyDocs" + + +def test_empty_source_without_row_creates_nothing( + db: Session, + tmp_path: Path, + monkeypatch: pytest.MonkeyPatch, + capsys: pytest.CaptureFixture[str], +) -> None: + empty = tmp_path / "EmptyDocs" + empty.mkdir() + llm = FakeEmbedder() + + rc, out = _run_main(monkeypatch, llm, ["--source", str(empty)], capsys) + + assert rc == 0 + assert "files=0" in out + assert out.rstrip().endswith("overview=skipped") + assert llm.chat_calls == [] # no KB → no outline, no wasted model call + assert _row(db) is None # nothing created diff --git a/tests/integration/test_kb_overview_api.py b/tests/integration/test_kb_overview_api.py new file mode 100644 index 0000000..7bb79ea --- /dev/null +++ b/tests/integration/test_kb_overview_api.py @@ -0,0 +1,249 @@ +"""Integration: KB overview (phase 31) — the ```` section +of the chat system prompt. + +Real Postgres (``podman compose up -d db``) seeded from +``tests/fixtures/docs/`` through the real importer; the chat path reuses +the deterministic capturing fake LLM from ``test_chat_api`` +(token-overlap embeddings), so the stored row's journey — +``kb_overview`` row → per-turn PK lookup → ```` section +of the **exact** captured system prompt (HIGH and LOW) — is verified +end-to-end without a network. + +The byte-identity contract (phase 15 convention): with no row, the +captured system prompt equals the pre-phase construction +(``build_high_prompt`` / ``build_deflect_prompt`` with +``kb_overview=None``) — asserted with ``==``, not ``in``. + +Requires: podman compose up -d db +""" +from __future__ import annotations + +import asyncio +import logging +from collections.abc import Iterator +from pathlib import Path + +import pytest +from fastapi.testclient import TestClient +from sqlalchemy import select, text +from test_chat_api import FakeRagLLM, _stream_chat, _token_vec + +from app.api import chat as chat_api +from app.main import app as fastapi_app +from app.models import Document, KbOverview +from app.rag.importer import import_sources +from app.rag.prompts import build_deflect_prompt, build_high_prompt +from app.rag.retriever import retrieve, weak_hit_titles + +FIXTURES = Path(__file__).resolve().parents[1] / "fixtures" / "docs" +QUESTION = "How is my Kubernetes cluster set up?" +OFF_TOPIC = "How do I bake sourdough bread?" + +#: A multi-line, multi-bullet outline: the section must carry it whole +#: (well within ``BOR_KB_OVERVIEW_MAX_CHARS``) and the per-turn log line +#: records its length. +OVERVIEW = ( + "- Kubernetes cluster and node maintenance notes\n" + "- Backup schedules and restore runbooks\n" + "- Networking: static DNS and kafkabridge" +) + + +@pytest.fixture(autouse=True) +def clean_kb_overview(db) -> Iterator[None]: + """The outline row + query log are global state: reset around every + test so no test inherits another test's row.""" + db.execute(text("TRUNCATE kb_overview, query_log")) + db.commit() + yield + db.execute(text("TRUNCATE kb_overview, query_log")) + db.commit() + + +@pytest.fixture() +def seeded_kb(db) -> Iterator[FakeRagLLM]: + """Fresh Postgres with the fixture docs imported (real pipeline).""" + db.execute(text("TRUNCATE chunks, documents, query_log, kb_overview")) + db.commit() + llm = FakeRagLLM() + summary = asyncio.run(import_sources([FIXTURES], llm, session=db)) + assert summary.added == 8 # A9 formats; .hidden/ skipped + yield llm + db.execute(text("TRUNCATE chunks, documents, query_log, kb_overview")) + db.commit() + + +def _seed_overview(db) -> None: + db.add(KbOverview(id=1, content=OVERVIEW)) + db.commit() + + +def _cited_docs(db, frames: list[dict]) -> list[Document]: + """The documents the done event cited, in citation order — the same + list ``plan_turn`` passed to the prompt builder.""" + docs = [] + for s in frames[-1]["sources"]: + doc = db.scalar(select(Document).where(Document.path == s["path"])) + assert doc is not None, f"done source {s['path']!r} missing from the KB" + docs.append(doc) + return docs + + +def _turn_log_lines(caplog: pytest.LogCaptureFixture) -> list[str]: + """The per-turn ``question=…`` log lines (PLAN §9) from this test.""" + return [r.getMessage() for r in caplog.records if "question=" in r.getMessage()] + + +# ---------- no row → byte-identical to the pre-phase prompts ---------- + + +def test_no_row_high_prompt_byte_identical_to_pre_phase( + client: TestClient, db, seeded_kb: FakeRagLLM, caplog: pytest.LogCaptureFixture +) -> None: + """No ``kb_overview`` row: the captured HIGH system prompt EQUALS the + pre-phase construction exactly — the section is absent, not empty.""" + caplog.set_level(logging.INFO, logger="app.chat") + fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb + try: + _, _, frames = _stream_chat(client, QUESTION) + finally: + fastapi_app.dependency_overrides.clear() + + assert frames[-1]["deflected"] is False + (system, user) = seeded_kb.seen_messages[0][0], seeded_kb.seen_messages[0][1] + assert user["content"] == QUESTION + expected = build_high_prompt(_cited_docs(db, frames), notes=[], kb_overview=None) + assert system["content"] == expected + assert "" not in system["content"] + + lines = _turn_log_lines(caplog) + assert lines and "kb_chars=0" in lines[-1] + + +def test_no_row_low_prompt_byte_identical_to_pre_phase( + client: TestClient, db, seeded_kb: FakeRagLLM, caplog: pytest.LogCaptureFixture +) -> None: + """No row, off-topic question: the captured LOW (deflection) prompt + EQUALS the pre-phase construction exactly.""" + caplog.set_level(logging.INFO, logger="app.chat") + fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb + try: + _, _, frames = _stream_chat(client, OFF_TOPIC) + finally: + fastapi_app.dependency_overrides.clear() + + assert frames[-1]["deflected"] is True + (system, user) = seeded_kb.seen_messages[0][0], seeded_kb.seen_messages[0][1] + assert user["content"] == OFF_TOPIC + # Reconstruct the LOW prompt the way plan_turn does — the pre-phase + # construction (kb_overview=None), the same deterministic retrieval. + chunks = retrieve(db, OFF_TOPIC, _token_vec(OFF_TOPIC)) + expected = build_deflect_prompt( + weak_hit_titles(chunks), notes=[], kb_overview=None + ) + assert system["content"] == expected + assert "" not in system["content"] + assert "DEFLECT_MODE" in system["content"] + + lines = _turn_log_lines(caplog) + assert lines and "kb_chars=0" in lines[-1] + + +# ---------- row present → section in BOTH prompts, exactly ---------- + + +def test_row_high_prompt_carries_kb_section_exactly( + client: TestClient, db, seeded_kb: FakeRagLLM, caplog: pytest.LogCaptureFixture +) -> None: + """Stored row: the captured HIGH prompt EQUALS the construction with + the outline — section present, ordered before ````.""" + _seed_overview(db) + caplog.set_level(logging.INFO, logger="app.chat") + fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb + try: + _, _, frames = _stream_chat(client, QUESTION) + finally: + fastapi_app.dependency_overrides.clear() + + assert frames[-1]["deflected"] is False + (system, _user) = seeded_kb.seen_messages[0][0], seeded_kb.seen_messages[0][1] + expected = build_high_prompt( + _cited_docs(db, frames), notes=[], kb_overview=OVERVIEW + ) + assert system["content"] == expected + + # Section shape + order: → → . + prompt = system["content"] + assert ( + prompt.index("HIGH") + < prompt.index("") + < prompt.index(OVERVIEW) + < prompt.index("") + < prompt.index("") + ) + + # The per-turn log line records the outline's length (PLAN §9). + lines = _turn_log_lines(caplog) + assert lines and f"kb_chars={len(OVERVIEW)}" in lines[-1] + + +def test_row_low_prompt_carries_kb_section_exactly( + client: TestClient, db, seeded_kb: FakeRagLLM, caplog: pytest.LogCaptureFixture +) -> None: + """Stored row, off-topic question: the LOW prompt EQUALS the + construction with the outline — the section is in the deflection + prompt too (real alternatives, not hallucinated ones).""" + _seed_overview(db) + caplog.set_level(logging.INFO, logger="app.chat") + fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb + try: + _, _, frames = _stream_chat(client, OFF_TOPIC) + finally: + fastapi_app.dependency_overrides.clear() + + assert frames[-1]["deflected"] is True + (system, _user) = seeded_kb.seen_messages[0][0], seeded_kb.seen_messages[0][1] + chunks = retrieve(db, OFF_TOPIC, _token_vec(OFF_TOPIC)) + expected = build_deflect_prompt( + weak_hit_titles(chunks), notes=[], kb_overview=OVERVIEW + ) + assert system["content"] == expected + + prompt = system["content"] + assert ( + prompt.index("LOW") + < prompt.index("") + < prompt.index(OVERVIEW) + < prompt.index("") + < prompt.index("DEFLECT_MODE") + ) + # Deflection still sees titles only — never document content. + assert "Talos Linux" not in prompt + + lines = _turn_log_lines(caplog) + assert lines and f"kb_chars={len(OVERVIEW)}" in lines[-1] + + +def test_row_reread_every_turn_and_deleted_row_stops_it( + client: TestClient, db, seeded_kb: FakeRagLLM +) -> None: + """The row is read per turn (not cached): it steers every turn until + it is deleted, and the following turn is section-free again.""" + _seed_overview(db) + fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb + try: + _stream_chat(client, QUESTION) + _stream_chat(client, QUESTION) + assert len(seeded_kb.seen_messages) == 2 + for messages in seeded_kb.seen_messages: + assert "" in messages[0]["content"] + assert OVERVIEW in messages[0]["content"] + + # Delete the row → the next turn's prompt drops the section. + db.execute(text("TRUNCATE kb_overview")) + db.commit() + _stream_chat(client, QUESTION) + assert len(seeded_kb.seen_messages) == 3 + assert "" not in seeded_kb.seen_messages[-1][0]["content"] + finally: + fastapi_app.dependency_overrides.clear() diff --git a/tests/integration/test_migration_0004.py b/tests/integration/test_migration_0004.py index 3a7b470..929c664 100644 --- a/tests/integration/test_migration_0004.py +++ b/tests/integration/test_migration_0004.py @@ -3,14 +3,15 @@ Drives the **real Alembic engine** against the live dev database (``podman compose up -d db``), mirroring the style of ``test_migration_0002.py`` (information_schema assertions on the state the -migration must leave): +migration must leave). The tests target revision ``0004`` explicitly so +later migrations (0005, …) cannot break them: -* upgrade to head → ``documents.summary`` (TEXT, nullable) and +* upgrade 0003 → 0004 → ``documents.summary`` (TEXT, nullable) and ``chunks.is_summary`` (BOOLEAN NOT NULL, default false) both exist, and a chunk inserted without the column gets ``is_summary = false`` (pre-0004 insert paths stay valid); * downgrade to 0003 → both columns are gone; -* upgrade to head again → both are back (round-trip). +* upgrade 0003 → 0004 again → both are back (round-trip). The ``alembic`` fixture guarantees the DB ends at head even if a test fails or the process is interrupted. @@ -65,13 +66,13 @@ def _version(db: Session) -> str | None: return db.execute(text("SELECT version_num FROM alembic_version")).scalar() -def test_upgrade_to_head_adds_summary_columns(db: Session, alembic: Config) -> None: - """Upgrade to head: both columns exist with the locked types/defaults.""" +def test_upgrade_to_0004_adds_summary_columns(db: Session, alembic: Config) -> None: + """Upgrade 0003 → 0004: both columns exist with the locked types/defaults.""" command.downgrade(alembic, "0003") # start from the pre-0004 state assert _version(db) == "0003" - command.upgrade(alembic, "head") - assert _version(db) == "0004", "alembic_version must be at 0004 (head)" + command.upgrade(alembic, "0004") + assert _version(db) == "0004", "alembic_version must be at 0004" summary = _column(db, "documents", "summary") assert summary is not None, "documents.summary is missing" @@ -130,9 +131,10 @@ def test_downgrade_to_0003_removes_columns(db: Session, alembic: Config) -> None def test_upgrade_round_trip_restores_columns(db: Session, alembic: Config) -> None: - """Upgrade back to head after the downgrade: both columns are back.""" - command.upgrade(alembic, "head") - assert _version(db) == "0004", "round-trip upgrade must land at 0004 (head)" + """Downgrade to 0003, then upgrade 0003 → 0004: both columns are back.""" + command.downgrade(alembic, "0003") + command.upgrade(alembic, "0004") + assert _version(db) == "0004", "round-trip upgrade must land at 0004" summary = _column(db, "documents", "summary") assert summary is not None and summary[1] == "YES", "documents.summary must be back" diff --git a/tests/integration/test_migration_0005.py b/tests/integration/test_migration_0005.py new file mode 100644 index 0000000..488baf8 --- /dev/null +++ b/tests/integration/test_migration_0005.py @@ -0,0 +1,152 @@ +"""Integration: migration 0005 (kb_overview) schema contract. + +Drives the **real Alembic engine** against the live dev database +(``podman compose up -d db``), mirroring the style of +``test_migration_0004.py`` (information_schema assertions on the state the +migration must leave): + +* upgrade to head → ``kb_overview`` exists with exactly the three columns + the phase locks in (``id INTEGER PK`` default 1, ``content TEXT NOT NULL`` + default ``''``, ``updated_at TIMESTAMPTZ NOT NULL`` default ``now()``), + and a bare insert lands the single-row defaults (id=1, content=''); +* downgrade to 0004 → the table is gone; +* upgrade to head again → it is back (round-trip). + +The ``alembic`` fixture guarantees the DB ends at head even if a test +fails or the process is interrupted. +""" +from __future__ import annotations + +from collections.abc import Iterator +from typing import Any + +import pytest +from alembic.config import Config +from sqlalchemy import text +from sqlalchemy.orm import Session + +from alembic import command +from app.db import db_available + + +@pytest.fixture() +def alembic(db: Session) -> Iterator[Config]: + """Real Alembic config bound to the dev DB (URL from app settings). + + Starts at head (repairs an interrupted earlier run); teardown upgrades + to head no matter what happened, so the dev DB is never left below + head. + """ + if not db_available(): + pytest.skip("Postgres not reachable — run `podman compose up -d db` first") + cfg = Config() # no alembic.ini file — env.py gets the URL from app config + cfg.set_main_option("script_location", "alembic") + command.upgrade(cfg, "head") + try: + yield cfg + finally: + command.upgrade(cfg, "head") + + +def _column(db: Session, column: str) -> tuple[Any, ...] | None: + """(data_type, is_nullable, column_default) for one kb_overview column.""" + row = db.execute( + text( + "SELECT data_type, is_nullable, column_default" + " FROM information_schema.columns" + " WHERE table_name = 'kb_overview' AND column_name = :c" + ), + {"c": column}, + ).fetchone() + return tuple(row) if row is not None else None + + +def _version(db: Session) -> str | None: + return db.execute(text("SELECT version_num FROM alembic_version")).scalar() + + +def test_upgrade_to_head_creates_kb_overview(db: Session, alembic: Config) -> None: + """Upgrade to head: the single-row table exists with the locked + column types, nullability, and server defaults.""" + command.downgrade(alembic, "0004") # start from the pre-0005 state + assert _version(db) == "0004" + + command.upgrade(alembic, "head") + assert _version(db) == "0005", "alembic_version must be at 0005 (head)" + + pk = db.execute( + text( + "SELECT column_name FROM information_schema.table_constraints tc" + " JOIN information_schema.key_column_usage kcu" + " ON tc.constraint_name = kcu.constraint_name" + " WHERE tc.table_name = 'kb_overview' AND tc.constraint_type = 'PRIMARY KEY'" + ) + ).scalar() + assert pk == "id", "kb_overview primary key must be id" + + id_col = _column(db, "id") + assert id_col is not None, "kb_overview.id is missing" + assert id_col[0] == "integer", "kb_overview.id must be INTEGER" + assert id_col[1] == "NO", "kb_overview.id must be NOT NULL" + assert id_col[2] == "1", "kb_overview.id must have server default 1" + + content = _column(db, "content") + assert content is not None, "kb_overview.content is missing" + assert content[0] == "text", "kb_overview.content must be TEXT" + assert content[1] == "NO", "kb_overview.content must be NOT NULL" + assert content[2] is not None and "''" in content[2], ( + "kb_overview.content must have server default ''" + ) + + updated = _column(db, "updated_at") + assert updated is not None, "kb_overview.updated_at is missing" + assert updated[0] == "timestamp with time zone", "kb_overview.updated_at must be TIMESTAMPTZ" + assert updated[1] == "NO", "kb_overview.updated_at must be NOT NULL" + assert updated[2] is not None and "now()" in updated[2], ( + "kb_overview.updated_at must have server default now()" + ) + + +def test_bare_insert_gets_single_row_defaults(db: Session, alembic: Config) -> None: + """A column-less insert lands the single-row shape the phase relies on: + ``id = 1``, ``content = ''``, server-stamped ``updated_at``.""" + command.upgrade(alembic, "head") + try: + db.execute(text("DELETE FROM kb_overview")) + db.execute(text("INSERT INTO kb_overview DEFAULT VALUES")) + db.commit() + row = db.execute( + text("SELECT id, content, updated_at IS NOT NULL FROM kb_overview") + ).fetchone() + assert row is not None, "the bare insert must land one row" + assert row[0] == 1, "kb_overview.id must default to 1" + assert row[1] == "", "kb_overview.content must default to the empty string" + assert row[2] is True, "kb_overview.updated_at must be stamped by the server" + finally: + db.execute(text("DELETE FROM kb_overview")) + db.commit() + + +def test_downgrade_to_0004_drops_table(db: Session, alembic: Config) -> None: + """Downgrade to 0004: the table is dropped (A13 — reversible).""" + command.downgrade(alembic, "0004") + assert _version(db) == "0004" + + exists = db.execute( + text("SELECT to_regclass('public.kb_overview') IS NOT NULL") + ).scalar() + assert exists is False, "kb_overview must be dropped by the downgrade" + + +def test_upgrade_round_trip_restores_table(db: Session, alembic: Config) -> None: + """Upgrade back to head after the downgrade: table + defaults are back.""" + command.upgrade(alembic, "head") + assert _version(db) == "0005", "round-trip upgrade must land at 0005 (head)" + + id_col = _column(db, "id") + assert id_col is not None and id_col[2] == "1", "kb_overview.id must be back with default 1" + + content = _column(db, "content") + assert content is not None and content[2] is not None and "''" in content[2], ( + "kb_overview.content must keep its '' default after the round-trip" + ) diff --git a/tests/unit/test_chat_gate.py b/tests/unit/test_chat_gate.py index 22f183d..043e9be 100644 --- a/tests/unit/test_chat_gate.py +++ b/tests/unit/test_chat_gate.py @@ -19,13 +19,16 @@ from fastapi.testclient import TestClient from app.api import chat as chat_api from app.config import Settings from app.main import app as fastapi_app -from app.models import Document, QueryLog +from app.models import Document, KbOverview, QueryLog from app.rag.llm import StreamPiece from app.rag.retriever import RetrievedChunk, weak_hit_titles from app.rag.suggestions import MAX_SUGGESTIONS, derive_suggestions ANSWER = "I haven't done anything like that — try one of these instead!" +#: A small KB outline standing in for the lite-generated one (phase 31). +KB_OVERVIEW = "- Homelab\n - Kubernetes (k3s)\n- Deployments\n - Borg backups" + def _settings(threshold: float = 0.30) -> Settings: return Settings( @@ -236,6 +239,87 @@ def test_no_summary_chunks_yields_zero_summary_hits() -> None: assert plan_low.summary_hits == 0 +# ---------- KB overview (phase 31: section + kb_chars) ---------- + + +def test_plan_turn_high_injects_kb_overview() -> None: + """HIGH branch: the stored outline lands in the prompt between + ```` and ```` (or the ```` body with no + notes), and ``kb_chars`` records the outline's length.""" + doc = _doc("Kubernetes Homelab Cluster", "TALOS_DOC_CONTENT") + plan = chat_api.plan_turn( + [_chunk(doc, 0.90)], _settings(threshold=0.30), kb_overview=KB_OVERVIEW + ) + assert plan.deflected is False + assert plan.kb_chars == len(KB_OVERVIEW) + prompt = plan.system_prompt + assert "" in prompt + assert KB_OVERVIEW in prompt + i_rel = prompt.index("HIGH") + i_kb = prompt.index("") + i_docs = prompt.index("") + assert i_rel < i_kb < i_docs + + +def test_plan_turn_high_kb_section_ordered_before_tuning() -> None: + """Both sections present: ```` → ```` → + ```` → ```` (the locked phase-31 order).""" + doc = _doc("Kubernetes Homelab Cluster", "TALOS_DOC_CONTENT") + plan = chat_api.plan_turn( + [_chunk(doc, 0.90)], + _settings(threshold=0.30), + notes=["be concise"], + kb_overview=KB_OVERVIEW, + ) + prompt = plan.system_prompt + i_rel = prompt.index("HIGH") + i_kb = prompt.index("") + i_kb_close = prompt.index("") + i_tuning = prompt.index("") + i_docs = prompt.index("") + assert i_rel < i_kb < i_kb_close < i_tuning < i_docs + assert plan.kb_chars == len(KB_OVERVIEW) + + +def test_plan_turn_low_injects_kb_overview() -> None: + """LOW (deflected) branch: the outline is injected there too, ahead + of the DEFLECT_MODE body, and document content stays excluded.""" + doc = _doc("Kubernetes Homelab Cluster", "TALOS_DOC_NEVER_SENT") + plan = chat_api.plan_turn( + [_chunk(doc, 0.10)], _settings(threshold=0.30), kb_overview=KB_OVERVIEW + ) + assert plan.deflected is True + assert plan.kb_chars == len(KB_OVERVIEW) + prompt = plan.system_prompt + assert "" in prompt + assert KB_OVERVIEW in prompt + i_rel = prompt.index("LOW") + i_kb = prompt.index("") + i_mode = prompt.index("DEFLECT_MODE") + assert i_rel < i_kb < i_mode + assert "TALOS_DOC_NEVER_SENT" not in prompt # titles only, still + + +def test_plan_turn_empty_overview_keeps_prompt_and_zero_kb_chars() -> None: + """No outline (None/empty/blank) → ``kb_chars == 0`` and a prompt + byte-identical to the no-overview build in both branches.""" + high_doc = _doc("Kubernetes Homelab Cluster", "TALOS_DOC_CONTENT") + low_doc = _doc("Backup Strategy", "BACKUP_DOC_CONTENT") + for chunks, kb in ( + ([_chunk(high_doc, 0.90)], None), + ([_chunk(high_doc, 0.90)], ""), + ([_chunk(high_doc, 0.90)], " \n\t "), + ([_chunk(low_doc, 0.10)], None), + ([_chunk(low_doc, 0.10)], ""), + ): + plan = chat_api.plan_turn(chunks, _settings(threshold=0.30), kb_overview=kb) + baseline = chat_api.plan_turn(chunks, _settings(threshold=0.30)) + assert plan.kb_chars == 0 + assert baseline.kb_chars == 0 + assert plan.system_prompt == baseline.system_prompt # byte-identical + assert "" not in plan.system_prompt + + # ---------- prompt content (LOW vs HIGH) ---------- @@ -360,12 +444,15 @@ class _FakeSession: """Stands in for the DB session: records the QueryLog row it is given. ``scalars`` always yields no steering notes (phase 15) so the chat - turn's ``load_steering_notes`` call stays a no-op here. + turn's ``load_steering_notes`` call stays a no-op here, and ``get`` + returns the single ``kb_overview`` row when one is configured + (phase 31) — ``None`` by default, i.e. no stored outline. """ - def __init__(self) -> None: + def __init__(self, kb_overview: str = "") -> None: self.added: list[Any] = [] self.commits = 0 + self.kb_overview = kb_overview def add(self, obj: Any) -> None: self.added.append(obj) @@ -376,6 +463,11 @@ class _FakeSession: def scalars(self, _stmt: Any) -> _FakeSteeringResult: return _FakeSteeringResult() + def get(self, model: Any, pk: Any) -> Any: + if model is KbOverview and self.kb_overview: + return KbOverview(id=1, content=self.kb_overview) + return None + @pytest.fixture() def gate_env(monkeypatch: pytest.MonkeyPatch) -> Iterator[tuple[_FakeSession, _CannedLLM]]: @@ -479,3 +571,55 @@ def test_endpoint_score_at_threshold_answers( assert isinstance(row, QueryLog) assert row.deflected is False assert row.top_score == pytest.approx(0.30) + + +# ---------- endpoint: KB overview row (phase 31) ---------- + + +def test_endpoint_stored_kb_row_injected_into_system_prompt( + client: TestClient, + gate_env: tuple[_FakeSession, _CannedLLM], + monkeypatch: pytest.MonkeyPatch, + caplog: pytest.LogCaptureFixture, +) -> None: + """A non-empty ``kb_overview`` row (read via one PK lookup) reaches + the LLM's system prompt in both modes, and the per-turn log line + records ``kb_chars=N`` (PLAN §9).""" + session, llm = gate_env + doc = _doc("Kubernetes Homelab Cluster", "TALOS_DOC_SENT") + monkeypatch.setattr(chat_api, "retrieve", _fake_retriever([_chunk(doc, 0.30)])) + + session.kb_overview = f" {KB_OVERVIEW} " # the loader trims it + with caplog.at_level("INFO", logger="app.chat"): + _ask(client, "How is my Kubernetes cluster set up?") + + (system, _user) = llm.seen[0][0], llm.seen[0][1] + assert "" in system["content"] + assert KB_OVERVIEW in system["content"] + assert system["content"].index("HIGH") < system["content"].index( + "" + ) < system["content"].index("") + log_lines = [r.getMessage() for r in caplog.records if "question=" in r.getMessage()] + assert log_lines and f"kb_chars={len(KB_OVERVIEW)}" in log_lines[-1] + + +def test_endpoint_no_kb_row_prompt_unchanged( + client: TestClient, + gate_env: tuple[_FakeSession, _CannedLLM], + monkeypatch: pytest.MonkeyPatch, + caplog: pytest.LogCaptureFixture, +) -> None: + """No ``kb_overview`` row → the section is absent (byte-identical to + the pre-phase prompt) and the log line records ``kb_chars=0``.""" + session, llm = gate_env + assert session.kb_overview == "" # fixture default: no stored row + doc = _doc("Kubernetes Homelab Cluster", "TALOS_DOC_SENT") + monkeypatch.setattr(chat_api, "retrieve", _fake_retriever([_chunk(doc, 0.30)])) + + with caplog.at_level("INFO", logger="app.chat"): + _ask(client, "How is my Kubernetes cluster set up?") + + (system, _user) = llm.seen[0][0], llm.seen[0][1] + assert "" not in system["content"] + log_lines = [r.getMessage() for r in caplog.records if "question=" in r.getMessage()] + assert log_lines and "kb_chars=0" in log_lines[-1] diff --git a/tests/unit/test_overview.py b/tests/unit/test_overview.py new file mode 100644 index 0000000..419f1f3 --- /dev/null +++ b/tests/unit/test_overview.py @@ -0,0 +1,387 @@ +"""Unit: KB overview generator (phase 31, task 02). + +The prompt tests are pure (no DB): ``KB_OVERVIEW_MODE`` system prompt, +per-document user lines (source/path/title/first summary line), and the +input cap with the shared ``[…truncated…]`` marker. The loader and +regenerator tests run against the local compose Postgres (preferred — +real single-row upsert), skipping with clear instructions when the stack +is not up — same pattern as ``test_importer.py``. +""" +from __future__ import annotations + +import asyncio +import logging +import uuid +from datetime import UTC, datetime + +import pytest +from sqlalchemy import text +from sqlalchemy.orm import Session + +from app.config import Settings, get_settings +from app.models import Document, KbOverview +from app.rag.llm import LLMError +from app.rag.overview import ( + KB_OVERVIEW_MODE, + OVERVIEW_INSTRUCTION, + SYSTEM_PROMPT, + build_overview_prompt, + load_kb_overview, + regenerate_overview, +) +from app.rag.retriever import TRUNCATION_MARKER + +REPLY = "- Homelab\n - Kubernetes (k3s)\n- Deployments\n - Borg backups" + + +class _FakeLLM: + """Duck-typed stand-in for ``LLMClient`` (``chat`` + ``settings``). + + Records the messages and the ``model`` kwarg it was called with; + returns a canned reply or raises (e.g. :class:`LLMError`). + """ + + def __init__(self, reply: str = REPLY, fail: Exception | None = None) -> None: + self._reply = reply + self._fail = fail + self.settings = Settings(_env_file=None) # pyright: ignore[reportCallIssue] + self.messages: list[dict[str, str]] = [] + self.model: str | None = None + self.calls = 0 + + async def chat( + self, messages: list[dict[str, str]], model: str | None = None + ) -> str: + self.calls += 1 + self.messages = list(messages) + self.model = model + if self._fail is not None: + raise self._fail + return self._reply + + +# ---------- build_overview_prompt: system ---------- + + +def test_system_prompt_has_marker_and_locked_instruction() -> None: + assert SYSTEM_PROMPT.startswith(KB_OVERVIEW_MODE) + assert OVERVIEW_INSTRUCTION in SYSTEM_PROMPT + for fragment in ( + "compact plain-text outline", + "basic categories and topics", + "Group by source", + "use `-` bullet lines", + "~1500 characters", + "no markdown headings", + "no topics not present in the list", + ): + assert fragment in SYSTEM_PROMPT + system, _ = build_overview_prompt([]) + assert system == SYSTEM_PROMPT + assert KB_OVERVIEW_MODE in system # the marker the E2E mock keys on + + +# ---------- build_overview_prompt: user lines ---------- + + +def test_user_lines_carry_source_path_title_and_first_summary_line() -> None: + rows = [ + ("Homelab", "kubernetes/k3s.md", "K3s Cluster", None), + ( + "Deployments", + "backups/borg.yaml", + "Borg Backups", + "Backups run nightly via the borg schedule.\nSource: Deployments/backups/borg.yaml", + ), + ] + system, user = build_overview_prompt(rows) + assert system == SYSTEM_PROMPT + assert user == ( + "Homelab — kubernetes/k3s.md — K3s Cluster\n" + "Deployments — backups/borg.yaml — Borg Backups — " + "Backups run nightly via the borg schedule." + ) + + +def test_user_line_omits_summary_field_when_absent() -> None: + _, user = build_overview_prompt([("Homelab", "a.md", "Title A", None)]) + assert user == "Homelab — a.md — Title A" + assert not user.endswith(" — ") # no dangling dash + + +def test_user_line_omits_summary_field_when_blank() -> None: + _, user = build_overview_prompt([("Homelab", "a.md", "Title A", " \n\t ")]) + assert user == "Homelab — a.md — Title A" + + +def test_user_line_uses_only_first_summary_line() -> None: + """Multi-line summaries contribute their first line only (the + model-written lead sentence; the ``Source: …`` pointer is last).""" + _, user = build_overview_prompt( + [ + ( + "Homelab", + "a.yaml", + "Title A", + "First line.\nSecond line.\nSource: Homelab/a.yaml", + ) + ] + ) + assert user == "Homelab — a.yaml — Title A — First line." + assert "Second line" not in user + assert "Source:" not in user + + +def test_zero_rows_yield_empty_user() -> None: + _, user = build_overview_prompt([]) + assert user == "" + + +# ---------- build_overview_prompt: input cap ---------- + + +def test_user_prompt_truncated_with_marker_when_over_custom_cap() -> None: + rows = [("S", f"p{i}.md", f"T{i}", None) for i in range(10)] + _, full = build_overview_prompt(rows, max_chars=10_000) + cap = 25 + _, user = build_overview_prompt(rows, max_chars=cap) + assert user == full[:cap] + "\n" + TRUNCATION_MARKER + assert user.endswith(TRUNCATION_MARKER) + assert len(user) > cap # the marker makes the cut visible past the cap + + +def test_user_prompt_at_exact_cap_not_truncated() -> None: + rows = [("S", "p.md", "T", None)] # "S — p.md — T" = 12 chars + _, user = build_overview_prompt(rows, max_chars=12) + assert user == "S — p.md — T" + assert TRUNCATION_MARKER not in user + + +def test_user_prompt_truncated_at_default_cap() -> None: + """No explicit cap → ``BOR_OVERVIEW_INPUT_MAX_CHARS`` (read from the + live settings, so the test holds for any configured value).""" + cap = get_settings().overview_input_max_chars + rows = [("S", f"p{i}.md", "T", None) for i in range(3_000)] + _, user = build_overview_prompt(rows) + assert user.endswith(TRUNCATION_MARKER) + body = user.removesuffix("\n" + TRUNCATION_MARKER) + assert len(body) == cap # cut exactly at the cap, marker on its own line + assert "p2999.md" not in body # the overflow never reaches the model + + +# ---------- load_kb_overview ---------- + + +def test_load_kb_overview_no_row_empty(db: Session) -> None: + db.execute(text("DELETE FROM kb_overview")) + db.commit() + assert load_kb_overview(db) == "" + + +def test_load_kb_overview_returns_trimmed_content(db: Session) -> None: + db.execute(text("DELETE FROM kb_overview")) + db.add(KbOverview(id=1, content=f" {REPLY} \n")) + db.commit() + assert load_kb_overview(db) == REPLY + + +def test_load_kb_overview_whitespace_only_row_empty(db: Session) -> None: + db.execute(text("DELETE FROM kb_overview")) + db.add(KbOverview(id=1, content=" \n\t ")) + db.commit() + assert load_kb_overview(db) == "" + + +# ---------- regenerate_overview ---------- + + +def _add_doc( + db: Session, source: str, path: str, title: str, summary: str | None = None +) -> Document: + doc = Document( + id=uuid.uuid4(), + source=source, + path=path, + full_path=f"/tmp/{path}", + title=title, + content="body", + content_hash="0" * 64, + summary=summary, + ) + db.add(doc) + db.commit() + return doc + + +def _truncate_documents(db: Session) -> None: + db.execute(text("TRUNCATE chunks, documents")) + db.commit() + + +def _overview_row(db: Session) -> KbOverview | None: + return db.get(KbOverview, 1) + + +@pytest.fixture() +def clean_overview(db: Session): + db.execute(text("DELETE FROM kb_overview")) + db.commit() + yield + db.execute(text("DELETE FROM kb_overview")) + db.commit() + + +def test_regenerate_happy_path_creates_single_row_id_1( + db: Session, clean_overview, caplog: pytest.LogCaptureFixture +) -> None: + _truncate_documents(db) + try: + _add_doc(db, "Homelab", "b.md", "B") + _add_doc(db, "Homelab", "a.md", "A", "A summary.\nSource: Homelab/a.md") + llm = _FakeLLM() + with caplog.at_level(logging.INFO, logger="app.rag.overview"): + ok = asyncio.run(regenerate_overview(llm, db)) + assert ok is True + assert llm.calls == 1 + row = db.get(KbOverview, 1) + assert row is not None, "the single row must be upserted" + assert row.id == 1 + assert row.content == REPLY + assert row.updated_at is not None + age = datetime.now(UTC) - row.updated_at + assert age.total_seconds() < 300, "updated_at must be a fresh UTC timestamp" + assert f"overview: regenerated docs=2 chars={len(REPLY)}" in caplog.text + finally: + _truncate_documents(db) + + +def test_regenerate_updates_existing_row_in_place(db: Session, clean_overview) -> None: + _truncate_documents(db) + past = datetime(2020, 1, 1, 12, 0, 0, tzinfo=UTC) + db.add(KbOverview(id=1, content="old outline", updated_at=past)) + db.commit() + try: + _add_doc(db, "Homelab", "a.md", "A") + ok = asyncio.run(regenerate_overview(_FakeLLM(), db)) + assert ok is True + count = db.scalar(text("SELECT count(*) FROM kb_overview")) + assert count == 1, "still exactly one row after the re-generation" + row = db.get(KbOverview, 1) + assert row is not None + assert row.content == REPLY # replaced, not appended + assert row.updated_at is not None and row.updated_at > past + finally: + _truncate_documents(db) + + +def test_regenerate_orders_documents_by_source_path(db: Session, clean_overview) -> None: + _truncate_documents(db) + try: + # Inserted out of order — the model must see a deterministic order. + _add_doc(db, "Deployments", "z.md", "Z") + _add_doc(db, "Homelab", "b.md", "B") + _add_doc(db, "Homelab", "a.md", "A") + llm = _FakeLLM() + assert asyncio.run(regenerate_overview(llm, db)) is True + finally: + _truncate_documents(db) + + user = next(m["content"] for m in llm.messages if m["role"] == "user") + assert user == ( + "Deployments — z.md — Z\nHomelab — a.md — A\nHomelab — b.md — B" + ) + + +def test_regenerate_calls_the_configured_summary_model( + db: Session, clean_overview +) -> None: + _truncate_documents(db) + try: + _add_doc(db, "Homelab", "a.md", "A") + llm = _FakeLLM() + assert asyncio.run(regenerate_overview(llm, db)) is True + finally: + _truncate_documents(db) + + assert llm.model == llm.settings.llm_summary_model # the ``lite`` default + assert llm.model == "lite" + assert [m["role"] for m in llm.messages] == ["system", "user"] + assert KB_OVERVIEW_MODE in llm.messages[0]["content"] + assert "A" in llm.messages[1]["content"] + + +def test_regenerate_zero_docs_leaves_existing_row_untouched( + db: Session, clean_overview +) -> None: + db.execute(text("DELETE FROM kb_overview")) + db.add(KbOverview(id=1, content="old outline")) + db.commit() + _truncate_documents(db) + llm = _FakeLLM() + ok = asyncio.run(regenerate_overview(llm, db)) + assert ok is False + assert llm.calls == 0, "no KB → no wasted lite-model call" + row = db.get(KbOverview, 1) + assert row is not None and row.content == "old outline" # untouched + + +def test_regenerate_zero_docs_without_row_creates_nothing(db: Session, clean_overview) -> None: + _truncate_documents(db) + llm = _FakeLLM() + ok = asyncio.run(regenerate_overview(llm, db)) + assert ok is False + assert llm.calls == 0 + assert _overview_row(db) is None, "no row must be invented for an empty KB" + + +def test_regenerate_llm_error_leaves_previous_row_intact( + db: Session, clean_overview, caplog: pytest.LogCaptureFixture +) -> None: + db.execute(text("DELETE FROM kb_overview")) + db.add(KbOverview(id=1, content="old outline")) + db.commit() + _truncate_documents(db) + try: + _add_doc(db, "Homelab", "a.md", "A") + llm = _FakeLLM(fail=LLMError("simulated lite-model failure")) + with caplog.at_level(logging.ERROR, logger="app.rag.overview"): + ok = asyncio.run(regenerate_overview(llm, db)) + assert ok is False, "fail-soft: the import must not be dragged down" + row = db.get(KbOverview, 1) + assert row is not None and row.content == "old outline", ( + "the previous outline stays — an old outline is better than none" + ) + assert "overview: regeneration failed —" in caplog.text + assert "simulated lite-model failure" in caplog.text + finally: + _truncate_documents(db) + + +def test_regenerate_llm_error_without_row_creates_nothing( + db: Session, clean_overview +) -> None: + _truncate_documents(db) + try: + _add_doc(db, "Homelab", "a.md", "A") + llm = _FakeLLM(fail=LLMError("simulated lite-model failure")) + ok = asyncio.run(regenerate_overview(llm, db)) + finally: + _truncate_documents(db) + + assert ok is False + assert _overview_row(db) is None + + +def test_regenerate_opens_own_session_when_none(db: Session, clean_overview) -> None: + """``session=None`` → a private ``SessionLocal`` session is opened, + committed, and closed (the import-script call path).""" + _truncate_documents(db) + try: + _add_doc(db, "Homelab", "a.md", "A") + ok = asyncio.run(regenerate_overview(_FakeLLM())) + finally: + _truncate_documents(db) + + assert ok is True + row = db.get(KbOverview, 1) # fresh lookup via the test session + assert row is not None and row.content == REPLY diff --git a/tests/unit/test_prompts.py b/tests/unit/test_prompts.py index ee6befb..9ac0aff 100644 --- a/tests/unit/test_prompts.py +++ b/tests/unit/test_prompts.py @@ -1,12 +1,34 @@ -"""Unit: locked persona prompt builder (PLAN §6 verbatim + both modes).""" +"""Unit: locked persona prompt builder (PLAN §6 verbatim + both modes). + +Also covers the phase-31 ```` section (the stored, +lite-generated KB outline): its own builder contract (empty → ``""``, +char budget + ``[…truncated…]`` marker, pathological budgets), its +placement between ```` and ```` in both modes, and +the byte-identical-when-absent convention (phase 15 precedent). +""" from __future__ import annotations import uuid import pytest +from app.config import Settings from app.models import Document -from app.rag.prompts import PERSONA, _base, build_deflect_prompt, build_high_prompt +from app.rag.prompts import ( + PERSONA, + _base, + build_deflect_prompt, + build_high_prompt, + build_kb_section, + build_steering_section, +) +from app.rag.retriever import TRUNCATION_MARKER + +#: A small multi-line outline standing in for the lite-generated one. +OVERVIEW = "- Homelab\n - Kubernetes (k3s)\n- Deployments\n - Borg backups" + +#: The locked one-line intro of the ```` section. +KB_INTRO = "The basic categories of everything in this knowledge base (generated at import time):" def _doc(path: str, content: str, title: str) -> Document: @@ -116,3 +138,195 @@ def test_zero_note_prompt_is_byte_identical_to_pre_steering() -> None: def test_relevance_placeholder_rejected_for_garbage() -> None: with pytest.raises(ValueError, match="HIGH or LOW"): _base("MEDIUM") + + +# ---------- section (phase 31) ---------- + + +def test_kb_section_empty_when_no_overview() -> None: + assert build_kb_section("") == "" + assert build_kb_section(" \n\t ") == "" + + +def test_kb_section_format_intro_and_content() -> None: + assert build_kb_section(OVERVIEW) == ( + f"\n{KB_INTRO}\n{OVERVIEW}\n" + ) + + +def test_kb_section_trims_overview_edges() -> None: + assert build_kb_section(f" {OVERVIEW} \n") == build_kb_section(OVERVIEW) + + +def test_kb_section_fits_budget_exactly_no_marker() -> None: + exact = f"\n{KB_INTRO}\n{OVERVIEW}\n" + section = build_kb_section(OVERVIEW, max_chars=len(exact)) + assert TRUNCATION_MARKER not in section + assert section == exact + + +def test_kb_section_over_budget_capped_with_marker() -> None: + text = "- " + "x" * 500 + # The section frame alone is 121 chars, so the cap must clear it for + # any outline prefix to fit (pathological budgets are tested below). + cap = 200 + section = build_kb_section(text, max_chars=cap) + assert len(section) <= cap # the budget is never exceeded + assert TRUNCATION_MARKER in section + assert section.startswith(f"\n{KB_INTRO}\n-") + assert section.endswith(f"{TRUNCATION_MARKER}\n") + # The body is the kept prefix + the marker on its own line, and the + # kept part must be a true prefix of the outline (longest-fitting). + body = section.removeprefix(f"\n{KB_INTRO}\n").removesuffix( + "\n" + ) + kept, marker = body.rsplit("\n", 1) + assert marker == TRUNCATION_MARKER + assert kept.startswith("- ") + assert text.startswith(kept), "the kept part must be a prefix of the outline" + # And it is the longest such prefix: one more char would not fit. + assert len(section) > cap - 2, "the cut must sit as close to the cap as possible" + + +def test_kb_section_default_budget_from_settings(monkeypatch: pytest.MonkeyPatch) -> None: + from app.rag import prompts as prompts_mod + + monkeypatch.setattr( + prompts_mod, "get_settings", lambda: Settings(_env_file=None) # pyright: ignore[reportCallIssue] + ) + text = "y" * 9_000 # > the 4 000-char default + section = build_kb_section(text) + assert TRUNCATION_MARKER in section + assert len(section) <= 4_000 + + +def test_kb_section_nonpositive_budget_is_empty() -> None: + assert build_kb_section(OVERVIEW, max_chars=0) == "" + assert build_kb_section(OVERVIEW, max_chars=-10) == "" + + +def test_kb_section_tiny_budget_never_exceeds_cap() -> None: + # Pathological budget (steering precedent, phase 15): the section must + # never exceed the cap — bare marker when it fits, no section at all + # when even that doesn't. + assert build_kb_section("a" * 500, max_chars=10) == "" # marker (13) > 10 + fits_marker = build_kb_section("a" * 500, max_chars=len(TRUNCATION_MARKER)) + assert fits_marker == TRUNCATION_MARKER + + +# ---------- placement (both modes) ---------- + + +def test_no_overview_prompt_is_byte_identical_to_pre_phase() -> None: + """Phase 31 contract: with no KB overview (None, empty, or blank) + every prompt is exactly what it was before the ```` + section existed — with or without steering notes.""" + doc = _doc("kubernetes.md", "Talos Linux on three nodes.", "Kubernetes Homelab Cluster") + block = ( + '\n' + "Talos Linux on three nodes.\n" + "" + ) + docs_block = "\n\n" + block + "\n" + high_plain = _base("HIGH") + docs_block + high_steered = _base("HIGH") + "\n" + build_steering_section(["be concise"]) + docs_block + low_plain = ( + _base("LOW") + + "\nDEFLECT_MODE: retrieval was weak — the titles below are the closest " + "your notes come to the question. They are titles only; do not pretend " + "they answer it. Use them to propose 2-3 alternative questions.\n" + + "- T1\n- T2" + ) + low_steered = ( + _base("LOW") + + "\n" + + build_steering_section(["be concise"]) + + "\nDEFLECT_MODE: retrieval was weak — the titles below are the closest " + "your notes come to the question. They are titles only; do not pretend " + "they answer it. Use them to propose 2-3 alternative questions.\n" + + "- T1\n- T2" + ) + for kb in (None, "", " \n\t "): + assert build_high_prompt([doc], kb_overview=kb) == high_plain + assert build_high_prompt([doc], notes=["be concise"], kb_overview=kb) == high_steered + assert build_deflect_prompt(["T1", "T2"], kb_overview=kb) == low_plain + assert build_deflect_prompt( + ["T1", "T2"], notes=["be concise"], kb_overview=kb + ) == low_steered + assert "" not in build_high_prompt( + [doc], notes=["be concise"], kb_overview=kb + ) + assert "" not in build_deflect_prompt( + ["T1"], notes=["be concise"], kb_overview=kb + ) + + +def test_high_prompt_kb_section_ordered_between_relevance_and_tuning() -> None: + doc = _doc("kubernetes.md", "TALOS_DOC_CONTENT", "Kubernetes Homelab Cluster") + prompt = build_high_prompt([doc], notes=["be concise"], kb_overview=OVERVIEW) + i_rel = prompt.index("HIGH") + i_kb_open = prompt.index("") + i_kb_close = prompt.index("") + i_tuning = prompt.index("") + i_docs = prompt.index("") + assert i_rel < i_kb_open < i_kb_close < i_tuning < i_docs + assert KB_INTRO in prompt + assert OVERVIEW in prompt # outline intact within the section + assert "1. be concise" in prompt # steering still there + assert "TALOS_DOC_CONTENT" in prompt # documents still full + + +def test_high_prompt_kb_section_without_steering() -> None: + doc = _doc("kubernetes.md", "TALOS_DOC_CONTENT", "Kubernetes Homelab Cluster") + prompt = build_high_prompt([doc], kb_overview=OVERVIEW) + i_rel = prompt.index("HIGH") + i_kb_close = prompt.index("") + i_docs = prompt.index("") + assert i_rel < i_kb_close < i_docs + assert "" not in prompt # no notes → no steering section + assert build_kb_section(OVERVIEW) in prompt + + +def test_deflect_prompt_kb_section_ordered_between_relevance_and_tuning() -> None: + prompt = build_deflect_prompt( + ["Title A", "Title B"], notes=["be concise"], kb_overview=OVERVIEW + ) + i_rel = prompt.index("LOW") + i_kb_open = prompt.index("") + i_kb_close = prompt.index("") + i_tuning = prompt.index("") + i_mode = prompt.index("DEFLECT_MODE") + assert i_rel < i_kb_open < i_kb_close < i_tuning < i_mode + assert KB_INTRO in prompt + assert OVERVIEW in prompt + assert "1. be concise" in prompt + assert "- Title A" in prompt # weak-hit titles still carried + + +def test_prompt_kb_section_over_settings_budget_capped_with_marker( + monkeypatch: pytest.MonkeyPatch, +) -> None: + """Prompt-level budget: an overview longer than + ``kb_overview_max_chars`` is capped with the shared marker — in both + modes.""" + from app.rag import prompts as prompts_mod + + monkeypatch.setattr( + prompts_mod, + "get_settings", + # 200 > the 121-char section frame, so a prefix + marker can fit. + lambda: Settings(_env_file=None, kb_overview_max_chars=200), # pyright: ignore[reportCallIssue] + ) + text = "- " + "z" * 500 + doc = _doc("kubernetes.md", "TALOS_DOC_CONTENT", "Kubernetes Homelab Cluster") + for prompt in ( + build_high_prompt([doc], kb_overview=text), + build_deflect_prompt(["Title A"], kb_overview=text), + ): + assert TRUNCATION_MARKER in prompt + assert prompt.index("") < prompt.index(TRUNCATION_MARKER) + # The capped section (open tag through close tag) fits the budget. + section = prompt[prompt.index("") :] + close = section.index("") + section = section[: close + len("")] + assert len(section) <= 200