Files
brain-of-reese/app/rag/importer.py
T
ducoterra 9820c361b0
Build and Push Containers / build-and-push-app (push) Successful in 2m2s
Build and Push Containers / build-and-push-db (push) Successful in 14s
phase: 118_summary_seed_context
**Phase 118 final verification pass — complete.** All criteria verified; 4 pre-existing defects found and fixed.

- **Verified:** summary-seed wiring (`select_suggested` top-5 no-floor → summary blocks, no full text in HIGH prompt), all-doc markdown summaries + NULL backfill (`summary_backfilled`, no `sources_meta` bump), `read` adds full text with `read_docs`-only dedupe, `done.sources` = suggested+read / durable record = suggested+related+read + `suggested=N` log line (seen live in E2E), byte-locked PERSONA/LOW/TOOLS_SECTION, battery gate PASS recorded in `TOOL_CALLING_TESTING.md` §10 (turbo 2026-09-16: 1/2/4 GREEN, cond-3 reported 9/10 per A7, contract 21/21, caps 0).
- **Defects fixed (all pre-existing, none phase-118):** ① `ChatMessage` schema missing the phase-113 `related` key → `extra="forbid"` 422'd every done-time auto-save of grounded turns with a related tier, leaving `message_count=1` (root cause of `test_share_chat` 3F; browser-level instrumentation proved the PUT 422) — added the field + unit/integration pins; ② `test_theme_semantic_completion` pins stale vs phase-117 debox (border/chip removed) — re-targeted to assert border/chip *absence*; ③ `test_header_consistency` `<26`px pin red on 26.125px native date-input line — bound relaxed to `<34` (wrap-detection intent kept); ④ `test_navbar_refresh` bor.chat.v1 key set updated for `related`.
- **Test/lint/coverage:** `uv run pytest --cov=app --cov-report=term-missing` → **2506 passed, app/ 99%** (>90%); `uv run ruff check . && uv run pyright` → clean, 0 errors.
- **E2E:** new story suite in isolation → **2 passed**; full 103-suite matrix sweep (each isolated) → **all 103 green** after the fixes; `test_share_chat` 4 passed, `test_theme_semantic_completion` 8 passed, `test_header_consistency` 3 passed, `test_navbar_refresh` 7 passed.
- **Deviations:** none from LOCKED decisions. Note: orphaned diagnostic uvicorn processes briefly made E2E sessions exercise stale code — killed and re-verified; a sweep-regenerated tracked screenshot was restored. No commits made (harness commits).
- **Completion criteria:** all 7 ✅ (commit/phase-move is the harness's step).
- **Next pending phase:** none — `todo/` holds only this phase's overview pending the harness move.
2026-09-16 06:57:49 -04:00

663 lines
29 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Knowledge-base importer (PLAN §5 / §9 / §11).
Walks the in-scope files (the A9 family by default — the original seven
plus the quadlet family and ``j2`` — ``BOR_IMPORT_EXTENSIONS``, which may
name any well-formed extension; case-insensitive), diffs by sha256
against ``documents.content_hash`` and, for every new or changed file, runs
the two-phase upsert:
1. upsert the document row and replace its chunk rows (embeddings NULL)
2. embed the new chunks in batches and attach the vectors
3. commit — one transaction per file, so a failed embedding leaves the
database untouched and the file is simply retried on the next run
4. every file (phase 30; phase 118, A2: markdown included — the
non-markdown-only scope is retired): generate a ``lite``-model summary
and, best-effort, store it on ``documents.summary`` plus one extra
embedded chunk (``is_summary``, position −1). The document row and its
content chunks are already committed at this point, so a summary
failure only means the file is indexed without a summary (logged and
counted in ``summary_errors``) — it is never lost.
Scope (A9, revised 2026-08-21): any path containing a dot-prefixed
component (hidden dirs — vendored caches like ``.esphome/.espressif/**`` —
or hidden files) is skipped, plus the well-known exclusion list — UNLESS
the source's phase-105 hidden-folders flag admits dot-prefixed paths; the
exclusion list always applies. A token may also name extensionless files
by their exact lowercased full filename (``Dockerfile`` under the
``dockerfile`` token — phase 102).
``prune=True`` deletes documents (of the imported sources only) whose files
no longer exist **or no longer match the format filter** — this is how
previously-imported junk (e.g. dot-dir READMEs) leaves the index. Per-file
logging uses the verbs ``added | updated | unchanged | pruned`` plus a
summary line with per-format counts (PLAN §9).
Document dates (phase 106, D2/D4): every import sources
``documents.created_at`` from the file's source — the per-file git
last-commit date when a ``doc_dates_by_root`` entry names the file,
else the file's mtime — normalized by
:func:`app.rag.doc_dates.normalize_doc_date` (undetermined or future →
today, D3) on every add and update. On the unchanged path the stored
date is REFRESHED from the same source (it may go OLDER — no monotonic
guard) and counted in ``summary.dates_updated`` — unless the row
carries the owner's manual correction (``created_at_manual``, D1), which
the sync never touches.
NULL-summary backfill (phase 118, A2): an UNCHANGED file (same
``content_hash``) whose ``documents.summary`` is still NULL — a
pre-phase-30 row, or an earlier fail-soft miss — gets the same
best-effort summary pass on every sync until it sticks. A success counts
``summary_backfilled`` (never ``summaries``) and touches nothing else: no
content re-embed, no added/updated/pruned count — so no
``sources_meta`` bump, no KB-overview/folder-summary regeneration. The
backfill runs BEFORE the ``created_at_manual`` early-return (the manual
flag protects the DATE only, D1) and the strict ``is None`` check leaves
owner-set summaries (even empty strings, phase 57) alone.
``import_sources`` accepts an optional per-file ``progress`` callback
(phase 64, task 01) reporting the file being processed right now.
"""
from __future__ import annotations
import hashlib
import logging
from collections.abc import Callable
from dataclasses import dataclass, field
from datetime import UTC, datetime
from pathlib import Path
from typing import Protocol
from sqlalchemy import select
from sqlalchemy.orm import Session
from app.config import Settings
from app.db import SessionLocal
from app.models import Chunk, Document
from app.rag.chunker import chunk_document, extract_title
from app.rag.doc_dates import file_mtime_datetime, normalize_doc_date
from app.rag.llm import EmbeddingError, LLMError
from app.rag.summarizer import generate_summary
logger = logging.getLogger("app.importer")
#: Non-content directories never imported (PLAN anchor A9).
EXCLUDED_DIRS: frozenset[str] = frozenset(
{".venv", "node_modules", ".git", "__pycache__", ".pytest_cache", "dist", "build"}
)
class Embedder(Protocol):
"""Everything the importer needs from the LLM client (duck-typed for tests)."""
settings: Settings
embed_batches: int
async def embed(self, texts: list[str]) -> list[list[float]]: ...
async def chat(self, messages: list[dict[str, str]], model: str | None = None) -> str: ...
# ^ the one-shot completion the summarizer uses for the ``lite`` model
# (phase 30, task 01); :class:`app.rag.llm.LLMClient` satisfies it.
@dataclass
class ImportSummary:
"""Counts for one import run (also printed by the CLI)."""
files: int = 0
added: int = 0
updated: int = 0
unchanged: int = 0
pruned: int = 0
errors: int = 0
chunks: int = 0
embed_batches: int = 0
#: Files whose lite summary was generated + indexed (phase 30; phase
#: 118, A2: every A9 format, markdown included). One ``is_summary``
#: chunk per success.
summaries: int = 0
#: Files whose summary generation failed (best-effort — the document
#: is still indexed, without a summary).
summary_errors: int = 0
#: Unchanged docs whose NULL summary was backfilled (phase 118, A2) —
#: one ``is_summary`` chunk per success; the content is untouched, so
#: a backfill NEVER counts added/updated/pruned (no
#: ``sources_meta`` bump, no overview/folder-summary regeneration).
summary_backfilled: int = 0
#: Files whose ``created_at`` was refreshed on the UNCHANGED path —
#: content untouched, date re-sourced (phase 106, D4: the date may
#: go OLDER; a date-only refresh NEVER counts added/updated/pruned,
#: so no ``sources_meta`` bump, no overview/folder-summary
#: regeneration).
dates_updated: int = 0
#: Files walked, keyed by lowercased extension (``md``, ``yaml``, …).
formats: dict[str, int] = field(default_factory=dict)
def format_counts(self) -> str:
"""``md:203,yaml:267,py:14`` — highest count first (PLAN §9)."""
if not self.formats:
return "none"
ordered = sorted(self.formats.items(), key=lambda kv: (-kv[1], kv[0]))
return ",".join(f"{ext}:{count}" for ext, count in ordered)
def log(self) -> None:
logger.info(
"import: summary files=%d added=%d updated=%d unchanged=%d pruned=%d "
"errors=%d chunks=%d embed_batches=%d summaries=%d summary_errors=%d "
"summary_backfilled=%d dates_updated=%d formats=%s",
self.files,
self.added,
self.updated,
self.unchanged,
self.pruned,
self.errors,
self.chunks,
self.embed_batches,
self.summaries,
self.summary_errors,
self.summary_backfilled,
self.dates_updated,
self.format_counts(),
)
def normalize_ignore_path(entry: str) -> str:
"""One ignore-path entry → canonical form (phase 89, A1).
Trim surrounding whitespace, then strip ALL leading/trailing
``/`` — so ``"/my/files/"``, ``"my/files/"`` and ``"my/files"``
all become ``"my/files"``. ``""`` / ``"//"``,
``" "`` normalize to ``""`` (callers drop empties).
"""
return entry.strip().strip("/")
def is_ignored(rel: str, prefixes: tuple[str, ...]) -> bool:
"""Phase 89, A1 — the pure prefix rule, nothing else.
``rel`` is the source-relative POSIX path WITHOUT a leading
slash (the ``documents.path`` string). Match = ``rel`` STARTS
WITH a normalized entry: raw string prefix — deliberately NO
component-boundary check (``"my/files"`` also matches
``"my/files2/x.md"``) and NO mid-path matching (``"myfile.txt"``
matches ``"myfile.txt"`` but not ``"some/path/myfile.txt"``).
"""
return any(rel.startswith(p) for p in prefixes)
def _ignore_for_root(
root: Path, ignore_by_root: dict[str, list[str]] | None
) -> tuple[str, ...]:
"""The normalized, non-empty prefix tuple for one root (phase 89).
Keyed by ``str(root)`` — the root string exactly as the caller
passed it in ``sources`` (unambiguous when two rows share a
source *name* but different dirs). Callers may pass RAW box
lines: the importer normalizes + drops empties here, the single
choke point — stored lists (already normalized) normalize to
themselves.
"""
raw = (ignore_by_root or {}).get(str(root)) or []
return tuple(p for p in (normalize_ignore_path(e) for e in raw) if p)
def _include_hidden_for_root(
root: Path, include_hidden_by_root: dict[str, bool] | None
) -> bool:
"""The per-root hidden-folders flag (phase 105, A1).
Keyed by ``str(root)`` — the root string exactly as the caller
passed it in ``sources`` (the ``_ignore_for_root`` convention,
phase 89): ``True`` only for roots the caller lists as True;
unlisted/``None`` roots are ``False`` — every existing caller
behaves byte-identically (A4).
"""
return bool((include_hidden_by_root or {}).get(str(root), False))
def match_extension(path: Path, extensions: frozenset[str]) -> str | None:
"""The bare lowercased token *path* imports under, or ``None``.
1. Non-empty lowercased dotted suffix in *extensions* (the A9 rule —
``kubernetes.md`` → ``md``).
2. No suffix: the lowercased FULL filename equals a bare token of
*extensions* (``Dockerfile`` → ``dockerfile``) — the phase-102
extensionless rule. Exact name only: ``mydockerfile`` never
matches the ``dockerfile`` token.
"""
suffix = path.suffix.lower()
if suffix and suffix in extensions:
return suffix.lstrip(".")
if not suffix and path.name.lower() in {e.lstrip(".") for e in extensions}:
return path.name.lower()
return None
def iter_importable_files(
root: Path,
extensions: frozenset[str],
excluded: frozenset[str] = EXCLUDED_DIRS,
ignore: tuple[str, ...] = (),
include_hidden: bool = False,
) -> list[Path]:
"""All importable files under *root* (sorted), per the A9 scope rules.
*extensions* is a set of lowercased dotted suffixes (``{'.md', '.py'}``).
Skips: when ``include_hidden`` is False (the default), any path with a
dot-prefixed component (hidden dirs/files — vendored caches like
``.esphome/.espressif/**``); when True, dot-prefixed components are
ADMITTED (files inside hidden folders, and hidden files) and only
*excluded* is consulted (A1 — caches/VCS internals are never content).
The well-known non-content directories in *excluded* are skipped in
BOTH states. *ignore* (phase 89, A1) is a
tuple of ALREADY-normalized, non-empty source-relative path prefixes
(the importer's ``_ignore_for_root`` is the normalization choke point
— raw box lines never reach this function): a file is skipped when its
source-relative POSIX path starts with any entry; the default ``()``
keeps every existing caller byte-identical. The *ignore* tuple
composes additively in both states.
"""
if not root.is_dir():
return []
files: list[Path] = []
for path in sorted(root.rglob("*")):
if not path.is_file():
continue
rel = path.relative_to(root)
if any(
(not include_hidden and part.startswith(".")) or part in excluded
for part in rel.parts
):
continue
if ignore and is_ignored(rel.as_posix(), ignore):
continue
if match_extension(path, extensions) is None:
continue
files.append(path)
return files
async def import_sources(
sources: list[Path],
llm: Embedder,
*,
prune: bool = False,
limit: int | None = None,
session: Session | None = None,
progress: Callable[[str, str, int, int], None] | None = None,
ignore_by_root: dict[str, list[str]] | None = None,
include_hidden_by_root: dict[str, bool] | None = None,
doc_dates_by_root: dict[str, dict[str, datetime]] | None = None,
) -> ImportSummary:
"""Import every A9-format file under *sources* (see module docstring).
``session`` may be supplied (tests); a private one is opened and closed
otherwise. ``limit`` caps the number of files processed (debug only) and
disables pruning, since an incomplete walk must not drive deletions.
``progress`` (phase 64, task 01) is an optional per-file hook called
once per importable file, immediately before that file's
``_index_file`` — with ``(source, rel_posix_path, done, total)``: the
same POSIX *rel* the document rows use, ``done`` = the 1-based index of
the current file **across all sources**, and ``total`` = the combined
pre-walk count of importable files across all *sources* roots. The
pre-walk (same extension/exclusion rules, directory stats only, no file
reads) happens **only when *progress* is provided**: callers passing
nothing pay no extra walk and behave exactly as before. Under
``limit``, the hook still fires per processed file only — ``done`` never
exceeds the limit, but ``total`` stays the full pre-walk count (an
incomplete walk must not misreport the denominator).
``ignore_by_root`` (phase 89, A1/A2) maps ``str(root)`` — the root path
string exactly as passed in *sources* — to that source's RAW ignore-path
lines (the importer normalizes them via ``_ignore_for_root``, the single
choke point): matching files are never walked, so they are never
embedded and never summarized, and the progress pre-walk uses the same
per-root tuple as the processing loop, so ``total`` never counts them.
A file that newly matches a pattern simply never enters ``seen``, so the
next ``prune=True`` run deletes its row automatically (A2 — the A9
junk-precedent). ``None`` (the default) changes nothing: the map is read
per root, unlisted roots get an empty tuple, and every existing caller
behaves byte-identically.
``include_hidden_by_root`` (phase 105, A1) maps ``str(root)`` to the
stored flag: ``True`` admits dot-prefixed components for that root
(``EXCLUDED_DIRS`` and the extension filter still apply; the ignore
tuple composes additively). Unlisted/``None`` roots are ``False`` —
byte-identical to pre-phase-105. A file that was indexed with the
flag ON and is walked again with it OFF simply never enters
``seen``, so the next ``prune=True`` run deletes its row
automatically (A2 — the A9/phase-89 precedent).
``doc_dates_by_root`` (phase 106, D2/D4) maps ``str(root)`` — the
root path string exactly as passed in *sources* — to that source's
RAW per-file source dates: source-relative POSIX path → the git
last-commit datetime (task 03's ``file_commit_dates``). ONLY git
roots are listed — unlisted roots (local dirs, unpacked uploads)
take the mtime fallback, and a path missing from its root's map
does too. The map entry beats the file's mtime when present. The
progress pre-walk is untouched (dates change no file count).
``None`` (the default) changes nothing for existing callers: the
mtime fallback applies to every file — which IS the behavior
change, D4: an unchanged file now refreshes its stored date from
its source on every run (the backfill-correction case).
"""
if limit is not None and limit <= 0:
raise ValueError("limit must be >= 1")
summary = ImportSummary()
owns_session = session is None
if session is None:
session = SessionLocal()
seen: set[tuple[str, str]] = set()
source_names: set[str] = set()
# phase 64 (task 01): the hook's combined denominator, walked with the
# exact same rules as the processing loop below (directory stats only,
# no file reads). Skipped entirely for ``progress=None`` callers — no
# extra pass, byte-identical behaviour and cost.
total = 0
if progress is not None:
for root in sources:
total += len(
iter_importable_files(
root,
llm.settings.import_extension_set,
ignore=_ignore_for_root(root, ignore_by_root),
include_hidden=_include_hidden_for_root(
root, include_hidden_by_root
),
)
)
try:
for root in sources:
if not root.is_dir():
logger.warning("import: source dir not found, skipping: %s", root)
continue
if limit is not None and summary.files >= limit:
break
source = root.name
source_names.add(source)
ignore = _ignore_for_root(root, ignore_by_root)
include_hidden = _include_hidden_for_root(root, include_hidden_by_root)
# Phase 106 (D2): the root's raw source dates (git last-commit
# for git roots, keyed by the same str(root) convention); {}
# for unlisted roots — every file then takes the mtime fallback.
dates_map = (doc_dates_by_root or {}).get(str(root), {})
for path in iter_importable_files(
root,
llm.settings.import_extension_set,
ignore=ignore,
include_hidden=include_hidden,
):
if limit is not None and summary.files >= limit:
break
rel = path.relative_to(root).as_posix()
seen.add((source, rel))
summary.files += 1
# Phase 102: the matched bare token (``dockerfile`` for an
# extensionless ``Dockerfile``), never ``unknown`` — the
# file is in scope, so the walk matched it.
ext = (
match_extension(path, llm.settings.import_extension_set)
or "unknown"
)
summary.formats[ext] = summary.formats.get(ext, 0) + 1
if progress is not None:
# phase 64: report the file *before* indexing it — a
# file that then errors or turns out unchanged was
# already the "current file". No try/except around the
# call: the hooks in this repo only assign fields.
progress(source, rel, summary.files, total)
try:
await _index_file(
session, source=source, rel=rel, full_path=path, llm=llm,
summary=summary, raw_date=dates_map.get(rel),
)
except EmbeddingError as e:
# A pathological file (e.g. content the embedding endpoint
# refuses) must not abort the whole KB: roll back its
# uncommitted rows, log loudly, and keep going. The next
# run retries it.
session.rollback()
summary.errors += 1
logger.error("import: error source=%s path=%s — %s", source, rel, e)
if prune:
if limit is not None:
logger.warning("import: --prune ignored because --limit was given")
else:
summary.pruned = _prune(session, source_names, seen)
summary.embed_batches = llm.embed_batches
summary.log()
return summary
finally:
if owns_session:
session.close()
async def _index_file(
session: Session,
*,
source: str,
rel: str,
full_path: Path,
llm: Embedder,
summary: ImportSummary,
raw_date: datetime | None = None,
) -> None:
"""Upsert one file: doc row + chunk rows + embeddings, one transaction.
``raw_date`` (phase 106, D2) is the file's RAW source date — the
git last-commit datetime from the caller's ``doc_dates_by_root``
map, or ``None`` (every non-git case): the file's mtime is read
here, once, and becomes the source date (the D2 fallback).
"""
settings = llm.settings
content = full_path.read_text(encoding="utf-8", errors="replace").replace("\x00", "")
digest = hashlib.sha256(content.encode("utf-8")).hexdigest()
doc = session.scalar(select(Document).where(Document.source == source, Document.path == rel))
if raw_date is None:
# D2 fallback: no source date in the map → the file's mtime
# (one stat). Read before the unchanged early-return — the
# unchanged path refreshes the stored date from the same source.
raw_date = file_mtime_datetime(full_path)
if doc is not None and doc.content_hash == digest:
summary.unchanged += 1
logger.info("import: unchanged source=%s path=%s", source, rel)
# Phase 118 (A2): an unchanged doc whose summary is still NULL
# (a pre-phase-30 row, or an earlier fail-soft miss) gets a
# summary-only backfill — one ``is_summary`` chunk, no content
# re-embed, and NEVER an added/updated/pruned count (so no
# ``sources_meta`` bump, no overview/folder-summary
# regeneration). Strict ``is None``: an empty-string summary is
# owner-set (phase 57) and is never overwritten. BEFORE the
# manual-date early-return: ``created_at_manual`` protects the
# DATE only (phase 106, D1), not the summary.
if doc.summary is None:
await _store_summary(
session, doc=doc, source=source, rel=rel, content=content,
llm=llm, summary=summary, backfill=True,
)
if doc.created_at_manual:
# D1/D4: the owner's correction survives the sync — no
# write at all (the phase-97 ``manually_edited`` precedent).
return
# D4: the date refreshes on every sync, including unchanged
# files, and may go OLDER (no monotonic guard). A date-only
# refresh is still counted ``unchanged`` — never added/updated/
# pruned, so no ``sources_meta`` bump and no regeneration.
target = normalize_doc_date(raw_date)
if target != doc.created_at:
doc.created_at = target
session.commit()
summary.dates_updated += 1
logger.info(
"import: date-refreshed source=%s path=%s date=%s",
source, rel,
doc.created_at.isoformat(),
)
return
verb = "updated" if doc is not None else "added"
# A ``#`` line is a real heading in markdown but a comment in every
# other format — titles for those come from the file stem.
if full_path.suffix.lower() in (".md", ".markdown"):
title = extract_title(content, fallback=full_path.stem)
else:
title = full_path.stem
if doc is None:
doc = Document(
source=source,
path=rel,
full_path=str(full_path),
title=title,
content=content,
content_hash=digest,
indexed_at=datetime.now(UTC),
# Phase 106 (D2/D3): the sourced creation date, normalized
# (undetermined or future → today). ``created_at_manual``
# stays the column default (False) — only the date-edit API
# (task 05) sets it.
created_at=normalize_doc_date(raw_date),
)
session.add(doc)
else:
doc.full_path = str(full_path)
doc.title = title
doc.content = content
doc.content_hash = digest
doc.indexed_at = datetime.now(UTC)
# Phase 106 (D4): a content change is a new document version —
# the date is re-sourced and a previous manual correction is
# reset (it referred to the old content).
doc.created_at = normalize_doc_date(raw_date)
doc.created_at_manual = False
session.flush() # guarantees doc.id even for brand-new rows
# Phase 1+2 — chunk, replace the chunk rows (embeddings NULL), embed,
# and commit the whole file atomically. If the endpoint rejects a chunk
# as over its input token cap (URL-dense paragraphs tokenize at ~1.1
# chars/token), halve this file's chunk target and retry — the global
# policy stays intact for the rest of the KB.
target = max(400, settings.chunk_target_chars)
while True:
chunks_text = chunk_document(content, rel, target, settings.chunk_overlap_chars)
doc.chunks = [
Chunk(document_id=doc.id, position=i, content=c) for i, c in enumerate(chunks_text)
]
session.flush() # delete-orphan cascade drops the previous rows
if not doc.chunks:
break
try:
vectors = await llm.embed([c.content for c in doc.chunks])
for row, vec in zip(doc.chunks, vectors, strict=True):
row.embedding = vec
break
except EmbeddingError as e:
if "token cap" not in str(e) or target <= 400:
raise
logger.info(
"import: re-chunking at %d chars after endpoint token cap: %s",
target // 2,
rel,
)
target //= 2
session.commit()
if verb == "added":
summary.added += 1
else:
summary.updated += 1
summary.chunks += len(chunks_text)
logger.info("import: %s source=%s path=%s chunks=%d", verb, source, rel, len(chunks_text))
# Phase 30; phase 118 (A2, 2026-09-15): EVERY new/changed document
# gets a ``lite``-model summary — markdown included. Phase 30's
# "markdown is already natural language" exclusion is retired: the
# summary is the retrieval seed context (the phase-118 suggestion
# blocks), not a formatting convenience.
await _store_summary(
session, doc=doc, source=source, rel=rel, content=content, llm=llm, summary=summary
)
async def _store_summary(
session: Session,
*,
doc: Document,
source: str,
rel: str,
content: str,
llm: Embedder,
summary: ImportSummary,
backfill: bool = False,
) -> None:
"""Best-effort ``lite`` summary for one already-committed document.
Generates the summary (task 03), stores it on ``documents.summary``
and indexes it as one extra embedded chunk (``is_summary``,
position −1) that hybrid search can hit instead of badly-formatted
raw text. Replacement is guaranteed: any pre-existing ``is_summary``
chunk of this document is deleted first, so at most one summary chunk
exists per document at a time.
Best-effort by contract: the doc row + content chunks are committed
by the caller before this runs, so an :class:`LLMError` /
:class:`EmbeddingError` only rolls back the summary rows — the file
stays indexed, without a summary, and the failure is counted in
``summary_errors`` (PLAN phase 30).
``backfill`` (phase 118, A2): the unchanged-doc NULL-summary path —
a success counts ``summary_backfilled`` instead of ``summaries``
(the doc content is untouched, so the import's KB-change signal must
not move); the rest of the mechanics are identical.
"""
try:
text = await generate_summary(llm, source=source, path=rel, content=content)
# Replacement: at most one summary chunk per document at a time.
# Removing from the collection is what the ``delete-orphan``
# cascade turns into a row delete on flush — and it keeps the
# in-memory collection consistent (this session runs with
# ``expire_on_commit=False``).
for old in [c for c in doc.chunks if c.is_summary]:
doc.chunks.remove(old)
chunk = Chunk(document_id=doc.id, position=-1, content=text, is_summary=True)
vector = (await llm.embed([text]))[0]
chunk.embedding = vector
doc.summary = text
# Append through the relationship (the ``all`` cascade persists the
# row) so the collection — live in this session because of
# ``expire_on_commit=False`` — reflects the committed state.
doc.chunks.append(chunk)
session.commit()
if backfill:
# Phase 118 (A2): the backfill counts itself apart from fresh
# imports — the doc content is unchanged, so ``summaries``
# (a KB-change signal) must not move.
summary.summary_backfilled += 1
logger.info(
"import: summary-backfill source=%s path=%s chars=%d",
source, rel, len(text),
)
else:
summary.summaries += 1
logger.info("import: summary source=%s path=%s chars=%d", source, rel, len(text))
except (LLMError, EmbeddingError) as e:
session.rollback()
summary.summary_errors += 1
logger.error("import: summary failed source=%s path=%s — %s", source, rel, e)
def _prune(session: Session, source_names: set[str], seen: set[tuple[str, str]]) -> int:
"""Delete documents of *source_names* whose file is no longer in *seen*."""
if not source_names:
return 0
pruned = 0
docs = session.scalars(select(Document).where(Document.source.in_(source_names))).all()
for doc in docs:
if (doc.source, doc.path) not in seen:
session.delete(doc)
pruned += 1
logger.info("import: pruned source=%s path=%s", doc.source, doc.path)
if pruned:
session.commit()
return pruned