**Phase 118 final verification pass — complete.** All criteria verified; 4 pre-existing defects found and fixed.
- **Verified:** summary-seed wiring (`select_suggested` top-5 no-floor → summary blocks, no full text in HIGH prompt), all-doc markdown summaries + NULL backfill (`summary_backfilled`, no `sources_meta` bump), `read` adds full text with `read_docs`-only dedupe, `done.sources` = suggested+read / durable record = suggested+related+read + `suggested=N` log line (seen live in E2E), byte-locked PERSONA/LOW/TOOLS_SECTION, battery gate PASS recorded in `TOOL_CALLING_TESTING.md` §10 (turbo 2026-09-16: 1/2/4 GREEN, cond-3 reported 9/10 per A7, contract 21/21, caps 0).
- **Defects fixed (all pre-existing, none phase-118):** ① `ChatMessage` schema missing the phase-113 `related` key → `extra="forbid"` 422'd every done-time auto-save of grounded turns with a related tier, leaving `message_count=1` (root cause of `test_share_chat` 3F; browser-level instrumentation proved the PUT 422) — added the field + unit/integration pins; ② `test_theme_semantic_completion` pins stale vs phase-117 debox (border/chip removed) — re-targeted to assert border/chip *absence*; ③ `test_header_consistency` `<26`px pin red on 26.125px native date-input line — bound relaxed to `<34` (wrap-detection intent kept); ④ `test_navbar_refresh` bor.chat.v1 key set updated for `related`.
- **Test/lint/coverage:** `uv run pytest --cov=app --cov-report=term-missing` → **2506 passed, app/ 99%** (>90%); `uv run ruff check . && uv run pyright` → clean, 0 errors.
- **E2E:** new story suite in isolation → **2 passed**; full 103-suite matrix sweep (each isolated) → **all 103 green** after the fixes; `test_share_chat` 4 passed, `test_theme_semantic_completion` 8 passed, `test_header_consistency` 3 passed, `test_navbar_refresh` 7 passed.
- **Deviations:** none from LOCKED decisions. Note: orphaned diagnostic uvicorn processes briefly made E2E sessions exercise stale code — killed and re-verified; a sweep-regenerated tracked screenshot was restored. No commits made (harness commits).
- **Completion criteria:** all 7 ✅ (commit/phase-move is the harness's step).
- **Next pending phase:** none — `todo/` holds only this phase's overview pending the harness move.
663 lines
29 KiB
Python
663 lines
29 KiB
Python
"""Knowledge-base importer (PLAN §5 / §9 / §11).
|
||
|
||
Walks the in-scope files (the A9 family by default — the original seven
|
||
plus the quadlet family and ``j2`` — ``BOR_IMPORT_EXTENSIONS``, which may
|
||
name any well-formed extension; case-insensitive), diffs by sha256
|
||
against ``documents.content_hash`` and, for every new or changed file, runs
|
||
the two-phase upsert:
|
||
|
||
1. upsert the document row and replace its chunk rows (embeddings NULL)
|
||
2. embed the new chunks in batches and attach the vectors
|
||
3. commit — one transaction per file, so a failed embedding leaves the
|
||
database untouched and the file is simply retried on the next run
|
||
4. every file (phase 30; phase 118, A2: markdown included — the
|
||
non-markdown-only scope is retired): generate a ``lite``-model summary
|
||
and, best-effort, store it on ``documents.summary`` plus one extra
|
||
embedded chunk (``is_summary``, position −1). The document row and its
|
||
content chunks are already committed at this point, so a summary
|
||
failure only means the file is indexed without a summary (logged and
|
||
counted in ``summary_errors``) — it is never lost.
|
||
|
||
Scope (A9, revised 2026-08-21): any path containing a dot-prefixed
|
||
component (hidden dirs — vendored caches like ``.esphome/.espressif/**`` —
|
||
or hidden files) is skipped, plus the well-known exclusion list — UNLESS
|
||
the source's phase-105 hidden-folders flag admits dot-prefixed paths; the
|
||
exclusion list always applies. A token may also name extensionless files
|
||
by their exact lowercased full filename (``Dockerfile`` under the
|
||
``dockerfile`` token — phase 102).
|
||
|
||
``prune=True`` deletes documents (of the imported sources only) whose files
|
||
no longer exist **or no longer match the format filter** — this is how
|
||
previously-imported junk (e.g. dot-dir READMEs) leaves the index. Per-file
|
||
logging uses the verbs ``added | updated | unchanged | pruned`` plus a
|
||
summary line with per-format counts (PLAN §9).
|
||
|
||
Document dates (phase 106, D2/D4): every import sources
|
||
``documents.created_at`` from the file's source — the per-file git
|
||
last-commit date when a ``doc_dates_by_root`` entry names the file,
|
||
else the file's mtime — normalized by
|
||
:func:`app.rag.doc_dates.normalize_doc_date` (undetermined or future →
|
||
today, D3) on every add and update. On the unchanged path the stored
|
||
date is REFRESHED from the same source (it may go OLDER — no monotonic
|
||
guard) and counted in ``summary.dates_updated`` — unless the row
|
||
carries the owner's manual correction (``created_at_manual``, D1), which
|
||
the sync never touches.
|
||
|
||
NULL-summary backfill (phase 118, A2): an UNCHANGED file (same
|
||
``content_hash``) whose ``documents.summary`` is still NULL — a
|
||
pre-phase-30 row, or an earlier fail-soft miss — gets the same
|
||
best-effort summary pass on every sync until it sticks. A success counts
|
||
``summary_backfilled`` (never ``summaries``) and touches nothing else: no
|
||
content re-embed, no added/updated/pruned count — so no
|
||
``sources_meta`` bump, no KB-overview/folder-summary regeneration. The
|
||
backfill runs BEFORE the ``created_at_manual`` early-return (the manual
|
||
flag protects the DATE only, D1) and the strict ``is None`` check leaves
|
||
owner-set summaries (even empty strings, phase 57) alone.
|
||
|
||
``import_sources`` accepts an optional per-file ``progress`` callback
|
||
(phase 64, task 01) reporting the file being processed right now.
|
||
"""
|
||
from __future__ import annotations
|
||
|
||
import hashlib
|
||
import logging
|
||
from collections.abc import Callable
|
||
from dataclasses import dataclass, field
|
||
from datetime import UTC, datetime
|
||
from pathlib import Path
|
||
from typing import Protocol
|
||
|
||
from sqlalchemy import select
|
||
from sqlalchemy.orm import Session
|
||
|
||
from app.config import Settings
|
||
from app.db import SessionLocal
|
||
from app.models import Chunk, Document
|
||
from app.rag.chunker import chunk_document, extract_title
|
||
from app.rag.doc_dates import file_mtime_datetime, normalize_doc_date
|
||
from app.rag.llm import EmbeddingError, LLMError
|
||
from app.rag.summarizer import generate_summary
|
||
|
||
logger = logging.getLogger("app.importer")
|
||
|
||
#: Non-content directories never imported (PLAN anchor A9).
|
||
EXCLUDED_DIRS: frozenset[str] = frozenset(
|
||
{".venv", "node_modules", ".git", "__pycache__", ".pytest_cache", "dist", "build"}
|
||
)
|
||
|
||
|
||
class Embedder(Protocol):
|
||
"""Everything the importer needs from the LLM client (duck-typed for tests)."""
|
||
|
||
settings: Settings
|
||
embed_batches: int
|
||
|
||
async def embed(self, texts: list[str]) -> list[list[float]]: ...
|
||
|
||
async def chat(self, messages: list[dict[str, str]], model: str | None = None) -> str: ...
|
||
# ^ the one-shot completion the summarizer uses for the ``lite`` model
|
||
# (phase 30, task 01); :class:`app.rag.llm.LLMClient` satisfies it.
|
||
|
||
|
||
@dataclass
|
||
class ImportSummary:
|
||
"""Counts for one import run (also printed by the CLI)."""
|
||
|
||
files: int = 0
|
||
added: int = 0
|
||
updated: int = 0
|
||
unchanged: int = 0
|
||
pruned: int = 0
|
||
errors: int = 0
|
||
chunks: int = 0
|
||
embed_batches: int = 0
|
||
#: Files whose lite summary was generated + indexed (phase 30; phase
|
||
#: 118, A2: every A9 format, markdown included). One ``is_summary``
|
||
#: chunk per success.
|
||
summaries: int = 0
|
||
#: Files whose summary generation failed (best-effort — the document
|
||
#: is still indexed, without a summary).
|
||
summary_errors: int = 0
|
||
#: Unchanged docs whose NULL summary was backfilled (phase 118, A2) —
|
||
#: one ``is_summary`` chunk per success; the content is untouched, so
|
||
#: a backfill NEVER counts added/updated/pruned (no
|
||
#: ``sources_meta`` bump, no overview/folder-summary regeneration).
|
||
summary_backfilled: int = 0
|
||
#: Files whose ``created_at`` was refreshed on the UNCHANGED path —
|
||
#: content untouched, date re-sourced (phase 106, D4: the date may
|
||
#: go OLDER; a date-only refresh NEVER counts added/updated/pruned,
|
||
#: so no ``sources_meta`` bump, no overview/folder-summary
|
||
#: regeneration).
|
||
dates_updated: int = 0
|
||
#: Files walked, keyed by lowercased extension (``md``, ``yaml``, …).
|
||
formats: dict[str, int] = field(default_factory=dict)
|
||
|
||
def format_counts(self) -> str:
|
||
"""``md:203,yaml:267,py:14`` — highest count first (PLAN §9)."""
|
||
if not self.formats:
|
||
return "none"
|
||
ordered = sorted(self.formats.items(), key=lambda kv: (-kv[1], kv[0]))
|
||
return ",".join(f"{ext}:{count}" for ext, count in ordered)
|
||
|
||
def log(self) -> None:
|
||
logger.info(
|
||
"import: summary files=%d added=%d updated=%d unchanged=%d pruned=%d "
|
||
"errors=%d chunks=%d embed_batches=%d summaries=%d summary_errors=%d "
|
||
"summary_backfilled=%d dates_updated=%d formats=%s",
|
||
self.files,
|
||
self.added,
|
||
self.updated,
|
||
self.unchanged,
|
||
self.pruned,
|
||
self.errors,
|
||
self.chunks,
|
||
self.embed_batches,
|
||
self.summaries,
|
||
self.summary_errors,
|
||
self.summary_backfilled,
|
||
self.dates_updated,
|
||
self.format_counts(),
|
||
)
|
||
|
||
|
||
def normalize_ignore_path(entry: str) -> str:
|
||
"""One ignore-path entry → canonical form (phase 89, A1).
|
||
|
||
Trim surrounding whitespace, then strip ALL leading/trailing
|
||
``/`` — so ``"/my/files/"``, ``"my/files/"`` and ``"my/files"``
|
||
all become ``"my/files"``. ``""`` / ``"//"``,
|
||
``" "`` normalize to ``""`` (callers drop empties).
|
||
"""
|
||
return entry.strip().strip("/")
|
||
|
||
|
||
def is_ignored(rel: str, prefixes: tuple[str, ...]) -> bool:
|
||
"""Phase 89, A1 — the pure prefix rule, nothing else.
|
||
|
||
``rel`` is the source-relative POSIX path WITHOUT a leading
|
||
slash (the ``documents.path`` string). Match = ``rel`` STARTS
|
||
WITH a normalized entry: raw string prefix — deliberately NO
|
||
component-boundary check (``"my/files"`` also matches
|
||
``"my/files2/x.md"``) and NO mid-path matching (``"myfile.txt"``
|
||
matches ``"myfile.txt"`` but not ``"some/path/myfile.txt"``).
|
||
"""
|
||
return any(rel.startswith(p) for p in prefixes)
|
||
|
||
|
||
def _ignore_for_root(
|
||
root: Path, ignore_by_root: dict[str, list[str]] | None
|
||
) -> tuple[str, ...]:
|
||
"""The normalized, non-empty prefix tuple for one root (phase 89).
|
||
|
||
Keyed by ``str(root)`` — the root string exactly as the caller
|
||
passed it in ``sources`` (unambiguous when two rows share a
|
||
source *name* but different dirs). Callers may pass RAW box
|
||
lines: the importer normalizes + drops empties here, the single
|
||
choke point — stored lists (already normalized) normalize to
|
||
themselves.
|
||
"""
|
||
raw = (ignore_by_root or {}).get(str(root)) or []
|
||
return tuple(p for p in (normalize_ignore_path(e) for e in raw) if p)
|
||
|
||
|
||
def _include_hidden_for_root(
|
||
root: Path, include_hidden_by_root: dict[str, bool] | None
|
||
) -> bool:
|
||
"""The per-root hidden-folders flag (phase 105, A1).
|
||
|
||
Keyed by ``str(root)`` — the root string exactly as the caller
|
||
passed it in ``sources`` (the ``_ignore_for_root`` convention,
|
||
phase 89): ``True`` only for roots the caller lists as True;
|
||
unlisted/``None`` roots are ``False`` — every existing caller
|
||
behaves byte-identically (A4).
|
||
"""
|
||
return bool((include_hidden_by_root or {}).get(str(root), False))
|
||
|
||
|
||
def match_extension(path: Path, extensions: frozenset[str]) -> str | None:
|
||
"""The bare lowercased token *path* imports under, or ``None``.
|
||
|
||
1. Non-empty lowercased dotted suffix in *extensions* (the A9 rule —
|
||
``kubernetes.md`` → ``md``).
|
||
2. No suffix: the lowercased FULL filename equals a bare token of
|
||
*extensions* (``Dockerfile`` → ``dockerfile``) — the phase-102
|
||
extensionless rule. Exact name only: ``mydockerfile`` never
|
||
matches the ``dockerfile`` token.
|
||
"""
|
||
suffix = path.suffix.lower()
|
||
if suffix and suffix in extensions:
|
||
return suffix.lstrip(".")
|
||
if not suffix and path.name.lower() in {e.lstrip(".") for e in extensions}:
|
||
return path.name.lower()
|
||
return None
|
||
|
||
|
||
def iter_importable_files(
|
||
root: Path,
|
||
extensions: frozenset[str],
|
||
excluded: frozenset[str] = EXCLUDED_DIRS,
|
||
ignore: tuple[str, ...] = (),
|
||
include_hidden: bool = False,
|
||
) -> list[Path]:
|
||
"""All importable files under *root* (sorted), per the A9 scope rules.
|
||
|
||
*extensions* is a set of lowercased dotted suffixes (``{'.md', '.py'}``).
|
||
Skips: when ``include_hidden`` is False (the default), any path with a
|
||
dot-prefixed component (hidden dirs/files — vendored caches like
|
||
``.esphome/.espressif/**``); when True, dot-prefixed components are
|
||
ADMITTED (files inside hidden folders, and hidden files) and only
|
||
*excluded* is consulted (A1 — caches/VCS internals are never content).
|
||
The well-known non-content directories in *excluded* are skipped in
|
||
BOTH states. *ignore* (phase 89, A1) is a
|
||
tuple of ALREADY-normalized, non-empty source-relative path prefixes
|
||
(the importer's ``_ignore_for_root`` is the normalization choke point
|
||
— raw box lines never reach this function): a file is skipped when its
|
||
source-relative POSIX path starts with any entry; the default ``()``
|
||
keeps every existing caller byte-identical. The *ignore* tuple
|
||
composes additively in both states.
|
||
"""
|
||
if not root.is_dir():
|
||
return []
|
||
files: list[Path] = []
|
||
for path in sorted(root.rglob("*")):
|
||
if not path.is_file():
|
||
continue
|
||
rel = path.relative_to(root)
|
||
if any(
|
||
(not include_hidden and part.startswith(".")) or part in excluded
|
||
for part in rel.parts
|
||
):
|
||
continue
|
||
if ignore and is_ignored(rel.as_posix(), ignore):
|
||
continue
|
||
if match_extension(path, extensions) is None:
|
||
continue
|
||
files.append(path)
|
||
return files
|
||
|
||
|
||
async def import_sources(
|
||
sources: list[Path],
|
||
llm: Embedder,
|
||
*,
|
||
prune: bool = False,
|
||
limit: int | None = None,
|
||
session: Session | None = None,
|
||
progress: Callable[[str, str, int, int], None] | None = None,
|
||
ignore_by_root: dict[str, list[str]] | None = None,
|
||
include_hidden_by_root: dict[str, bool] | None = None,
|
||
doc_dates_by_root: dict[str, dict[str, datetime]] | None = None,
|
||
) -> ImportSummary:
|
||
"""Import every A9-format file under *sources* (see module docstring).
|
||
|
||
``session`` may be supplied (tests); a private one is opened and closed
|
||
otherwise. ``limit`` caps the number of files processed (debug only) and
|
||
disables pruning, since an incomplete walk must not drive deletions.
|
||
|
||
``progress`` (phase 64, task 01) is an optional per-file hook called
|
||
once per importable file, immediately before that file's
|
||
``_index_file`` — with ``(source, rel_posix_path, done, total)``: the
|
||
same POSIX *rel* the document rows use, ``done`` = the 1-based index of
|
||
the current file **across all sources**, and ``total`` = the combined
|
||
pre-walk count of importable files across all *sources* roots. The
|
||
pre-walk (same extension/exclusion rules, directory stats only, no file
|
||
reads) happens **only when *progress* is provided**: callers passing
|
||
nothing pay no extra walk and behave exactly as before. Under
|
||
``limit``, the hook still fires per processed file only — ``done`` never
|
||
exceeds the limit, but ``total`` stays the full pre-walk count (an
|
||
incomplete walk must not misreport the denominator).
|
||
|
||
``ignore_by_root`` (phase 89, A1/A2) maps ``str(root)`` — the root path
|
||
string exactly as passed in *sources* — to that source's RAW ignore-path
|
||
lines (the importer normalizes them via ``_ignore_for_root``, the single
|
||
choke point): matching files are never walked, so they are never
|
||
embedded and never summarized, and the progress pre-walk uses the same
|
||
per-root tuple as the processing loop, so ``total`` never counts them.
|
||
A file that newly matches a pattern simply never enters ``seen``, so the
|
||
next ``prune=True`` run deletes its row automatically (A2 — the A9
|
||
junk-precedent). ``None`` (the default) changes nothing: the map is read
|
||
per root, unlisted roots get an empty tuple, and every existing caller
|
||
behaves byte-identically.
|
||
|
||
``include_hidden_by_root`` (phase 105, A1) maps ``str(root)`` to the
|
||
stored flag: ``True`` admits dot-prefixed components for that root
|
||
(``EXCLUDED_DIRS`` and the extension filter still apply; the ignore
|
||
tuple composes additively). Unlisted/``None`` roots are ``False`` —
|
||
byte-identical to pre-phase-105. A file that was indexed with the
|
||
flag ON and is walked again with it OFF simply never enters
|
||
``seen``, so the next ``prune=True`` run deletes its row
|
||
automatically (A2 — the A9/phase-89 precedent).
|
||
|
||
``doc_dates_by_root`` (phase 106, D2/D4) maps ``str(root)`` — the
|
||
root path string exactly as passed in *sources* — to that source's
|
||
RAW per-file source dates: source-relative POSIX path → the git
|
||
last-commit datetime (task 03's ``file_commit_dates``). ONLY git
|
||
roots are listed — unlisted roots (local dirs, unpacked uploads)
|
||
take the mtime fallback, and a path missing from its root's map
|
||
does too. The map entry beats the file's mtime when present. The
|
||
progress pre-walk is untouched (dates change no file count).
|
||
``None`` (the default) changes nothing for existing callers: the
|
||
mtime fallback applies to every file — which IS the behavior
|
||
change, D4: an unchanged file now refreshes its stored date from
|
||
its source on every run (the backfill-correction case).
|
||
"""
|
||
if limit is not None and limit <= 0:
|
||
raise ValueError("limit must be >= 1")
|
||
summary = ImportSummary()
|
||
owns_session = session is None
|
||
if session is None:
|
||
session = SessionLocal()
|
||
seen: set[tuple[str, str]] = set()
|
||
source_names: set[str] = set()
|
||
# phase 64 (task 01): the hook's combined denominator, walked with the
|
||
# exact same rules as the processing loop below (directory stats only,
|
||
# no file reads). Skipped entirely for ``progress=None`` callers — no
|
||
# extra pass, byte-identical behaviour and cost.
|
||
total = 0
|
||
if progress is not None:
|
||
for root in sources:
|
||
total += len(
|
||
iter_importable_files(
|
||
root,
|
||
llm.settings.import_extension_set,
|
||
ignore=_ignore_for_root(root, ignore_by_root),
|
||
include_hidden=_include_hidden_for_root(
|
||
root, include_hidden_by_root
|
||
),
|
||
)
|
||
)
|
||
try:
|
||
for root in sources:
|
||
if not root.is_dir():
|
||
logger.warning("import: source dir not found, skipping: %s", root)
|
||
continue
|
||
if limit is not None and summary.files >= limit:
|
||
break
|
||
source = root.name
|
||
source_names.add(source)
|
||
ignore = _ignore_for_root(root, ignore_by_root)
|
||
include_hidden = _include_hidden_for_root(root, include_hidden_by_root)
|
||
# Phase 106 (D2): the root's raw source dates (git last-commit
|
||
# for git roots, keyed by the same str(root) convention); {}
|
||
# for unlisted roots — every file then takes the mtime fallback.
|
||
dates_map = (doc_dates_by_root or {}).get(str(root), {})
|
||
for path in iter_importable_files(
|
||
root,
|
||
llm.settings.import_extension_set,
|
||
ignore=ignore,
|
||
include_hidden=include_hidden,
|
||
):
|
||
if limit is not None and summary.files >= limit:
|
||
break
|
||
rel = path.relative_to(root).as_posix()
|
||
seen.add((source, rel))
|
||
summary.files += 1
|
||
# Phase 102: the matched bare token (``dockerfile`` for an
|
||
# extensionless ``Dockerfile``), never ``unknown`` — the
|
||
# file is in scope, so the walk matched it.
|
||
ext = (
|
||
match_extension(path, llm.settings.import_extension_set)
|
||
or "unknown"
|
||
)
|
||
summary.formats[ext] = summary.formats.get(ext, 0) + 1
|
||
if progress is not None:
|
||
# phase 64: report the file *before* indexing it — a
|
||
# file that then errors or turns out unchanged was
|
||
# already the "current file". No try/except around the
|
||
# call: the hooks in this repo only assign fields.
|
||
progress(source, rel, summary.files, total)
|
||
try:
|
||
await _index_file(
|
||
session, source=source, rel=rel, full_path=path, llm=llm,
|
||
summary=summary, raw_date=dates_map.get(rel),
|
||
)
|
||
except EmbeddingError as e:
|
||
# A pathological file (e.g. content the embedding endpoint
|
||
# refuses) must not abort the whole KB: roll back its
|
||
# uncommitted rows, log loudly, and keep going. The next
|
||
# run retries it.
|
||
session.rollback()
|
||
summary.errors += 1
|
||
logger.error("import: error source=%s path=%s — %s", source, rel, e)
|
||
if prune:
|
||
if limit is not None:
|
||
logger.warning("import: --prune ignored because --limit was given")
|
||
else:
|
||
summary.pruned = _prune(session, source_names, seen)
|
||
summary.embed_batches = llm.embed_batches
|
||
summary.log()
|
||
return summary
|
||
finally:
|
||
if owns_session:
|
||
session.close()
|
||
|
||
|
||
async def _index_file(
|
||
session: Session,
|
||
*,
|
||
source: str,
|
||
rel: str,
|
||
full_path: Path,
|
||
llm: Embedder,
|
||
summary: ImportSummary,
|
||
raw_date: datetime | None = None,
|
||
) -> None:
|
||
"""Upsert one file: doc row + chunk rows + embeddings, one transaction.
|
||
|
||
``raw_date`` (phase 106, D2) is the file's RAW source date — the
|
||
git last-commit datetime from the caller's ``doc_dates_by_root``
|
||
map, or ``None`` (every non-git case): the file's mtime is read
|
||
here, once, and becomes the source date (the D2 fallback).
|
||
"""
|
||
settings = llm.settings
|
||
content = full_path.read_text(encoding="utf-8", errors="replace").replace("\x00", "")
|
||
digest = hashlib.sha256(content.encode("utf-8")).hexdigest()
|
||
doc = session.scalar(select(Document).where(Document.source == source, Document.path == rel))
|
||
if raw_date is None:
|
||
# D2 fallback: no source date in the map → the file's mtime
|
||
# (one stat). Read before the unchanged early-return — the
|
||
# unchanged path refreshes the stored date from the same source.
|
||
raw_date = file_mtime_datetime(full_path)
|
||
if doc is not None and doc.content_hash == digest:
|
||
summary.unchanged += 1
|
||
logger.info("import: unchanged source=%s path=%s", source, rel)
|
||
# Phase 118 (A2): an unchanged doc whose summary is still NULL
|
||
# (a pre-phase-30 row, or an earlier fail-soft miss) gets a
|
||
# summary-only backfill — one ``is_summary`` chunk, no content
|
||
# re-embed, and NEVER an added/updated/pruned count (so no
|
||
# ``sources_meta`` bump, no overview/folder-summary
|
||
# regeneration). Strict ``is None``: an empty-string summary is
|
||
# owner-set (phase 57) and is never overwritten. BEFORE the
|
||
# manual-date early-return: ``created_at_manual`` protects the
|
||
# DATE only (phase 106, D1), not the summary.
|
||
if doc.summary is None:
|
||
await _store_summary(
|
||
session, doc=doc, source=source, rel=rel, content=content,
|
||
llm=llm, summary=summary, backfill=True,
|
||
)
|
||
if doc.created_at_manual:
|
||
# D1/D4: the owner's correction survives the sync — no
|
||
# write at all (the phase-97 ``manually_edited`` precedent).
|
||
return
|
||
# D4: the date refreshes on every sync, including unchanged
|
||
# files, and may go OLDER (no monotonic guard). A date-only
|
||
# refresh is still counted ``unchanged`` — never added/updated/
|
||
# pruned, so no ``sources_meta`` bump and no regeneration.
|
||
target = normalize_doc_date(raw_date)
|
||
if target != doc.created_at:
|
||
doc.created_at = target
|
||
session.commit()
|
||
summary.dates_updated += 1
|
||
logger.info(
|
||
"import: date-refreshed source=%s path=%s date=%s",
|
||
source, rel,
|
||
doc.created_at.isoformat(),
|
||
)
|
||
return
|
||
|
||
verb = "updated" if doc is not None else "added"
|
||
# A ``#`` line is a real heading in markdown but a comment in every
|
||
# other format — titles for those come from the file stem.
|
||
if full_path.suffix.lower() in (".md", ".markdown"):
|
||
title = extract_title(content, fallback=full_path.stem)
|
||
else:
|
||
title = full_path.stem
|
||
if doc is None:
|
||
doc = Document(
|
||
source=source,
|
||
path=rel,
|
||
full_path=str(full_path),
|
||
title=title,
|
||
content=content,
|
||
content_hash=digest,
|
||
indexed_at=datetime.now(UTC),
|
||
# Phase 106 (D2/D3): the sourced creation date, normalized
|
||
# (undetermined or future → today). ``created_at_manual``
|
||
# stays the column default (False) — only the date-edit API
|
||
# (task 05) sets it.
|
||
created_at=normalize_doc_date(raw_date),
|
||
)
|
||
session.add(doc)
|
||
else:
|
||
doc.full_path = str(full_path)
|
||
doc.title = title
|
||
doc.content = content
|
||
doc.content_hash = digest
|
||
doc.indexed_at = datetime.now(UTC)
|
||
# Phase 106 (D4): a content change is a new document version —
|
||
# the date is re-sourced and a previous manual correction is
|
||
# reset (it referred to the old content).
|
||
doc.created_at = normalize_doc_date(raw_date)
|
||
doc.created_at_manual = False
|
||
|
||
session.flush() # guarantees doc.id even for brand-new rows
|
||
|
||
# Phase 1+2 — chunk, replace the chunk rows (embeddings NULL), embed,
|
||
# and commit the whole file atomically. If the endpoint rejects a chunk
|
||
# as over its input token cap (URL-dense paragraphs tokenize at ~1.1
|
||
# chars/token), halve this file's chunk target and retry — the global
|
||
# policy stays intact for the rest of the KB.
|
||
target = max(400, settings.chunk_target_chars)
|
||
while True:
|
||
chunks_text = chunk_document(content, rel, target, settings.chunk_overlap_chars)
|
||
doc.chunks = [
|
||
Chunk(document_id=doc.id, position=i, content=c) for i, c in enumerate(chunks_text)
|
||
]
|
||
session.flush() # delete-orphan cascade drops the previous rows
|
||
if not doc.chunks:
|
||
break
|
||
try:
|
||
vectors = await llm.embed([c.content for c in doc.chunks])
|
||
for row, vec in zip(doc.chunks, vectors, strict=True):
|
||
row.embedding = vec
|
||
break
|
||
except EmbeddingError as e:
|
||
if "token cap" not in str(e) or target <= 400:
|
||
raise
|
||
logger.info(
|
||
"import: re-chunking at %d chars after endpoint token cap: %s",
|
||
target // 2,
|
||
rel,
|
||
)
|
||
target //= 2
|
||
session.commit()
|
||
|
||
if verb == "added":
|
||
summary.added += 1
|
||
else:
|
||
summary.updated += 1
|
||
summary.chunks += len(chunks_text)
|
||
logger.info("import: %s source=%s path=%s chunks=%d", verb, source, rel, len(chunks_text))
|
||
|
||
# Phase 30; phase 118 (A2, 2026-09-15): EVERY new/changed document
|
||
# gets a ``lite``-model summary — markdown included. Phase 30's
|
||
# "markdown is already natural language" exclusion is retired: the
|
||
# summary is the retrieval seed context (the phase-118 suggestion
|
||
# blocks), not a formatting convenience.
|
||
await _store_summary(
|
||
session, doc=doc, source=source, rel=rel, content=content, llm=llm, summary=summary
|
||
)
|
||
|
||
|
||
async def _store_summary(
|
||
session: Session,
|
||
*,
|
||
doc: Document,
|
||
source: str,
|
||
rel: str,
|
||
content: str,
|
||
llm: Embedder,
|
||
summary: ImportSummary,
|
||
backfill: bool = False,
|
||
) -> None:
|
||
"""Best-effort ``lite`` summary for one already-committed document.
|
||
|
||
Generates the summary (task 03), stores it on ``documents.summary``
|
||
and indexes it as one extra embedded chunk (``is_summary``,
|
||
position −1) that hybrid search can hit instead of badly-formatted
|
||
raw text. Replacement is guaranteed: any pre-existing ``is_summary``
|
||
chunk of this document is deleted first, so at most one summary chunk
|
||
exists per document at a time.
|
||
|
||
Best-effort by contract: the doc row + content chunks are committed
|
||
by the caller before this runs, so an :class:`LLMError` /
|
||
:class:`EmbeddingError` only rolls back the summary rows — the file
|
||
stays indexed, without a summary, and the failure is counted in
|
||
``summary_errors`` (PLAN phase 30).
|
||
|
||
``backfill`` (phase 118, A2): the unchanged-doc NULL-summary path —
|
||
a success counts ``summary_backfilled`` instead of ``summaries``
|
||
(the doc content is untouched, so the import's KB-change signal must
|
||
not move); the rest of the mechanics are identical.
|
||
"""
|
||
try:
|
||
text = await generate_summary(llm, source=source, path=rel, content=content)
|
||
# Replacement: at most one summary chunk per document at a time.
|
||
# Removing from the collection is what the ``delete-orphan``
|
||
# cascade turns into a row delete on flush — and it keeps the
|
||
# in-memory collection consistent (this session runs with
|
||
# ``expire_on_commit=False``).
|
||
for old in [c for c in doc.chunks if c.is_summary]:
|
||
doc.chunks.remove(old)
|
||
chunk = Chunk(document_id=doc.id, position=-1, content=text, is_summary=True)
|
||
vector = (await llm.embed([text]))[0]
|
||
chunk.embedding = vector
|
||
doc.summary = text
|
||
# Append through the relationship (the ``all`` cascade persists the
|
||
# row) so the collection — live in this session because of
|
||
# ``expire_on_commit=False`` — reflects the committed state.
|
||
doc.chunks.append(chunk)
|
||
session.commit()
|
||
if backfill:
|
||
# Phase 118 (A2): the backfill counts itself apart from fresh
|
||
# imports — the doc content is unchanged, so ``summaries``
|
||
# (a KB-change signal) must not move.
|
||
summary.summary_backfilled += 1
|
||
logger.info(
|
||
"import: summary-backfill source=%s path=%s chars=%d",
|
||
source, rel, len(text),
|
||
)
|
||
else:
|
||
summary.summaries += 1
|
||
logger.info("import: summary source=%s path=%s chars=%d", source, rel, len(text))
|
||
except (LLMError, EmbeddingError) as e:
|
||
session.rollback()
|
||
summary.summary_errors += 1
|
||
logger.error("import: summary failed source=%s path=%s — %s", source, rel, e)
|
||
|
||
|
||
def _prune(session: Session, source_names: set[str], seen: set[tuple[str, str]]) -> int:
|
||
"""Delete documents of *source_names* whose file is no longer in *seen*."""
|
||
if not source_names:
|
||
return 0
|
||
pruned = 0
|
||
docs = session.scalars(select(Document).where(Document.source.in_(source_names))).all()
|
||
for doc in docs:
|
||
if (doc.source, doc.path) not in seen:
|
||
session.delete(doc)
|
||
pruned += 1
|
||
logger.info("import: pruned source=%s path=%s", doc.source, doc.path)
|
||
if pruned:
|
||
session.commit()
|
||
return pruned
|