phase: 122_image_documents
**Phase 122 (image documents) — final verification pass: all green. No code changes were needed; defects found: none.**
**Verified (implementation already complete in working tree, reviewed end-to-end):**
- Toggle (`BOR_IMAGES`/`BOR_IMAGE_EXTENSIONS`/`BOR_IMAGE_DIR`, off by default) + `GET /api/config` `images` flag
- Ingest: bytes digest, `image_dir` persistent copy, `content = summary = vision description` (chat-model call; only text embedded), fail-soft skip + `images_failed` counter
- Serve/display: `/api/documents/{id}/image` route (404 matrix), viewer `<img>` + description, Sources 48px lazy thumbnails, chat inline source figure (alt = summary), agent `read` marker
- Prune guard: images-off syncs never prune `is_image` docs
**Test / lint / coverage (exact commands & outcomes):**
- `uv run pytest` → exit 0 (green; note: pytest 9.1.1 `-q` omits the final count line in output — exit code authoritative)
- `uv run pytest --cov=app --cov-report=term-missing` → **2715 passed, exit 0, TOTAL 99%** (>90% gate)
- `uv run ruff check . && uv run pyright` → "All checks passed!" / "0 errors, 0 warnings, 0 informations"
- `uv run pytest tests/e2e/test_image_documents.py -v --no-cov` → **4 passed, exit 0** (isolation)
**Completion criteria:** (1) images=true → described/embedded/displayed docs: ✅ (E2E + integration) · (2) images=false byte-identical + image docs survive sync: ✅ (E2E negative app + unit/integration) · (3) viewer + chat rendering with alt text; failed description skips + logs, sync completes: ✅ · (4) test/lint/coverage gates: ✅ · (5) commit + phase move: deferred to harness per this pass's rules (working tree left uncommitted).
**Notable deviation (pre-existing, documented in code):** image route uses `require_user` (phase-79 posture, same gate as the document content endpoint) rather than the phase text's "public" parenthetical — matches the endpoint it mirrors.
**Next pending phase:** `123_chat_image_questions`.
This commit is contained in:
+374
-13
@@ -30,7 +30,12 @@ by their exact lowercased full filename (``Dockerfile`` under the
|
||||
no longer exist **or no longer match the format filter** — this is how
|
||||
previously-imported junk (e.g. dot-dir READMEs) leaves the index. Per-file
|
||||
logging uses the verbs ``added | updated | unchanged | pruned`` plus a
|
||||
summary line with per-format counts (PLAN §9).
|
||||
summary line with per-format counts (PLAN §9). Phase 122 prune guard
|
||||
(LOCKED, derived from A3/A4): while the ``images`` toggle is OFF, an
|
||||
``is_image`` doc is INVISIBLE to the walk, not a deleted file — prune
|
||||
skips it (turning the toggle off and syncing must never destroy image
|
||||
documents); a toggle-ON run prunes a deleted image file normally and
|
||||
deletes its ``image_dir`` copy with the row.
|
||||
|
||||
Document dates (phase 106, D2/D4): every import sources
|
||||
``documents.created_at`` from the file's source — the per-file git
|
||||
@@ -54,6 +59,32 @@ backfill runs BEFORE the ``created_at_manual`` early-return (the manual
|
||||
flag protects the DATE only, D1) and the strict ``is None`` check leaves
|
||||
owner-set summaries (even empty strings, phase 57) alone.
|
||||
|
||||
Standalone images (phase 122, LOCKED A3): with the ``images`` toggle
|
||||
(``BOR_IMAGES``) ON, the walk also admits the image extension set
|
||||
(``BOR_IMAGE_EXTENSIONS`` — a SEPARATE set from ``import_extension_set``;
|
||||
images are never user-added via ``BOR_IMPORT_EXTENSIONS``, the toggle is
|
||||
the single knob). Such a file takes the binary index path
|
||||
(:func:`_index_image_file`): the sha256 digest is over the raw BYTES
|
||||
(content identity — the digest rule is unchanged), the bytes are copied
|
||||
to the persistent home ``settings.image_dir/<doc-id>.<ext>`` (dir created
|
||||
on demand; the copy is written ONLY after a successful description, so a
|
||||
failure never leaves an orphan; a changed image deletes the stale copy
|
||||
first; a pruned image doc deletes its copy), the row carries
|
||||
``is_image=True`` + ``image_path``, and ``content`` is the vision
|
||||
description — the ONLY embedded text of the document (the embedding model
|
||||
never sees pixels; ``read_text`` is never called for an image). The
|
||||
description comes through the single seam :func:`_describe_or_skip`
|
||||
(task 03: :func:`app.rag.summarizer.describe_image` — ONE CHAT-model
|
||||
(vision) call with the image bytes as a base64 data URL; the ``lite``
|
||||
summary model is NOT assumed vision-capable, LOCKED A3); a failed/empty
|
||||
description SKIPS the doc entirely (no row, no copy) — counted in
|
||||
``images_failed`` + a warning, the sync continues (fail-soft). The normal
|
||||
chunk pipeline then embeds ``content`` and the phase-30 summary path runs
|
||||
on it — image-aware (task 03): for an image doc the description IS the
|
||||
summary (stored verbatim, no ``lite`` call, no pointer line), so the
|
||||
``is_summary`` position −1 chunk mirrors ``Document.summary``, which
|
||||
equals ``Document.content``.
|
||||
|
||||
``import_sources`` accepts an optional per-file ``progress`` callback
|
||||
(phase 64, task 01) reporting the file being processed right now.
|
||||
"""
|
||||
@@ -61,11 +92,12 @@ from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import logging
|
||||
import uuid
|
||||
from collections.abc import Callable
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Protocol
|
||||
from typing import Any, Protocol
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session
|
||||
@@ -76,7 +108,12 @@ from app.models import Chunk, Document
|
||||
from app.rag.chunker import chunk_document, extract_title
|
||||
from app.rag.doc_dates import file_mtime_datetime, normalize_doc_date
|
||||
from app.rag.llm import EmbeddingError, LLMError
|
||||
from app.rag.summarizer import generate_summary
|
||||
from app.rag.summarizer import (
|
||||
IMAGE_FALLBACK_MIME,
|
||||
IMAGE_MIMES,
|
||||
describe_image,
|
||||
generate_summary,
|
||||
)
|
||||
|
||||
logger = logging.getLogger("app.importer")
|
||||
|
||||
@@ -94,9 +131,12 @@ class Embedder(Protocol):
|
||||
|
||||
async def embed(self, texts: list[str]) -> list[list[float]]: ...
|
||||
|
||||
async def chat(self, messages: list[dict[str, str]], model: str | None = None) -> str: ...
|
||||
# ^ the one-shot completion the summarizer uses for the ``lite`` model
|
||||
# (phase 30, task 01); :class:`app.rag.llm.LLMClient` satisfies it.
|
||||
async def chat(self, messages: list[dict[str, Any]], model: str | None = None) -> str: ...
|
||||
# ^ the one-shot completion the summarizer uses — the ``lite`` model
|
||||
# for text summaries (phase 30, task 01) and the CHAT (vision)
|
||||
# model for the phase-122 image description (multimodal content:
|
||||
# a string or a list of OpenAI-compatible parts); :class:`app.rag.
|
||||
# llm.LLMClient` satisfies it.
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -129,6 +169,12 @@ class ImportSummary:
|
||||
#: so no ``sources_meta`` bump, no overview/folder-summary
|
||||
#: regeneration).
|
||||
dates_updated: int = 0
|
||||
#: Image docs (phase 122, LOCKED A3) whose vision description failed
|
||||
#: or came back empty — the doc is SKIPPED entirely (no row, no
|
||||
#: ``image_dir`` copy): an undescribed image is unsearchable noise.
|
||||
#: Fail-soft: the sync continues, this counter + the warning line
|
||||
#: are the signal.
|
||||
images_failed: int = 0
|
||||
#: Files walked, keyed by lowercased extension (``md``, ``yaml``, …).
|
||||
formats: dict[str, int] = field(default_factory=dict)
|
||||
|
||||
@@ -143,7 +189,7 @@ class ImportSummary:
|
||||
logger.info(
|
||||
"import: summary files=%d added=%d updated=%d unchanged=%d pruned=%d "
|
||||
"errors=%d chunks=%d embed_batches=%d summaries=%d summary_errors=%d "
|
||||
"summary_backfilled=%d dates_updated=%d formats=%s",
|
||||
"summary_backfilled=%d dates_updated=%d images_failed=%d formats=%s",
|
||||
self.files,
|
||||
self.added,
|
||||
self.updated,
|
||||
@@ -156,6 +202,7 @@ class ImportSummary:
|
||||
self.summary_errors,
|
||||
self.summary_backfilled,
|
||||
self.dates_updated,
|
||||
self.images_failed,
|
||||
self.format_counts(),
|
||||
)
|
||||
|
||||
@@ -238,6 +285,7 @@ def iter_importable_files(
|
||||
excluded: frozenset[str] = EXCLUDED_DIRS,
|
||||
ignore: tuple[str, ...] = (),
|
||||
include_hidden: bool = False,
|
||||
image_extensions: frozenset[str] = frozenset(),
|
||||
) -> list[Path]:
|
||||
"""All importable files under *root* (sorted), per the A9 scope rules.
|
||||
|
||||
@@ -255,6 +303,13 @@ def iter_importable_files(
|
||||
source-relative POSIX path starts with any entry; the default ``()``
|
||||
keeps every existing caller byte-identical. The *ignore* tuple
|
||||
composes additively in both states.
|
||||
|
||||
*image_extensions* (phase 122) is the lowercased dotted
|
||||
image-extension set admitted IN ADDITION to *extensions* — passed by
|
||||
:func:`import_sources` only while the ``images`` toggle is on (it
|
||||
reads ``llm.settings``; the image set is never merged into
|
||||
*extensions*). The empty default admits nothing: every existing caller
|
||||
(and the toggle-off walk) stays byte-identical to pre-phase.
|
||||
"""
|
||||
if not root.is_dir():
|
||||
return []
|
||||
@@ -270,7 +325,12 @@ def iter_importable_files(
|
||||
continue
|
||||
if ignore and is_ignored(rel.as_posix(), ignore):
|
||||
continue
|
||||
if match_extension(path, extensions) is None:
|
||||
matched = match_extension(path, extensions)
|
||||
if matched is None and image_extensions:
|
||||
# Phase 122: the image set is admitted IN ADDITION to the
|
||||
# import set (toggle on only — the caller passes it in).
|
||||
matched = match_extension(path, image_extensions)
|
||||
if matched is None:
|
||||
continue
|
||||
files.append(path)
|
||||
return files
|
||||
@@ -340,10 +400,23 @@ async def import_sources(
|
||||
mtime fallback applies to every file — which IS the behavior
|
||||
change, D4: an unchanged file now refreshes its stored date from
|
||||
its source on every run (the backfill-correction case).
|
||||
|
||||
Images (phase 122): when ``llm.settings.images`` is on, BOTH walks
|
||||
(the progress pre-walk and the processing loop — same rules, so
|
||||
``total`` counts images) also admit ``llm.settings.image_extension_set``
|
||||
files, each indexed through the binary image path (see the module
|
||||
docstring). ``prune=True`` with the toggle ON prunes a deleted image
|
||||
file normally (row + ``image_dir`` copy); with the toggle OFF the
|
||||
prune skips ``is_image`` docs (the prune guard — the image is
|
||||
invisible to the walk, not a deleted file).
|
||||
"""
|
||||
if limit is not None and limit <= 0:
|
||||
raise ValueError("limit must be >= 1")
|
||||
summary = ImportSummary()
|
||||
# Phase 122: the image set is admitted by the walks ONLY while the
|
||||
# toggle is on — the empty set admits nothing, so the toggle-off run
|
||||
# (walk, counts, prune) stays byte-identical to pre-phase.
|
||||
image_exts = llm.settings.image_extension_set if llm.settings.images else frozenset()
|
||||
owns_session = session is None
|
||||
if session is None:
|
||||
session = SessionLocal()
|
||||
@@ -364,6 +437,7 @@ async def import_sources(
|
||||
include_hidden=_include_hidden_for_root(
|
||||
root, include_hidden_by_root
|
||||
),
|
||||
image_extensions=image_exts,
|
||||
)
|
||||
)
|
||||
try:
|
||||
@@ -386,6 +460,7 @@ async def import_sources(
|
||||
llm.settings.import_extension_set,
|
||||
ignore=ignore,
|
||||
include_hidden=include_hidden,
|
||||
image_extensions=image_exts,
|
||||
):
|
||||
if limit is not None and summary.files >= limit:
|
||||
break
|
||||
@@ -394,9 +469,11 @@ async def import_sources(
|
||||
summary.files += 1
|
||||
# Phase 102: the matched bare token (``dockerfile`` for an
|
||||
# extensionless ``Dockerfile``), never ``unknown`` — the
|
||||
# file is in scope, so the walk matched it.
|
||||
# file is in scope, so the walk matched it. Phase 122: an
|
||||
# image file matches the image set, not the import set.
|
||||
ext = (
|
||||
match_extension(path, llm.settings.import_extension_set)
|
||||
or match_extension(path, image_exts)
|
||||
or "unknown"
|
||||
)
|
||||
summary.formats[ext] = summary.formats.get(ext, 0) + 1
|
||||
@@ -423,7 +500,9 @@ async def import_sources(
|
||||
if limit is not None:
|
||||
logger.warning("import: --prune ignored because --limit was given")
|
||||
else:
|
||||
summary.pruned = _prune(session, source_names, seen)
|
||||
summary.pruned = _prune(
|
||||
session, source_names, seen, images=llm.settings.images
|
||||
)
|
||||
summary.embed_batches = llm.embed_batches
|
||||
summary.log()
|
||||
return summary
|
||||
@@ -448,8 +527,24 @@ async def _index_file(
|
||||
git last-commit datetime from the caller's ``doc_dates_by_root``
|
||||
map, or ``None`` (every non-git case): the file's mtime is read
|
||||
here, once, and becomes the source date (the D2 fallback).
|
||||
|
||||
Image files (phase 122, toggle on) delegate to
|
||||
:func:`_index_image_file` — the binary path (bytes digest,
|
||||
persistent copy, ``content`` = the vision description) — BEFORE
|
||||
any text read: ``read_text`` is never called for an image.
|
||||
"""
|
||||
settings = llm.settings
|
||||
# Phase 122 (task 02): the image branch FIRST. Only reachable while
|
||||
# the ``images`` toggle is on — the walk never admits image files
|
||||
# while it is off, and with it off this check is a no-op (the text
|
||||
# path below stays byte-identical to pre-phase).
|
||||
if settings.images:
|
||||
image_set = settings.image_extension_set
|
||||
if match_extension(full_path, image_set) is not None:
|
||||
return await _index_image_file(
|
||||
session, source=source, rel=rel, full_path=full_path, llm=llm,
|
||||
summary=summary, raw_date=raw_date,
|
||||
)
|
||||
content = full_path.read_text(encoding="utf-8", errors="replace").replace("\x00", "")
|
||||
digest = hashlib.sha256(content.encode("utf-8")).hexdigest()
|
||||
doc = session.scalar(select(Document).where(Document.source == source, Document.path == rel))
|
||||
@@ -609,9 +704,24 @@ async def _store_summary(
|
||||
a success counts ``summary_backfilled`` instead of ``summaries``
|
||||
(the doc content is untouched, so the import's KB-change signal must
|
||||
not move); the rest of the mechanics are identical.
|
||||
|
||||
Image docs (phase 122, task 03, LOCKED A3): for an ``is_image`` doc
|
||||
the vision description — ``content`` (which equals ``doc.content``
|
||||
on the backfill path) — IS the summary: no ``lite`` call, no
|
||||
pointer line (the summary mirrors the description verbatim, so
|
||||
``doc.summary`` == ``doc.content``). The phase-30 chunk mechanics
|
||||
(one ``is_summary`` position −1 chunk, replacement, best-effort
|
||||
rollback) are unchanged; the only remaining failure class is the
|
||||
summary chunk's embed (the ``doc`` row + content chunks survive —
|
||||
fail-soft, same as the text path).
|
||||
"""
|
||||
try:
|
||||
text = await generate_summary(llm, source=source, path=rel, content=content)
|
||||
if doc.is_image:
|
||||
# Phase 122 (task 03): the description IS the summary —
|
||||
# stored verbatim (no ``lite`` call, no pointer line).
|
||||
text = content
|
||||
else:
|
||||
text = await generate_summary(llm, source=source, path=rel, content=content)
|
||||
# Replacement: at most one summary chunk per document at a time.
|
||||
# Removing from the collection is what the ``delete-orphan``
|
||||
# cascade turns into a row delete on flush — and it keeps the
|
||||
@@ -646,14 +756,265 @@ async def _store_summary(
|
||||
logger.error("import: summary failed source=%s path=%s — %s", source, rel, e)
|
||||
|
||||
|
||||
def _prune(session: Session, source_names: set[str], seen: set[tuple[str, str]]) -> int:
|
||||
"""Delete documents of *source_names* whose file is no longer in *seen*."""
|
||||
async def _describe_or_skip(
|
||||
llm: Embedder, *, data: bytes, source: str, rel: str, full_path: Path
|
||||
) -> str | None:
|
||||
"""Phase 122 — the SINGLE seam for the image description (task 03).
|
||||
|
||||
Returns the vision description that becomes the image document's
|
||||
``content`` (and ``summary`` — the ONLY embedded text of the doc),
|
||||
or ``None`` when the description failed or came back empty — the
|
||||
caller then SKIPS the doc entirely (no row, no copy) and counts
|
||||
``summary.images_failed`` (LOCKED A3, fail-soft: the sync continues,
|
||||
the warning + counter are the signal).
|
||||
|
||||
The seam is one line by design (the importer's tests patch exactly
|
||||
this function): :func:`app.rag.summarizer.describe_image` — ONE
|
||||
CHAT-model (vision) call (LOCKED A3) with the bytes as a data URL
|
||||
whose mime comes from :data:`app.rag.summarizer.IMAGE_MIMES`
|
||||
(dotted extension; an unlisted ``BOR_IMAGE_EXTENSIONS`` token takes
|
||||
the generic fallback — a rejection there fails soft like any other
|
||||
description error). ``source``/``rel`` stay on the signature so the
|
||||
caller (and the patch) reads like the document being described;
|
||||
the failure's doc identity is logged by the caller's warning.
|
||||
"""
|
||||
mime = IMAGE_MIMES.get(full_path.suffix.lower(), IMAGE_FALLBACK_MIME)
|
||||
return await describe_image(llm, data=data, mime=mime)
|
||||
|
||||
|
||||
def _delete_image_copy(image_path: str | None) -> None:
|
||||
"""Best-effort removal of a stale image copy (phase 122).
|
||||
|
||||
A missing path is a no-op (already gone — e.g. the owner cleaned
|
||||
the image dir); an unreadable one is logged, never raised — copy
|
||||
cleanup must not break the sync (the doc row's fate is decided by
|
||||
the upsert/prune logic, not by filesystem hygiene).
|
||||
"""
|
||||
if not image_path:
|
||||
return
|
||||
try:
|
||||
Path(image_path).unlink(missing_ok=True)
|
||||
except OSError as e:
|
||||
logger.warning("import: could not delete image copy %s — %s", image_path, e)
|
||||
|
||||
|
||||
async def _index_image_file(
|
||||
session: Session,
|
||||
*,
|
||||
source: str,
|
||||
rel: str,
|
||||
full_path: Path,
|
||||
llm: Embedder,
|
||||
summary: ImportSummary,
|
||||
raw_date: datetime | None = None,
|
||||
) -> None:
|
||||
"""The phase-122 image branch of :func:`_index_file` — a standalone
|
||||
image is indexed from its BYTES, never its text:
|
||||
|
||||
* the sha256 digest is over the raw bytes (the digest rule is
|
||||
content identity — the same bytes are the same document);
|
||||
* the PERSISTENT copy lands in ``settings.image_dir`` as
|
||||
``<doc-id>.<ext>`` (the dir is created on demand; the copy is
|
||||
written only AFTER a successful description, so a failure never
|
||||
leaves an orphan; a changed image deletes the stale copy before
|
||||
replacing it);
|
||||
* ``content`` is the vision description — the ONLY embedded text of
|
||||
the document (the embedding model never sees pixels) — and the
|
||||
normal chunk pipeline then embeds it, with the phase-30 summary
|
||||
path running on it (the ``is_summary`` position −1 chunk mirrors
|
||||
``Document.summary``).
|
||||
|
||||
``raw_date`` follows the text path exactly (the phase-106 D2
|
||||
fallback: no source date in the map → the file's mtime, read before
|
||||
the unchanged early-return because the unchanged path refreshes the
|
||||
stored date from the same source; the D1 manual-date lock and the
|
||||
D4 refresh apply unmodified).
|
||||
|
||||
Fail-soft (LOCKED A3): a failed/empty description SKIPS the doc
|
||||
entirely (no row, no copy) — ``summary.images_failed`` + a warning,
|
||||
the sync continues.
|
||||
"""
|
||||
settings = llm.settings
|
||||
data = full_path.read_bytes()
|
||||
digest = hashlib.sha256(data).hexdigest()
|
||||
doc = session.scalar(select(Document).where(Document.source == source, Document.path == rel))
|
||||
if raw_date is None:
|
||||
# D2 fallback (same as the text path): no source date in the map
|
||||
# → the file's mtime (one stat).
|
||||
raw_date = file_mtime_datetime(full_path)
|
||||
|
||||
if doc is not None and doc.content_hash == digest:
|
||||
# Unchanged image (byte digest) — the text path's unchanged
|
||||
# branch, unmodified in shape.
|
||||
summary.unchanged += 1
|
||||
logger.info("import: unchanged source=%s path=%s", source, rel)
|
||||
# Phase 118 (A2) backfill, image flavour: an unchanged image doc
|
||||
# whose summary is still NULL (an earlier fail-soft summary miss
|
||||
# — for an image, the one remaining failure class: the summary
|
||||
# chunk's embed) gets the same best-effort summary pass. For an
|
||||
# image the summary IS the stored description (``doc.content``),
|
||||
# so the image-aware ``_store_summary`` (task 03) re-stores it
|
||||
# verbatim with one ``is_summary`` chunk; a failure (the
|
||||
# embed) keeps the doc as-is (no row mutation) — fail-soft,
|
||||
# same as the text path.
|
||||
if doc.summary is None:
|
||||
await _store_summary(
|
||||
session, doc=doc, source=source, rel=rel, content=doc.content,
|
||||
llm=llm, summary=summary, backfill=True,
|
||||
)
|
||||
if doc.created_at_manual:
|
||||
# D1/D4: the owner's correction survives the sync — no write
|
||||
# at all (the text path's manual-date early-return).
|
||||
return
|
||||
# D4: the date refreshes on every sync, including unchanged
|
||||
# files, and may go OLDER (no monotonic guard).
|
||||
target = normalize_doc_date(raw_date)
|
||||
if target != doc.created_at:
|
||||
doc.created_at = target
|
||||
session.commit()
|
||||
summary.dates_updated += 1
|
||||
logger.info(
|
||||
"import: date-refreshed source=%s path=%s date=%s",
|
||||
source, rel, doc.created_at.isoformat(),
|
||||
)
|
||||
return
|
||||
|
||||
verb = "updated" if doc is not None else "added"
|
||||
# The description is the doc's content (task 03: and its summary) —
|
||||
# it is generated BEFORE anything is written, so a failure skips the
|
||||
# doc with no row and no copy (the copy is only made after a
|
||||
# successful description — a failure never leaves an orphan).
|
||||
content = await _describe_or_skip(
|
||||
llm, data=data, source=source, rel=rel, full_path=full_path
|
||||
)
|
||||
if content is None:
|
||||
# LOCKED A3 fail-soft: an undescribed image is unsearchable
|
||||
# noise — skip the doc entirely (no row, no copy).
|
||||
summary.images_failed += 1
|
||||
logger.warning("import: image description failed source=%s path=%s", source, rel)
|
||||
return
|
||||
|
||||
# The persistent copy: uploads are replaced on every upload, git
|
||||
# checkouts are re-cloned, local dirs are user-edited — the served
|
||||
# bytes must outlive the source file. Named by the doc id: a new
|
||||
# doc's id is the uuid4 chosen here (row and copy agree); a changed
|
||||
# doc keeps its id (the copy path is stable).
|
||||
doc_id = doc.id if doc is not None else uuid.uuid4()
|
||||
image_dir = Path(settings.image_dir).expanduser()
|
||||
image_dir.mkdir(parents=True, exist_ok=True)
|
||||
if doc is not None:
|
||||
# A CHANGED image (hash differs): the stale copy is deleted
|
||||
# before replacement.
|
||||
_delete_image_copy(doc.image_path)
|
||||
copy_path = image_dir / f"{doc_id}{full_path.suffix.lower()}"
|
||||
copy_path.write_bytes(data)
|
||||
|
||||
# The non-markdown title rule (the image's content is prose, but the
|
||||
# doc IS the image — the file stem is the title).
|
||||
title = full_path.stem
|
||||
if doc is None:
|
||||
doc = Document(
|
||||
id=doc_id,
|
||||
source=source,
|
||||
path=rel,
|
||||
full_path=str(full_path),
|
||||
title=title,
|
||||
content=content,
|
||||
content_hash=digest,
|
||||
indexed_at=datetime.now(UTC),
|
||||
created_at=normalize_doc_date(raw_date),
|
||||
is_image=True,
|
||||
image_path=str(copy_path),
|
||||
)
|
||||
session.add(doc)
|
||||
else:
|
||||
doc.full_path = str(full_path)
|
||||
doc.title = title
|
||||
doc.content = content
|
||||
doc.content_hash = digest
|
||||
doc.indexed_at = datetime.now(UTC)
|
||||
# Phase 106 (D4): a content change is a new document version —
|
||||
# the date is re-sourced and a previous manual correction is
|
||||
# reset (it referred to the old content).
|
||||
doc.created_at = normalize_doc_date(raw_date)
|
||||
doc.created_at_manual = False
|
||||
doc.is_image = True
|
||||
doc.image_path = str(copy_path)
|
||||
|
||||
session.flush() # guarantees doc.id even for brand-new rows
|
||||
|
||||
# Phase 1+2 — the UNCHANGED pipeline on the description: chunk,
|
||||
# replace the chunk rows (embeddings NULL), embed, and commit the
|
||||
# whole file atomically (one transaction per file). The token-cap
|
||||
# retry loop is copied from the text path; a description is short,
|
||||
# so it never fires in practice.
|
||||
target = max(400, settings.chunk_target_chars)
|
||||
while True:
|
||||
chunks_text = chunk_document(content, rel, target, settings.chunk_overlap_chars)
|
||||
doc.chunks = [
|
||||
Chunk(document_id=doc.id, position=i, content=c) for i, c in enumerate(chunks_text)
|
||||
]
|
||||
session.flush() # delete-orphan cascade drops the previous rows
|
||||
if not doc.chunks:
|
||||
break
|
||||
try:
|
||||
vectors = await llm.embed([c.content for c in doc.chunks])
|
||||
for row, vec in zip(doc.chunks, vectors, strict=True):
|
||||
row.embedding = vec
|
||||
break
|
||||
except EmbeddingError as e:
|
||||
if "token cap" not in str(e) or target <= 400:
|
||||
raise
|
||||
logger.info(
|
||||
"import: re-chunking at %d chars after endpoint token cap: %s",
|
||||
target // 2,
|
||||
rel,
|
||||
)
|
||||
target //= 2
|
||||
session.commit()
|
||||
|
||||
if verb == "added":
|
||||
summary.added += 1
|
||||
else:
|
||||
summary.updated += 1
|
||||
summary.chunks += len(chunks_text)
|
||||
logger.info("import: %s source=%s path=%s chunks=%d", verb, source, rel, len(chunks_text))
|
||||
|
||||
# Phase 30 shape on the description — image-aware (task 03, LOCKED
|
||||
# A3): the description IS the summary (stored verbatim, no
|
||||
# ``lite`` call), so ``doc.summary`` == ``doc.content`` and the
|
||||
# ``is_summary`` position −1 chunk mirrors it.
|
||||
await _store_summary(
|
||||
session, doc=doc, source=source, rel=rel, content=content, llm=llm, summary=summary
|
||||
)
|
||||
|
||||
|
||||
def _prune(
|
||||
session: Session,
|
||||
source_names: set[str],
|
||||
seen: set[tuple[str, str]],
|
||||
images: bool = False,
|
||||
) -> int:
|
||||
"""Delete documents of *source_names* whose file is no longer in *seen*.
|
||||
|
||||
*images* (phase 122 prune guard, LOCKED, derived from A3/A4): while
|
||||
the image toggle is OFF (``images`` False), every ``is_image`` doc is
|
||||
SKIPPED — the image is invisible to an images-off walk, not a deleted
|
||||
file, so pruning it would silently destroy image documents on the
|
||||
first images-off sync. Toggle ON → normal semantics: a deleted image
|
||||
file prunes its doc, and the pruned image's ``image_dir`` copy is
|
||||
deleted with it.
|
||||
"""
|
||||
if not source_names:
|
||||
return 0
|
||||
pruned = 0
|
||||
docs = session.scalars(select(Document).where(Document.source.in_(source_names))).all()
|
||||
for doc in docs:
|
||||
if (doc.source, doc.path) not in seen:
|
||||
if doc.is_image and not images:
|
||||
# Prune guard: invisible to the walk, not deleted.
|
||||
continue
|
||||
_delete_image_copy(doc.image_path)
|
||||
session.delete(doc)
|
||||
pruned += 1
|
||||
logger.info("import: pruned source=%s path=%s", doc.source, doc.path)
|
||||
|
||||
Reference in New Issue
Block a user