phase: 122_image_documents
**Phase 122 (image documents) — final verification pass: all green. No code changes were needed; defects found: none.**
**Verified (implementation already complete in working tree, reviewed end-to-end):**
- Toggle (`BOR_IMAGES`/`BOR_IMAGE_EXTENSIONS`/`BOR_IMAGE_DIR`, off by default) + `GET /api/config` `images` flag
- Ingest: bytes digest, `image_dir` persistent copy, `content = summary = vision description` (chat-model call; only text embedded), fail-soft skip + `images_failed` counter
- Serve/display: `/api/documents/{id}/image` route (404 matrix), viewer `<img>` + description, Sources 48px lazy thumbnails, chat inline source figure (alt = summary), agent `read` marker
- Prune guard: images-off syncs never prune `is_image` docs
**Test / lint / coverage (exact commands & outcomes):**
- `uv run pytest` → exit 0 (green; note: pytest 9.1.1 `-q` omits the final count line in output — exit code authoritative)
- `uv run pytest --cov=app --cov-report=term-missing` → **2715 passed, exit 0, TOTAL 99%** (>90% gate)
- `uv run ruff check . && uv run pyright` → "All checks passed!" / "0 errors, 0 warnings, 0 informations"
- `uv run pytest tests/e2e/test_image_documents.py -v --no-cov` → **4 passed, exit 0** (isolation)
**Completion criteria:** (1) images=true → described/embedded/displayed docs: ✅ (E2E + integration) · (2) images=false byte-identical + image docs survive sync: ✅ (E2E negative app + unit/integration) · (3) viewer + chat rendering with alt text; failed description skips + logs, sync completes: ✅ · (4) test/lint/coverage gates: ✅ · (5) commit + phase move: deferred to harness per this pass's rules (working tree left uncommitted).
**Notable deviation (pre-existing, documented in code):** image route uses `require_user` (phase-79 posture, same gate as the document content endpoint) rather than the phase text's "public" parenthetical — matches the endpoint it mirrors.
**Next pending phase:** `123_chat_image_questions`.
This commit is contained in:
+23
-2
@@ -427,6 +427,18 @@ READ_TRUNCATION_NOTICE = (
|
||||
"document."
|
||||
)
|
||||
|
||||
#: The image-document marker (phase 122, task 05): the line prefixed to
|
||||
#: the vision DESCRIPTION a ``read`` of a standalone-image document
|
||||
#: returns — the model must reason about what it is reading (the text
|
||||
#: below is a description GENERATED from the image, not the image's
|
||||
#: own words). It sits on the result's THIRD line: the ``Document …``
|
||||
#: header and the phase-106 date line stay byte-identical (the E2E
|
||||
#: mock's ``_READ_RESULT_PREFIX`` header contract), and a NON-image
|
||||
#: doc's result carries no marker at all (byte-identical to pre-122).
|
||||
IMAGE_DOC_MARKER = (
|
||||
"Image document — the text below is a description generated from the image:"
|
||||
)
|
||||
|
||||
#: The no-source ``ls`` refusal with the teaching parenthetical
|
||||
#: appended (phase 72): used when a stripped scope has no ``/`` and
|
||||
#: matches no registered source (the incident's ``ls(path='.')``). The
|
||||
@@ -1270,6 +1282,12 @@ def _execute_tool(
|
||||
return _no_document_refusal(db, arg)
|
||||
holder.read_docs.append(doc)
|
||||
holder.tool_calls += 1
|
||||
# Phase 122 (task 05): an image document's content IS the vision
|
||||
# description — the marker line (a third line between the
|
||||
# byte-identical header/date lines and the text) tells the model
|
||||
# what it is reading. A text doc's ``marker`` is "" — the result
|
||||
# stays byte-identical to pre-122.
|
||||
marker = f"{IMAGE_DOC_MARKER}\n" if doc.is_image else ""
|
||||
cap = settings.read_max_chars
|
||||
if len(doc.content) > cap:
|
||||
# Phase 95 (owner permission 2026-09-10, ``TODO.md`` L5): the
|
||||
@@ -1293,17 +1311,20 @@ def _execute_tool(
|
||||
return (
|
||||
f"Document {doc.source}/{doc.path}:\n"
|
||||
f"date: {doc.created_at:%Y-%m-%d}\n"
|
||||
f"{marker}"
|
||||
f"{doc.content[:cap]}\n"
|
||||
f"{TRUNCATION_MARKER}\n"
|
||||
f"{READ_TRUNCATION_NOTICE.format(shown=cap, total=len(doc.content))}"
|
||||
)
|
||||
# At or under the cap: the pre-phase-95 result plus the
|
||||
# phase-106 D5 date line (first line byte-identical — the
|
||||
# mock's header contract; no marker, no notice, no holder
|
||||
# entry, no ToolResultPiece).
|
||||
# mock's header contract; no truncation marker, no notice, no
|
||||
# holder entry, no ToolResultPiece) — and, phase 122, the
|
||||
# image-document marker line for image docs only.
|
||||
return (
|
||||
f"Document {doc.source}/{doc.path}:\n"
|
||||
f"date: {doc.created_at:%Y-%m-%d}\n"
|
||||
f"{marker}"
|
||||
f"{doc.content}"
|
||||
)
|
||||
if call.name == "grep":
|
||||
|
||||
+374
-13
@@ -30,7 +30,12 @@ by their exact lowercased full filename (``Dockerfile`` under the
|
||||
no longer exist **or no longer match the format filter** — this is how
|
||||
previously-imported junk (e.g. dot-dir READMEs) leaves the index. Per-file
|
||||
logging uses the verbs ``added | updated | unchanged | pruned`` plus a
|
||||
summary line with per-format counts (PLAN §9).
|
||||
summary line with per-format counts (PLAN §9). Phase 122 prune guard
|
||||
(LOCKED, derived from A3/A4): while the ``images`` toggle is OFF, an
|
||||
``is_image`` doc is INVISIBLE to the walk, not a deleted file — prune
|
||||
skips it (turning the toggle off and syncing must never destroy image
|
||||
documents); a toggle-ON run prunes a deleted image file normally and
|
||||
deletes its ``image_dir`` copy with the row.
|
||||
|
||||
Document dates (phase 106, D2/D4): every import sources
|
||||
``documents.created_at`` from the file's source — the per-file git
|
||||
@@ -54,6 +59,32 @@ backfill runs BEFORE the ``created_at_manual`` early-return (the manual
|
||||
flag protects the DATE only, D1) and the strict ``is None`` check leaves
|
||||
owner-set summaries (even empty strings, phase 57) alone.
|
||||
|
||||
Standalone images (phase 122, LOCKED A3): with the ``images`` toggle
|
||||
(``BOR_IMAGES``) ON, the walk also admits the image extension set
|
||||
(``BOR_IMAGE_EXTENSIONS`` — a SEPARATE set from ``import_extension_set``;
|
||||
images are never user-added via ``BOR_IMPORT_EXTENSIONS``, the toggle is
|
||||
the single knob). Such a file takes the binary index path
|
||||
(:func:`_index_image_file`): the sha256 digest is over the raw BYTES
|
||||
(content identity — the digest rule is unchanged), the bytes are copied
|
||||
to the persistent home ``settings.image_dir/<doc-id>.<ext>`` (dir created
|
||||
on demand; the copy is written ONLY after a successful description, so a
|
||||
failure never leaves an orphan; a changed image deletes the stale copy
|
||||
first; a pruned image doc deletes its copy), the row carries
|
||||
``is_image=True`` + ``image_path``, and ``content`` is the vision
|
||||
description — the ONLY embedded text of the document (the embedding model
|
||||
never sees pixels; ``read_text`` is never called for an image). The
|
||||
description comes through the single seam :func:`_describe_or_skip`
|
||||
(task 03: :func:`app.rag.summarizer.describe_image` — ONE CHAT-model
|
||||
(vision) call with the image bytes as a base64 data URL; the ``lite``
|
||||
summary model is NOT assumed vision-capable, LOCKED A3); a failed/empty
|
||||
description SKIPS the doc entirely (no row, no copy) — counted in
|
||||
``images_failed`` + a warning, the sync continues (fail-soft). The normal
|
||||
chunk pipeline then embeds ``content`` and the phase-30 summary path runs
|
||||
on it — image-aware (task 03): for an image doc the description IS the
|
||||
summary (stored verbatim, no ``lite`` call, no pointer line), so the
|
||||
``is_summary`` position −1 chunk mirrors ``Document.summary``, which
|
||||
equals ``Document.content``.
|
||||
|
||||
``import_sources`` accepts an optional per-file ``progress`` callback
|
||||
(phase 64, task 01) reporting the file being processed right now.
|
||||
"""
|
||||
@@ -61,11 +92,12 @@ from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import logging
|
||||
import uuid
|
||||
from collections.abc import Callable
|
||||
from dataclasses import dataclass, field
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Protocol
|
||||
from typing import Any, Protocol
|
||||
|
||||
from sqlalchemy import select
|
||||
from sqlalchemy.orm import Session
|
||||
@@ -76,7 +108,12 @@ from app.models import Chunk, Document
|
||||
from app.rag.chunker import chunk_document, extract_title
|
||||
from app.rag.doc_dates import file_mtime_datetime, normalize_doc_date
|
||||
from app.rag.llm import EmbeddingError, LLMError
|
||||
from app.rag.summarizer import generate_summary
|
||||
from app.rag.summarizer import (
|
||||
IMAGE_FALLBACK_MIME,
|
||||
IMAGE_MIMES,
|
||||
describe_image,
|
||||
generate_summary,
|
||||
)
|
||||
|
||||
logger = logging.getLogger("app.importer")
|
||||
|
||||
@@ -94,9 +131,12 @@ class Embedder(Protocol):
|
||||
|
||||
async def embed(self, texts: list[str]) -> list[list[float]]: ...
|
||||
|
||||
async def chat(self, messages: list[dict[str, str]], model: str | None = None) -> str: ...
|
||||
# ^ the one-shot completion the summarizer uses for the ``lite`` model
|
||||
# (phase 30, task 01); :class:`app.rag.llm.LLMClient` satisfies it.
|
||||
async def chat(self, messages: list[dict[str, Any]], model: str | None = None) -> str: ...
|
||||
# ^ the one-shot completion the summarizer uses — the ``lite`` model
|
||||
# for text summaries (phase 30, task 01) and the CHAT (vision)
|
||||
# model for the phase-122 image description (multimodal content:
|
||||
# a string or a list of OpenAI-compatible parts); :class:`app.rag.
|
||||
# llm.LLMClient` satisfies it.
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -129,6 +169,12 @@ class ImportSummary:
|
||||
#: so no ``sources_meta`` bump, no overview/folder-summary
|
||||
#: regeneration).
|
||||
dates_updated: int = 0
|
||||
#: Image docs (phase 122, LOCKED A3) whose vision description failed
|
||||
#: or came back empty — the doc is SKIPPED entirely (no row, no
|
||||
#: ``image_dir`` copy): an undescribed image is unsearchable noise.
|
||||
#: Fail-soft: the sync continues, this counter + the warning line
|
||||
#: are the signal.
|
||||
images_failed: int = 0
|
||||
#: Files walked, keyed by lowercased extension (``md``, ``yaml``, …).
|
||||
formats: dict[str, int] = field(default_factory=dict)
|
||||
|
||||
@@ -143,7 +189,7 @@ class ImportSummary:
|
||||
logger.info(
|
||||
"import: summary files=%d added=%d updated=%d unchanged=%d pruned=%d "
|
||||
"errors=%d chunks=%d embed_batches=%d summaries=%d summary_errors=%d "
|
||||
"summary_backfilled=%d dates_updated=%d formats=%s",
|
||||
"summary_backfilled=%d dates_updated=%d images_failed=%d formats=%s",
|
||||
self.files,
|
||||
self.added,
|
||||
self.updated,
|
||||
@@ -156,6 +202,7 @@ class ImportSummary:
|
||||
self.summary_errors,
|
||||
self.summary_backfilled,
|
||||
self.dates_updated,
|
||||
self.images_failed,
|
||||
self.format_counts(),
|
||||
)
|
||||
|
||||
@@ -238,6 +285,7 @@ def iter_importable_files(
|
||||
excluded: frozenset[str] = EXCLUDED_DIRS,
|
||||
ignore: tuple[str, ...] = (),
|
||||
include_hidden: bool = False,
|
||||
image_extensions: frozenset[str] = frozenset(),
|
||||
) -> list[Path]:
|
||||
"""All importable files under *root* (sorted), per the A9 scope rules.
|
||||
|
||||
@@ -255,6 +303,13 @@ def iter_importable_files(
|
||||
source-relative POSIX path starts with any entry; the default ``()``
|
||||
keeps every existing caller byte-identical. The *ignore* tuple
|
||||
composes additively in both states.
|
||||
|
||||
*image_extensions* (phase 122) is the lowercased dotted
|
||||
image-extension set admitted IN ADDITION to *extensions* — passed by
|
||||
:func:`import_sources` only while the ``images`` toggle is on (it
|
||||
reads ``llm.settings``; the image set is never merged into
|
||||
*extensions*). The empty default admits nothing: every existing caller
|
||||
(and the toggle-off walk) stays byte-identical to pre-phase.
|
||||
"""
|
||||
if not root.is_dir():
|
||||
return []
|
||||
@@ -270,7 +325,12 @@ def iter_importable_files(
|
||||
continue
|
||||
if ignore and is_ignored(rel.as_posix(), ignore):
|
||||
continue
|
||||
if match_extension(path, extensions) is None:
|
||||
matched = match_extension(path, extensions)
|
||||
if matched is None and image_extensions:
|
||||
# Phase 122: the image set is admitted IN ADDITION to the
|
||||
# import set (toggle on only — the caller passes it in).
|
||||
matched = match_extension(path, image_extensions)
|
||||
if matched is None:
|
||||
continue
|
||||
files.append(path)
|
||||
return files
|
||||
@@ -340,10 +400,23 @@ async def import_sources(
|
||||
mtime fallback applies to every file — which IS the behavior
|
||||
change, D4: an unchanged file now refreshes its stored date from
|
||||
its source on every run (the backfill-correction case).
|
||||
|
||||
Images (phase 122): when ``llm.settings.images`` is on, BOTH walks
|
||||
(the progress pre-walk and the processing loop — same rules, so
|
||||
``total`` counts images) also admit ``llm.settings.image_extension_set``
|
||||
files, each indexed through the binary image path (see the module
|
||||
docstring). ``prune=True`` with the toggle ON prunes a deleted image
|
||||
file normally (row + ``image_dir`` copy); with the toggle OFF the
|
||||
prune skips ``is_image`` docs (the prune guard — the image is
|
||||
invisible to the walk, not a deleted file).
|
||||
"""
|
||||
if limit is not None and limit <= 0:
|
||||
raise ValueError("limit must be >= 1")
|
||||
summary = ImportSummary()
|
||||
# Phase 122: the image set is admitted by the walks ONLY while the
|
||||
# toggle is on — the empty set admits nothing, so the toggle-off run
|
||||
# (walk, counts, prune) stays byte-identical to pre-phase.
|
||||
image_exts = llm.settings.image_extension_set if llm.settings.images else frozenset()
|
||||
owns_session = session is None
|
||||
if session is None:
|
||||
session = SessionLocal()
|
||||
@@ -364,6 +437,7 @@ async def import_sources(
|
||||
include_hidden=_include_hidden_for_root(
|
||||
root, include_hidden_by_root
|
||||
),
|
||||
image_extensions=image_exts,
|
||||
)
|
||||
)
|
||||
try:
|
||||
@@ -386,6 +460,7 @@ async def import_sources(
|
||||
llm.settings.import_extension_set,
|
||||
ignore=ignore,
|
||||
include_hidden=include_hidden,
|
||||
image_extensions=image_exts,
|
||||
):
|
||||
if limit is not None and summary.files >= limit:
|
||||
break
|
||||
@@ -394,9 +469,11 @@ async def import_sources(
|
||||
summary.files += 1
|
||||
# Phase 102: the matched bare token (``dockerfile`` for an
|
||||
# extensionless ``Dockerfile``), never ``unknown`` — the
|
||||
# file is in scope, so the walk matched it.
|
||||
# file is in scope, so the walk matched it. Phase 122: an
|
||||
# image file matches the image set, not the import set.
|
||||
ext = (
|
||||
match_extension(path, llm.settings.import_extension_set)
|
||||
or match_extension(path, image_exts)
|
||||
or "unknown"
|
||||
)
|
||||
summary.formats[ext] = summary.formats.get(ext, 0) + 1
|
||||
@@ -423,7 +500,9 @@ async def import_sources(
|
||||
if limit is not None:
|
||||
logger.warning("import: --prune ignored because --limit was given")
|
||||
else:
|
||||
summary.pruned = _prune(session, source_names, seen)
|
||||
summary.pruned = _prune(
|
||||
session, source_names, seen, images=llm.settings.images
|
||||
)
|
||||
summary.embed_batches = llm.embed_batches
|
||||
summary.log()
|
||||
return summary
|
||||
@@ -448,8 +527,24 @@ async def _index_file(
|
||||
git last-commit datetime from the caller's ``doc_dates_by_root``
|
||||
map, or ``None`` (every non-git case): the file's mtime is read
|
||||
here, once, and becomes the source date (the D2 fallback).
|
||||
|
||||
Image files (phase 122, toggle on) delegate to
|
||||
:func:`_index_image_file` — the binary path (bytes digest,
|
||||
persistent copy, ``content`` = the vision description) — BEFORE
|
||||
any text read: ``read_text`` is never called for an image.
|
||||
"""
|
||||
settings = llm.settings
|
||||
# Phase 122 (task 02): the image branch FIRST. Only reachable while
|
||||
# the ``images`` toggle is on — the walk never admits image files
|
||||
# while it is off, and with it off this check is a no-op (the text
|
||||
# path below stays byte-identical to pre-phase).
|
||||
if settings.images:
|
||||
image_set = settings.image_extension_set
|
||||
if match_extension(full_path, image_set) is not None:
|
||||
return await _index_image_file(
|
||||
session, source=source, rel=rel, full_path=full_path, llm=llm,
|
||||
summary=summary, raw_date=raw_date,
|
||||
)
|
||||
content = full_path.read_text(encoding="utf-8", errors="replace").replace("\x00", "")
|
||||
digest = hashlib.sha256(content.encode("utf-8")).hexdigest()
|
||||
doc = session.scalar(select(Document).where(Document.source == source, Document.path == rel))
|
||||
@@ -609,9 +704,24 @@ async def _store_summary(
|
||||
a success counts ``summary_backfilled`` instead of ``summaries``
|
||||
(the doc content is untouched, so the import's KB-change signal must
|
||||
not move); the rest of the mechanics are identical.
|
||||
|
||||
Image docs (phase 122, task 03, LOCKED A3): for an ``is_image`` doc
|
||||
the vision description — ``content`` (which equals ``doc.content``
|
||||
on the backfill path) — IS the summary: no ``lite`` call, no
|
||||
pointer line (the summary mirrors the description verbatim, so
|
||||
``doc.summary`` == ``doc.content``). The phase-30 chunk mechanics
|
||||
(one ``is_summary`` position −1 chunk, replacement, best-effort
|
||||
rollback) are unchanged; the only remaining failure class is the
|
||||
summary chunk's embed (the ``doc`` row + content chunks survive —
|
||||
fail-soft, same as the text path).
|
||||
"""
|
||||
try:
|
||||
text = await generate_summary(llm, source=source, path=rel, content=content)
|
||||
if doc.is_image:
|
||||
# Phase 122 (task 03): the description IS the summary —
|
||||
# stored verbatim (no ``lite`` call, no pointer line).
|
||||
text = content
|
||||
else:
|
||||
text = await generate_summary(llm, source=source, path=rel, content=content)
|
||||
# Replacement: at most one summary chunk per document at a time.
|
||||
# Removing from the collection is what the ``delete-orphan``
|
||||
# cascade turns into a row delete on flush — and it keeps the
|
||||
@@ -646,14 +756,265 @@ async def _store_summary(
|
||||
logger.error("import: summary failed source=%s path=%s — %s", source, rel, e)
|
||||
|
||||
|
||||
def _prune(session: Session, source_names: set[str], seen: set[tuple[str, str]]) -> int:
|
||||
"""Delete documents of *source_names* whose file is no longer in *seen*."""
|
||||
async def _describe_or_skip(
|
||||
llm: Embedder, *, data: bytes, source: str, rel: str, full_path: Path
|
||||
) -> str | None:
|
||||
"""Phase 122 — the SINGLE seam for the image description (task 03).
|
||||
|
||||
Returns the vision description that becomes the image document's
|
||||
``content`` (and ``summary`` — the ONLY embedded text of the doc),
|
||||
or ``None`` when the description failed or came back empty — the
|
||||
caller then SKIPS the doc entirely (no row, no copy) and counts
|
||||
``summary.images_failed`` (LOCKED A3, fail-soft: the sync continues,
|
||||
the warning + counter are the signal).
|
||||
|
||||
The seam is one line by design (the importer's tests patch exactly
|
||||
this function): :func:`app.rag.summarizer.describe_image` — ONE
|
||||
CHAT-model (vision) call (LOCKED A3) with the bytes as a data URL
|
||||
whose mime comes from :data:`app.rag.summarizer.IMAGE_MIMES`
|
||||
(dotted extension; an unlisted ``BOR_IMAGE_EXTENSIONS`` token takes
|
||||
the generic fallback — a rejection there fails soft like any other
|
||||
description error). ``source``/``rel`` stay on the signature so the
|
||||
caller (and the patch) reads like the document being described;
|
||||
the failure's doc identity is logged by the caller's warning.
|
||||
"""
|
||||
mime = IMAGE_MIMES.get(full_path.suffix.lower(), IMAGE_FALLBACK_MIME)
|
||||
return await describe_image(llm, data=data, mime=mime)
|
||||
|
||||
|
||||
def _delete_image_copy(image_path: str | None) -> None:
|
||||
"""Best-effort removal of a stale image copy (phase 122).
|
||||
|
||||
A missing path is a no-op (already gone — e.g. the owner cleaned
|
||||
the image dir); an unreadable one is logged, never raised — copy
|
||||
cleanup must not break the sync (the doc row's fate is decided by
|
||||
the upsert/prune logic, not by filesystem hygiene).
|
||||
"""
|
||||
if not image_path:
|
||||
return
|
||||
try:
|
||||
Path(image_path).unlink(missing_ok=True)
|
||||
except OSError as e:
|
||||
logger.warning("import: could not delete image copy %s — %s", image_path, e)
|
||||
|
||||
|
||||
async def _index_image_file(
|
||||
session: Session,
|
||||
*,
|
||||
source: str,
|
||||
rel: str,
|
||||
full_path: Path,
|
||||
llm: Embedder,
|
||||
summary: ImportSummary,
|
||||
raw_date: datetime | None = None,
|
||||
) -> None:
|
||||
"""The phase-122 image branch of :func:`_index_file` — a standalone
|
||||
image is indexed from its BYTES, never its text:
|
||||
|
||||
* the sha256 digest is over the raw bytes (the digest rule is
|
||||
content identity — the same bytes are the same document);
|
||||
* the PERSISTENT copy lands in ``settings.image_dir`` as
|
||||
``<doc-id>.<ext>`` (the dir is created on demand; the copy is
|
||||
written only AFTER a successful description, so a failure never
|
||||
leaves an orphan; a changed image deletes the stale copy before
|
||||
replacing it);
|
||||
* ``content`` is the vision description — the ONLY embedded text of
|
||||
the document (the embedding model never sees pixels) — and the
|
||||
normal chunk pipeline then embeds it, with the phase-30 summary
|
||||
path running on it (the ``is_summary`` position −1 chunk mirrors
|
||||
``Document.summary``).
|
||||
|
||||
``raw_date`` follows the text path exactly (the phase-106 D2
|
||||
fallback: no source date in the map → the file's mtime, read before
|
||||
the unchanged early-return because the unchanged path refreshes the
|
||||
stored date from the same source; the D1 manual-date lock and the
|
||||
D4 refresh apply unmodified).
|
||||
|
||||
Fail-soft (LOCKED A3): a failed/empty description SKIPS the doc
|
||||
entirely (no row, no copy) — ``summary.images_failed`` + a warning,
|
||||
the sync continues.
|
||||
"""
|
||||
settings = llm.settings
|
||||
data = full_path.read_bytes()
|
||||
digest = hashlib.sha256(data).hexdigest()
|
||||
doc = session.scalar(select(Document).where(Document.source == source, Document.path == rel))
|
||||
if raw_date is None:
|
||||
# D2 fallback (same as the text path): no source date in the map
|
||||
# → the file's mtime (one stat).
|
||||
raw_date = file_mtime_datetime(full_path)
|
||||
|
||||
if doc is not None and doc.content_hash == digest:
|
||||
# Unchanged image (byte digest) — the text path's unchanged
|
||||
# branch, unmodified in shape.
|
||||
summary.unchanged += 1
|
||||
logger.info("import: unchanged source=%s path=%s", source, rel)
|
||||
# Phase 118 (A2) backfill, image flavour: an unchanged image doc
|
||||
# whose summary is still NULL (an earlier fail-soft summary miss
|
||||
# — for an image, the one remaining failure class: the summary
|
||||
# chunk's embed) gets the same best-effort summary pass. For an
|
||||
# image the summary IS the stored description (``doc.content``),
|
||||
# so the image-aware ``_store_summary`` (task 03) re-stores it
|
||||
# verbatim with one ``is_summary`` chunk; a failure (the
|
||||
# embed) keeps the doc as-is (no row mutation) — fail-soft,
|
||||
# same as the text path.
|
||||
if doc.summary is None:
|
||||
await _store_summary(
|
||||
session, doc=doc, source=source, rel=rel, content=doc.content,
|
||||
llm=llm, summary=summary, backfill=True,
|
||||
)
|
||||
if doc.created_at_manual:
|
||||
# D1/D4: the owner's correction survives the sync — no write
|
||||
# at all (the text path's manual-date early-return).
|
||||
return
|
||||
# D4: the date refreshes on every sync, including unchanged
|
||||
# files, and may go OLDER (no monotonic guard).
|
||||
target = normalize_doc_date(raw_date)
|
||||
if target != doc.created_at:
|
||||
doc.created_at = target
|
||||
session.commit()
|
||||
summary.dates_updated += 1
|
||||
logger.info(
|
||||
"import: date-refreshed source=%s path=%s date=%s",
|
||||
source, rel, doc.created_at.isoformat(),
|
||||
)
|
||||
return
|
||||
|
||||
verb = "updated" if doc is not None else "added"
|
||||
# The description is the doc's content (task 03: and its summary) —
|
||||
# it is generated BEFORE anything is written, so a failure skips the
|
||||
# doc with no row and no copy (the copy is only made after a
|
||||
# successful description — a failure never leaves an orphan).
|
||||
content = await _describe_or_skip(
|
||||
llm, data=data, source=source, rel=rel, full_path=full_path
|
||||
)
|
||||
if content is None:
|
||||
# LOCKED A3 fail-soft: an undescribed image is unsearchable
|
||||
# noise — skip the doc entirely (no row, no copy).
|
||||
summary.images_failed += 1
|
||||
logger.warning("import: image description failed source=%s path=%s", source, rel)
|
||||
return
|
||||
|
||||
# The persistent copy: uploads are replaced on every upload, git
|
||||
# checkouts are re-cloned, local dirs are user-edited — the served
|
||||
# bytes must outlive the source file. Named by the doc id: a new
|
||||
# doc's id is the uuid4 chosen here (row and copy agree); a changed
|
||||
# doc keeps its id (the copy path is stable).
|
||||
doc_id = doc.id if doc is not None else uuid.uuid4()
|
||||
image_dir = Path(settings.image_dir).expanduser()
|
||||
image_dir.mkdir(parents=True, exist_ok=True)
|
||||
if doc is not None:
|
||||
# A CHANGED image (hash differs): the stale copy is deleted
|
||||
# before replacement.
|
||||
_delete_image_copy(doc.image_path)
|
||||
copy_path = image_dir / f"{doc_id}{full_path.suffix.lower()}"
|
||||
copy_path.write_bytes(data)
|
||||
|
||||
# The non-markdown title rule (the image's content is prose, but the
|
||||
# doc IS the image — the file stem is the title).
|
||||
title = full_path.stem
|
||||
if doc is None:
|
||||
doc = Document(
|
||||
id=doc_id,
|
||||
source=source,
|
||||
path=rel,
|
||||
full_path=str(full_path),
|
||||
title=title,
|
||||
content=content,
|
||||
content_hash=digest,
|
||||
indexed_at=datetime.now(UTC),
|
||||
created_at=normalize_doc_date(raw_date),
|
||||
is_image=True,
|
||||
image_path=str(copy_path),
|
||||
)
|
||||
session.add(doc)
|
||||
else:
|
||||
doc.full_path = str(full_path)
|
||||
doc.title = title
|
||||
doc.content = content
|
||||
doc.content_hash = digest
|
||||
doc.indexed_at = datetime.now(UTC)
|
||||
# Phase 106 (D4): a content change is a new document version —
|
||||
# the date is re-sourced and a previous manual correction is
|
||||
# reset (it referred to the old content).
|
||||
doc.created_at = normalize_doc_date(raw_date)
|
||||
doc.created_at_manual = False
|
||||
doc.is_image = True
|
||||
doc.image_path = str(copy_path)
|
||||
|
||||
session.flush() # guarantees doc.id even for brand-new rows
|
||||
|
||||
# Phase 1+2 — the UNCHANGED pipeline on the description: chunk,
|
||||
# replace the chunk rows (embeddings NULL), embed, and commit the
|
||||
# whole file atomically (one transaction per file). The token-cap
|
||||
# retry loop is copied from the text path; a description is short,
|
||||
# so it never fires in practice.
|
||||
target = max(400, settings.chunk_target_chars)
|
||||
while True:
|
||||
chunks_text = chunk_document(content, rel, target, settings.chunk_overlap_chars)
|
||||
doc.chunks = [
|
||||
Chunk(document_id=doc.id, position=i, content=c) for i, c in enumerate(chunks_text)
|
||||
]
|
||||
session.flush() # delete-orphan cascade drops the previous rows
|
||||
if not doc.chunks:
|
||||
break
|
||||
try:
|
||||
vectors = await llm.embed([c.content for c in doc.chunks])
|
||||
for row, vec in zip(doc.chunks, vectors, strict=True):
|
||||
row.embedding = vec
|
||||
break
|
||||
except EmbeddingError as e:
|
||||
if "token cap" not in str(e) or target <= 400:
|
||||
raise
|
||||
logger.info(
|
||||
"import: re-chunking at %d chars after endpoint token cap: %s",
|
||||
target // 2,
|
||||
rel,
|
||||
)
|
||||
target //= 2
|
||||
session.commit()
|
||||
|
||||
if verb == "added":
|
||||
summary.added += 1
|
||||
else:
|
||||
summary.updated += 1
|
||||
summary.chunks += len(chunks_text)
|
||||
logger.info("import: %s source=%s path=%s chunks=%d", verb, source, rel, len(chunks_text))
|
||||
|
||||
# Phase 30 shape on the description — image-aware (task 03, LOCKED
|
||||
# A3): the description IS the summary (stored verbatim, no
|
||||
# ``lite`` call), so ``doc.summary`` == ``doc.content`` and the
|
||||
# ``is_summary`` position −1 chunk mirrors it.
|
||||
await _store_summary(
|
||||
session, doc=doc, source=source, rel=rel, content=content, llm=llm, summary=summary
|
||||
)
|
||||
|
||||
|
||||
def _prune(
|
||||
session: Session,
|
||||
source_names: set[str],
|
||||
seen: set[tuple[str, str]],
|
||||
images: bool = False,
|
||||
) -> int:
|
||||
"""Delete documents of *source_names* whose file is no longer in *seen*.
|
||||
|
||||
*images* (phase 122 prune guard, LOCKED, derived from A3/A4): while
|
||||
the image toggle is OFF (``images`` False), every ``is_image`` doc is
|
||||
SKIPPED — the image is invisible to an images-off walk, not a deleted
|
||||
file, so pruning it would silently destroy image documents on the
|
||||
first images-off sync. Toggle ON → normal semantics: a deleted image
|
||||
file prunes its doc, and the pruned image's ``image_dir`` copy is
|
||||
deleted with it.
|
||||
"""
|
||||
if not source_names:
|
||||
return 0
|
||||
pruned = 0
|
||||
docs = session.scalars(select(Document).where(Document.source.in_(source_names))).all()
|
||||
for doc in docs:
|
||||
if (doc.source, doc.path) not in seen:
|
||||
if doc.is_image and not images:
|
||||
# Prune guard: invisible to the walk, not deleted.
|
||||
continue
|
||||
_delete_image_copy(doc.image_path)
|
||||
session.delete(doc)
|
||||
pruned += 1
|
||||
logger.info("import: pruned source=%s path=%s", doc.source, doc.path)
|
||||
|
||||
@@ -92,6 +92,7 @@ from sqlalchemy.orm import Session
|
||||
|
||||
from app.config import get_settings
|
||||
from app.models import Chunk, Document
|
||||
from app.schemas import SourceRef
|
||||
|
||||
#: Shared overflow marker (phase 15; imported by ``app.rag.prompts``)
|
||||
#: — used by the steering (<tuning>) section, the phase-118 NULL-summary
|
||||
@@ -939,3 +940,27 @@ def select_related(
|
||||
continue
|
||||
out.append(doc)
|
||||
return out
|
||||
|
||||
|
||||
def source_ref_with_image(doc: Document) -> SourceRef:
|
||||
"""One :class:`~app.schemas.SourceRef` wire frame for *doc* — the
|
||||
phase-122 (task 05) SHARED frame builder: the chat API's cited
|
||||
tier (the agent's read docs) and the related tier both run through
|
||||
it, so the per-doc ref shape has exactly one construction site.
|
||||
|
||||
The ref carries the chip identity (``source`` / ``path`` /
|
||||
``title``) and, ONLY for a standalone-image document (``is_image``),
|
||||
the optional ``image_url`` — the image BYTES route
|
||||
``/api/documents/<id>/image`` (the chat's sources block renders the
|
||||
compact inline image from it, the summary as alt + caption —
|
||||
"shown in the chat nicely", TODO L6). For a TEXT document the field
|
||||
stays ``None`` and is DROPPED by the model's serializer (never
|
||||
``null`` — the omission rule): a text-doc frame is byte-identical
|
||||
to pre-phase. The frame's doc id rides the path — the same way the
|
||||
document content endpoint's ``(source, path)`` lookup does (no new
|
||||
id leak beyond what the frame already carries).
|
||||
"""
|
||||
ref = SourceRef(source=doc.source, path=doc.path, title=doc.title)
|
||||
if doc.is_image:
|
||||
ref.image_url = f"/api/documents/{doc.id}/image"
|
||||
return ref
|
||||
|
||||
+131
-6
@@ -1,4 +1,5 @@
|
||||
"""Document summarizer (phase 30, task 03).
|
||||
"""Document summarizer (phase 30, task 03) + image descriptions
|
||||
(phase 122, task 03).
|
||||
|
||||
Builds the ``SUMMARY_MODE`` prompt for one document, calls the aipi
|
||||
``lite`` model through the one-shot ``LLMClient.chat`` (phase 30,
|
||||
@@ -22,18 +23,33 @@ Quality contracts enforced here:
|
||||
summarizer re-asserts defensively and never hands the importer a
|
||||
pointer-only row).
|
||||
|
||||
The ``SUMMARY_MODE`` marker follows the ``DEFLECT_MODE`` convention:
|
||||
the deterministic E2E mock LLM keys on it in the system prompt
|
||||
(``tests/e2e/mock_llm.py`` — wired in task 06).
|
||||
Image descriptions (phase 122, LOCKED A3): :func:`describe_image` is
|
||||
this module's second one-shot generation path — a SINGLE CHAT-model
|
||||
(vision) call describing one image's bytes as a base64 data URL. The
|
||||
description becomes the image document's ``content`` AND ``summary``
|
||||
(it is the ONLY embedded text of the doc — the embedding model never
|
||||
sees pixels, and the ``lite`` summary model is deliberately NOT used:
|
||||
it is not assumed vision-capable). Fail-soft by contract: any client
|
||||
error, empty reply, or non-2xx yields ``None`` — the importer skips
|
||||
the doc, counts ``images_failed``, and the sync continues.
|
||||
|
||||
The ``SUMMARY_MODE`` / ``IMAGE_DESCRIPTION_MODE`` markers follow the
|
||||
``DEFLECT_MODE`` convention: the deterministic E2E mock LLM keys on
|
||||
them (``tests/e2e/mock_llm.py`` — the image branch is wired by task
|
||||
06's story suite).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Protocol
|
||||
import base64
|
||||
import logging
|
||||
from typing import Any, Protocol
|
||||
|
||||
from app.config import Settings, get_settings
|
||||
from app.rag.llm import LLMError
|
||||
from app.rag.retriever import TRUNCATION_MARKER
|
||||
|
||||
logger = logging.getLogger("app.summarizer")
|
||||
|
||||
#: System-prompt marker for summary generation — the E2E mock LLM keys on
|
||||
#: it (same convention as ``DEFLECT_MODE``, PLAN §6).
|
||||
SUMMARY_MODE = "SUMMARY_MODE"
|
||||
@@ -51,6 +67,55 @@ SUMMARY_INSTRUCTION = (
|
||||
#: Full system prompt: marker first (the mock's key), then the instruction.
|
||||
SYSTEM_PROMPT = f"{SUMMARY_MODE}: {SUMMARY_INSTRUCTION}"
|
||||
|
||||
#: Image-description marker (phase 122, task 03) — the deterministic
|
||||
#: E2E mock LLM keys on it (same convention as ``SUMMARY_MODE`` /
|
||||
#: ``DEFLECT_MODE``, PLAN §6; the story suite wires the mock's branch
|
||||
#: in task 06). It heads the text part of the multimodal describe
|
||||
#: message, so it rides the user message, not a system prompt.
|
||||
IMAGE_DESCRIPTION_MODE = "IMAGE_DESCRIPTION_MODE"
|
||||
|
||||
#: Locked instruction for the CHAT (vision) model (phase 122, LOCKED
|
||||
#: A3): the description is the ONLY retrievable text of the image
|
||||
#: document (the embedding model never sees pixels), so it must be a
|
||||
#: faithful, retrieval-oriented account that carries the image's full
|
||||
#: meaning — what is depicted, any visible text/labels/titles,
|
||||
#: diagram/table structure, salient details.
|
||||
DESCRIBE_INSTRUCTION = (
|
||||
"Describe this image faithfully, in plain text, for a search index. "
|
||||
"State what is depicted, transcribe any visible text, labels, or "
|
||||
"titles, describe the structure of any diagram, table, or layout, and "
|
||||
"call out the most salient details. Write 2-4 sentences of substance. "
|
||||
"Do not use markdown. Do not invent anything that is not visible in "
|
||||
"the image. Your description is the ONLY text that will ever be "
|
||||
"retrieved for this image — it must carry the image's full meaning."
|
||||
)
|
||||
|
||||
#: The full describe prompt: marker first (the mock's key), then the
|
||||
#: instruction — the single text part of the multimodal user message.
|
||||
DESCRIBE_PROMPT = f"{IMAGE_DESCRIPTION_MODE}: {DESCRIBE_INSTRUCTION}"
|
||||
|
||||
#: Extension → MIME type for the image family (phase 122). Dotted,
|
||||
#: lowercase keys — the ``image_extension_set`` shape. Task 03 uses it
|
||||
#: for the describe call's data-URL mime; task 04's serve route reuses
|
||||
#: it for the image bytes' ``Content-Type`` (one map, one truth). A
|
||||
#: ``BOR_IMAGE_EXTENSIONS`` token outside this map (a custom format)
|
||||
#: takes the :data:`IMAGE_FALLBACK_MIME` data-URL mime in the describe
|
||||
#: call — the vision endpoint may reject it, and the fail-soft skip
|
||||
#: (``images_failed``) is the honest outcome.
|
||||
IMAGE_MIMES: dict[str, str] = {
|
||||
".png": "image/png",
|
||||
".jpg": "image/jpeg",
|
||||
".jpeg": "image/jpeg",
|
||||
".webp": "image/webp",
|
||||
".gif": "image/gif",
|
||||
".bmp": "image/bmp",
|
||||
}
|
||||
|
||||
#: The data-URL mime for an image extension :data:`IMAGE_MIMES` does not
|
||||
#: name (phase 122) — the best-effort generic, never a guess at a
|
||||
#: specific type.
|
||||
IMAGE_FALLBACK_MIME = "application/octet-stream"
|
||||
|
||||
|
||||
class SummaryLLM(Protocol):
|
||||
"""The one-shot chat surface the summarizer needs.
|
||||
@@ -58,12 +123,17 @@ class SummaryLLM(Protocol):
|
||||
:class:`app.rag.llm.LLMClient` satisfies it; unit tests pass a
|
||||
duck-typed fake (``chat`` + ``settings``) instead — same pattern as
|
||||
the importer's ``Embedder`` protocol.
|
||||
|
||||
``content`` may be a string (text calls — ``generate_summary``) or
|
||||
a list of OpenAI-compatible parts (phase 122 multimodal image
|
||||
descriptions — ``{type: "text", …}`` + ``{type: "image_url", …}``);
|
||||
the client passes message dicts through untouched.
|
||||
"""
|
||||
|
||||
settings: Settings
|
||||
|
||||
async def chat(
|
||||
self, messages: list[dict[str, str]], model: str | None = None
|
||||
self, messages: list[dict[str, Any]], model: str | None = None
|
||||
) -> str: ...
|
||||
|
||||
|
||||
@@ -121,3 +191,58 @@ async def generate_summary(
|
||||
"refusing to store a silent summary"
|
||||
)
|
||||
return f"{summary}\nSource: {source}/{path}"
|
||||
|
||||
|
||||
async def describe_image(
|
||||
llm: SummaryLLM,
|
||||
*,
|
||||
data: bytes,
|
||||
mime: str,
|
||||
settings: Settings | None = None,
|
||||
) -> str | None:
|
||||
"""One-shot CHAT-model (vision) description of one image (phase 122,
|
||||
LOCKED A3) — the text that becomes the image document's ``content``
|
||||
AND ``summary`` (the ONLY embedded text of the doc; the embedding
|
||||
model never sees pixels).
|
||||
|
||||
ONE chat-model call against ``settings.llm_chat_model`` (the vision
|
||||
model — the ``lite`` summary model is NOT assumed vision-capable)
|
||||
with the multimodal user message the OpenAI-compatible API expects:
|
||||
``[{type: "text", text: DESCRIBE_PROMPT}, {type: "image_url",
|
||||
image_url: {url: <data URL from *data* + *mime*>}}]`` — no system
|
||||
prompt, no tools, no app-level retries beyond the client's own
|
||||
(SDK-level + the house one-shot empty-content policy) — a
|
||||
description failure must not stall a sync.
|
||||
|
||||
Returns the stripped reply cut exactly at
|
||||
``(settings or llm.settings).summary_max_chars`` (the phase-30 cap —
|
||||
the description IS the summary, so it keeps the same uniform
|
||||
ceiling). Returns ``None`` on any client error, empty reply, or
|
||||
non-2xx (the client raises :class:`LLMError` for all three classes)
|
||||
— the caller (the importer's ``_describe_or_skip`` seam) fails soft:
|
||||
the doc is skipped, counted in ``images_failed``, and the sync
|
||||
continues (LOCKED A3).
|
||||
"""
|
||||
data_url = f"data:{mime};base64,{base64.b64encode(data).decode('ascii')}"
|
||||
messages: list[dict[str, Any]] = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "text", "text": DESCRIBE_PROMPT},
|
||||
{"type": "image_url", "image_url": {"url": data_url}},
|
||||
],
|
||||
}
|
||||
]
|
||||
model = llm.settings.llm_chat_model
|
||||
try:
|
||||
raw = await llm.chat(messages, model=model)
|
||||
except LLMError as e:
|
||||
logger.warning("image description failed (model=%s): %s", model, e)
|
||||
return None
|
||||
text = raw.strip()
|
||||
if not text:
|
||||
# The client already rejects empty content; this is the
|
||||
# defensive re-assert (the duck-typed fakes may return it).
|
||||
return None
|
||||
limit = (settings or llm.settings).summary_max_chars
|
||||
return text[:limit]
|
||||
|
||||
Reference in New Issue
Block a user