feat(import): index quadlet unit files and jinja templates (A9 revision)

Phase 47 (owner permission 2026-08-27, TODO.md L10–11, roadmap R1): the
full Podman quadlet family (.container, .network, .volume, .image,
.pod, .kube, .swap, .os, .endpoint) and .j2 Jinja templates join the
allowed + default A9 import formats, chunked as plain text (owner
decision — no TOML/Jinja-aware splitter). No env configuration needed:
a default import now indexes them.

- app/config.py: _ALLOWED_IMPORT_EXTENSIONS + the default
  import_extensions CSV gain the ten names (the original seven first);
  the never-widen BOR_IMPORT_EXTENSIONS validator is untouched and
  still rejects truly unknown extensions.
- app/rag/chunker.py: ten _FORMAT_CHUNKERS entries -> chunk_text
  (HARD_MAX_CHARS 1200 honored, unknown-suffix fallback unchanged);
  docstring/comments cite the A9 revision 2026-08-27.
- tests/fixtures/docs/homelab/: quadlet/compose.container (realistic
  quadlet TOML, >1500 chars, [Unit]/[Service]/[Container] sections,
  RESE-QUADLET-SENTINEL-77aa), quadlet/lan.network,
  quadlet/cache.volume, templates/deploy.j2 (for/set/if Jinja
  constructs + RESE-JINJA-SENTINEL-33dd). Every suite that seeds the
  fixture tree updates its 9 -> 13 document-count constants.
- tests/unit/test_config.py: allowed set carries all seventeen formats,
  default CSV + dotted import_extension_set include the ten, the
  validator accepts the new names and still rejects unknowns.
- tests/unit/test_chunker.py: dispatch parity with chunk_text for every
  new suffix (parametrized), the .container fixture chunks >=2 under
  the cap with the sentinel surviving, the .j2 fixture keeps {{ }}
  verbatim, the unknown-suffix fallback is unchanged.
- tests/unit/test_importer.py: a default-extensions walk over a temp
  tree indexes exactly the ten new files (unknown/hidden/excluded
  filtered), the original seven still walk, stem-title fallback holds.
- tests/integration/test_import_quadlet_jinja.py (new): import_sources
  over a temp tree with .container/.volume/.j2 -> documents + chunks
  rows with stem titles; delta re-import updates only the changed .j2
  doc; prune drops the deleted .volume doc with cascade.
- tests/e2e/test_quadlet_jinja_import.py (new, story suite, mock-only,
  isolation): GET /api/docs (admin session) lists the four new-format
  docs with non-zero chunk counts and stem titles; the Sources table
  renders a row + .doc-link per file; the phase-26 modal shows the
  .container TOML ([Container] section + sentinel) with stem title and
  the container format badge; a RESE-JINJA-SENTINEL-33dd question
  FTS-matches the .j2 chunk -> honest-positive (A8: LOW requires zero
  FTS hits) — the bubble is not .is-deflected and a source chip names
  templates/deploy.j2.
- README.md + .env.example: the extended default format set (A9
  revised 2026-08-27, plain-text chunking, narrow-only rule intact).
- .agent/PLAN.md: the A9 revision (owner-locked R1) — A9 row status,
  the revision note under the anchors table, and the §5 chunking-policy
  + §11 workflow lines. The only PLAN edit this phase.

Gates: uv run pytest 795 passed; app/ coverage TOTAL 99% (>90%);
ruff check + pyright clean; story E2E 4/4 in isolation (DB up);
regression E2E suites test_import_documents (3) / test_sync_button
(3) / test_git_sources_admin (6) green in isolation.

Also records the 47_quadlet_jinja_import task-file moves (01–03)
todo/ -> complete/.
This commit is contained in:
2026-08-28 07:02:24 -04:00
parent 6be692d999
commit 5d679f5184
40 changed files with 811 additions and 73 deletions
+60
View File
@@ -7,6 +7,7 @@ unchanged); the per-format tests cover the phase-09 dispatcher
from __future__ import annotations
from itertools import pairwise
from pathlib import Path
import pytest
@@ -361,6 +362,63 @@ def test_txt_long_doc_packs_with_overlap() -> None:
assert all(f"para {i}" in "\n".join(chunks) for i in range(6))
# ---------------------------------------------------------------------------
# quadlet family + j2 (A9 revised 2026-08-27) — plain-text dispatch
# ---------------------------------------------------------------------------
#: The ten suffixes added 2026-08-27 (owner decision R1): the full Podman
#: quadlet family + Jinja templates, all chunked as plain text.
NEW_FORMAT_SUFFIXES = [
".container", ".network", ".volume", ".image", ".pod",
".kube", ".swap", ".os", ".endpoint", ".j2",
]
REPO_ROOT = Path(__file__).resolve().parents[2]
FIXTURE_DOCS = REPO_ROOT / "tests" / "fixtures" / "docs" / "homelab"
@pytest.mark.parametrize("suffix", NEW_FORMAT_SUFFIXES)
def test_dispatch_new_suffixes_match_plain_text(suffix: str) -> None:
"""Every new suffix dispatches to plain-text paragraph packing — the
chunks are identical to a direct ``chunk_text`` call."""
content = "[Unit]\nDescription=one\n\n[Service]\nRestart=always\n\n[Container]\nImage=busybox\n"
name = suffix.lstrip(".")
assert chunk_document(content, f"x/{name}{suffix}") == chunk_text(content)
def test_container_fixture_chunks_past_single_paragraph_pack() -> None:
"""The realistic quadlet TOML fixture (≥1 500 chars) splits into ≥2
chunks, every chunk stays under the hard cap, and the sentinel token
survives the split."""
rel = "quadlet/compose.container"
content = (FIXTURE_DOCS / rel).read_text(encoding="utf-8")
assert len(content) >= 1500 # the fixture must exercise sub-splitting
chunks = chunk_document(content, rel)
assert len(chunks) >= 2
assert all(len(c) <= HARD_MAX_CHARS for c in chunks)
assert any("RESE-QUADLET-SENTINEL-77aa" in c for c in chunks)
joined = "\n".join(chunks)
assert "[Container]" in joined and "Restart=always" in joined
def test_j2_fixture_chunks_braces_as_plain_text() -> None:
"""Jinja braces are just text — no special handling; the ``{{ … }}``
line and the ``{% set %}`` sentinel appear verbatim in the chunks."""
rel = "templates/deploy.j2"
content = (FIXTURE_DOCS / rel).read_text(encoding="utf-8")
chunks = chunk_document(content, rel)
assert chunks
assert all(len(c) <= HARD_MAX_CHARS for c in chunks)
joined = "\n".join(chunks)
assert "RESE-JINJA-SENTINEL-33dd" in joined
assert any("[{{ svc.name }}]" in c for c in chunks)
def test_dispatch_still_falls_back_for_unknown_suffix() -> None:
content = "alpha\n\nbeta\n"
assert chunk_document(content, "x/notes.whatever") == chunk_text(content)
# ---------------------------------------------------------------------------
# Hard cap across every format (aipi ~1024-token request cap)
# ---------------------------------------------------------------------------
@@ -374,6 +432,8 @@ def test_txt_long_doc_packs_with_overlap() -> None:
('{"blob": "' + "z" * 5000 + '"}', "big.json"),
("def f():\n" + " x = 1\n" * 1000, "big.py"),
("line of text\n\n" * 800, "big.txt"),
("[Unit]\n" + "key=" + "v" * 5000 + "\n", "big.container"), # phase 47
("{% for x in y %}" + "{{ x }} " * 400 + "{% endfor %}", "big.j2"), # phase 47
],
)
def test_hard_cap_holds_for_every_format(content: str, path: str) -> None: