# Story: Import quadlet + jinja files **Phase:** `47_quadlet_jinja_import` · **Source:** `TODO.md` L10–L11 · **E2E:** `tests/e2e/test_quadlet_jinja_import.py` ## Bug reports (verbatim, `TODO.md` L10–L11) > "Add \".container\", \".network\", \".volume\" and other quadlet files > to the list of allowed/parsed files" > "Add \".j2\" jinja files to the list of allowed/parsed files" ## Narrative As **the owner**, my homelab notes increasingly live in Podman quadlet unit files (`.container`, `.network`, `.volume`, …) and Jinja templates (`.j2`). Neither is in the A9 import format list, so the KB is blind to exactly the config files I ask questions about. Both families join the allowed + default import formats and are parsed (chunked) by the importer. - **Given** a source directory containing quadlet and/or `.j2` files - **When** the importer (CLI or sync) runs - **Then** those files are indexed (chunked, embedded, upserted) like any other A9-format file, and their content is retrievable. ## Acceptance criteria 1. **Allowed set:** `_ALLOWED_IMPORT_EXTENSIONS` in `app/config.py` gains `container, network, volume, image, pod, kube, swap, os, endpoint` (the full Podman quadlet family) and `j2`; a `BOR_IMPORT_EXTENSIONS` env value may name any of them (the never-widen validator keeps rejecting truly unknown extensions). 2. **Default set:** the default `import_extensions` CSV includes all ten new formats after the existing seven — a default import now picks them up with no env configuration. 3. **Parsing:** `chunker.py` dispatches every new suffix to plain-text paragraph packing (`chunk_text`) — quadlet files are TOML unit files and `.j2` files are templates; no format-specific splitter (owner decision). Every chunk still honors `HARD_MAX_CHARS` (1200). 4. **Title:** no H1 → `extract_title` falls back to the file stem, as with other non-markdown formats (no change needed, verified). 5. **Behavior parity:** hidden (dot) directories are still skipped, the exclusion list is unchanged, sha256 delta detection / prune work unchanged for the new formats. 6. **Docs:** `.env.example`'s `BOR_IMPORT_EXTENSIONS` comment, the README's format list, and an **A9 revision note** in `.agent/PLAN.md` (owner permission 2026-08-27, `TODO.md` L10–L11) record the extended set. ## Owner-confirmed (2026-08-27, roadmap R1) 1. **Full quadlet family:** `container, network, volume, image, pod, kube, swap, os, endpoint` — plus `j2`. 2. **Chunked as plain text** — no TOML/Jinja-aware splitting. 3. **A9 is revised** with a PLAN.md revision note (the established owner-permission pattern). ## UI Visualization & Structure - **Config** (`app/config.py`): extend the `_ALLOWED_IMPORT_EXTENSIONS` frozenset + the `import_extensions` default string (comment cites the A9 revision 2026-08-27). No validator change — it already normalizes/dedups and rejects unknowns. - **Chunker** (`app/rag/chunker.py`): ten new `_FORMAT_CHUNKERS` entries → `chunk_text`; module docstring's format list updated. - **Fixtures:** `tests/fixtures/docs/homelab/quadlet/compose.container` (realistic quadlet TOML: `[Unit]` / `[Service]` / `[Container]` sections, one unique sentinel token, >1200 chars to exercise sub-splitting), `…/quadlet/lan.network`, `…/quadlet/cache.volume`, and `tests/fixtures/docs/homelab/templates/deploy.j2` (Jinja snippet with `{{ … }}` / `{% … %}` tags + its own sentinel). - **Tests:** unit (`test_config.py` allowed-set/default/validator; `test_chunker.py` dispatch for every new suffix + fixture chunking; `test_importer.py` directory walk picks the new files up); integration (import_sources over a temp dir with quadlet+j2 files → documents + chunks rows). - **Non-goals:** no new DB column, no format badge change (the viewer shows the extension it already shows), no summary-model changes. ## Playwright Mapping Rule **Test Scenario → `tests/e2e/test_quadlet_jinja_import.py`** (mock LLM; DB up): 1. `test_quadlet_and_jinja_indexed` — truncate + import the fixture tree (the `test_import_documents.py` pattern): `GET /api/docs` lists the `.container` / `.network` / `.volume` / `.j2` files with non-zero chunk counts. 2. `test_sources_table_shows_them` (admin) — the Sources table renders rows for the new files; their path links open the document modal. 3. `test_container_content_viewable` — the viewer modal shows the `.container` file's TOML content (sentinel token present) with its stem as the title. 4. `test_jinja_retrievable_not_deflected` — ask a question containing the `.j2` file's sentinel word: the FTS hit keeps the honesty gate honest-positive (A8) — the answer bubble is **not** `.is-deflected` and the source chip names the `.j2` document. 5. `test_default_walk_includes_new_formats` (regression, unit-backed) — a default-extensions walk over a temp tree with all ten new extensions indexes every file; hidden directories + the exclusion list still filter (covered by `test_importer.py` in the regression pass).