feat(rag): index markdown KB — chunker, embed client, delta importer, Sources page

Phase 02 (story: import documents):

- fence-aware markdown chunker (heading sections, 200-char overlap,
  heading anchor on every chunk, 1200-char hard cap, fence blocks
  kept atomic and split under the cap)
- LLMClient over aipi (LiteLLM) reusing the openai client's httpx
  transport to send a clean {model, input} payload — the openai SDK
  injects encoding_format, which aipi's openai_like group rejects;
  token-budget batching + halving retry for the endpoint's
  ~1024-token per-request input cap
- two-phase per-file upsert importer: sha256 delta (unchanged skip),
  atomic commit, A9 exclusion walk, per-source prune, per-file error
  tolerance (rollback + log + continue, non-zero CLI exit), adaptive
  re-chunk at half target for URL-dense files the endpoint rejects
- scripts/import_docs CLI (repeatable --source, --prune, --limit,
  defaults ~/Homelab + ~/Deployments)
- GET /api/docs with per-doc chunk counts; Sources page wired to the
  real endpoint (stat cards, full-width a11y table, designed empty
  state, DOM-built rows — no innerHTML)
- tests: 63 passed (chunker/llm/importer units, docs API + importer
  integration), story E2E 3/3 (real endpoints, in-thread import);
  app/ coverage 98%
- real KB imported: 672 docs / 8969 chunks in ~3m, idempotent
  re-run (672 unchanged, 0 batches)
- harness: .agent/validate.sh now gates through uv (pytest +
  coverage >90% + ruff + pyright) instead of system python3
This commit is contained in:
2026-08-21 16:24:45 -04:00
parent dce6d0d6f1
commit 99c48cbe06
19 changed files with 1822 additions and 11 deletions
+19 -11
View File
@@ -1,6 +1,9 @@
/* Brain of Reese — Sources page (knowledge base index view).
* Scaffolding-stage: fetches /api/docs (implemented in the import phase);
* until then it renders the empty state.
*
* Wires the real `GET /api/docs` endpoint (import phase): stat cards +
* full-width document table, or the designed empty state when nothing is
* indexed yet. Cells are built with DOM APIs (textContent) — never
* innerHTML with document-derived data (XSS-safe by construction).
*/
const tbody = document.querySelector("#docs-tbody");
@@ -36,20 +39,13 @@ async function loadDocs() {
return;
}
tbody.innerHTML = "";
tbody.replaceChildren();
let totalChunks = 0;
let last = "";
for (const d of documents) {
totalChunks += d.chunks;
if (d.indexed_at > last) last = d.indexed_at;
const tr = document.createElement("tr");
tr.innerHTML = `
<td>${d.source}</td>
<td title="${d.path}">${d.path}</td>
<td>${d.title}</td>
<td>${d.chunks}</td>
<td>${fmtDate(d.indexed_at)}</td>`;
tbody.appendChild(tr);
tbody.appendChild(makeRow(d));
}
statDocs.textContent = String(documents.length);
statChunks.textContent = String(totalChunks);
@@ -58,6 +54,18 @@ async function loadDocs() {
tableWrap.hidden = false;
}
function makeRow(d) {
const tr = document.createElement("tr");
const cells = [d.source, d.path, d.title, String(d.chunks), fmtDate(d.indexed_at)];
for (const value of cells) {
const td = document.createElement("td");
td.textContent = value; // document-derived text — never innerHTML
tr.appendChild(td);
}
tr.children[1].title = d.path; // full path on hover (column is ellipsized)
return tr;
}
function showEmpty() {
statDocs.textContent = "0";
statChunks.textContent = "0";