feat(rag): index markdown KB — chunker, embed client, delta importer, Sources page
Phase 02 (story: import documents):
- fence-aware markdown chunker (heading sections, 200-char overlap,
heading anchor on every chunk, 1200-char hard cap, fence blocks
kept atomic and split under the cap)
- LLMClient over aipi (LiteLLM) reusing the openai client's httpx
transport to send a clean {model, input} payload — the openai SDK
injects encoding_format, which aipi's openai_like group rejects;
token-budget batching + halving retry for the endpoint's
~1024-token per-request input cap
- two-phase per-file upsert importer: sha256 delta (unchanged skip),
atomic commit, A9 exclusion walk, per-source prune, per-file error
tolerance (rollback + log + continue, non-zero CLI exit), adaptive
re-chunk at half target for URL-dense files the endpoint rejects
- scripts/import_docs CLI (repeatable --source, --prune, --limit,
defaults ~/Homelab + ~/Deployments)
- GET /api/docs with per-doc chunk counts; Sources page wired to the
real endpoint (stat cards, full-width a11y table, designed empty
state, DOM-built rows — no innerHTML)
- tests: 63 passed (chunker/llm/importer units, docs API + importer
integration), story E2E 3/3 (real endpoints, in-thread import);
app/ coverage 98%
- real KB imported: 672 docs / 8969 chunks in ~3m, idempotent
re-run (672 unchanged, 0 batches)
- harness: .agent/validate.sh now gates through uv (pytest +
coverage >90% + ruff + pyright) instead of system python3
This commit is contained in:
@@ -77,6 +77,10 @@ uv run python -m scripts.import_docs --source ~/SomeOtherDocs
|
||||
- Only **`*.md`** files are indexed. Directories like `.venv`,
|
||||
`node_modules`, `.git`, `__pycache__`, `.pytest_cache`, `dist`, `build`
|
||||
are skipped (see `.agent/PLAN.md` anchor A9).
|
||||
- Every file is logged on its own line (`import: added|updated|unchanged|
|
||||
pruned …`), and the run ends with a one-line summary (`import: summary
|
||||
files=… added=… updated=… unchanged=… pruned=… chunks=… embed_batches=…`)
|
||||
so the counts are greppable in logs.
|
||||
- Unchanged files are **not re-embedded** — only new/changed ones, so
|
||||
refreshes are cheap.
|
||||
- To sanity-check the LLM backend (models + embedding dimension) after any
|
||||
@@ -194,6 +198,14 @@ served locally (no CDN), `BOR_ENVIRONMENT=production`.
|
||||
## Troubleshooting
|
||||
|
||||
- **`401` from aipi** — set `BOR_LLM_API_KEY` (or `$AIPI_KEY`).
|
||||
- **`litellm.UnsupportedParamsError … encoding_format` from aipi** — the
|
||||
aipi proxy (litellm `openai_like`) rejects the `encoding_format` parameter
|
||||
that the `openai` SDK injects into every embeddings request. The app
|
||||
already works around this by POSTing a minimal `{model, input}` payload
|
||||
through the openai client's own httpx transport (`app/rag/llm.py` →
|
||||
`LLMClient._embed_batch`). If you see this, you are likely calling the
|
||||
endpoint with a different client — drop the parameter (or set
|
||||
`litellm.drop_params = True` on the proxy).
|
||||
- **Embedding dimension mismatch** — aipi changed models; run
|
||||
`uv run python -m scripts.llm_probe`, update `BOR_EMBEDDING_DIM`, then
|
||||
drop + recreate the chunks table (new migration or manual `TRUNCATE
|
||||
|
||||
Reference in New Issue
Block a user