phase: 112_honesty_gate_weak_hits
**Phase 112 — final verification pass (all 4 tasks already complete in `complete/`):** - Verified gate fix: `app/api/chat.py::plan_turn` — HIGH iff `best_cosine >= relevance_threshold` OR (`fts_hits > 0` AND `best_cosine >= lexical_support_floor`); `lexical_support_floor` (default 0.35, `BOR_LEXICAL_SUPPORT_FLOOR`, bounds-validated) in `app/config.py` + `.env.example`; A8 revision note (2026-09-14) in `.agents/PLAN.md`. - Verified prompt contract: `app/rag/prompts.py` diff is docstring-only (dated owner-decision-iii entry); `tests/unit/test_prompt_lock.py` byte-pins PERSONA/TOOLS_SECTION/DEFLECT body (sha256+length). - Verified README: L11 + L575 deflection copy refreshed; `grep "haven't done anything" README.md` → no hits; disclosed-answer behavior documented. - Tests: `uv run pytest --cov=app --cov-report=term-missing` → **2378 passed, 99% coverage (>90%)**; includes Mongolia-quadrant unit pins (fts>0 + cosine<floor → LOW). - E2E in isolation: `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → **4 passed** (out-of-KB question: `deflected=true`, `sources==[]`, 2–3 suggestions); regression `test_chat_rag.py` + `test_retrieval_quality.py` → **7 passed**. - Lint/types: `uv run ruff check .` → clean; `uv run pyright` → 0 errors. **Completion criteria:** weak-FTS→LOW unit-pinned ✅ · no false citations + 2–3 alternatives E2E ✅ · prompts byte-identical (test-pinned) + README matches ✅ · suite/coverage/e2e/lint all green ✅ · commit + phase move → left to harness (no `git commit` run, per rules; changes in working tree). **Deviations:** none. Next pending phase: `113_source_chip_quality`.
This commit is contained in:
+1
-1
@@ -46,7 +46,7 @@ accounts (one admin + hand-out tokens is the model).
|
|||||||
| A5 | LLM backend | OpenAI-compatible self-hosted endpoint `https://aipi.reeseapps.com/v1` via the `openai` async client: **`turbo`** (chat, streams `reasoning_content` thinking), **`embed`** (embeddings), **`lite`** (one-shot: document summaries, KB overview, folder summaries) |
|
| A5 | LLM backend | OpenAI-compatible self-hosted endpoint `https://aipi.reeseapps.com/v1` via the `openai` async client: **`turbo`** (chat, streams `reasoning_content` thinking), **`embed`** (embeddings), **`lite`** (one-shot: document summaries, KB overview, folder summaries) |
|
||||||
| A6 | Embedding dim | **768** (verified against the live endpoint); `chunks.embedding` is fixed at table creation — a dim mismatch must **fail loudly**, never silently re-embed |
|
| A6 | Embedding dim | **768** (verified against the live endpoint); `chunks.embedding` is fixed at table creation — a dim mismatch must **fail loudly**, never silently re-embed |
|
||||||
| A7 | Retrieval→context | **Hybrid:** cosine top-100 ∪ Postgres FTS top-30 (OR tsquery, `ts_rank`), fused with **RRF (k=60)** → parent docs ranked by best fused chunk → **full text of top-N=2 documents, never truncated** on the retrieval path (owner 2026-08-24: "this should never happen"; the agent `read`-tool cap is a separate owner-permitted path) |
|
| A7 | Retrieval→context | **Hybrid:** cosine top-100 ∪ Postgres FTS top-30 (OR tsquery, `ts_rank`), fused with **RRF (k=60)** → parent docs ranked by best fused chunk → **full text of top-N=2 documents, never truncated** on the retrieval path (owner 2026-08-24: "this should never happen"; the agent `read`-tool cap is a separate owner-permitted path) |
|
||||||
| A8 | Honesty gate | Deflect (LOW mode) **only when** best cosine < `BOR_RELEVANCE_THRESHOLD` (0.62) **and** zero FTS hits; LOW prompt carries weak-hit *titles only* + the `DEFLECT_MODE` marker (the E2E mock keys on its presence) + the plain-text no-tools line |
|
| A8 | Honesty gate | Deflect (LOW mode) **only when** best cosine < `BOR_RELEVANCE_THRESHOLD` (0.62) **and** (zero FTS hits **or** best cosine < `BOR_LEXICAL_SUPPORT_FLOOR` (0.35)); an FTS hit flips HIGH only when `best_cosine >= lexical_support_floor` — the vector signal must corroborate the lexical match (A8 revised 2026-09-14, owner-confirmed, TODO L2a: lexical-only hits without vector support deflect). LOW prompt carries weak-hit *titles only* + the `DEFLECT_MODE` marker (the E2E mock keys on its presence) + the plain-text no-tools line |
|
||||||
| A9 | Content scope | Default `md, markdown, txt, yaml, yml, json, py` + Podman quadlet family + `j2`; `BOR_IMPORT_EXTENSIONS` may name **any** well-formed extension or narrow the set; hidden (dot) paths + exclusion list + per-source `ignore_paths` (raw prefixes, no globs) always apply |
|
| A9 | Content scope | Default `md, markdown, txt, yaml, yml, json, py` + Podman quadlet family + `j2`; `BOR_IMPORT_EXTENSIONS` may name **any** well-formed extension or narrow the set; hidden (dot) paths + exclusion list + per-source `ignore_paths` (raw prefixes, no globs) always apply |
|
||||||
| A10 | State & auth | `POST /api/chat` is **stateless** (client-provided `history` only, budget-trimmed — nothing stored per conversation). Auth = single admin, signed `bor_session` cookie (Starlette `SessionMiddleware` + itsdangerous; no server-side session store); admin-issued SHA-256-hashed access tokens are the only other identity; **only** shared chats stay anonymous. Fail-loud at boot while `BOR_ADMIN_PASSWORD`/`BOR_SESSION_SECRET` are empty |
|
| A10 | State & auth | `POST /api/chat` is **stateless** (client-provided `history` only, budget-trimmed — nothing stored per conversation). Auth = single admin, signed `bor_session` cookie (Starlette `SessionMiddleware` + itsdangerous; no server-side session store); admin-issued SHA-256-hashed access tokens are the only other identity; **only** shared chats stay anonymous. Fail-loud at boot while `BOR_ADMIN_PASSWORD`/`BOR_SESSION_SECRET` are empty |
|
||||||
| A11 | Frontend | Vanilla HTML/CSS/JS in git, **no CDN** — everything served by FastAPI `StaticFiles`; system font stack; the navbar views are views of ONE shell document (`frontend/index.html` + `router.js` deep-links) |
|
| A11 | Frontend | Vanilla HTML/CSS/JS in git, **no CDN** — everything served by FastAPI `StaticFiles`; system font stack; the navbar views are views of ONE shell document (`frontend/index.html` + `router.js` deep-links) |
|
||||||
|
|||||||
@@ -0,0 +1,12 @@
|
|||||||
|
**Phase 112 — final verification pass (all 4 tasks already complete in `complete/`):**
|
||||||
|
|
||||||
|
- Verified gate fix: `app/api/chat.py::plan_turn` — HIGH iff `best_cosine >= relevance_threshold` OR (`fts_hits > 0` AND `best_cosine >= lexical_support_floor`); `lexical_support_floor` (default 0.35, `BOR_LEXICAL_SUPPORT_FLOOR`, bounds-validated) in `app/config.py` + `.env.example`; A8 revision note (2026-09-14) in `.agents/PLAN.md`.
|
||||||
|
- Verified prompt contract: `app/rag/prompts.py` diff is docstring-only (dated owner-decision-iii entry); `tests/unit/test_prompt_lock.py` byte-pins PERSONA/TOOLS_SECTION/DEFLECT body (sha256+length).
|
||||||
|
- Verified README: L11 + L575 deflection copy refreshed; `grep "haven't done anything" README.md` → no hits; disclosed-answer behavior documented.
|
||||||
|
- Tests: `uv run pytest --cov=app --cov-report=term-missing` → **2378 passed, 99% coverage (>90%)**; includes Mongolia-quadrant unit pins (fts>0 + cosine<floor → LOW).
|
||||||
|
- E2E in isolation: `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → **4 passed** (out-of-KB question: `deflected=true`, `sources==[]`, 2–3 suggestions); regression `test_chat_rag.py` + `test_retrieval_quality.py` → **7 passed**.
|
||||||
|
- Lint/types: `uv run ruff check .` → clean; `uv run pyright` → 0 errors.
|
||||||
|
|
||||||
|
**Completion criteria:** weak-FTS→LOW unit-pinned ✅ · no false citations + 2–3 alternatives E2E ✅ · prompts byte-identical (test-pinned) + README matches ✅ · suite/coverage/e2e/lint all green ✅ · commit + phase move → left to harness (no `git commit` run, per rules; changes in working tree).
|
||||||
|
|
||||||
|
**Deviations:** none. Next pending phase: `113_source_chip_quality`.
|
||||||
+102
@@ -0,0 +1,102 @@
|
|||||||
|
........................................................................ [ 3%]
|
||||||
|
........................................................................ [ 6%]
|
||||||
|
........................................................................ [ 9%]
|
||||||
|
........................................................................ [ 12%]
|
||||||
|
........................................................................ [ 15%]
|
||||||
|
........................................................................ [ 18%]
|
||||||
|
........................................................................ [ 21%]
|
||||||
|
........................................................................ [ 24%]
|
||||||
|
........................................................................ [ 27%]
|
||||||
|
........................................................................ [ 30%]
|
||||||
|
........................................................................ [ 33%]
|
||||||
|
........................................................................ [ 36%]
|
||||||
|
........................................................................ [ 39%]
|
||||||
|
........................................................................ [ 42%]
|
||||||
|
........................................................................ [ 45%]
|
||||||
|
........................................................................ [ 48%]
|
||||||
|
........................................................................ [ 51%]
|
||||||
|
........................................................................ [ 54%]
|
||||||
|
........................................................................ [ 57%]
|
||||||
|
........................................................................ [ 60%]
|
||||||
|
........................................................................ [ 63%]
|
||||||
|
........................................................................ [ 66%]
|
||||||
|
........................................................................ [ 69%]
|
||||||
|
........................................................................ [ 72%]
|
||||||
|
........................................................................ [ 75%]
|
||||||
|
........................................................................ [ 78%]
|
||||||
|
........................................................................ [ 81%]
|
||||||
|
........................................................................ [ 84%]
|
||||||
|
........................................................................ [ 87%]
|
||||||
|
........................................................................ [ 90%]
|
||||||
|
........................................................................ [ 93%]
|
||||||
|
........................................................................ [ 96%]
|
||||||
|
........................................................................ [ 99%]
|
||||||
|
.. [100%]
|
||||||
|
=============================== warnings summary ===============================
|
||||||
|
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
|
||||||
|
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||||
|
from starlette.testclient import TestClient as TestClient # noqa
|
||||||
|
|
||||||
|
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||||
|
================================ tests coverage ================================
|
||||||
|
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
|
||||||
|
|
||||||
|
Name Stmts Miss Cover
|
||||||
|
--------------------------------------------------
|
||||||
|
app/__init__.py 1 0 100%
|
||||||
|
app/api/__init__.py 0 0 100%
|
||||||
|
app/api/auth.py 52 0 100%
|
||||||
|
app/api/chat.py 208 1 99%
|
||||||
|
app/api/chats.py 110 0 100%
|
||||||
|
app/api/config.py 13 0 100%
|
||||||
|
app/api/doc_drafts.py 94 0 100%
|
||||||
|
app/api/docs.py 156 1 99%
|
||||||
|
app/api/git_sources.py 232 0 100%
|
||||||
|
app/api/health.py 10 0 100%
|
||||||
|
app/api/steering.py 42 0 100%
|
||||||
|
app/api/suggestions.py 33 0 100%
|
||||||
|
app/api/sync.py 139 0 100%
|
||||||
|
app/api/tokens.py 40 0 100%
|
||||||
|
app/api/ui_settings.py 55 0 100%
|
||||||
|
app/config.py 186 0 100%
|
||||||
|
app/core/__init__.py 0 0 100%
|
||||||
|
app/core/auth.py 45 0 100%
|
||||||
|
app/core/caching.py 124 0 100%
|
||||||
|
app/core/debugging.py 29 2 93%
|
||||||
|
app/core/docs_push.py 39 0 100%
|
||||||
|
app/core/errors.py 5 0 100%
|
||||||
|
app/core/logging.py 13 0 100%
|
||||||
|
app/core/rate_limit.py 44 0 100%
|
||||||
|
app/core/security_headers.py 20 0 100%
|
||||||
|
app/core/theming.py 38 0 100%
|
||||||
|
app/core/tokens.py 44 0 100%
|
||||||
|
app/db.py 22 0 100%
|
||||||
|
app/main.py 66 0 100%
|
||||||
|
app/models.py 128 0 100%
|
||||||
|
app/rag/__init__.py 0 0 100%
|
||||||
|
app/rag/agent.py 317 1 99%
|
||||||
|
app/rag/archive_upload.py 134 0 100%
|
||||||
|
app/rag/chunker.py 206 4 98%
|
||||||
|
app/rag/doc_dates.py 18 0 100%
|
||||||
|
app/rag/folder_summaries.py 123 0 100%
|
||||||
|
app/rag/git_sources.py 14 0 100%
|
||||||
|
app/rag/importer.py 215 3 99%
|
||||||
|
app/rag/llm.py 243 1 99%
|
||||||
|
app/rag/overview.py 71 0 100%
|
||||||
|
app/rag/prompts.py 88 0 100%
|
||||||
|
app/rag/retriever.py 172 3 98%
|
||||||
|
app/rag/scaffolding.py 55 0 100%
|
||||||
|
app/rag/source_removal.py 41 0 100%
|
||||||
|
app/rag/sources_meta.py 16 0 100%
|
||||||
|
app/rag/suggestions.py 27 0 100%
|
||||||
|
app/rag/summarizer.py 24 0 100%
|
||||||
|
app/schemas.py 327 0 100%
|
||||||
|
--------------------------------------------------
|
||||||
|
TOTAL 4079 16 99%
|
||||||
|
coverage gate: app/ 99% (>90%) OK
|
||||||
|
All checks passed!
|
||||||
|
0 errors, 0 warnings, 0 informations
|
||||||
|
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||||
|
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||||
|
|
||||||
|
validation OK
|
||||||
+1
@@ -0,0 +1 @@
|
|||||||
|
All 2373 tests pass with 99% coverage. Now let me run the linter:
|
||||||
+101
@@ -0,0 +1,101 @@
|
|||||||
|
........................................................................ [ 3%]
|
||||||
|
........................................................................ [ 6%]
|
||||||
|
........................................................................ [ 9%]
|
||||||
|
........................................................................ [ 12%]
|
||||||
|
........................................................................ [ 15%]
|
||||||
|
........................................................................ [ 18%]
|
||||||
|
........................................................................ [ 21%]
|
||||||
|
........................................................................ [ 24%]
|
||||||
|
........................................................................ [ 27%]
|
||||||
|
........................................................................ [ 30%]
|
||||||
|
........................................................................ [ 33%]
|
||||||
|
........................................................................ [ 36%]
|
||||||
|
........................................................................ [ 39%]
|
||||||
|
........................................................................ [ 42%]
|
||||||
|
........................................................................ [ 45%]
|
||||||
|
........................................................................ [ 48%]
|
||||||
|
........................................................................ [ 51%]
|
||||||
|
........................................................................ [ 54%]
|
||||||
|
........................................................................ [ 57%]
|
||||||
|
........................................................................ [ 60%]
|
||||||
|
........................................................................ [ 63%]
|
||||||
|
........................................................................ [ 66%]
|
||||||
|
........................................................................ [ 69%]
|
||||||
|
........................................................................ [ 72%]
|
||||||
|
........................................................................ [ 75%]
|
||||||
|
........................................................................ [ 78%]
|
||||||
|
........................................................................ [ 81%]
|
||||||
|
........................................................................ [ 84%]
|
||||||
|
........................................................................ [ 87%]
|
||||||
|
........................................................................ [ 91%]
|
||||||
|
........................................................................ [ 94%]
|
||||||
|
........................................................................ [ 97%]
|
||||||
|
..................................................................... [100%]
|
||||||
|
=============================== warnings summary ===============================
|
||||||
|
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
|
||||||
|
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||||
|
from starlette.testclient import TestClient as TestClient # noqa
|
||||||
|
|
||||||
|
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||||
|
================================ tests coverage ================================
|
||||||
|
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
|
||||||
|
|
||||||
|
Name Stmts Miss Cover
|
||||||
|
--------------------------------------------------
|
||||||
|
app/__init__.py 1 0 100%
|
||||||
|
app/api/__init__.py 0 0 100%
|
||||||
|
app/api/auth.py 52 0 100%
|
||||||
|
app/api/chat.py 205 1 99%
|
||||||
|
app/api/chats.py 110 0 100%
|
||||||
|
app/api/config.py 13 0 100%
|
||||||
|
app/api/doc_drafts.py 94 0 100%
|
||||||
|
app/api/docs.py 156 1 99%
|
||||||
|
app/api/git_sources.py 232 0 100%
|
||||||
|
app/api/health.py 10 0 100%
|
||||||
|
app/api/steering.py 42 0 100%
|
||||||
|
app/api/suggestions.py 33 0 100%
|
||||||
|
app/api/sync.py 139 0 100%
|
||||||
|
app/api/tokens.py 40 0 100%
|
||||||
|
app/api/ui_settings.py 55 0 100%
|
||||||
|
app/config.py 186 0 100%
|
||||||
|
app/core/__init__.py 0 0 100%
|
||||||
|
app/core/auth.py 45 0 100%
|
||||||
|
app/core/caching.py 124 0 100%
|
||||||
|
app/core/debugging.py 29 2 93%
|
||||||
|
app/core/docs_push.py 39 0 100%
|
||||||
|
app/core/errors.py 5 0 100%
|
||||||
|
app/core/logging.py 13 0 100%
|
||||||
|
app/core/rate_limit.py 44 0 100%
|
||||||
|
app/core/security_headers.py 20 0 100%
|
||||||
|
app/core/theming.py 38 0 100%
|
||||||
|
app/core/tokens.py 44 0 100%
|
||||||
|
app/db.py 22 0 100%
|
||||||
|
app/main.py 66 0 100%
|
||||||
|
app/models.py 128 0 100%
|
||||||
|
app/rag/__init__.py 0 0 100%
|
||||||
|
app/rag/agent.py 317 1 99%
|
||||||
|
app/rag/archive_upload.py 134 0 100%
|
||||||
|
app/rag/chunker.py 206 4 98%
|
||||||
|
app/rag/doc_dates.py 18 0 100%
|
||||||
|
app/rag/folder_summaries.py 123 0 100%
|
||||||
|
app/rag/git_sources.py 14 0 100%
|
||||||
|
app/rag/importer.py 215 3 99%
|
||||||
|
app/rag/llm.py 243 1 99%
|
||||||
|
app/rag/overview.py 71 0 100%
|
||||||
|
app/rag/prompts.py 88 0 100%
|
||||||
|
app/rag/retriever.py 172 3 98%
|
||||||
|
app/rag/scaffolding.py 55 0 100%
|
||||||
|
app/rag/source_removal.py 41 0 100%
|
||||||
|
app/rag/sources_meta.py 16 0 100%
|
||||||
|
app/rag/suggestions.py 27 0 100%
|
||||||
|
app/rag/summarizer.py 24 0 100%
|
||||||
|
app/schemas.py 327 0 100%
|
||||||
|
--------------------------------------------------
|
||||||
|
TOTAL 4076 16 99%
|
||||||
|
coverage gate: app/ 99% (>90%) OK
|
||||||
|
All checks passed!
|
||||||
|
0 errors, 0 warnings, 0 informations
|
||||||
|
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||||
|
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||||
|
|
||||||
|
validation OK
|
||||||
+1
@@ -0,0 +1 @@
|
|||||||
|
All 2373 tests pass with 99% coverage. Now let me run the linter:
|
||||||
+101
@@ -0,0 +1,101 @@
|
|||||||
|
........................................................................ [ 3%]
|
||||||
|
........................................................................ [ 6%]
|
||||||
|
........................................................................ [ 9%]
|
||||||
|
........................................................................ [ 12%]
|
||||||
|
........................................................................ [ 15%]
|
||||||
|
........................................................................ [ 18%]
|
||||||
|
........................................................................ [ 21%]
|
||||||
|
........................................................................ [ 24%]
|
||||||
|
........................................................................ [ 27%]
|
||||||
|
........................................................................ [ 30%]
|
||||||
|
........................................................................ [ 33%]
|
||||||
|
........................................................................ [ 36%]
|
||||||
|
........................................................................ [ 39%]
|
||||||
|
........................................................................ [ 42%]
|
||||||
|
........................................................................ [ 45%]
|
||||||
|
........................................................................ [ 48%]
|
||||||
|
........................................................................ [ 51%]
|
||||||
|
........................................................................ [ 54%]
|
||||||
|
........................................................................ [ 57%]
|
||||||
|
........................................................................ [ 60%]
|
||||||
|
........................................................................ [ 63%]
|
||||||
|
........................................................................ [ 66%]
|
||||||
|
........................................................................ [ 69%]
|
||||||
|
........................................................................ [ 72%]
|
||||||
|
........................................................................ [ 75%]
|
||||||
|
........................................................................ [ 78%]
|
||||||
|
........................................................................ [ 81%]
|
||||||
|
........................................................................ [ 84%]
|
||||||
|
........................................................................ [ 87%]
|
||||||
|
........................................................................ [ 91%]
|
||||||
|
........................................................................ [ 94%]
|
||||||
|
........................................................................ [ 97%]
|
||||||
|
..................................................................... [100%]
|
||||||
|
=============================== warnings summary ===============================
|
||||||
|
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
|
||||||
|
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||||
|
from starlette.testclient import TestClient as TestClient # noqa
|
||||||
|
|
||||||
|
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||||
|
================================ tests coverage ================================
|
||||||
|
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
|
||||||
|
|
||||||
|
Name Stmts Miss Cover
|
||||||
|
--------------------------------------------------
|
||||||
|
app/__init__.py 1 0 100%
|
||||||
|
app/api/__init__.py 0 0 100%
|
||||||
|
app/api/auth.py 52 0 100%
|
||||||
|
app/api/chat.py 205 1 99%
|
||||||
|
app/api/chats.py 110 0 100%
|
||||||
|
app/api/config.py 13 0 100%
|
||||||
|
app/api/doc_drafts.py 94 0 100%
|
||||||
|
app/api/docs.py 156 1 99%
|
||||||
|
app/api/git_sources.py 232 0 100%
|
||||||
|
app/api/health.py 10 0 100%
|
||||||
|
app/api/steering.py 42 0 100%
|
||||||
|
app/api/suggestions.py 33 0 100%
|
||||||
|
app/api/sync.py 139 0 100%
|
||||||
|
app/api/tokens.py 40 0 100%
|
||||||
|
app/api/ui_settings.py 55 0 100%
|
||||||
|
app/config.py 186 0 100%
|
||||||
|
app/core/__init__.py 0 0 100%
|
||||||
|
app/core/auth.py 45 0 100%
|
||||||
|
app/core/caching.py 124 0 100%
|
||||||
|
app/core/debugging.py 29 2 93%
|
||||||
|
app/core/docs_push.py 39 0 100%
|
||||||
|
app/core/errors.py 5 0 100%
|
||||||
|
app/core/logging.py 13 0 100%
|
||||||
|
app/core/rate_limit.py 44 0 100%
|
||||||
|
app/core/security_headers.py 20 0 100%
|
||||||
|
app/core/theming.py 38 0 100%
|
||||||
|
app/core/tokens.py 44 0 100%
|
||||||
|
app/db.py 22 0 100%
|
||||||
|
app/main.py 66 0 100%
|
||||||
|
app/models.py 128 0 100%
|
||||||
|
app/rag/__init__.py 0 0 100%
|
||||||
|
app/rag/agent.py 317 1 99%
|
||||||
|
app/rag/archive_upload.py 134 0 100%
|
||||||
|
app/rag/chunker.py 206 4 98%
|
||||||
|
app/rag/doc_dates.py 18 0 100%
|
||||||
|
app/rag/folder_summaries.py 123 0 100%
|
||||||
|
app/rag/git_sources.py 14 0 100%
|
||||||
|
app/rag/importer.py 215 3 99%
|
||||||
|
app/rag/llm.py 243 1 99%
|
||||||
|
app/rag/overview.py 71 0 100%
|
||||||
|
app/rag/prompts.py 88 0 100%
|
||||||
|
app/rag/retriever.py 172 3 98%
|
||||||
|
app/rag/scaffolding.py 55 0 100%
|
||||||
|
app/rag/source_removal.py 41 0 100%
|
||||||
|
app/rag/sources_meta.py 16 0 100%
|
||||||
|
app/rag/suggestions.py 27 0 100%
|
||||||
|
app/rag/summarizer.py 24 0 100%
|
||||||
|
app/schemas.py 327 0 100%
|
||||||
|
--------------------------------------------------
|
||||||
|
TOTAL 4076 16 99%
|
||||||
|
coverage gate: app/ 99% (>90%) OK
|
||||||
|
All checks passed!
|
||||||
|
0 errors, 0 warnings, 0 informations
|
||||||
|
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||||
|
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||||
|
|
||||||
|
validation OK
|
||||||
+19
@@ -0,0 +1,19 @@
|
|||||||
|
All green. Final verification complete.
|
||||||
|
|
||||||
|
**Task 02 (gate tests) — complete**
|
||||||
|
|
||||||
|
Implemented:
|
||||||
|
- Unit quadrants in `tests/unit/test_chat_gate.py`: added the code-defaults quadrant (fts>0 + cosine 0.50 ≥ floor 0.35 → HIGH) and pinned the new done-frame contract on the endpoint (deflected → `done["sources"] == []`, `query_log.sources` still records retrieval)
|
||||||
|
- E2E `test_honest_deflection.py`: new `test_out_of_kb_question_deflects_without_citations` — "Who composed Rhapsody in Blue?" (LLM-known, absent from fixture KB; fts_hits≥1, cosine 0.124 < floor 0.15 — the fts>0/cosine<floor quadrant pinned end-to-end) → `deflected: true`, `sources: []`, 2–3 suggestions
|
||||||
|
- Server: deflected turns' `done` frame now sends `sources=[]` (no false citations — the L2 acceptance; UI chips nothing); `query_log.sources` + per-turn log line keep recording the retrieval (observability unchanged, phase-113 A3 precedent)
|
||||||
|
- Updated the phase-04 integration test that pinned the old contract; reworded `test_retrieval_quality`'s kafkabridge question ("How does kafkabridge handle DNS?") — the old one (cos 0.134) now correctly falls in the new deflection quadrant
|
||||||
|
|
||||||
|
Results:
|
||||||
|
- `uv run pytest` → 2374 passed; `uv run pytest --cov=app --cov-report=term-missing` → 99% on `app/` (>90%)
|
||||||
|
- `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → 4 passed (in isolation); `test_chat_rag.py` → 3 passed; `test_retrieval_quality.py` → 4 passed
|
||||||
|
- At-risk audit: scanned every E2E question literal for gate flips (only the intended one + branch-independent/empty-KB suites); ran 13 gate-sensitive E2E suites (38 tests) → all passed
|
||||||
|
- `uv run ruff check .` → clean; `uv run pyright` → 0 errors, 0 warnings
|
||||||
|
|
||||||
|
Decisions: the task's example question scores 0.303 (> E2E threshold 0.30) under the mock, so it would ground — replaced with Rhapsody in Blue, which also exercises the new quadrant; deflected-done-`sources=[]` is a deliberate contract change required by the task's acceptance (matches phase 113's assumed end state).
|
||||||
|
|
||||||
|
Next pending task: `112_honesty_gate_weak_hits/03_prompt_contract_documentation.md`
|
||||||
+101
@@ -0,0 +1,101 @@
|
|||||||
|
........................................................................ [ 3%]
|
||||||
|
........................................................................ [ 6%]
|
||||||
|
........................................................................ [ 9%]
|
||||||
|
........................................................................ [ 12%]
|
||||||
|
........................................................................ [ 15%]
|
||||||
|
........................................................................ [ 18%]
|
||||||
|
........................................................................ [ 21%]
|
||||||
|
........................................................................ [ 24%]
|
||||||
|
........................................................................ [ 27%]
|
||||||
|
........................................................................ [ 30%]
|
||||||
|
........................................................................ [ 33%]
|
||||||
|
........................................................................ [ 36%]
|
||||||
|
........................................................................ [ 39%]
|
||||||
|
........................................................................ [ 42%]
|
||||||
|
........................................................................ [ 45%]
|
||||||
|
........................................................................ [ 48%]
|
||||||
|
........................................................................ [ 51%]
|
||||||
|
........................................................................ [ 54%]
|
||||||
|
........................................................................ [ 57%]
|
||||||
|
........................................................................ [ 60%]
|
||||||
|
........................................................................ [ 63%]
|
||||||
|
........................................................................ [ 66%]
|
||||||
|
........................................................................ [ 69%]
|
||||||
|
........................................................................ [ 72%]
|
||||||
|
........................................................................ [ 75%]
|
||||||
|
........................................................................ [ 78%]
|
||||||
|
........................................................................ [ 81%]
|
||||||
|
........................................................................ [ 84%]
|
||||||
|
........................................................................ [ 87%]
|
||||||
|
........................................................................ [ 90%]
|
||||||
|
........................................................................ [ 94%]
|
||||||
|
........................................................................ [ 97%]
|
||||||
|
...................................................................... [100%]
|
||||||
|
=============================== warnings summary ===============================
|
||||||
|
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
|
||||||
|
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||||
|
from starlette.testclient import TestClient as TestClient # noqa
|
||||||
|
|
||||||
|
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||||
|
================================ tests coverage ================================
|
||||||
|
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
|
||||||
|
|
||||||
|
Name Stmts Miss Cover
|
||||||
|
--------------------------------------------------
|
||||||
|
app/__init__.py 1 0 100%
|
||||||
|
app/api/__init__.py 0 0 100%
|
||||||
|
app/api/auth.py 52 0 100%
|
||||||
|
app/api/chat.py 208 1 99%
|
||||||
|
app/api/chats.py 110 0 100%
|
||||||
|
app/api/config.py 13 0 100%
|
||||||
|
app/api/doc_drafts.py 94 0 100%
|
||||||
|
app/api/docs.py 156 1 99%
|
||||||
|
app/api/git_sources.py 232 0 100%
|
||||||
|
app/api/health.py 10 0 100%
|
||||||
|
app/api/steering.py 42 0 100%
|
||||||
|
app/api/suggestions.py 33 0 100%
|
||||||
|
app/api/sync.py 139 0 100%
|
||||||
|
app/api/tokens.py 40 0 100%
|
||||||
|
app/api/ui_settings.py 55 0 100%
|
||||||
|
app/config.py 186 0 100%
|
||||||
|
app/core/__init__.py 0 0 100%
|
||||||
|
app/core/auth.py 45 0 100%
|
||||||
|
app/core/caching.py 124 0 100%
|
||||||
|
app/core/debugging.py 29 2 93%
|
||||||
|
app/core/docs_push.py 39 0 100%
|
||||||
|
app/core/errors.py 5 0 100%
|
||||||
|
app/core/logging.py 13 0 100%
|
||||||
|
app/core/rate_limit.py 44 0 100%
|
||||||
|
app/core/security_headers.py 20 0 100%
|
||||||
|
app/core/theming.py 38 0 100%
|
||||||
|
app/core/tokens.py 44 0 100%
|
||||||
|
app/db.py 22 0 100%
|
||||||
|
app/main.py 66 0 100%
|
||||||
|
app/models.py 128 0 100%
|
||||||
|
app/rag/__init__.py 0 0 100%
|
||||||
|
app/rag/agent.py 317 1 99%
|
||||||
|
app/rag/archive_upload.py 134 0 100%
|
||||||
|
app/rag/chunker.py 206 4 98%
|
||||||
|
app/rag/doc_dates.py 18 0 100%
|
||||||
|
app/rag/folder_summaries.py 123 0 100%
|
||||||
|
app/rag/git_sources.py 14 0 100%
|
||||||
|
app/rag/importer.py 215 3 99%
|
||||||
|
app/rag/llm.py 243 1 99%
|
||||||
|
app/rag/overview.py 71 0 100%
|
||||||
|
app/rag/prompts.py 88 0 100%
|
||||||
|
app/rag/retriever.py 172 3 98%
|
||||||
|
app/rag/scaffolding.py 55 0 100%
|
||||||
|
app/rag/source_removal.py 41 0 100%
|
||||||
|
app/rag/sources_meta.py 16 0 100%
|
||||||
|
app/rag/suggestions.py 27 0 100%
|
||||||
|
app/rag/summarizer.py 24 0 100%
|
||||||
|
app/schemas.py 327 0 100%
|
||||||
|
--------------------------------------------------
|
||||||
|
TOTAL 4079 16 99%
|
||||||
|
coverage gate: app/ 99% (>90%) OK
|
||||||
|
All checks passed!
|
||||||
|
0 errors, 0 warnings, 0 informations
|
||||||
|
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||||
|
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||||
|
|
||||||
|
validation OK
|
||||||
+10
@@ -0,0 +1,10 @@
|
|||||||
|
All green. Task 03 complete.
|
||||||
|
|
||||||
|
**Report**
|
||||||
|
|
||||||
|
- **prompts.py docstring**: dated entry recording the 2026-09-15 interactive-test finding (stochastic "Ulaanbaatar" disclosed answer vs. clean deflection on re-run, Mongolia case) and owner decision (iii, 2026-09-14): disclosed answers acceptable + documented; gate (A8 revised) is the deterministic lever; options (i)/(ii) left open. Prompt strings untouched (diff is docstring-only).
|
||||||
|
- **Prompt-lock pin**: new `tests/unit/test_prompt_lock.py` (4 tests) — sha256 + exact prefix/suffix + total length pinning `PERSONA` (706 B), `TOOLS_SECTION` (2273 B), the inline LOW `DEFLECT_MODE` body (259 B), and both `_base` lengths; any byte change fails.
|
||||||
|
- **README**: added the rare disclosed-answer sentence to the Honest-deflection bullet (work item 2). Defect fix (task 01 leftovers): refreshed the now-stale A8 gate copy in the retrieval paragraph, env table (added `BOR_LEXICAL_SUPPORT_FLOOR` row), tuning bullet, and two spots in `app/rag/retriever.py`.
|
||||||
|
- **Results**: `uv run pytest --cov=app --cov-report=term-missing` → 2378 passed, **99%** on `app/` (>90%); `uv run ruff check .` clean; `uv run pyright` 0 errors; `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → 4 passed in isolation.
|
||||||
|
- **Decisions**: kept the quoted *"I haven't done anything like that"* openings intact — task 04 rewrites those; my gate-copy fixes there will fold into it.
|
||||||
|
- **Next pending task**: `04_readme_deflection_copy.md`.
|
||||||
+102
@@ -0,0 +1,102 @@
|
|||||||
|
........................................................................ [ 3%]
|
||||||
|
........................................................................ [ 6%]
|
||||||
|
........................................................................ [ 9%]
|
||||||
|
........................................................................ [ 12%]
|
||||||
|
........................................................................ [ 15%]
|
||||||
|
........................................................................ [ 18%]
|
||||||
|
........................................................................ [ 21%]
|
||||||
|
........................................................................ [ 24%]
|
||||||
|
........................................................................ [ 27%]
|
||||||
|
........................................................................ [ 30%]
|
||||||
|
........................................................................ [ 33%]
|
||||||
|
........................................................................ [ 36%]
|
||||||
|
........................................................................ [ 39%]
|
||||||
|
........................................................................ [ 42%]
|
||||||
|
........................................................................ [ 45%]
|
||||||
|
........................................................................ [ 48%]
|
||||||
|
........................................................................ [ 51%]
|
||||||
|
........................................................................ [ 54%]
|
||||||
|
........................................................................ [ 57%]
|
||||||
|
........................................................................ [ 60%]
|
||||||
|
........................................................................ [ 63%]
|
||||||
|
........................................................................ [ 66%]
|
||||||
|
........................................................................ [ 69%]
|
||||||
|
........................................................................ [ 72%]
|
||||||
|
........................................................................ [ 75%]
|
||||||
|
........................................................................ [ 78%]
|
||||||
|
........................................................................ [ 81%]
|
||||||
|
........................................................................ [ 84%]
|
||||||
|
........................................................................ [ 87%]
|
||||||
|
........................................................................ [ 90%]
|
||||||
|
........................................................................ [ 93%]
|
||||||
|
........................................................................ [ 96%]
|
||||||
|
........................................................................ [ 99%]
|
||||||
|
.. [100%]
|
||||||
|
=============================== warnings summary ===============================
|
||||||
|
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
|
||||||
|
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||||
|
from starlette.testclient import TestClient as TestClient # noqa
|
||||||
|
|
||||||
|
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||||
|
================================ tests coverage ================================
|
||||||
|
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
|
||||||
|
|
||||||
|
Name Stmts Miss Cover
|
||||||
|
--------------------------------------------------
|
||||||
|
app/__init__.py 1 0 100%
|
||||||
|
app/api/__init__.py 0 0 100%
|
||||||
|
app/api/auth.py 52 0 100%
|
||||||
|
app/api/chat.py 208 1 99%
|
||||||
|
app/api/chats.py 110 0 100%
|
||||||
|
app/api/config.py 13 0 100%
|
||||||
|
app/api/doc_drafts.py 94 0 100%
|
||||||
|
app/api/docs.py 156 1 99%
|
||||||
|
app/api/git_sources.py 232 0 100%
|
||||||
|
app/api/health.py 10 0 100%
|
||||||
|
app/api/steering.py 42 0 100%
|
||||||
|
app/api/suggestions.py 33 0 100%
|
||||||
|
app/api/sync.py 139 0 100%
|
||||||
|
app/api/tokens.py 40 0 100%
|
||||||
|
app/api/ui_settings.py 55 0 100%
|
||||||
|
app/config.py 186 0 100%
|
||||||
|
app/core/__init__.py 0 0 100%
|
||||||
|
app/core/auth.py 45 0 100%
|
||||||
|
app/core/caching.py 124 0 100%
|
||||||
|
app/core/debugging.py 29 2 93%
|
||||||
|
app/core/docs_push.py 39 0 100%
|
||||||
|
app/core/errors.py 5 0 100%
|
||||||
|
app/core/logging.py 13 0 100%
|
||||||
|
app/core/rate_limit.py 44 0 100%
|
||||||
|
app/core/security_headers.py 20 0 100%
|
||||||
|
app/core/theming.py 38 0 100%
|
||||||
|
app/core/tokens.py 44 0 100%
|
||||||
|
app/db.py 22 0 100%
|
||||||
|
app/main.py 66 0 100%
|
||||||
|
app/models.py 128 0 100%
|
||||||
|
app/rag/__init__.py 0 0 100%
|
||||||
|
app/rag/agent.py 317 1 99%
|
||||||
|
app/rag/archive_upload.py 134 0 100%
|
||||||
|
app/rag/chunker.py 206 4 98%
|
||||||
|
app/rag/doc_dates.py 18 0 100%
|
||||||
|
app/rag/folder_summaries.py 123 0 100%
|
||||||
|
app/rag/git_sources.py 14 0 100%
|
||||||
|
app/rag/importer.py 215 3 99%
|
||||||
|
app/rag/llm.py 243 1 99%
|
||||||
|
app/rag/overview.py 71 0 100%
|
||||||
|
app/rag/prompts.py 88 0 100%
|
||||||
|
app/rag/retriever.py 172 3 98%
|
||||||
|
app/rag/scaffolding.py 55 0 100%
|
||||||
|
app/rag/source_removal.py 41 0 100%
|
||||||
|
app/rag/sources_meta.py 16 0 100%
|
||||||
|
app/rag/suggestions.py 27 0 100%
|
||||||
|
app/rag/summarizer.py 24 0 100%
|
||||||
|
app/schemas.py 327 0 100%
|
||||||
|
--------------------------------------------------
|
||||||
|
TOTAL 4079 16 99%
|
||||||
|
coverage gate: app/ 99% (>90%) OK
|
||||||
|
All checks passed!
|
||||||
|
0 errors, 0 warnings, 0 informations
|
||||||
|
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||||
|
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||||
|
|
||||||
|
validation OK
|
||||||
+12
@@ -0,0 +1,12 @@
|
|||||||
|
**Task 04 complete — README deflection copy refreshed (docs-only, no code changes).**
|
||||||
|
|
||||||
|
- Rewrote README intro (L11): now "it admits it — and offers 2-3 concrete alternative questions about things it **does** have notes on" (no exact-copy promise; the mandated opening was removed in the 2026-08-22 prompt revision).
|
||||||
|
- Rewrote Troubleshooting "Honest deflection" entry (L579): amber-bubble reference no longer quotes the removed opening; kept the accurate amber-banner + "Maybe try" chips description.
|
||||||
|
- Third stale occurrence (L588, the "reply opens with…" clause) found via the required README grep — updated to admit + 2-3 concrete alternatives. `grep -rn "haven't done anything" README.md` → no hits (exit 1).
|
||||||
|
- Tests grep: all `haven't done anything` hits in `tests/` are mock-LLM/stub-LLM *fixture* copies (the mock's own reply, which already matches the new contract; `DEFLECT_PHRASE` is documented as "the mock's deflection answer must match this") — none assert the app/prompt copy, so no test changes were needed.
|
||||||
|
- `uv run pytest` → 2378 passed.
|
||||||
|
- `uv run pytest --cov=app` → TOTAL 99% (>90% gate).
|
||||||
|
- `uv run ruff check .` → all checks passed; `uv run pyright` → 0 errors.
|
||||||
|
- `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → 4 passed (phase E2E green in isolation; last task of phase 112).
|
||||||
|
- No git add/commit (left in working tree for the harness phase commit).
|
||||||
|
- Next pending task: `.agents/phases/todo/113_source_chip_quality/01_usefulness_bar_sources.md`
|
||||||
+102
@@ -0,0 +1,102 @@
|
|||||||
|
........................................................................ [ 3%]
|
||||||
|
........................................................................ [ 6%]
|
||||||
|
........................................................................ [ 9%]
|
||||||
|
........................................................................ [ 12%]
|
||||||
|
........................................................................ [ 15%]
|
||||||
|
........................................................................ [ 18%]
|
||||||
|
........................................................................ [ 21%]
|
||||||
|
........................................................................ [ 24%]
|
||||||
|
........................................................................ [ 27%]
|
||||||
|
........................................................................ [ 30%]
|
||||||
|
........................................................................ [ 33%]
|
||||||
|
........................................................................ [ 36%]
|
||||||
|
........................................................................ [ 39%]
|
||||||
|
........................................................................ [ 42%]
|
||||||
|
........................................................................ [ 45%]
|
||||||
|
........................................................................ [ 48%]
|
||||||
|
........................................................................ [ 51%]
|
||||||
|
........................................................................ [ 54%]
|
||||||
|
........................................................................ [ 57%]
|
||||||
|
........................................................................ [ 60%]
|
||||||
|
........................................................................ [ 63%]
|
||||||
|
........................................................................ [ 66%]
|
||||||
|
........................................................................ [ 69%]
|
||||||
|
........................................................................ [ 72%]
|
||||||
|
........................................................................ [ 75%]
|
||||||
|
........................................................................ [ 78%]
|
||||||
|
........................................................................ [ 81%]
|
||||||
|
........................................................................ [ 84%]
|
||||||
|
........................................................................ [ 87%]
|
||||||
|
........................................................................ [ 90%]
|
||||||
|
........................................................................ [ 93%]
|
||||||
|
........................................................................ [ 96%]
|
||||||
|
........................................................................ [ 99%]
|
||||||
|
.. [100%]
|
||||||
|
=============================== warnings summary ===============================
|
||||||
|
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
|
||||||
|
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
|
||||||
|
from starlette.testclient import TestClient as TestClient # noqa
|
||||||
|
|
||||||
|
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
|
||||||
|
================================ tests coverage ================================
|
||||||
|
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
|
||||||
|
|
||||||
|
Name Stmts Miss Cover
|
||||||
|
--------------------------------------------------
|
||||||
|
app/__init__.py 1 0 100%
|
||||||
|
app/api/__init__.py 0 0 100%
|
||||||
|
app/api/auth.py 52 0 100%
|
||||||
|
app/api/chat.py 208 1 99%
|
||||||
|
app/api/chats.py 110 0 100%
|
||||||
|
app/api/config.py 13 0 100%
|
||||||
|
app/api/doc_drafts.py 94 0 100%
|
||||||
|
app/api/docs.py 156 1 99%
|
||||||
|
app/api/git_sources.py 232 0 100%
|
||||||
|
app/api/health.py 10 0 100%
|
||||||
|
app/api/steering.py 42 0 100%
|
||||||
|
app/api/suggestions.py 33 0 100%
|
||||||
|
app/api/sync.py 139 0 100%
|
||||||
|
app/api/tokens.py 40 0 100%
|
||||||
|
app/api/ui_settings.py 55 0 100%
|
||||||
|
app/config.py 186 0 100%
|
||||||
|
app/core/__init__.py 0 0 100%
|
||||||
|
app/core/auth.py 45 0 100%
|
||||||
|
app/core/caching.py 124 0 100%
|
||||||
|
app/core/debugging.py 29 2 93%
|
||||||
|
app/core/docs_push.py 39 0 100%
|
||||||
|
app/core/errors.py 5 0 100%
|
||||||
|
app/core/logging.py 13 0 100%
|
||||||
|
app/core/rate_limit.py 44 0 100%
|
||||||
|
app/core/security_headers.py 20 0 100%
|
||||||
|
app/core/theming.py 38 0 100%
|
||||||
|
app/core/tokens.py 44 0 100%
|
||||||
|
app/db.py 22 0 100%
|
||||||
|
app/main.py 66 0 100%
|
||||||
|
app/models.py 128 0 100%
|
||||||
|
app/rag/__init__.py 0 0 100%
|
||||||
|
app/rag/agent.py 317 1 99%
|
||||||
|
app/rag/archive_upload.py 134 0 100%
|
||||||
|
app/rag/chunker.py 206 4 98%
|
||||||
|
app/rag/doc_dates.py 18 0 100%
|
||||||
|
app/rag/folder_summaries.py 123 0 100%
|
||||||
|
app/rag/git_sources.py 14 0 100%
|
||||||
|
app/rag/importer.py 215 3 99%
|
||||||
|
app/rag/llm.py 243 1 99%
|
||||||
|
app/rag/overview.py 71 0 100%
|
||||||
|
app/rag/prompts.py 88 0 100%
|
||||||
|
app/rag/retriever.py 172 3 98%
|
||||||
|
app/rag/scaffolding.py 55 0 100%
|
||||||
|
app/rag/source_removal.py 41 0 100%
|
||||||
|
app/rag/sources_meta.py 16 0 100%
|
||||||
|
app/rag/suggestions.py 27 0 100%
|
||||||
|
app/rag/summarizer.py 24 0 100%
|
||||||
|
app/schemas.py 327 0 100%
|
||||||
|
--------------------------------------------------
|
||||||
|
TOTAL 4079 16 99%
|
||||||
|
coverage gate: app/ 99% (>90%) OK
|
||||||
|
All checks passed!
|
||||||
|
0 errors, 0 warnings, 0 informations
|
||||||
|
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
|
||||||
|
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
|
||||||
|
|
||||||
|
validation OK
|
||||||
+2
-1
@@ -34,7 +34,8 @@ BOR_STREAM_THINKING=1 # stream the model's thinking as `thinking` SS
|
|||||||
|
|
||||||
# --- RAG tuning ---
|
# --- RAG tuning ---
|
||||||
BOR_TOP_N_DOCS=2
|
BOR_TOP_N_DOCS=2
|
||||||
BOR_RELEVANCE_THRESHOLD=0.62 # answer when best cosine >= this OR an FTS hit; else honest deflection
|
BOR_RELEVANCE_THRESHOLD=0.62 # answer when best cosine >= this OR an FTS hit corroborated by cosine >= lexical_support_floor; else honest deflection
|
||||||
|
# BOR_LEXICAL_SUPPORT_FLOOR=0.35 # cosine floor for FTS hits to flip HIGH (A8 revised 2026-09-14); 0 <= floor <= relevance_threshold
|
||||||
BOR_MAX_OUTPUT_TOKENS=32768 # max answer length in tokens (answers must not be cut off)
|
BOR_MAX_OUTPUT_TOKENS=32768 # max answer length in tokens (answers must not be cut off)
|
||||||
BOR_STEERING_MAX_CHARS=8000 # char budget for the <tuning> (steering notes) prompt section
|
BOR_STEERING_MAX_CHARS=8000 # char budget for the <tuning> (steering notes) prompt section
|
||||||
BOR_SUMMARY_MAX_CHARS=12000 # cap on document content sent to the lite summary model (phase 30)
|
BOR_SUMMARY_MAX_CHARS=12000 # cap on document content sent to the lite summary model (phase 30)
|
||||||
|
|||||||
@@ -8,8 +8,8 @@ Reciprocal Rank Fusion), feeds the **whole relevant document** to a
|
|||||||
**self-hosted LLM** (`turbo` via `https://aipi.reeseapps.com/v1`), and
|
**self-hosted LLM** (`turbo` via `https://aipi.reeseapps.com/v1`), and
|
||||||
streams a grounded answer back.
|
streams a grounded answer back.
|
||||||
|
|
||||||
If it doesn't have notes for your question, it admits it: *"I haven't done
|
If it doesn't have notes for your question, it admits it — and offers
|
||||||
anything like that"* — plus suggestions for what it **does** know.
|
2-3 concrete alternative questions about things it **does** have notes on.
|
||||||
|
|
||||||
> **Updated your notes?** Re-run the import — it's idempotent and only
|
> **Updated your notes?** Re-run the import — it's idempotent and only
|
||||||
> re-embeds what changed:
|
> re-embeds what changed:
|
||||||
@@ -345,9 +345,12 @@ nearly double, which lets a name-your-tool question ("gitlab") find its own
|
|||||||
document even when the question embeds close to generic templates.
|
document even when the question embeds close to generic templates.
|
||||||
|
|
||||||
The **honesty gate** (A8) then answers (HIGH) when the best cosine is ≥
|
The **honesty gate** (A8) then answers (HIGH) when the best cosine is ≥
|
||||||
`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **or** at least one chunk matched
|
`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **or** when at least one chunk
|
||||||
lexically (`fts_hits > 0`) — it deflects (LOW) only when *both* signals are
|
matched lexically (`fts_hits > 0`) **and** the best cosine clears
|
||||||
absent. The top `BOR_TOP_N_DOCS` full documents are what the LLM sees.
|
`BOR_LEXICAL_SUPPORT_FLOOR` (default `0.35`) — the vector signal must
|
||||||
|
corroborate a lexical hit (A8 revised 2026-09-14). It deflects (LOW)
|
||||||
|
otherwise, including a weak single-token hit with vector-unsupported
|
||||||
|
docs. The top `BOR_TOP_N_DOCS` full documents are what the LLM sees.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -412,7 +415,8 @@ asset references of the known pages in flight.
|
|||||||
| `BOR_LLM_SUMMARY_MODEL` | `lite` | One-shot completions: document summaries at import, KB overview |
|
| `BOR_LLM_SUMMARY_MODEL` | `lite` | One-shot completions: document summaries at import, KB overview |
|
||||||
| `BOR_EMBEDDING_DIM` | `768` | Vector dimension (fixed at table creation) |
|
| `BOR_EMBEDDING_DIM` | `768` | Vector dimension (fixed at table creation) |
|
||||||
| `BOR_TOP_N_DOCS` | `2` | Full documents fed to the LLM |
|
| `BOR_TOP_N_DOCS` | `2` | Full documents fed to the LLM |
|
||||||
| `BOR_RELEVANCE_THRESHOLD` | `0.62` | Answer when best cosine ≥ this **or** an FTS hit; below + no FTS ⇒ honest deflection |
|
| `BOR_RELEVANCE_THRESHOLD` | `0.62` | Answer when best cosine ≥ this **or** a lexical hit corroborated by cosine ≥ `BOR_LEXICAL_SUPPORT_FLOOR`; otherwise honest deflection |
|
||||||
|
| `BOR_LEXICAL_SUPPORT_FLOOR` | `0.35` | Best-cosine floor an FTS hit must clear to flip the gate HIGH (A8 revised 2026-09-14) |
|
||||||
| `BOR_HYBRID_VECTOR_CANDIDATES` | `100` | Cosine list width for the RRF fusion |
|
| `BOR_HYBRID_VECTOR_CANDIDATES` | `100` | Cosine list width for the RRF fusion |
|
||||||
| `BOR_HYBRID_LEXICAL_CANDIDATES` | `30` | FTS list width for the RRF fusion |
|
| `BOR_HYBRID_LEXICAL_CANDIDATES` | `30` | FTS list width for the RRF fusion |
|
||||||
| `BOR_RRF_K` | `60` | RRF damping constant (`1/(k + rank)`) |
|
| `BOR_RRF_K` | `60` | RRF damping constant (`1/(k + rank)`) |
|
||||||
@@ -572,22 +576,32 @@ served locally (no CDN), `BOR_ENVIRONMENT=production`.
|
|||||||
- **Embedding dimension mismatch** — aipi changed models; run
|
- **Embedding dimension mismatch** — aipi changed models; run
|
||||||
`uv run python -m scripts.llm_probe`, update `BOR_EMBEDDING_DIM`, then
|
`uv run python -m scripts.llm_probe`, update `BOR_EMBEDDING_DIM`, then
|
||||||
drop + recreate the chunks table.
|
drop + recreate the chunks table.
|
||||||
- **Honest deflection** (the amber *"I haven't done anything like that"*
|
- **Honest deflection** (the amber deflection bubble) — every question
|
||||||
bubble) — every question passes the honesty gate: deflection happens only
|
passes the honesty gate: deflection happens
|
||||||
when the best cosine similarity is below `BOR_RELEVANCE_THRESHOLD`
|
when the best cosine similarity is below `BOR_RELEVANCE_THRESHOLD`
|
||||||
(default `0.62`) **and** no chunk matched lexically (`fts_hits = 0`).
|
(default `0.62`) **and** the lexical signal is absent or not
|
||||||
A weak cosine with a lexical hit (name-your-tool questions) still gets a
|
vector-corroborated (`fts_hits = 0`, or best cosine below
|
||||||
grounded answer. When it does deflect, the LLM prompt carries weak-hit
|
`BOR_LEXICAL_SUPPORT_FLOOR` — default `0.35` — A8 revised 2026-09-14).
|
||||||
*titles only* (no document content), the reply opens with *"I haven't
|
A weak cosine with a lexical hit (name-your-tool questions) still gets
|
||||||
done anything like that"*, the bubble renders amber with *"Maybe try"*
|
a grounded answer, provided the best cosine clears the floor. When it
|
||||||
chips derived from the closest indexed titles, and the `query_log` row
|
does deflect, the LLM prompt carries weak-hit *titles only* (no document
|
||||||
records `deflected=true`. This is a feature, not a bug.
|
content), the reply admits it has no notes on that and offers 2-3
|
||||||
|
concrete alternative questions about things it **does** have notes on,
|
||||||
|
the bubble renders amber with *"Maybe try"* chips derived from the
|
||||||
|
closest indexed titles, and the `query_log` row records
|
||||||
|
`deflected=true`. This is a feature, not a bug. With a small local
|
||||||
|
model, a rare turn may still answer from general knowledge with an
|
||||||
|
explicit disclosure when retrieval was borderline — the gate (A8
|
||||||
|
revised) minimizes this by keeping vector-unsupported docs out of the
|
||||||
|
grounded prompt, and the disclosure is surfaced, never silent.
|
||||||
- **Answers deflect too often / too rarely** — tune
|
- **Answers deflect too often / too rarely** — tune
|
||||||
`BOR_RELEVANCE_THRESHOLD` (lower = answers more, higher = more honest
|
`BOR_RELEVANCE_THRESHOLD` (lower = answers more, higher = more honest
|
||||||
deflection): `0.0` ⇒ the gate leans entirely on FTS hits; `1.0` ⇒
|
deflection) and `BOR_LEXICAL_SUPPORT_FLOOR` (the best-cosine floor a
|
||||||
everything deflects unless a chunk matches lexically. The `embed` model's
|
lexical hit must clear to flip HIGH — lower = lexical hits ground more
|
||||||
cosines cluster in a ~0.6–0.85 band on the live KB, so the default is
|
easily): `0.0` ⇒ the gate leans entirely on FTS hits; `1.0` ⇒
|
||||||
`0.62`. Check real scores:
|
everything deflects unless a lexical hit is vector-corroborated. The
|
||||||
|
`embed` model's cosines cluster in a ~0.6–0.85 band on the live KB, so
|
||||||
|
the default is `0.62`. Check real scores:
|
||||||
```sql
|
```sql
|
||||||
SELECT question, top_score, fts_hits, deflected
|
SELECT question, top_score, fts_hits, deflected
|
||||||
FROM query_log ORDER BY created_at DESC LIMIT 20;
|
FROM query_log ORDER BY created_at DESC LIMIT 20;
|
||||||
|
|||||||
+49
-26
@@ -15,14 +15,22 @@ turn's thinking is counted in the per-turn log line
|
|||||||
(``thinking_chars=N``); ``BOR_STREAM_THINKING=0`` suppresses the
|
(``thinking_chars=N``); ``BOR_STREAM_THINKING=0`` suppresses the
|
||||||
``thinking`` frames (the pieces are still counted).
|
``thinking`` frames (the pieces are still counted).
|
||||||
|
|
||||||
Honesty gate (A8, revised 2026-08-21): LOW — deflection — only when the
|
Honesty gate (A8, revised 2026-09-14): LOW — deflection — only when
|
||||||
best cosine is strictly below ``BOR_RELEVANCE_THRESHOLD`` **and** no
|
``best_cosine < BOR_RELEVANCE_THRESHOLD`` **and** (``fts_hits == 0``
|
||||||
candidate chunk FTS-matches the question (``fts_hits == 0``). A
|
or ``best_cosine < BOR_LEXICAL_SUPPORT_FLOOR``). A lexical hit alone
|
||||||
name-your-tool question with weak vector overlap but a lexical hit still
|
(cosine < lexical_support_floor) no longer promotes to HIGH — the vector
|
||||||
gets a grounded answer. Deflection mode carries weak-hit *titles only*
|
signal must corroborate the lexical match (A8 revised 2026-09-14, the
|
||||||
(never document content) plus deterministic "Maybe try" chips, and the
|
"Mongolia" fix). HIGH fires when ``best_cosine >= threshold`` OR
|
||||||
``done`` event / ``query_log`` row record ``deflected=true``, the weak
|
(fts>0 AND cosine >= lexical_support_floor). Deflection mode carries
|
||||||
score and the ``fts_hits`` count.
|
weak-hit *titles only* (never document content) plus deterministic
|
||||||
|
"Maybe try" chips, and the ``done`` event / ``query_log`` row record
|
||||||
|
``deflected=true``, the weak score and the ``fts_hits`` count. A
|
||||||
|
deflected turn cites nothing: the ``done`` frame's ``sources`` is
|
||||||
|
``[]`` (``done.sources`` is the citation surface — the UI chips every
|
||||||
|
entry as "the answer used this" — and weak hits are scored docs, not
|
||||||
|
citations, TODO L2), while ``query_log.sources`` and the per-turn log
|
||||||
|
line keep recording the retrieval for tuning (phase 112, A8 revised).
|
||||||
|
|
||||||
|
|
||||||
Steering (phase 15): the owner's stored tuning notes are loaded per turn
|
Steering (phase 15): the owner's stored tuning notes are loaded per turn
|
||||||
(oldest first) and injected into the system prompt as a ``<tuning>``
|
(oldest first) and injected into the system prompt as a ``<tuning>``
|
||||||
@@ -270,13 +278,15 @@ def plan_turn(
|
|||||||
"""Apply the honesty gate (A8, revised) and assemble prompt + context.
|
"""Apply the honesty gate (A8, revised) and assemble prompt + context.
|
||||||
|
|
||||||
* **HIGH (grounded)** when ``best_cosine >= threshold`` **or**
|
* **HIGH (grounded)** when ``best_cosine >= threshold`` **or**
|
||||||
``fts_hits > 0``: HIGH prompt with the full top-N documents, no
|
(``fts_hits > 0`` **and** ``best_cosine >= lexical_support_floor``):
|
||||||
suggestions. A cosine exactly at the threshold is an answer — the
|
HIGH prompt with the full top-N documents, no suggestions. A cosine
|
||||||
gate is strict (``< threshold``).
|
exactly at the threshold is an answer — the gate is strict
|
||||||
* **LOW (deflected)** only when ``best_cosine < threshold`` **and**
|
(``< threshold``). An FTS hit alone, without vector corroboration
|
||||||
``fts_hits == 0`` (or no hits at all): LOW prompt (``DEFLECT_MODE``)
|
(cosine < lexical_support_floor), stays LOW (A8 revised 2026-09-14).
|
||||||
with weak-hit titles only — never document content — plus
|
* **LOW (deflected)** otherwise — including the fts>0 / cosine < floor
|
||||||
deterministic alternative-question chips derived from those titles.
|
case (the "Mongolia" case): LOW prompt (``DEFLECT_MODE``) with
|
||||||
|
weak-hit titles only — never document content — plus deterministic
|
||||||
|
alternative-question chips derived from those titles.
|
||||||
|
|
||||||
``top_score`` (stored in ``query_log``) is the best cosine, so the
|
``top_score`` (stored in ``query_log``) is the best cosine, so the
|
||||||
gate input is always a pure vector-similarity number; the lexical
|
gate input is always a pure vector-similarity number; the lexical
|
||||||
@@ -305,7 +315,8 @@ def plan_turn(
|
|||||||
docs = select_documents(chunks, n=settings.top_n_docs)
|
docs = select_documents(chunks, n=settings.top_n_docs)
|
||||||
selected_ids = {d.id for d in docs}
|
selected_ids = {d.id for d in docs}
|
||||||
summary_hits = sum(1 for c in chunks if c.is_summary and c.document.id in selected_ids)
|
summary_hits = sum(1 for c in chunks if c.is_summary and c.document.id in selected_ids)
|
||||||
if best_cosine >= settings.relevance_threshold or fts_hits > 0:
|
lexical_supported = fts_hits > 0 and best_cosine >= settings.lexical_support_floor
|
||||||
|
if best_cosine >= settings.relevance_threshold or lexical_supported:
|
||||||
return TurnPlan(
|
return TurnPlan(
|
||||||
best_cosine,
|
best_cosine,
|
||||||
fts_hits,
|
fts_hits,
|
||||||
@@ -736,10 +747,14 @@ async def chat(
|
|||||||
|
|
||||||
# 4. Durable record + required per-turn log line (PLAN §9).
|
# 4. Durable record + required per-turn log line (PLAN §9).
|
||||||
# Phase 37: the agent's read documents join the
|
# Phase 37: the agent's read documents join the
|
||||||
# retrieval's — deduped by (source, path), order preserved
|
# retrieval's — deduped by (source, path), order preserved.
|
||||||
# — and the same combined list feeds done.sources,
|
# The combined list feeds query_log.sources and the log
|
||||||
# query_log.sources and the log line (empty on deflected
|
# line (retrieval docs even on deflected turns —
|
||||||
# turns: the agent never runs). A cancelled turn (the
|
# observability, the phase-113 A3 precedent). Phase 112:
|
||||||
|
# done.sources is the CITATION surface — it carries the
|
||||||
|
# combined list on grounded turns and [] on deflected
|
||||||
|
# ones (a deflected answer cites nothing; the weak hits
|
||||||
|
# stay in the durable record). A cancelled turn (the
|
||||||
# generator closed by the consumer) never reaches this
|
# generator closed by the consumer) never reaches this
|
||||||
# step — no query_log row.
|
# step — no query_log row.
|
||||||
cited_docs: list[Document] = []
|
cited_docs: list[Document] = []
|
||||||
@@ -790,15 +805,23 @@ async def chat(
|
|||||||
scaffold_stripped,
|
scaffold_stripped,
|
||||||
)
|
)
|
||||||
settled = True # terminal: the done frame settles the turn
|
settled = True # terminal: the done frame settles the turn
|
||||||
|
# Phase 112 (A8 revised, TODO L2): a deflected turn cites
|
||||||
|
# nothing — done.sources is [] (the UI chips every entry
|
||||||
|
# as a citation; the weak hits are scored docs, not
|
||||||
|
# citations). The retrieval stays durably recorded above
|
||||||
|
# (query_log.sources + the log line — observability
|
||||||
|
# unchanged); the phase-113 related-doc tier is the home
|
||||||
|
# for the weak hits' visibility.
|
||||||
|
cited_refs: list[SourceRef] = []
|
||||||
|
if not plan.deflected:
|
||||||
|
cited_refs = [
|
||||||
|
SourceRef(source=d.source, path=d.path, title=d.title)
|
||||||
|
for d in cited_docs
|
||||||
|
]
|
||||||
yield sse_event(
|
yield sse_event(
|
||||||
ChatDoneEvent(
|
ChatDoneEvent(
|
||||||
deflected=plan.deflected,
|
deflected=plan.deflected,
|
||||||
sources=[
|
sources=cited_refs,
|
||||||
SourceRef(
|
|
||||||
source=d.source, path=d.path, title=d.title
|
|
||||||
)
|
|
||||||
for d in cited_docs
|
|
||||||
],
|
|
||||||
suggestions=plan.suggestions,
|
suggestions=plan.suggestions,
|
||||||
).model_dump()
|
).model_dump()
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -117,6 +117,15 @@ class Settings(BaseSettings):
|
|||||||
# default never discriminated. LOW only fires when best cosine < this
|
# default never discriminated. LOW only fires when best cosine < this
|
||||||
# AND no candidate chunk matches the question lexically (see A8).
|
# AND no candidate chunk matches the question lexically (see A8).
|
||||||
relevance_threshold: float = 0.62
|
relevance_threshold: float = 0.62
|
||||||
|
#: Lexical support floor (A8 revised 2026-09-14): an FTS hit promotes
|
||||||
|
# a turn to HIGH only when the best cosine is >= this value — it
|
||||||
|
# requires the vector signal to corroborate the lexical match. A
|
||||||
|
# single weak token hit with vector-unsupported docs (cosine < floor)
|
||||||
|
# stays LOW (deflected). Default 0.35 ≈ half the relevance threshold;
|
||||||
|
# tunable via ``BOR_LEXICAL_SUPPORT_FLOOR``. Must be <=
|
||||||
|
# ``relevance_threshold`` (a floor above the threshold is a typo that
|
||||||
|
# would make every FTS hit require a HIGH cosine anyway).
|
||||||
|
lexical_support_floor: float = 0.35
|
||||||
#: Maximum output tokens a chat answer may use (owner instruction
|
#: Maximum output tokens a chat answer may use (owner instruction
|
||||||
#: 2026-08-22: answers must run to their natural end — the old hard
|
#: 2026-08-22: answers must run to their natural end — the old hard
|
||||||
#: 700-token cap cut long answers off mid-sentence).
|
#: 700-token cap cut long answers off mid-sentence).
|
||||||
@@ -302,6 +311,22 @@ class Settings(BaseSettings):
|
|||||||
#: separate from ``sources_dir`` (the source checkouts).
|
#: separate from ``sources_dir`` (the source checkouts).
|
||||||
docs_work_dir: str = "~/bor-docs"
|
docs_work_dir: str = "~/bor-docs"
|
||||||
|
|
||||||
|
@field_validator("lexical_support_floor")
|
||||||
|
@classmethod
|
||||||
|
def _lexical_support_floor_bounds(cls, v: float, info: ValidationInfo) -> float:
|
||||||
|
"""The lexical support floor must be in [0, relevance_threshold].
|
||||||
|
A value above the relevance threshold would be a typo — it would
|
||||||
|
make every FTS hit require a HIGH cosine anyway, defeating the
|
||||||
|
purpose of the floor (A8 revised 2026-09-14)."""
|
||||||
|
if v < 0:
|
||||||
|
raise ValueError("lexical_support_floor must be >= 0")
|
||||||
|
threshold = info.data.get("relevance_threshold")
|
||||||
|
if isinstance(threshold, float) and v > threshold:
|
||||||
|
raise ValueError(
|
||||||
|
f"lexical_support_floor ({v}) must be <= relevance_threshold ({threshold})"
|
||||||
|
)
|
||||||
|
return v
|
||||||
|
|
||||||
@field_validator("import_extensions")
|
@field_validator("import_extensions")
|
||||||
@classmethod
|
@classmethod
|
||||||
def _import_extensions_known(cls, v: str) -> str:
|
def _import_extensions_known(cls, v: str) -> str:
|
||||||
|
|||||||
@@ -58,6 +58,27 @@ recovery in :mod:`app.rag.scaffolding` / :mod:`app.rag.agent` is the
|
|||||||
backstop). The ``DEFLECT_MODE`` marker and everything else in the
|
backstop). The ``DEFLECT_MODE`` marker and everything else in the
|
||||||
prompt stay put — the E2E mock LLM keys on the marker's *presence*,
|
prompt stay put — the E2E mock LLM keys on the marker's *presence*,
|
||||||
not the wording, so that contract is unchanged.
|
not the wording, so that contract is unchanged.
|
||||||
|
|
||||||
|
Disclosed-answer contract (phase 112, task 03; owner decision
|
||||||
|
2026-09-14, roadmap confirmation — option (iii) of the three the
|
||||||
|
2026-09-15 interactive deflection test raised): that test found the
|
||||||
|
HONESTY GATE's compliance is **stochastic** across runs when
|
||||||
|
misleading context is injected — the "What is the capital of
|
||||||
|
Mongolia?" question had two *irrelevant* docs promoted into the HIGH
|
||||||
|
prompt by a weak single-token FTS hit (the pre-phase A8 rule: any
|
||||||
|
``fts_hits > 0`` flipped HIGH), and run 1 answered parametrically —
|
||||||
|
"Ulaanbaatar" — *with an explicit disclosure*, while the identical
|
||||||
|
one-tap re-run produced a clean, textbook deflection. The owner's
|
||||||
|
decision: the disclosed general-knowledge answer is treated as
|
||||||
|
**acceptable and documented** — a small local model cannot be relied
|
||||||
|
on to obey Rules 1/3 100% when handed misleading context, so this
|
||||||
|
prompt text stays byte-identical (LOCKED verbatim; options (i) tighten
|
||||||
|
the copy / (ii) amend this locked prompt via the plan to explicitly
|
||||||
|
permit disclosed general-knowledge answers remain open to a future
|
||||||
|
owner decision). The deterministic lever is the honesty gate itself
|
||||||
|
(A8 revised 2026-09-14: an FTS hit flips HIGH only when
|
||||||
|
``best_cosine >= lexical_support_floor``, so the misleading docs are
|
||||||
|
no longer injected — see :mod:`app.api.chat`).
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
|||||||
@@ -21,8 +21,9 @@
|
|||||||
match a document's ``qwen3``/``8``/``27b`` tokens, while every
|
match a document's ``qwen3``/``8``/``27b`` tokens, while every
|
||||||
unrelated llama.cpp quadlet out-ranks the target on the shared
|
unrelated llama.cpp quadlet out-ranks the target on the shared
|
||||||
``llama``/``cpp`` tokens). The name-hit rows carry ``fts_hit=True``
|
``llama``/``cpp`` tokens). The name-hit rows carry ``fts_hit=True``
|
||||||
(they ARE the lexical signal — the A8 honesty gate then answers
|
(they ARE the lexical signal — the A8 honesty gate answers on them
|
||||||
instead of deflecting) and ``cosine=0.0``; the RRF fusion is
|
only when the best cosine clears ``BOR_LEXICAL_SUPPORT_FLOOR``,
|
||||||
|
A8 revised 2026-09-14) and ``cosine=0.0``; the RRF fusion is
|
||||||
unchanged (same lists, same ``1/(k+rank)`` terms).
|
unchanged (same lists, same ``1/(k+rank)`` terms).
|
||||||
* **Fusion** — Reciprocal Rank Fusion (``score = Σ 1/(k + rank)`` over the
|
* **Fusion** — Reciprocal Rank Fusion (``score = Σ 1/(k + rank)`` over the
|
||||||
lists a chunk appears in; chunks hit by both lists get both terms). The
|
lists a chunk appears in; chunks hit by both lists get both terms). The
|
||||||
@@ -445,7 +446,7 @@ def _name_hit_chunks(db: Session, question: str) -> list[RetrievedChunk]:
|
|||||||
score=0.0, # filled in by :func:`fuse`
|
score=0.0, # filled in by :func:`fuse`
|
||||||
document=doc,
|
document=doc,
|
||||||
cosine=0.0, # no vector rank — name-only hit
|
cosine=0.0, # no vector rank — name-only hit
|
||||||
fts_hit=True, # the lexical signal — the A8 gate answers
|
fts_hit=True, # lexical signal — A8 answers if cosine corroborates
|
||||||
is_summary=bool(row.is_summary),
|
is_summary=bool(row.is_summary),
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -15,6 +15,8 @@ from sqlalchemy.orm import Session
|
|||||||
# conftest.py. Must be set before ``app.main`` (below) caches settings.
|
# conftest.py. Must be set before ``app.main`` (below) caches settings.
|
||||||
# The production default stays 0.62 (app/config.py, A8 revised).
|
# The production default stays 0.62 (app/config.py, A8 revised).
|
||||||
os.environ.setdefault("BOR_RELEVANCE_THRESHOLD", "0.30")
|
os.environ.setdefault("BOR_RELEVANCE_THRESHOLD", "0.30")
|
||||||
|
# A8 revised 2026-09-14: lexical_support_floor must be <= relevance_threshold.
|
||||||
|
os.environ.setdefault("BOR_LEXICAL_SUPPORT_FLOOR", "0.15")
|
||||||
|
|
||||||
# Phase 16: single-admin auth is fail-loud — create_app() refuses to boot
|
# Phase 16: single-admin auth is fail-loud — create_app() refuses to boot
|
||||||
# without both vars, and app.main (imported below) builds the app at
|
# without both vars, and app.main (imported below) builds the app at
|
||||||
|
|||||||
@@ -34,6 +34,17 @@ from e2e.auth_helpers import ADMIN_PASSWORD, login
|
|||||||
REPO = Path(__file__).resolve().parents[2]
|
REPO = Path(__file__).resolve().parents[2]
|
||||||
FIXTURES = REPO / "tests" / "fixtures" / "docs"
|
FIXTURES = REPO / "tests" / "fixtures" / "docs"
|
||||||
OFF_TOPIC = "How do I bake sourdough bread?"
|
OFF_TOPIC = "How do I bake sourdough bread?"
|
||||||
|
# Phase 112 (A8 revised 2026-09-14, TODO L2a — the "Mongolia case"): a
|
||||||
|
# question the LLM knows (Gershwin) but the fixture KB does not cover.
|
||||||
|
# Unlike the plain no-hit deflection above, it carries WEAK FTS hits
|
||||||
|
# (the "compos" stem matches the compose fixture docs — fts_hits >= 1)
|
||||||
|
# while its best mock token-overlap cosine (~0.12) sits BELOW
|
||||||
|
# lexical_support_floor (0.15, the mock-calibrated conftest value). The
|
||||||
|
# pre-phase gate (fts>0 → HIGH) grounded it and injected irrelevant docs
|
||||||
|
# into the prompt; the revised gate (cosine corroboration) must keep it
|
||||||
|
# LOW. The mock keys on DEFLECT_MODE, so the test pins the gate, not
|
||||||
|
# model compliance.
|
||||||
|
OUT_OF_KB = "Who composed Rhapsody in Blue?"
|
||||||
# The mock's deflection answer (tests/e2e/mock_llm.py) must match this.
|
# The mock's deflection answer (tests/e2e/mock_llm.py) must match this.
|
||||||
DEFLECT_PHRASE = r"haven't done anything like that"
|
DEFLECT_PHRASE = r"haven't done anything like that"
|
||||||
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
|
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
|
||||||
@@ -201,3 +212,67 @@ def test_deflected_done_event_and_query_log(app_url: str, mock_llm: int, db_read
|
|||||||
assert row.deflected is True
|
assert row.deflected is True
|
||||||
assert 0.0 < row.top_score < get_settings().relevance_threshold
|
assert 0.0 < row.top_score < get_settings().relevance_threshold
|
||||||
assert row.chunk_hits >= 1
|
assert row.chunk_hits >= 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_out_of_kb_question_deflects_without_citations(
|
||||||
|
app_url: str, mock_llm: int, db_ready: None
|
||||||
|
) -> None:
|
||||||
|
"""Phase 112 acceptance (TODO L2): a known-out-of-KB question whose
|
||||||
|
weak lexical hits NO LONGER promote (the fts>0 / cosine<floor
|
||||||
|
quadrant, pinned end-to-end) deflects with ZERO source citations —
|
||||||
|
done.sources is empty (the UI chips nothing under a deflected
|
||||||
|
answer) and 2-3 concrete alternative questions are offered.
|
||||||
|
|
||||||
|
Raw SSE (like the done-event test above): the done frame is the
|
||||||
|
contract surface; the mock's DEFLECT_MODE phrasing proves the
|
||||||
|
server sent the LOW prompt (the gate, not the model, decides).
|
||||||
|
"""
|
||||||
|
_reset_db(mock_llm, seed=True)
|
||||||
|
|
||||||
|
client = httpx.Client(timeout=60.0)
|
||||||
|
r = client.post(f"{app_url}/api/login", json={"password": ADMIN_PASSWORD})
|
||||||
|
assert r.status_code == 204
|
||||||
|
|
||||||
|
frames: list[dict[str, Any]] = []
|
||||||
|
with client.stream(
|
||||||
|
"POST", f"{app_url}/api/chat", json={"message": OUT_OF_KB}, timeout=60.0
|
||||||
|
) as r:
|
||||||
|
assert r.status_code == 200
|
||||||
|
assert r.headers["content-type"].startswith("text/event-stream")
|
||||||
|
buf = ""
|
||||||
|
for part in r.iter_text():
|
||||||
|
buf += part
|
||||||
|
while "\n\n" in buf:
|
||||||
|
frame, buf = buf.split("\n\n", 1)
|
||||||
|
if frame.strip().startswith("data:"):
|
||||||
|
frames.append(json.loads(frame.strip().removeprefix("data:").strip()))
|
||||||
|
assert buf.strip() == "" # stream ends cleanly on a frame boundary
|
||||||
|
|
||||||
|
deltas = [f for f in frames if f.get("type") == "delta"]
|
||||||
|
answer = "".join(d["text"] for d in deltas)
|
||||||
|
# The DEFLECT_MODE phrasing streamed ⇒ the LOW prompt reached the
|
||||||
|
# model (the mock answers it only for the deflection system prompt).
|
||||||
|
assert re.search(DEFLECT_PHRASE, answer, re.IGNORECASE)
|
||||||
|
|
||||||
|
done = frames[-1]
|
||||||
|
assert done["type"] == "done"
|
||||||
|
assert done["deflected"] is True
|
||||||
|
# No false citations (TODO L2): the weak hits never ride the wire as
|
||||||
|
# sources — a deflected answer cites nothing.
|
||||||
|
assert done["sources"] == []
|
||||||
|
# 2-3 concrete alternative questions, all non-empty.
|
||||||
|
assert 2 <= len(done["suggestions"]) <= 3
|
||||||
|
assert all(s.strip() for s in done["suggestions"])
|
||||||
|
|
||||||
|
# Durable record: the quadrant pinned end-to-end — the lexical leg
|
||||||
|
# FIRED (fts_hits > 0, the pre-phase gate's promotion trigger) while
|
||||||
|
# the vector signal never cleared lexical_support_floor, so the
|
||||||
|
# revised gate deflected. The retrieval itself stays recorded
|
||||||
|
# (query_log = observability, not citations).
|
||||||
|
with SessionLocal() as db:
|
||||||
|
row = db.scalars(select(QueryLog)).one()
|
||||||
|
assert row.question == OUT_OF_KB
|
||||||
|
assert row.deflected is True
|
||||||
|
assert (row.fts_hits or 0) >= 1
|
||||||
|
assert row.top_score < get_settings().lexical_support_floor
|
||||||
|
assert row.sources
|
||||||
|
|||||||
@@ -13,7 +13,8 @@ The four tests map the story's acceptance criteria:
|
|||||||
2. "How did I install gitlab?" — grounded (not deflected), gitlab chip,
|
2. "How did I install gitlab?" — grounded (not deflected), gitlab chip,
|
||||||
``query_log`` row with the gitlab doc in ``sources``
|
``query_log`` row with the gitlab doc in ``sources``
|
||||||
3. keyword-only question ("kafkabridge") beats the vector ranking — the
|
3. keyword-only question ("kafkabridge") beats the vector ranking — the
|
||||||
FTS-OR gate grounds it end to end despite weak cosine
|
corroborated-lexical gate (A8 revised 2026-09-14) grounds it end to
|
||||||
|
end: weak cosine, but an FTS hit AND cosine >= lexical_support_floor
|
||||||
4. "sourdough" — deflected bubble + ≥2 "Maybe try" chips
|
4. "sourdough" — deflected bubble + ≥2 "Maybe try" chips
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
@@ -37,7 +38,15 @@ from e2e.auth_helpers import login
|
|||||||
REPO = Path(__file__).resolve().parents[2]
|
REPO = Path(__file__).resolve().parents[2]
|
||||||
FIXTURES = REPO / "tests" / "fixtures" / "docs"
|
FIXTURES = REPO / "tests" / "fixtures" / "docs"
|
||||||
GITLAB_QUESTION = "How did I install gitlab?"
|
GITLAB_QUESTION = "How did I install gitlab?"
|
||||||
KEYWORD_QUESTION = "How does kafkabridge work?"
|
# Phase 112 (A8 revised): the pre-phase question ("How does kafkabridge
|
||||||
|
# work?" — mock cosine 0.134) now sits BELOW lexical_support_floor
|
||||||
|
# (0.15, mock-calibrated) with fts>0 — the new gate's deflection
|
||||||
|
# quadrant, so it can no longer demonstrate the grounded lexical path.
|
||||||
|
# "handle DNS" adds static-dns.json's own tokens: cosine ≈0.24 — still
|
||||||
|
# weak (below the 0.30 threshold) yet corroborated (>= floor) with the
|
||||||
|
# same single-doc FTS hit, and the FTS-matched doc still tops the fused
|
||||||
|
# ranking (the test's actual assertion).
|
||||||
|
KEYWORD_QUESTION = "How does kafkabridge handle DNS?"
|
||||||
OFF_TOPIC = "sourdough starter"
|
OFF_TOPIC = "sourdough starter"
|
||||||
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
|
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
|
||||||
|
|
||||||
@@ -162,9 +171,11 @@ def test_gitlab_question_is_grounded_with_gitlab_chip(
|
|||||||
def test_keyword_only_question_beats_vector_ranking(
|
def test_keyword_only_question_beats_vector_ranking(
|
||||||
page: Page, app_url: str, mock_llm: int, db_ready: None
|
page: Page, app_url: str, mock_llm: int, db_ready: None
|
||||||
) -> None:
|
) -> None:
|
||||||
"""The FTS-OR gate end to end: "kafkabridge" appears in exactly one
|
"""The corroborated-lexical gate end to end (A8 revised 2026-09-14):
|
||||||
fixture doc (static-dns.json) and the question's cosine overlap is
|
"kafkabridge" appears in exactly one fixture doc (static-dns.json)
|
||||||
weak — the lexical branch is what grounds the answer."""
|
and the question's cosine overlap is weak (below the threshold) —
|
||||||
|
the lexical hit plus cosine >= lexical_support_floor is what grounds
|
||||||
|
the answer (a lexical-only hit below the floor would deflect)."""
|
||||||
_reset_db(mock_llm, seed=True)
|
_reset_db(mock_llm, seed=True)
|
||||||
page.set_default_timeout(30_000)
|
page.set_default_timeout(30_000)
|
||||||
login(page, app_url, next="/") # phase 79: chat is require_user-gated
|
login(page, app_url, next="/") # phase 79: chat is require_user-gated
|
||||||
@@ -182,9 +193,11 @@ def test_keyword_only_question_beats_vector_ranking(
|
|||||||
|
|
||||||
with SessionLocal() as db:
|
with SessionLocal() as db:
|
||||||
row = db.scalars(select(QueryLog)).one()
|
row = db.scalars(select(QueryLog)).one()
|
||||||
# Weak vector score…
|
# Weak vector score — below the answer threshold…
|
||||||
assert row.top_score < get_settings().relevance_threshold
|
assert row.top_score < get_settings().relevance_threshold
|
||||||
# …but a lexical hit grounded it (the FTS-OR branch).
|
# …but cleared the lexical support floor, and a lexical hit fired —
|
||||||
|
# the corroborated-lexical path (A8 revised 2026-09-14) grounded it.
|
||||||
|
assert row.top_score >= get_settings().lexical_support_floor
|
||||||
assert (row.fts_hits or 0) >= 1
|
assert (row.fts_hits or 0) >= 1
|
||||||
assert row.deflected is False
|
assert row.deflected is False
|
||||||
assert "homelab/networking/static-dns.json" in row.sources
|
assert "homelab/networking/static-dns.json" in row.sources
|
||||||
|
|||||||
@@ -413,7 +413,11 @@ def test_off_topic_question_deflects_honestly(client, db, seeded_kb: FakeRagLLM)
|
|||||||
assert any(
|
assert any(
|
||||||
"Deploying a New Service" in s for s in done["suggestions"]
|
"Deploying a New Service" in s for s in done["suggestions"]
|
||||||
), "the best weak-hit title must be offered as a chip"
|
), "the best weak-hit title must be offered as a chip"
|
||||||
assert done["sources"], "weak hits are still reported as the closest sources"
|
# Phase 112 (A8 revised, TODO L2): a deflected turn cites nothing —
|
||||||
|
# done.sources is the citation surface (the UI chips every entry as
|
||||||
|
# "the answer used this"), and the weak hits are scored docs, not
|
||||||
|
# citations. (Pre-phase: they rode the wire as sources.)
|
||||||
|
assert done["sources"] == []
|
||||||
|
|
||||||
# The LLM saw the LOW prompt: DEFLECT_MODE + titles, never doc content.
|
# The LLM saw the LOW prompt: DEFLECT_MODE + titles, never doc content.
|
||||||
(system, user) = seeded_kb.seen_messages[0][0], seeded_kb.seen_messages[0][1]
|
(system, user) = seeded_kb.seen_messages[0][0], seeded_kb.seen_messages[0][1]
|
||||||
@@ -433,19 +437,34 @@ def test_off_topic_question_deflects_honestly(client, db, seeded_kb: FakeRagLLM)
|
|||||||
assert 0.0 < row.top_score < get_settings().relevance_threshold
|
assert 0.0 < row.top_score < get_settings().relevance_threshold
|
||||||
assert row.fts_hits == 0
|
assert row.fts_hits == 0
|
||||||
assert row.chunk_hits >= 1
|
assert row.chunk_hits >= 1
|
||||||
|
# The retrieval stays durably recorded for threshold tuning
|
||||||
|
# (observability unchanged — query_log records retrieval, not
|
||||||
|
# citations; the done frame's [] above is the citation surface).
|
||||||
|
assert row.sources
|
||||||
|
|
||||||
|
|
||||||
def test_keyword_question_grounded_by_lexical_hit_despite_weak_cosine(
|
def test_keyword_question_grounded_by_lexical_hit_despite_weak_cosine(
|
||||||
client, db, seeded_kb: FakeRagLLM
|
client, db, seeded_kb: FakeRagLLM,
|
||||||
|
monkeypatch: pytest.MonkeyPatch,
|
||||||
) -> None:
|
) -> None:
|
||||||
"""Phase 09: a name-your-tool question the vector model barely ranks
|
"""Phase 09: a name-your-tool question the vector model barely ranks
|
||||||
("kafkabridge" only appears in static-dns.json) must still be grounded
|
("kafkabridge" only appears in static-dns.json) must still be grounded
|
||||||
via the FTS branch — LOW only fires at weak cosine AND zero hits."""
|
via the FTS branch — HIGH when cosine >= lexical_support_floor AND
|
||||||
|
fts_hits > 0 (A8 revised 2026-09-14).
|
||||||
|
|
||||||
|
The conftest floor (0.15) is above the mock's cosine (~0.134), so we
|
||||||
|
lower the floor here so the corroborated-lexical path fires."""
|
||||||
|
from app.config import get_settings # noqa: E402
|
||||||
|
|
||||||
|
monkeypatch.setenv("BOR_LEXICAL_SUPPORT_FLOOR", "0.10")
|
||||||
|
# get_settings is lru_cached — clear the cache so the new env var takes effect.
|
||||||
|
get_settings.cache_clear()
|
||||||
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb
|
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb
|
||||||
try:
|
try:
|
||||||
_, _, frames = _stream_chat(client, "How does kafkabridge work?")
|
_, _, frames = _stream_chat(client, "How does kafkabridge work?")
|
||||||
finally:
|
finally:
|
||||||
fastapi_app.dependency_overrides.clear()
|
fastapi_app.dependency_overrides.clear()
|
||||||
|
get_settings.cache_clear()
|
||||||
|
|
||||||
done = frames[-1]
|
done = frames[-1]
|
||||||
assert done["type"] == "done"
|
assert done["type"] == "done"
|
||||||
|
|||||||
@@ -36,10 +36,13 @@ ANSWER = "I haven't done anything like that — try one of these instead!"
|
|||||||
KB_OVERVIEW = "- Homelab\n - Kubernetes (k3s)\n- Deployments\n - Borg backups"
|
KB_OVERVIEW = "- Homelab\n - Kubernetes (k3s)\n- Deployments\n - Borg backups"
|
||||||
|
|
||||||
|
|
||||||
def _settings(threshold: float = 0.30) -> Settings:
|
def _settings(threshold: float = 0.30, floor: float | None = None) -> Settings:
|
||||||
|
if floor is None:
|
||||||
|
floor = threshold * 0.5 # half the threshold — keeps existing tests green
|
||||||
return Settings(
|
return Settings(
|
||||||
_env_file=None, # pyright: ignore[reportCallIssue]
|
_env_file=None, # pyright: ignore[reportCallIssue]
|
||||||
relevance_threshold=threshold,
|
relevance_threshold=threshold,
|
||||||
|
lexical_support_floor=floor,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
@@ -118,12 +121,16 @@ def test_gate_is_env_tunable_via_settings() -> None:
|
|||||||
|
|
||||||
|
|
||||||
def test_gate_weak_cosine_with_fts_hit_still_answers() -> None:
|
def test_gate_weak_cosine_with_fts_hit_still_answers() -> None:
|
||||||
"""cosine < threshold but a lexical hit ⇒ HIGH — the FTS-OR branch.
|
"""cosine < threshold but a lexical hit corroborated by cosine >= floor
|
||||||
This is the name-your-tool case: "kafkabridge" grounds despite weak
|
⇒ HIGH — the FTS-OR branch. This is the name-your-tool case:
|
||||||
vector overlap."""
|
"kafkabridge" grounds despite weak vector overlap.
|
||||||
|
|
||||||
|
A8 revised 2026-09-14: FTS alone no longer promotes; cosine must also
|
||||||
|
clear lexical_support_floor (here 0.15 = half of threshold 0.30)."""
|
||||||
doc = _doc("Static DNS", "DNS_DOC_CONTENT")
|
doc = _doc("Static DNS", "DNS_DOC_CONTENT")
|
||||||
plan = chat_api.plan_turn(
|
plan = chat_api.plan_turn(
|
||||||
[_chunk(doc, 0.02, cosine=0.10, fts_hit=True)], _settings(threshold=0.30)
|
[_chunk(doc, 0.02, cosine=0.10, fts_hit=True)],
|
||||||
|
_settings(threshold=0.30, floor=0.05), # floor=0.05 so 0.10 >= floor
|
||||||
)
|
)
|
||||||
assert plan.deflected is False
|
assert plan.deflected is False
|
||||||
assert plan.top_score == pytest.approx(0.10) # gate input is the cosine
|
assert plan.top_score == pytest.approx(0.10) # gate input is the cosine
|
||||||
@@ -155,7 +162,7 @@ def test_gate_fts_hits_counts_all_lexical_candidates() -> None:
|
|||||||
_chunk(a, 0.02, cosine=0.04, fts_hit=True), # same doc, second chunk
|
_chunk(a, 0.02, cosine=0.04, fts_hit=True), # same doc, second chunk
|
||||||
_chunk(b, 0.01, cosine=0.03),
|
_chunk(b, 0.01, cosine=0.03),
|
||||||
]
|
]
|
||||||
plan = chat_api.plan_turn(chunks, _settings(threshold=0.30))
|
plan = chat_api.plan_turn(chunks, _settings(threshold=0.30, floor=0.03))
|
||||||
assert plan.deflected is False
|
assert plan.deflected is False
|
||||||
assert plan.fts_hits == 2 # per chunk, not per doc
|
assert plan.fts_hits == 2 # per chunk, not per doc
|
||||||
|
|
||||||
@@ -176,6 +183,184 @@ def test_gate_lexical_only_chunk_does_not_inflate_cosine() -> None:
|
|||||||
assert plan.docs[0].title == "Beta"
|
assert plan.docs[0].title == "Beta"
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- lexical support floor (A8 revised 2026-09-14) ----------
|
||||||
|
|
||||||
|
|
||||||
|
def test_gate_fts_hit_below_floor_deflects() -> None:
|
||||||
|
"""The Mongolia case: FTS hit with cosine below lexical_support_floor
|
||||||
|
→ LOW (deflected). The lexical-only hit no longer promotes to HIGH.
|
||||||
|
This is the regression pin for phase 112."""
|
||||||
|
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
|
||||||
|
plan = chat_api.plan_turn(
|
||||||
|
[_chunk(doc, 0.05, cosine=0.10, fts_hit=True)],
|
||||||
|
_settings(threshold=0.62),
|
||||||
|
)
|
||||||
|
assert plan.deflected is True
|
||||||
|
assert plan.top_score == pytest.approx(0.10)
|
||||||
|
assert plan.fts_hits == 1
|
||||||
|
assert "DEFLECT_MODE" in plan.system_prompt
|
||||||
|
assert "QUEST_DOC_CONTENT" not in plan.system_prompt
|
||||||
|
assert "Capital Quest" in plan.system_prompt # title only
|
||||||
|
assert plan.suggestions # derived from weak-hit titles
|
||||||
|
|
||||||
|
|
||||||
|
def test_gate_fts_hit_at_floor_answers() -> None:
|
||||||
|
"""FTS hit with cosine exactly at lexical_support_floor → HIGH.
|
||||||
|
The floor is inclusive (>=), not strict (<)."""
|
||||||
|
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
|
||||||
|
plan = chat_api.plan_turn(
|
||||||
|
[_chunk(doc, 0.40, cosine=0.35, fts_hit=True)],
|
||||||
|
_settings(threshold=0.62),
|
||||||
|
)
|
||||||
|
assert plan.deflected is False
|
||||||
|
assert plan.top_score == pytest.approx(0.35)
|
||||||
|
assert plan.fts_hits == 1
|
||||||
|
assert "QUEST_DOC_CONTENT" in plan.system_prompt
|
||||||
|
assert plan.suggestions == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_gate_fts_hit_above_floor_below_threshold_answers() -> None:
|
||||||
|
"""FTS hit with cosine between floor and threshold → HIGH.
|
||||||
|
The corroborated-lexical path fires."""
|
||||||
|
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
|
||||||
|
plan = chat_api.plan_turn(
|
||||||
|
[_chunk(doc, 0.50, cosine=0.50, fts_hit=True)],
|
||||||
|
_settings(threshold=0.62),
|
||||||
|
)
|
||||||
|
assert plan.deflected is False
|
||||||
|
assert plan.top_score == pytest.approx(0.50)
|
||||||
|
assert plan.fts_hits == 1
|
||||||
|
assert "QUEST_DOC_CONTENT" in plan.system_prompt
|
||||||
|
assert plan.suggestions == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_gate_fts_hit_above_code_default_floor_answers() -> None:
|
||||||
|
"""The quadrant table's "0.50 with default settings" row: the CODE
|
||||||
|
defaults (``test_lexical_support_floor_validation_default`` pins them:
|
||||||
|
threshold 0.62 / floor 0.35) — fts>0 + cosine 0.50 >= 0.35 → HIGH.
|
||||||
|
Named literally (not via the helper's half-threshold floor) so the
|
||||||
|
production-default path is pinned on its own."""
|
||||||
|
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
|
||||||
|
plan = chat_api.plan_turn(
|
||||||
|
[_chunk(doc, 0.50, cosine=0.50, fts_hit=True)],
|
||||||
|
_settings(threshold=0.62, floor=0.35),
|
||||||
|
)
|
||||||
|
assert plan.deflected is False
|
||||||
|
assert plan.top_score == pytest.approx(0.50)
|
||||||
|
assert plan.fts_hits == 1
|
||||||
|
assert "QUEST_DOC_CONTENT" in plan.system_prompt
|
||||||
|
assert plan.suggestions == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_gate_high_cosine_overrides_fts_deflection() -> None:
|
||||||
|
"""Strong cosine (>= threshold) → HIGH regardless of FTS status.
|
||||||
|
The cosine-primary path is unchanged."""
|
||||||
|
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
|
||||||
|
plan = chat_api.plan_turn(
|
||||||
|
[_chunk(doc, 0.90, cosine=0.80, fts_hit=True)],
|
||||||
|
_settings(threshold=0.62),
|
||||||
|
)
|
||||||
|
assert plan.deflected is False
|
||||||
|
assert plan.top_score == pytest.approx(0.80)
|
||||||
|
assert plan.fts_hits == 1
|
||||||
|
assert "QUEST_DOC_CONTENT" in plan.system_prompt
|
||||||
|
assert plan.suggestions == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_gate_fts_no_cosine_deflects() -> None:
|
||||||
|
"""FTS hit with cosine = 0.0 → LOW (the extreme Mongolia case)."""
|
||||||
|
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
|
||||||
|
plan = chat_api.plan_turn(
|
||||||
|
[_chunk(doc, 0.90, cosine=0.0, fts_hit=True)],
|
||||||
|
_settings(threshold=0.62),
|
||||||
|
)
|
||||||
|
assert plan.deflected is True
|
||||||
|
assert plan.top_score == 0.0
|
||||||
|
assert plan.fts_hits == 1
|
||||||
|
assert "DEFLECT_MODE" in plan.system_prompt
|
||||||
|
|
||||||
|
|
||||||
|
def test_gate_multiple_fts_below_floor_deflects() -> None:
|
||||||
|
"""Multiple FTS hits, all below lexical_support_floor → LOW.
|
||||||
|
The gate requires the BEST cosine to clear the floor, not just any hit."""
|
||||||
|
a = _doc("Alpha Quest", "ALPHA_CONTENT")
|
||||||
|
b = _doc("Beta Quest", "BETA_CONTENT")
|
||||||
|
chunks = [
|
||||||
|
_chunk(a, 0.30, cosine=0.20, fts_hit=True),
|
||||||
|
_chunk(b, 0.25, cosine=0.15, fts_hit=True),
|
||||||
|
]
|
||||||
|
plan = chat_api.plan_turn(chunks, _settings(threshold=0.62))
|
||||||
|
assert plan.deflected is True
|
||||||
|
assert plan.fts_hits == 2
|
||||||
|
assert "DEFLECT_MODE" in plan.system_prompt
|
||||||
|
|
||||||
|
|
||||||
|
def test_gate_one_fts_above_floor_answers() -> None:
|
||||||
|
"""Multiple chunks, one FTS hit above floor → HIGH.
|
||||||
|
The best cosine (from the corroborated hit) clears the floor."""
|
||||||
|
a = _doc("Alpha Quest", "ALPHA_CONTENT")
|
||||||
|
b = _doc("Beta Quest", "BETA_CONTENT")
|
||||||
|
chunks = [
|
||||||
|
_chunk(a, 0.30, cosine=0.20, fts_hit=True), # below floor
|
||||||
|
_chunk(b, 0.25, cosine=0.40, fts_hit=True), # above floor
|
||||||
|
]
|
||||||
|
plan = chat_api.plan_turn(chunks, _settings(threshold=0.62))
|
||||||
|
assert plan.deflected is False
|
||||||
|
assert plan.fts_hits == 2
|
||||||
|
assert "ALPHA_CONTENT" in plan.system_prompt
|
||||||
|
assert "BETA_CONTENT" in plan.system_prompt
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- config validation (lexical_support_floor) ----------
|
||||||
|
|
||||||
|
|
||||||
|
def test_lexical_support_floor_validation_floor_above_threshold_fails() -> None:
|
||||||
|
"""lexical_support_floor > relevance_threshold is rejected at startup."""
|
||||||
|
with pytest.raises(ValueError, match="lexical_support_floor"):
|
||||||
|
Settings(
|
||||||
|
_env_file=None, # pyright: ignore[reportCallIssue]
|
||||||
|
relevance_threshold=0.62,
|
||||||
|
lexical_support_floor=0.70,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_lexical_support_floor_validation_negative_fails() -> None:
|
||||||
|
"""Negative lexical_support_floor is rejected."""
|
||||||
|
with pytest.raises(ValueError, match="lexical_support_floor"):
|
||||||
|
Settings(
|
||||||
|
_env_file=None, # pyright: ignore[reportCallIssue]
|
||||||
|
lexical_support_floor=-0.1,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_lexical_support_floor_validation_at_threshold_succeeds() -> None:
|
||||||
|
"""lexical_support_floor == relevance_threshold is legal."""
|
||||||
|
s = Settings(
|
||||||
|
_env_file=None, # pyright: ignore[reportCallIssue]
|
||||||
|
relevance_threshold=0.62,
|
||||||
|
lexical_support_floor=0.62,
|
||||||
|
)
|
||||||
|
assert s.lexical_support_floor == 0.62
|
||||||
|
|
||||||
|
|
||||||
|
def test_lexical_support_floor_validation_default() -> None:
|
||||||
|
"""Default lexical_support_floor is 0.35."""
|
||||||
|
import os
|
||||||
|
# Conftest sets BOR_RELEVANCE_THRESHOLD=0.30 and BOR_LEXICAL_SUPPORT_FLOOR=0.15.
|
||||||
|
# We need the CODE defaults, so clear both and let the class defaults apply.
|
||||||
|
saved_relevance = os.environ.pop("BOR_RELEVANCE_THRESHOLD", None)
|
||||||
|
saved_floor = os.environ.pop("BOR_LEXICAL_SUPPORT_FLOOR", None)
|
||||||
|
try:
|
||||||
|
s = Settings(_env_file=None) # pyright: ignore[reportCallIssue]
|
||||||
|
assert s.lexical_support_floor == 0.35
|
||||||
|
assert s.relevance_threshold == 0.62
|
||||||
|
finally:
|
||||||
|
if saved_relevance is not None:
|
||||||
|
os.environ["BOR_RELEVANCE_THRESHOLD"] = saved_relevance
|
||||||
|
if saved_floor is not None:
|
||||||
|
os.environ["BOR_LEXICAL_SUPPORT_FLOOR"] = saved_floor
|
||||||
|
|
||||||
|
|
||||||
def test_gate_zero_chunks_deflects_with_fallback_chips() -> None:
|
def test_gate_zero_chunks_deflects_with_fallback_chips() -> None:
|
||||||
plan = chat_api.plan_turn([], _settings())
|
plan = chat_api.plan_turn([], _settings())
|
||||||
assert plan.deflected is True
|
assert plan.deflected is True
|
||||||
@@ -581,6 +766,11 @@ def test_endpoint_just_below_threshold_deflects(
|
|||||||
assert 2 <= len(done["suggestions"]) <= MAX_SUGGESTIONS # title chip + fallback
|
assert 2 <= len(done["suggestions"]) <= MAX_SUGGESTIONS # title chip + fallback
|
||||||
assert all(s.strip() for s in done["suggestions"])
|
assert all(s.strip() for s in done["suggestions"])
|
||||||
assert any("Deploying a New Service" in s for s in done["suggestions"])
|
assert any("Deploying a New Service" in s for s in done["suggestions"])
|
||||||
|
# Phase 112 (A8 revised, TODO L2): a deflected turn cites nothing —
|
||||||
|
# done.sources is the citation surface (the UI chips every entry as
|
||||||
|
# "the answer used this"), and the weak hits are scored docs, not
|
||||||
|
# citations.
|
||||||
|
assert done["sources"] == []
|
||||||
|
|
||||||
# The LLM saw the LOW prompt: DEFLECT_MODE + titles, never doc content.
|
# The LLM saw the LOW prompt: DEFLECT_MODE + titles, never doc content.
|
||||||
(system, user) = llm.seen[0][0], llm.seen[0][1]
|
(system, user) = llm.seen[0][0], llm.seen[0][1]
|
||||||
@@ -588,11 +778,14 @@ def test_endpoint_just_below_threshold_deflects(
|
|||||||
assert "DEFLECT_MODE" in system["content"]
|
assert "DEFLECT_MODE" in system["content"]
|
||||||
assert "DOC_CONTENT_NEVER_SENT" not in system["content"]
|
assert "DOC_CONTENT_NEVER_SENT" not in system["content"]
|
||||||
|
|
||||||
# Durable record: deflected + the weak score.
|
# Durable record: deflected + the weak score. The retrieval itself
|
||||||
|
# stays recorded (observability unchanged — query_log records
|
||||||
|
# retrieval, not citations; the phase-113 A3 precedent).
|
||||||
(row,) = session.added
|
(row,) = session.added
|
||||||
assert isinstance(row, QueryLog)
|
assert isinstance(row, QueryLog)
|
||||||
assert row.deflected is True
|
assert row.deflected is True
|
||||||
assert row.top_score == pytest.approx(0.2999)
|
assert row.top_score == pytest.approx(0.2999)
|
||||||
|
assert row.sources # the weak-hit doc's path, for threshold tuning
|
||||||
assert session.commits == 1
|
assert session.commits == 1
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,122 @@
|
|||||||
|
"""Prompt-lock pin (phase 112, task 03 — owner decision iii, 2026-09-14).
|
||||||
|
|
||||||
|
The persona + HONESTY GATE text is **locked verbatim** (PLAN §6): it
|
||||||
|
changes through the plan, never in code. This module byte-pins the
|
||||||
|
locked prompt constants against their pre-phase-112 anchor values —
|
||||||
|
sha256 + exact prefix/suffix + total length, so *any* byte change
|
||||||
|
(option (i)'s copy tightening, option (ii)'s plan amendment, or an
|
||||||
|
accidental edit) fails loudly until the anchors are re-cut as part of
|
||||||
|
the same plan revision. ``tests.unit.test_prompts`` pins the
|
||||||
|
assembled-prompt structure and the behavioral contracts on top of
|
||||||
|
these constants; this file pins the constants themselves.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import hashlib
|
||||||
|
|
||||||
|
from app.rag.prompts import (
|
||||||
|
PERSONA,
|
||||||
|
TOOLS_SECTION,
|
||||||
|
_base,
|
||||||
|
build_deflect_prompt,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _sha256(text: str) -> str:
|
||||||
|
return hashlib.sha256(text.encode("utf-8")).hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- PERSONA (the locked base of BOTH the HIGH and LOW prompts) ----------
|
||||||
|
|
||||||
|
#: Pre-phase-112 anchors for ``PERSONA`` (the ``{relevance}`` placeholder
|
||||||
|
#: and the line wrapping included).
|
||||||
|
PERSONA_SHA256 = "e31792a73e64c53853097e0f7b6df8b96c5f2d05c286944c3edba16dde7777fe"
|
||||||
|
PERSONA_LEN = 706
|
||||||
|
PERSONA_PREFIX = (
|
||||||
|
'You are "Brain of Reese" — the digital brain of Reese, a self-hoster and\n'
|
||||||
|
"homelab tinkerer. Personality: chippy, upbeat, warm, and genuinely\n"
|
||||||
|
"optimistic about the user's ability to do things.\n"
|
||||||
|
"\n"
|
||||||
|
"Rules:\n"
|
||||||
|
)
|
||||||
|
PERSONA_SUFFIX = (
|
||||||
|
"4. Never invent facts, hosts, or steps that are not in the context.\n"
|
||||||
|
"5. Keep answers tight: short paragraphs, bullets where helpful.\n"
|
||||||
|
"\n"
|
||||||
|
"<relevance>{relevance}</relevance>"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_persona_byte_locked() -> None:
|
||||||
|
"""Any byte change to the locked persona (opening, any rule line, the
|
||||||
|
``<relevance>`` placeholder) fails on the sha256; the prefix/suffix
|
||||||
|
anchors name the damaged region for the diff."""
|
||||||
|
assert len(PERSONA) == PERSONA_LEN
|
||||||
|
assert _sha256(PERSONA) == PERSONA_SHA256
|
||||||
|
assert PERSONA.startswith(PERSONA_PREFIX)
|
||||||
|
assert PERSONA.endswith(PERSONA_SUFFIX)
|
||||||
|
|
||||||
|
|
||||||
|
def test_high_and_low_bases_byte_locked() -> None:
|
||||||
|
"""``_base`` only substitutes ``{relevance}`` — the per-mode base
|
||||||
|
lengths pin the substitution against a moved or re-spelled
|
||||||
|
placeholder in the locked text."""
|
||||||
|
assert len(_base("HIGH")) == 699 # PERSONA_LEN - 11 + 4
|
||||||
|
assert len(_base("LOW")) == 698 # PERSONA_LEN - 11 + 3
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- TOOLS_SECTION (the HIGH prompt's locked ``<tools>`` copy) ----------
|
||||||
|
|
||||||
|
#: Pre-phase-112 anchors for ``TOOLS_SECTION``.
|
||||||
|
TOOLS_SECTION_SHA256 = "b834cbe368055e65da82ae3e37a91e6c658c713954703fc79b849a6ebdf4aa53"
|
||||||
|
TOOLS_SECTION_LEN = 2273
|
||||||
|
TOOLS_SECTION_PREFIX = (
|
||||||
|
"<tools>\n"
|
||||||
|
"You may extend your context with three tools. `ls` lists the "
|
||||||
|
"knowledge base as a tree, one level at a time: "
|
||||||
|
)
|
||||||
|
TOOLS_SECTION_SUFFIX = (
|
||||||
|
"Never repeat a call that was refused or already succeeded — the refusal "
|
||||||
|
"already told you the correct form. Answer as soon as you have what "
|
||||||
|
"you need.\n</tools>"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_tools_section_byte_locked() -> None:
|
||||||
|
"""The ``<tools>`` teaching is LOCKED verbatim too (the E2E mock keys
|
||||||
|
on the ``<tools>`` marker's presence; the wording is the owner's):
|
||||||
|
sha256 + exact prefix/suffix + total length."""
|
||||||
|
assert len(TOOLS_SECTION) == TOOLS_SECTION_LEN
|
||||||
|
assert _sha256(TOOLS_SECTION) == TOOLS_SECTION_SHA256
|
||||||
|
assert TOOLS_SECTION.startswith(TOOLS_SECTION_PREFIX)
|
||||||
|
assert TOOLS_SECTION.endswith(TOOLS_SECTION_SUFFIX)
|
||||||
|
|
||||||
|
|
||||||
|
# ---------- the LOW prompt's locked DEFLECT_MODE body ----------
|
||||||
|
|
||||||
|
#: The ``DEFLECT_MODE`` body exactly as it was pre-phase-112 (the
|
||||||
|
#: phase-71 plain-text line included) — inline in
|
||||||
|
#: :func:`app.rag.prompts.build_deflect_prompt`, so it is pinned through
|
||||||
|
#: the built prompt rather than a module constant.
|
||||||
|
LOW_BODY_SHA256 = "c9868cfccdd0ff79d7c1de5df5f6dea0182d3726ca912c9ef609ab54b563a0a4"
|
||||||
|
LOW_BODY_LEN = 259
|
||||||
|
LOW_BODY = (
|
||||||
|
"DEFLECT_MODE: retrieval was weak — the titles below are the closest "
|
||||||
|
"your notes come to the question. They are titles only; do not pretend "
|
||||||
|
"they answer it. Use them to propose 2-3 alternative questions.\n"
|
||||||
|
"Reply in plain text only — you have no tools in this mode."
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_deflect_body_byte_locked() -> None:
|
||||||
|
"""The LOW build = locked base + exactly the locked DEFLECT_MODE body
|
||||||
|
+ the weak-hit title list — byte for byte (the mock keys on the
|
||||||
|
``DEFLECT_MODE`` marker's presence; the body wording is locked)."""
|
||||||
|
assert len(LOW_BODY) == LOW_BODY_LEN
|
||||||
|
assert _sha256(LOW_BODY) == LOW_BODY_SHA256
|
||||||
|
assert build_deflect_prompt(["T1", "T2"]) == _base("LOW") + "\n" + LOW_BODY + "\n- T1\n- T2"
|
||||||
|
prompt = build_deflect_prompt(["T1", "T2"])
|
||||||
|
assert prompt.count(LOW_BODY) == 1
|
||||||
|
assert prompt.index("DEFLECT_MODE") < prompt.index(
|
||||||
|
"Reply in plain text only"
|
||||||
|
) # the marker precedes the plain-text line
|
||||||
Reference in New Issue
Block a user