phase: 112_honesty_gate_weak_hits
Build and Push Containers / build-and-push-app (push) Successful in 2m15s
Build and Push Containers / build-and-push-db (push) Successful in 14s

**Phase 112 — final verification pass (all 4 tasks already complete in `complete/`):**

- Verified gate fix: `app/api/chat.py::plan_turn` — HIGH iff `best_cosine >= relevance_threshold` OR (`fts_hits > 0` AND `best_cosine >= lexical_support_floor`); `lexical_support_floor` (default 0.35, `BOR_LEXICAL_SUPPORT_FLOOR`, bounds-validated) in `app/config.py` + `.env.example`; A8 revision note (2026-09-14) in `.agents/PLAN.md`.
- Verified prompt contract: `app/rag/prompts.py` diff is docstring-only (dated owner-decision-iii entry); `tests/unit/test_prompt_lock.py` byte-pins PERSONA/TOOLS_SECTION/DEFLECT body (sha256+length).
- Verified README: L11 + L575 deflection copy refreshed; `grep "haven't done anything" README.md` → no hits; disclosed-answer behavior documented.
- Tests: `uv run pytest --cov=app --cov-report=term-missing` → **2378 passed, 99% coverage (>90%)**; includes Mongolia-quadrant unit pins (fts>0 + cosine<floor → LOW).
- E2E in isolation: `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → **4 passed** (out-of-KB question: `deflected=true`, `sources==[]`, 2–3 suggestions); regression `test_chat_rag.py` + `test_retrieval_quality.py` → **7 passed**.
- Lint/types: `uv run ruff check .` → clean; `uv run pyright` → 0 errors.

**Completion criteria:** weak-FTS→LOW unit-pinned ✅ · no false citations + 2–3 alternatives E2E ✅ · prompts byte-identical (test-pinned) + README matches ✅ · suite/coverage/e2e/lint all green ✅ · commit + phase move → left to harness (no `git commit` run, per rules; changes in working tree).

**Deviations:** none. Next pending phase: `113_source_chip_quality`.
This commit is contained in:
2026-09-15 00:37:38 -04:00
parent 2683128876
commit 1374faf136
36 changed files with 1240 additions and 67 deletions
+1 -1
View File
@@ -46,7 +46,7 @@ accounts (one admin + hand-out tokens is the model).
| A5 | LLM backend | OpenAI-compatible self-hosted endpoint `https://aipi.reeseapps.com/v1` via the `openai` async client: **`turbo`** (chat, streams `reasoning_content` thinking), **`embed`** (embeddings), **`lite`** (one-shot: document summaries, KB overview, folder summaries) |
| A6 | Embedding dim | **768** (verified against the live endpoint); `chunks.embedding` is fixed at table creation — a dim mismatch must **fail loudly**, never silently re-embed |
| A7 | Retrieval→context | **Hybrid:** cosine top-100 ∪ Postgres FTS top-30 (OR tsquery, `ts_rank`), fused with **RRF (k=60)** → parent docs ranked by best fused chunk → **full text of top-N=2 documents, never truncated** on the retrieval path (owner 2026-08-24: "this should never happen"; the agent `read`-tool cap is a separate owner-permitted path) |
| A8 | Honesty gate | Deflect (LOW mode) **only when** best cosine < `BOR_RELEVANCE_THRESHOLD` (0.62) **and** zero FTS hits; LOW prompt carries weak-hit *titles only* + the `DEFLECT_MODE` marker (the E2E mock keys on its presence) + the plain-text no-tools line |
| A8 | Honesty gate | Deflect (LOW mode) **only when** best cosine < `BOR_RELEVANCE_THRESHOLD` (0.62) **and** (zero FTS hits **or** best cosine < `BOR_LEXICAL_SUPPORT_FLOOR` (0.35)); an FTS hit flips HIGH only when `best_cosine >= lexical_support_floor` — the vector signal must corroborate the lexical match (A8 revised 2026-09-14, owner-confirmed, TODO L2a: lexical-only hits without vector support deflect). LOW prompt carries weak-hit *titles only* + the `DEFLECT_MODE` marker (the E2E mock keys on its presence) + the plain-text no-tools line |
| A9 | Content scope | Default `md, markdown, txt, yaml, yml, json, py` + Podman quadlet family + `j2`; `BOR_IMPORT_EXTENSIONS` may name **any** well-formed extension or narrow the set; hidden (dot) paths + exclusion list + per-source `ignore_paths` (raw prefixes, no globs) always apply |
| A10 | State & auth | `POST /api/chat` is **stateless** (client-provided `history` only, budget-trimmed — nothing stored per conversation). Auth = single admin, signed `bor_session` cookie (Starlette `SessionMiddleware` + itsdangerous; no server-side session store); admin-issued SHA-256-hashed access tokens are the only other identity; **only** shared chats stay anonymous. Fail-loud at boot while `BOR_ADMIN_PASSWORD`/`BOR_SESSION_SECRET` are empty |
| A11 | Frontend | Vanilla HTML/CSS/JS in git, **no CDN** — everything served by FastAPI `StaticFiles`; system font stack; the navbar views are views of ONE shell document (`frontend/index.html` + `router.js` deep-links) |
@@ -0,0 +1,12 @@
**Phase 112 — final verification pass (all 4 tasks already complete in `complete/`):**
- Verified gate fix: `app/api/chat.py::plan_turn` — HIGH iff `best_cosine >= relevance_threshold` OR (`fts_hits > 0` AND `best_cosine >= lexical_support_floor`); `lexical_support_floor` (default 0.35, `BOR_LEXICAL_SUPPORT_FLOOR`, bounds-validated) in `app/config.py` + `.env.example`; A8 revision note (2026-09-14) in `.agents/PLAN.md`.
- Verified prompt contract: `app/rag/prompts.py` diff is docstring-only (dated owner-decision-iii entry); `tests/unit/test_prompt_lock.py` byte-pins PERSONA/TOOLS_SECTION/DEFLECT body (sha256+length).
- Verified README: L11 + L575 deflection copy refreshed; `grep "haven't done anything" README.md` → no hits; disclosed-answer behavior documented.
- Tests: `uv run pytest --cov=app --cov-report=term-missing` → **2378 passed, 99% coverage (>90%)**; includes Mongolia-quadrant unit pins (fts>0 + cosine<floor → LOW).
- E2E in isolation: `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → **4 passed** (out-of-KB question: `deflected=true`, `sources==[]`, 2–3 suggestions); regression `test_chat_rag.py` + `test_retrieval_quality.py` → **7 passed**.
- Lint/types: `uv run ruff check .` → clean; `uv run pyright` → 0 errors.
**Completion criteria:** weak-FTS→LOW unit-pinned ✅ · no false citations + 2–3 alternatives E2E ✅ · prompts byte-identical (test-pinned) + README matches ✅ · suite/coverage/e2e/lint all green ✅ · commit + phase move → left to harness (no `git commit` run, per rules; changes in working tree).
**Deviations:** none. Next pending phase: `113_source_chip_quality`.
@@ -0,0 +1,102 @@
........................................................................ [ 3%]
........................................................................ [ 6%]
........................................................................ [ 9%]
........................................................................ [ 12%]
........................................................................ [ 15%]
........................................................................ [ 18%]
........................................................................ [ 21%]
........................................................................ [ 24%]
........................................................................ [ 27%]
........................................................................ [ 30%]
........................................................................ [ 33%]
........................................................................ [ 36%]
........................................................................ [ 39%]
........................................................................ [ 42%]
........................................................................ [ 45%]
........................................................................ [ 48%]
........................................................................ [ 51%]
........................................................................ [ 54%]
........................................................................ [ 57%]
........................................................................ [ 60%]
........................................................................ [ 63%]
........................................................................ [ 66%]
........................................................................ [ 69%]
........................................................................ [ 72%]
........................................................................ [ 75%]
........................................................................ [ 78%]
........................................................................ [ 81%]
........................................................................ [ 84%]
........................................................................ [ 87%]
........................................................................ [ 90%]
........................................................................ [ 93%]
........................................................................ [ 96%]
........................................................................ [ 99%]
.. [100%]
=============================== warnings summary ===============================
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
Name Stmts Miss Cover
--------------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 52 0 100%
app/api/chat.py 208 1 99%
app/api/chats.py 110 0 100%
app/api/config.py 13 0 100%
app/api/doc_drafts.py 94 0 100%
app/api/docs.py 156 1 99%
app/api/git_sources.py 232 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 33 0 100%
app/api/sync.py 139 0 100%
app/api/tokens.py 40 0 100%
app/api/ui_settings.py 55 0 100%
app/config.py 186 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 45 0 100%
app/core/caching.py 124 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/errors.py 5 0 100%
app/core/logging.py 13 0 100%
app/core/rate_limit.py 44 0 100%
app/core/security_headers.py 20 0 100%
app/core/theming.py 38 0 100%
app/core/tokens.py 44 0 100%
app/db.py 22 0 100%
app/main.py 66 0 100%
app/models.py 128 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 317 1 99%
app/rag/archive_upload.py 134 0 100%
app/rag/chunker.py 206 4 98%
app/rag/doc_dates.py 18 0 100%
app/rag/folder_summaries.py 123 0 100%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 215 3 99%
app/rag/llm.py 243 1 99%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 88 0 100%
app/rag/retriever.py 172 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 327 0 100%
--------------------------------------------------
TOTAL 4079 16 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
validation OK
@@ -0,0 +1 @@
All 2373 tests pass with 99% coverage. Now let me run the linter:
@@ -0,0 +1,101 @@
........................................................................ [ 3%]
........................................................................ [ 6%]
........................................................................ [ 9%]
........................................................................ [ 12%]
........................................................................ [ 15%]
........................................................................ [ 18%]
........................................................................ [ 21%]
........................................................................ [ 24%]
........................................................................ [ 27%]
........................................................................ [ 30%]
........................................................................ [ 33%]
........................................................................ [ 36%]
........................................................................ [ 39%]
........................................................................ [ 42%]
........................................................................ [ 45%]
........................................................................ [ 48%]
........................................................................ [ 51%]
........................................................................ [ 54%]
........................................................................ [ 57%]
........................................................................ [ 60%]
........................................................................ [ 63%]
........................................................................ [ 66%]
........................................................................ [ 69%]
........................................................................ [ 72%]
........................................................................ [ 75%]
........................................................................ [ 78%]
........................................................................ [ 81%]
........................................................................ [ 84%]
........................................................................ [ 87%]
........................................................................ [ 91%]
........................................................................ [ 94%]
........................................................................ [ 97%]
..................................................................... [100%]
=============================== warnings summary ===============================
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
Name Stmts Miss Cover
--------------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 52 0 100%
app/api/chat.py 205 1 99%
app/api/chats.py 110 0 100%
app/api/config.py 13 0 100%
app/api/doc_drafts.py 94 0 100%
app/api/docs.py 156 1 99%
app/api/git_sources.py 232 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 33 0 100%
app/api/sync.py 139 0 100%
app/api/tokens.py 40 0 100%
app/api/ui_settings.py 55 0 100%
app/config.py 186 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 45 0 100%
app/core/caching.py 124 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/errors.py 5 0 100%
app/core/logging.py 13 0 100%
app/core/rate_limit.py 44 0 100%
app/core/security_headers.py 20 0 100%
app/core/theming.py 38 0 100%
app/core/tokens.py 44 0 100%
app/db.py 22 0 100%
app/main.py 66 0 100%
app/models.py 128 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 317 1 99%
app/rag/archive_upload.py 134 0 100%
app/rag/chunker.py 206 4 98%
app/rag/doc_dates.py 18 0 100%
app/rag/folder_summaries.py 123 0 100%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 215 3 99%
app/rag/llm.py 243 1 99%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 88 0 100%
app/rag/retriever.py 172 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 327 0 100%
--------------------------------------------------
TOTAL 4076 16 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
validation OK
@@ -0,0 +1 @@
All 2373 tests pass with 99% coverage. Now let me run the linter:
@@ -0,0 +1,101 @@
........................................................................ [ 3%]
........................................................................ [ 6%]
........................................................................ [ 9%]
........................................................................ [ 12%]
........................................................................ [ 15%]
........................................................................ [ 18%]
........................................................................ [ 21%]
........................................................................ [ 24%]
........................................................................ [ 27%]
........................................................................ [ 30%]
........................................................................ [ 33%]
........................................................................ [ 36%]
........................................................................ [ 39%]
........................................................................ [ 42%]
........................................................................ [ 45%]
........................................................................ [ 48%]
........................................................................ [ 51%]
........................................................................ [ 54%]
........................................................................ [ 57%]
........................................................................ [ 60%]
........................................................................ [ 63%]
........................................................................ [ 66%]
........................................................................ [ 69%]
........................................................................ [ 72%]
........................................................................ [ 75%]
........................................................................ [ 78%]
........................................................................ [ 81%]
........................................................................ [ 84%]
........................................................................ [ 87%]
........................................................................ [ 91%]
........................................................................ [ 94%]
........................................................................ [ 97%]
..................................................................... [100%]
=============================== warnings summary ===============================
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
Name Stmts Miss Cover
--------------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 52 0 100%
app/api/chat.py 205 1 99%
app/api/chats.py 110 0 100%
app/api/config.py 13 0 100%
app/api/doc_drafts.py 94 0 100%
app/api/docs.py 156 1 99%
app/api/git_sources.py 232 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 33 0 100%
app/api/sync.py 139 0 100%
app/api/tokens.py 40 0 100%
app/api/ui_settings.py 55 0 100%
app/config.py 186 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 45 0 100%
app/core/caching.py 124 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/errors.py 5 0 100%
app/core/logging.py 13 0 100%
app/core/rate_limit.py 44 0 100%
app/core/security_headers.py 20 0 100%
app/core/theming.py 38 0 100%
app/core/tokens.py 44 0 100%
app/db.py 22 0 100%
app/main.py 66 0 100%
app/models.py 128 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 317 1 99%
app/rag/archive_upload.py 134 0 100%
app/rag/chunker.py 206 4 98%
app/rag/doc_dates.py 18 0 100%
app/rag/folder_summaries.py 123 0 100%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 215 3 99%
app/rag/llm.py 243 1 99%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 88 0 100%
app/rag/retriever.py 172 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 327 0 100%
--------------------------------------------------
TOTAL 4076 16 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
validation OK
@@ -0,0 +1,19 @@
All green. Final verification complete.
**Task 02 (gate tests) — complete**
Implemented:
- Unit quadrants in `tests/unit/test_chat_gate.py`: added the code-defaults quadrant (fts>0 + cosine 0.50 ≥ floor 0.35 → HIGH) and pinned the new done-frame contract on the endpoint (deflected → `done["sources"] == []`, `query_log.sources` still records retrieval)
- E2E `test_honest_deflection.py`: new `test_out_of_kb_question_deflects_without_citations` — "Who composed Rhapsody in Blue?" (LLM-known, absent from fixture KB; fts_hits≥1, cosine 0.124 < floor 0.15 — the fts>0/cosine<floor quadrant pinned end-to-end) → `deflected: true`, `sources: []`, 2–3 suggestions
- Server: deflected turns' `done` frame now sends `sources=[]` (no false citations — the L2 acceptance; UI chips nothing); `query_log.sources` + per-turn log line keep recording the retrieval (observability unchanged, phase-113 A3 precedent)
- Updated the phase-04 integration test that pinned the old contract; reworded `test_retrieval_quality`'s kafkabridge question ("How does kafkabridge handle DNS?") — the old one (cos 0.134) now correctly falls in the new deflection quadrant
Results:
- `uv run pytest` → 2374 passed; `uv run pytest --cov=app --cov-report=term-missing` → 99% on `app/` (>90%)
- `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → 4 passed (in isolation); `test_chat_rag.py` → 3 passed; `test_retrieval_quality.py` → 4 passed
- At-risk audit: scanned every E2E question literal for gate flips (only the intended one + branch-independent/empty-KB suites); ran 13 gate-sensitive E2E suites (38 tests) → all passed
- `uv run ruff check .` → clean; `uv run pyright` → 0 errors, 0 warnings
Decisions: the task's example question scores 0.303 (> E2E threshold 0.30) under the mock, so it would ground — replaced with Rhapsody in Blue, which also exercises the new quadrant; deflected-done-`sources=[]` is a deliberate contract change required by the task's acceptance (matches phase 113's assumed end state).
Next pending task: `112_honesty_gate_weak_hits/03_prompt_contract_documentation.md`
@@ -0,0 +1,101 @@
........................................................................ [ 3%]
........................................................................ [ 6%]
........................................................................ [ 9%]
........................................................................ [ 12%]
........................................................................ [ 15%]
........................................................................ [ 18%]
........................................................................ [ 21%]
........................................................................ [ 24%]
........................................................................ [ 27%]
........................................................................ [ 30%]
........................................................................ [ 33%]
........................................................................ [ 36%]
........................................................................ [ 39%]
........................................................................ [ 42%]
........................................................................ [ 45%]
........................................................................ [ 48%]
........................................................................ [ 51%]
........................................................................ [ 54%]
........................................................................ [ 57%]
........................................................................ [ 60%]
........................................................................ [ 63%]
........................................................................ [ 66%]
........................................................................ [ 69%]
........................................................................ [ 72%]
........................................................................ [ 75%]
........................................................................ [ 78%]
........................................................................ [ 81%]
........................................................................ [ 84%]
........................................................................ [ 87%]
........................................................................ [ 90%]
........................................................................ [ 94%]
........................................................................ [ 97%]
...................................................................... [100%]
=============================== warnings summary ===============================
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
Name Stmts Miss Cover
--------------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 52 0 100%
app/api/chat.py 208 1 99%
app/api/chats.py 110 0 100%
app/api/config.py 13 0 100%
app/api/doc_drafts.py 94 0 100%
app/api/docs.py 156 1 99%
app/api/git_sources.py 232 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 33 0 100%
app/api/sync.py 139 0 100%
app/api/tokens.py 40 0 100%
app/api/ui_settings.py 55 0 100%
app/config.py 186 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 45 0 100%
app/core/caching.py 124 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/errors.py 5 0 100%
app/core/logging.py 13 0 100%
app/core/rate_limit.py 44 0 100%
app/core/security_headers.py 20 0 100%
app/core/theming.py 38 0 100%
app/core/tokens.py 44 0 100%
app/db.py 22 0 100%
app/main.py 66 0 100%
app/models.py 128 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 317 1 99%
app/rag/archive_upload.py 134 0 100%
app/rag/chunker.py 206 4 98%
app/rag/doc_dates.py 18 0 100%
app/rag/folder_summaries.py 123 0 100%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 215 3 99%
app/rag/llm.py 243 1 99%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 88 0 100%
app/rag/retriever.py 172 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 327 0 100%
--------------------------------------------------
TOTAL 4079 16 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
validation OK
@@ -0,0 +1,10 @@
All green. Task 03 complete.
**Report**
- **prompts.py docstring**: dated entry recording the 2026-09-15 interactive-test finding (stochastic "Ulaanbaatar" disclosed answer vs. clean deflection on re-run, Mongolia case) and owner decision (iii, 2026-09-14): disclosed answers acceptable + documented; gate (A8 revised) is the deterministic lever; options (i)/(ii) left open. Prompt strings untouched (diff is docstring-only).
- **Prompt-lock pin**: new `tests/unit/test_prompt_lock.py` (4 tests) — sha256 + exact prefix/suffix + total length pinning `PERSONA` (706 B), `TOOLS_SECTION` (2273 B), the inline LOW `DEFLECT_MODE` body (259 B), and both `_base` lengths; any byte change fails.
- **README**: added the rare disclosed-answer sentence to the Honest-deflection bullet (work item 2). Defect fix (task 01 leftovers): refreshed the now-stale A8 gate copy in the retrieval paragraph, env table (added `BOR_LEXICAL_SUPPORT_FLOOR` row), tuning bullet, and two spots in `app/rag/retriever.py`.
- **Results**: `uv run pytest --cov=app --cov-report=term-missing` → 2378 passed, **99%** on `app/` (>90%); `uv run ruff check .` clean; `uv run pyright` 0 errors; `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → 4 passed in isolation.
- **Decisions**: kept the quoted *"I haven't done anything like that"* openings intact — task 04 rewrites those; my gate-copy fixes there will fold into it.
- **Next pending task**: `04_readme_deflection_copy.md`.
@@ -0,0 +1,102 @@
........................................................................ [ 3%]
........................................................................ [ 6%]
........................................................................ [ 9%]
........................................................................ [ 12%]
........................................................................ [ 15%]
........................................................................ [ 18%]
........................................................................ [ 21%]
........................................................................ [ 24%]
........................................................................ [ 27%]
........................................................................ [ 30%]
........................................................................ [ 33%]
........................................................................ [ 36%]
........................................................................ [ 39%]
........................................................................ [ 42%]
........................................................................ [ 45%]
........................................................................ [ 48%]
........................................................................ [ 51%]
........................................................................ [ 54%]
........................................................................ [ 57%]
........................................................................ [ 60%]
........................................................................ [ 63%]
........................................................................ [ 66%]
........................................................................ [ 69%]
........................................................................ [ 72%]
........................................................................ [ 75%]
........................................................................ [ 78%]
........................................................................ [ 81%]
........................................................................ [ 84%]
........................................................................ [ 87%]
........................................................................ [ 90%]
........................................................................ [ 93%]
........................................................................ [ 96%]
........................................................................ [ 99%]
.. [100%]
=============================== warnings summary ===============================
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
Name Stmts Miss Cover
--------------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 52 0 100%
app/api/chat.py 208 1 99%
app/api/chats.py 110 0 100%
app/api/config.py 13 0 100%
app/api/doc_drafts.py 94 0 100%
app/api/docs.py 156 1 99%
app/api/git_sources.py 232 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 33 0 100%
app/api/sync.py 139 0 100%
app/api/tokens.py 40 0 100%
app/api/ui_settings.py 55 0 100%
app/config.py 186 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 45 0 100%
app/core/caching.py 124 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/errors.py 5 0 100%
app/core/logging.py 13 0 100%
app/core/rate_limit.py 44 0 100%
app/core/security_headers.py 20 0 100%
app/core/theming.py 38 0 100%
app/core/tokens.py 44 0 100%
app/db.py 22 0 100%
app/main.py 66 0 100%
app/models.py 128 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 317 1 99%
app/rag/archive_upload.py 134 0 100%
app/rag/chunker.py 206 4 98%
app/rag/doc_dates.py 18 0 100%
app/rag/folder_summaries.py 123 0 100%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 215 3 99%
app/rag/llm.py 243 1 99%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 88 0 100%
app/rag/retriever.py 172 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 327 0 100%
--------------------------------------------------
TOTAL 4079 16 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
validation OK
@@ -0,0 +1,12 @@
**Task 04 complete — README deflection copy refreshed (docs-only, no code changes).**
- Rewrote README intro (L11): now "it admits it — and offers 2-3 concrete alternative questions about things it **does** have notes on" (no exact-copy promise; the mandated opening was removed in the 2026-08-22 prompt revision).
- Rewrote Troubleshooting "Honest deflection" entry (L579): amber-bubble reference no longer quotes the removed opening; kept the accurate amber-banner + "Maybe try" chips description.
- Third stale occurrence (L588, the "reply opens with…" clause) found via the required README grep — updated to admit + 2-3 concrete alternatives. `grep -rn "haven't done anything" README.md` → no hits (exit 1).
- Tests grep: all `haven't done anything` hits in `tests/` are mock-LLM/stub-LLM *fixture* copies (the mock's own reply, which already matches the new contract; `DEFLECT_PHRASE` is documented as "the mock's deflection answer must match this") — none assert the app/prompt copy, so no test changes were needed.
- `uv run pytest` → 2378 passed.
- `uv run pytest --cov=app` → TOTAL 99% (>90% gate).
- `uv run ruff check .` → all checks passed; `uv run pyright` → 0 errors.
- `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → 4 passed (phase E2E green in isolation; last task of phase 112).
- No git add/commit (left in working tree for the harness phase commit).
- Next pending task: `.agents/phases/todo/113_source_chip_quality/01_usefulness_bar_sources.md`
@@ -0,0 +1,102 @@
........................................................................ [ 3%]
........................................................................ [ 6%]
........................................................................ [ 9%]
........................................................................ [ 12%]
........................................................................ [ 15%]
........................................................................ [ 18%]
........................................................................ [ 21%]
........................................................................ [ 24%]
........................................................................ [ 27%]
........................................................................ [ 30%]
........................................................................ [ 33%]
........................................................................ [ 36%]
........................................................................ [ 39%]
........................................................................ [ 42%]
........................................................................ [ 45%]
........................................................................ [ 48%]
........................................................................ [ 51%]
........................................................................ [ 54%]
........................................................................ [ 57%]
........................................................................ [ 60%]
........................................................................ [ 63%]
........................................................................ [ 66%]
........................................................................ [ 69%]
........................................................................ [ 72%]
........................................................................ [ 75%]
........................................................................ [ 78%]
........................................................................ [ 81%]
........................................................................ [ 84%]
........................................................................ [ 87%]
........................................................................ [ 90%]
........................................................................ [ 93%]
........................................................................ [ 96%]
........................................................................ [ 99%]
.. [100%]
=============================== warnings summary ===============================
.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1
/var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead.
from starlette.testclient import TestClient as TestClient # noqa
-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
================================ tests coverage ================================
_______________ coverage: platform linux, python 3.14.7-final-0 ________________
Name Stmts Miss Cover
--------------------------------------------------
app/__init__.py 1 0 100%
app/api/__init__.py 0 0 100%
app/api/auth.py 52 0 100%
app/api/chat.py 208 1 99%
app/api/chats.py 110 0 100%
app/api/config.py 13 0 100%
app/api/doc_drafts.py 94 0 100%
app/api/docs.py 156 1 99%
app/api/git_sources.py 232 0 100%
app/api/health.py 10 0 100%
app/api/steering.py 42 0 100%
app/api/suggestions.py 33 0 100%
app/api/sync.py 139 0 100%
app/api/tokens.py 40 0 100%
app/api/ui_settings.py 55 0 100%
app/config.py 186 0 100%
app/core/__init__.py 0 0 100%
app/core/auth.py 45 0 100%
app/core/caching.py 124 0 100%
app/core/debugging.py 29 2 93%
app/core/docs_push.py 39 0 100%
app/core/errors.py 5 0 100%
app/core/logging.py 13 0 100%
app/core/rate_limit.py 44 0 100%
app/core/security_headers.py 20 0 100%
app/core/theming.py 38 0 100%
app/core/tokens.py 44 0 100%
app/db.py 22 0 100%
app/main.py 66 0 100%
app/models.py 128 0 100%
app/rag/__init__.py 0 0 100%
app/rag/agent.py 317 1 99%
app/rag/archive_upload.py 134 0 100%
app/rag/chunker.py 206 4 98%
app/rag/doc_dates.py 18 0 100%
app/rag/folder_summaries.py 123 0 100%
app/rag/git_sources.py 14 0 100%
app/rag/importer.py 215 3 99%
app/rag/llm.py 243 1 99%
app/rag/overview.py 71 0 100%
app/rag/prompts.py 88 0 100%
app/rag/retriever.py 172 3 98%
app/rag/scaffolding.py 55 0 100%
app/rag/source_removal.py 41 0 100%
app/rag/sources_meta.py 16 0 100%
app/rag/suggestions.py 27 0 100%
app/rag/summarizer.py 24 0 100%
app/schemas.py 327 0 100%
--------------------------------------------------
TOTAL 4079 16 99%
coverage gate: app/ 99% (>90%) OK
All checks passed!
0 errors, 0 warnings, 0 informations
WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414).
Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest`
validation OK
+2 -1
View File
@@ -34,7 +34,8 @@ BOR_STREAM_THINKING=1 # stream the model's thinking as `thinking` SS
# --- RAG tuning ---
BOR_TOP_N_DOCS=2
BOR_RELEVANCE_THRESHOLD=0.62 # answer when best cosine >= this OR an FTS hit; else honest deflection
BOR_RELEVANCE_THRESHOLD=0.62 # answer when best cosine >= this OR an FTS hit corroborated by cosine >= lexical_support_floor; else honest deflection
# BOR_LEXICAL_SUPPORT_FLOOR=0.35 # cosine floor for FTS hits to flip HIGH (A8 revised 2026-09-14); 0 <= floor <= relevance_threshold
BOR_MAX_OUTPUT_TOKENS=32768 # max answer length in tokens (answers must not be cut off)
BOR_STEERING_MAX_CHARS=8000 # char budget for the <tuning> (steering notes) prompt section
BOR_SUMMARY_MAX_CHARS=12000 # cap on document content sent to the lite summary model (phase 30)
+33 -19
View File
@@ -8,8 +8,8 @@ Reciprocal Rank Fusion), feeds the **whole relevant document** to a
**self-hosted LLM** (`turbo` via `https://aipi.reeseapps.com/v1`), and
streams a grounded answer back.
If it doesn't have notes for your question, it admits it: *"I haven't done
anything like that"* — plus suggestions for what it **does** know.
If it doesn't have notes for your question, it admits it — and offers
2-3 concrete alternative questions about things it **does** have notes on.
> **Updated your notes?** Re-run the import — it's idempotent and only
> re-embeds what changed:
@@ -345,9 +345,12 @@ nearly double, which lets a name-your-tool question ("gitlab") find its own
document even when the question embeds close to generic templates.
The **honesty gate** (A8) then answers (HIGH) when the best cosine is ≥
`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **or** at least one chunk matched
lexically (`fts_hits > 0`) — it deflects (LOW) only when *both* signals are
absent. The top `BOR_TOP_N_DOCS` full documents are what the LLM sees.
`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **or** when at least one chunk
matched lexically (`fts_hits > 0`) **and** the best cosine clears
`BOR_LEXICAL_SUPPORT_FLOOR` (default `0.35`) — the vector signal must
corroborate a lexical hit (A8 revised 2026-09-14). It deflects (LOW)
otherwise, including a weak single-token hit with vector-unsupported
docs. The top `BOR_TOP_N_DOCS` full documents are what the LLM sees.
---
@@ -412,7 +415,8 @@ asset references of the known pages in flight.
| `BOR_LLM_SUMMARY_MODEL` | `lite` | One-shot completions: document summaries at import, KB overview |
| `BOR_EMBEDDING_DIM` | `768` | Vector dimension (fixed at table creation) |
| `BOR_TOP_N_DOCS` | `2` | Full documents fed to the LLM |
| `BOR_RELEVANCE_THRESHOLD` | `0.62` | Answer when best cosine ≥ this **or** an FTS hit; below + no FTS ⇒ honest deflection |
| `BOR_RELEVANCE_THRESHOLD` | `0.62` | Answer when best cosine ≥ this **or** a lexical hit corroborated by cosine ≥ `BOR_LEXICAL_SUPPORT_FLOOR`; otherwise honest deflection |
| `BOR_LEXICAL_SUPPORT_FLOOR` | `0.35` | Best-cosine floor an FTS hit must clear to flip the gate HIGH (A8 revised 2026-09-14) |
| `BOR_HYBRID_VECTOR_CANDIDATES` | `100` | Cosine list width for the RRF fusion |
| `BOR_HYBRID_LEXICAL_CANDIDATES` | `30` | FTS list width for the RRF fusion |
| `BOR_RRF_K` | `60` | RRF damping constant (`1/(k + rank)`) |
@@ -572,22 +576,32 @@ served locally (no CDN), `BOR_ENVIRONMENT=production`.
- **Embedding dimension mismatch** — aipi changed models; run
`uv run python -m scripts.llm_probe`, update `BOR_EMBEDDING_DIM`, then
drop + recreate the chunks table.
- **Honest deflection** (the amber *"I haven't done anything like that"*
bubble) — every question passes the honesty gate: deflection happens only
- **Honest deflection** (the amber deflection bubble) — every question
passes the honesty gate: deflection happens
when the best cosine similarity is below `BOR_RELEVANCE_THRESHOLD`
(default `0.62`) **and** no chunk matched lexically (`fts_hits = 0`).
A weak cosine with a lexical hit (name-your-tool questions) still gets a
grounded answer. When it does deflect, the LLM prompt carries weak-hit
*titles only* (no document content), the reply opens with *"I haven't
done anything like that"*, the bubble renders amber with *"Maybe try"*
chips derived from the closest indexed titles, and the `query_log` row
records `deflected=true`. This is a feature, not a bug.
(default `0.62`) **and** the lexical signal is absent or not
vector-corroborated (`fts_hits = 0`, or best cosine below
`BOR_LEXICAL_SUPPORT_FLOOR` — default `0.35` — A8 revised 2026-09-14).
A weak cosine with a lexical hit (name-your-tool questions) still gets
a grounded answer, provided the best cosine clears the floor. When it
does deflect, the LLM prompt carries weak-hit *titles only* (no document
content), the reply admits it has no notes on that and offers 2-3
concrete alternative questions about things it **does** have notes on,
the bubble renders amber with *"Maybe try"* chips derived from the
closest indexed titles, and the `query_log` row records
`deflected=true`. This is a feature, not a bug. With a small local
model, a rare turn may still answer from general knowledge with an
explicit disclosure when retrieval was borderline — the gate (A8
revised) minimizes this by keeping vector-unsupported docs out of the
grounded prompt, and the disclosure is surfaced, never silent.
- **Answers deflect too often / too rarely** — tune
`BOR_RELEVANCE_THRESHOLD` (lower = answers more, higher = more honest
deflection): `0.0` ⇒ the gate leans entirely on FTS hits; `1.0` ⇒
everything deflects unless a chunk matches lexically. The `embed` model's
cosines cluster in a ~0.6–0.85 band on the live KB, so the default is
`0.62`. Check real scores:
deflection) and `BOR_LEXICAL_SUPPORT_FLOOR` (the best-cosine floor a
lexical hit must clear to flip HIGH — lower = lexical hits ground more
easily): `0.0` ⇒ the gate leans entirely on FTS hits; `1.0` ⇒
everything deflects unless a lexical hit is vector-corroborated. The
`embed` model's cosines cluster in a ~0.6–0.85 band on the live KB, so
the default is `0.62`. Check real scores:
```sql
SELECT question, top_score, fts_hits, deflected
FROM query_log ORDER BY created_at DESC LIMIT 20;
+49 -26
View File
@@ -15,14 +15,22 @@ turn's thinking is counted in the per-turn log line
(``thinking_chars=N``); ``BOR_STREAM_THINKING=0`` suppresses the
``thinking`` frames (the pieces are still counted).
Honesty gate (A8, revised 2026-08-21): LOW — deflection — only when the
best cosine is strictly below ``BOR_RELEVANCE_THRESHOLD`` **and** no
candidate chunk FTS-matches the question (``fts_hits == 0``). A
name-your-tool question with weak vector overlap but a lexical hit still
gets a grounded answer. Deflection mode carries weak-hit *titles only*
(never document content) plus deterministic "Maybe try" chips, and the
``done`` event / ``query_log`` row record ``deflected=true``, the weak
score and the ``fts_hits`` count.
Honesty gate (A8, revised 2026-09-14): LOW — deflection — only when
``best_cosine < BOR_RELEVANCE_THRESHOLD`` **and** (``fts_hits == 0``
or ``best_cosine < BOR_LEXICAL_SUPPORT_FLOOR``). A lexical hit alone
(cosine < lexical_support_floor) no longer promotes to HIGH — the vector
signal must corroborate the lexical match (A8 revised 2026-09-14, the
"Mongolia" fix). HIGH fires when ``best_cosine >= threshold`` OR
(fts>0 AND cosine >= lexical_support_floor). Deflection mode carries
weak-hit *titles only* (never document content) plus deterministic
"Maybe try" chips, and the ``done`` event / ``query_log`` row record
``deflected=true``, the weak score and the ``fts_hits`` count. A
deflected turn cites nothing: the ``done`` frame's ``sources`` is
``[]`` (``done.sources`` is the citation surface — the UI chips every
entry as "the answer used this" — and weak hits are scored docs, not
citations, TODO L2), while ``query_log.sources`` and the per-turn log
line keep recording the retrieval for tuning (phase 112, A8 revised).
Steering (phase 15): the owner's stored tuning notes are loaded per turn
(oldest first) and injected into the system prompt as a ``<tuning>``
@@ -270,13 +278,15 @@ def plan_turn(
"""Apply the honesty gate (A8, revised) and assemble prompt + context.
* **HIGH (grounded)** when ``best_cosine >= threshold`` **or**
``fts_hits > 0``: HIGH prompt with the full top-N documents, no
suggestions. A cosine exactly at the threshold is an answer — the
gate is strict (``< threshold``).
* **LOW (deflected)** only when ``best_cosine < threshold`` **and**
``fts_hits == 0`` (or no hits at all): LOW prompt (``DEFLECT_MODE``)
with weak-hit titles only — never document content — plus
deterministic alternative-question chips derived from those titles.
(``fts_hits > 0`` **and** ``best_cosine >= lexical_support_floor``):
HIGH prompt with the full top-N documents, no suggestions. A cosine
exactly at the threshold is an answer — the gate is strict
(``< threshold``). An FTS hit alone, without vector corroboration
(cosine < lexical_support_floor), stays LOW (A8 revised 2026-09-14).
* **LOW (deflected)** otherwise — including the fts>0 / cosine < floor
case (the "Mongolia" case): LOW prompt (``DEFLECT_MODE``) with
weak-hit titles only — never document content — plus deterministic
alternative-question chips derived from those titles.
``top_score`` (stored in ``query_log``) is the best cosine, so the
gate input is always a pure vector-similarity number; the lexical
@@ -305,7 +315,8 @@ def plan_turn(
docs = select_documents(chunks, n=settings.top_n_docs)
selected_ids = {d.id for d in docs}
summary_hits = sum(1 for c in chunks if c.is_summary and c.document.id in selected_ids)
if best_cosine >= settings.relevance_threshold or fts_hits > 0:
lexical_supported = fts_hits > 0 and best_cosine >= settings.lexical_support_floor
if best_cosine >= settings.relevance_threshold or lexical_supported:
return TurnPlan(
best_cosine,
fts_hits,
@@ -736,10 +747,14 @@ async def chat(
# 4. Durable record + required per-turn log line (PLAN §9).
# Phase 37: the agent's read documents join the
# retrieval's — deduped by (source, path), order preserved
# — and the same combined list feeds done.sources,
# query_log.sources and the log line (empty on deflected
# turns: the agent never runs). A cancelled turn (the
# retrieval's — deduped by (source, path), order preserved.
# The combined list feeds query_log.sources and the log
# line (retrieval docs even on deflected turns —
# observability, the phase-113 A3 precedent). Phase 112:
# done.sources is the CITATION surface — it carries the
# combined list on grounded turns and [] on deflected
# ones (a deflected answer cites nothing; the weak hits
# stay in the durable record). A cancelled turn (the
# generator closed by the consumer) never reaches this
# step — no query_log row.
cited_docs: list[Document] = []
@@ -790,15 +805,23 @@ async def chat(
scaffold_stripped,
)
settled = True # terminal: the done frame settles the turn
# Phase 112 (A8 revised, TODO L2): a deflected turn cites
# nothing — done.sources is [] (the UI chips every entry
# as a citation; the weak hits are scored docs, not
# citations). The retrieval stays durably recorded above
# (query_log.sources + the log line — observability
# unchanged); the phase-113 related-doc tier is the home
# for the weak hits' visibility.
cited_refs: list[SourceRef] = []
if not plan.deflected:
cited_refs = [
SourceRef(source=d.source, path=d.path, title=d.title)
for d in cited_docs
]
yield sse_event(
ChatDoneEvent(
deflected=plan.deflected,
sources=[
SourceRef(
source=d.source, path=d.path, title=d.title
)
for d in cited_docs
],
sources=cited_refs,
suggestions=plan.suggestions,
).model_dump()
)
+25
View File
@@ -117,6 +117,15 @@ class Settings(BaseSettings):
# default never discriminated. LOW only fires when best cosine < this
# AND no candidate chunk matches the question lexically (see A8).
relevance_threshold: float = 0.62
#: Lexical support floor (A8 revised 2026-09-14): an FTS hit promotes
# a turn to HIGH only when the best cosine is >= this value — it
# requires the vector signal to corroborate the lexical match. A
# single weak token hit with vector-unsupported docs (cosine < floor)
# stays LOW (deflected). Default 0.35 ≈ half the relevance threshold;
# tunable via ``BOR_LEXICAL_SUPPORT_FLOOR``. Must be <=
# ``relevance_threshold`` (a floor above the threshold is a typo that
# would make every FTS hit require a HIGH cosine anyway).
lexical_support_floor: float = 0.35
#: Maximum output tokens a chat answer may use (owner instruction
#: 2026-08-22: answers must run to their natural end — the old hard
#: 700-token cap cut long answers off mid-sentence).
@@ -302,6 +311,22 @@ class Settings(BaseSettings):
#: separate from ``sources_dir`` (the source checkouts).
docs_work_dir: str = "~/bor-docs"
@field_validator("lexical_support_floor")
@classmethod
def _lexical_support_floor_bounds(cls, v: float, info: ValidationInfo) -> float:
"""The lexical support floor must be in [0, relevance_threshold].
A value above the relevance threshold would be a typo — it would
make every FTS hit require a HIGH cosine anyway, defeating the
purpose of the floor (A8 revised 2026-09-14)."""
if v < 0:
raise ValueError("lexical_support_floor must be >= 0")
threshold = info.data.get("relevance_threshold")
if isinstance(threshold, float) and v > threshold:
raise ValueError(
f"lexical_support_floor ({v}) must be <= relevance_threshold ({threshold})"
)
return v
@field_validator("import_extensions")
@classmethod
def _import_extensions_known(cls, v: str) -> str:
+21
View File
@@ -58,6 +58,27 @@ recovery in :mod:`app.rag.scaffolding` / :mod:`app.rag.agent` is the
backstop). The ``DEFLECT_MODE`` marker and everything else in the
prompt stay put — the E2E mock LLM keys on the marker's *presence*,
not the wording, so that contract is unchanged.
Disclosed-answer contract (phase 112, task 03; owner decision
2026-09-14, roadmap confirmation — option (iii) of the three the
2026-09-15 interactive deflection test raised): that test found the
HONESTY GATE's compliance is **stochastic** across runs when
misleading context is injected — the "What is the capital of
Mongolia?" question had two *irrelevant* docs promoted into the HIGH
prompt by a weak single-token FTS hit (the pre-phase A8 rule: any
``fts_hits > 0`` flipped HIGH), and run 1 answered parametrically —
"Ulaanbaatar" — *with an explicit disclosure*, while the identical
one-tap re-run produced a clean, textbook deflection. The owner's
decision: the disclosed general-knowledge answer is treated as
**acceptable and documented** — a small local model cannot be relied
on to obey Rules 1/3 100% when handed misleading context, so this
prompt text stays byte-identical (LOCKED verbatim; options (i) tighten
the copy / (ii) amend this locked prompt via the plan to explicitly
permit disclosed general-knowledge answers remain open to a future
owner decision). The deterministic lever is the honesty gate itself
(A8 revised 2026-09-14: an FTS hit flips HIGH only when
``best_cosine >= lexical_support_floor``, so the misleading docs are
no longer injected — see :mod:`app.api.chat`).
"""
from __future__ import annotations
+4 -3
View File
@@ -21,8 +21,9 @@
match a document's ``qwen3``/``8``/``27b`` tokens, while every
unrelated llama.cpp quadlet out-ranks the target on the shared
``llama``/``cpp`` tokens). The name-hit rows carry ``fts_hit=True``
(they ARE the lexical signal — the A8 honesty gate then answers
instead of deflecting) and ``cosine=0.0``; the RRF fusion is
(they ARE the lexical signal — the A8 honesty gate answers on them
only when the best cosine clears ``BOR_LEXICAL_SUPPORT_FLOOR``,
A8 revised 2026-09-14) and ``cosine=0.0``; the RRF fusion is
unchanged (same lists, same ``1/(k+rank)`` terms).
* **Fusion** — Reciprocal Rank Fusion (``score = Σ 1/(k + rank)`` over the
lists a chunk appears in; chunks hit by both lists get both terms). The
@@ -445,7 +446,7 @@ def _name_hit_chunks(db: Session, question: str) -> list[RetrievedChunk]:
score=0.0, # filled in by :func:`fuse`
document=doc,
cosine=0.0, # no vector rank — name-only hit
fts_hit=True, # the lexical signal — the A8 gate answers
fts_hit=True, # lexical signal — A8 answers if cosine corroborates
is_summary=bool(row.is_summary),
)
)
+2
View File
@@ -15,6 +15,8 @@ from sqlalchemy.orm import Session
# conftest.py. Must be set before ``app.main`` (below) caches settings.
# The production default stays 0.62 (app/config.py, A8 revised).
os.environ.setdefault("BOR_RELEVANCE_THRESHOLD", "0.30")
# A8 revised 2026-09-14: lexical_support_floor must be <= relevance_threshold.
os.environ.setdefault("BOR_LEXICAL_SUPPORT_FLOOR", "0.15")
# Phase 16: single-admin auth is fail-loud — create_app() refuses to boot
# without both vars, and app.main (imported below) builds the app at
+75
View File
@@ -34,6 +34,17 @@ from e2e.auth_helpers import ADMIN_PASSWORD, login
REPO = Path(__file__).resolve().parents[2]
FIXTURES = REPO / "tests" / "fixtures" / "docs"
OFF_TOPIC = "How do I bake sourdough bread?"
# Phase 112 (A8 revised 2026-09-14, TODO L2a — the "Mongolia case"): a
# question the LLM knows (Gershwin) but the fixture KB does not cover.
# Unlike the plain no-hit deflection above, it carries WEAK FTS hits
# (the "compos" stem matches the compose fixture docs — fts_hits >= 1)
# while its best mock token-overlap cosine (~0.12) sits BELOW
# lexical_support_floor (0.15, the mock-calibrated conftest value). The
# pre-phase gate (fts>0 → HIGH) grounded it and injected irrelevant docs
# into the prompt; the revised gate (cosine corroboration) must keep it
# LOW. The mock keys on DEFLECT_MODE, so the test pins the gate, not
# model compliance.
OUT_OF_KB = "Who composed Rhapsody in Blue?"
# The mock's deflection answer (tests/e2e/mock_llm.py) must match this.
DEFLECT_PHRASE = r"haven't done anything like that"
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
@@ -201,3 +212,67 @@ def test_deflected_done_event_and_query_log(app_url: str, mock_llm: int, db_read
assert row.deflected is True
assert 0.0 < row.top_score < get_settings().relevance_threshold
assert row.chunk_hits >= 1
def test_out_of_kb_question_deflects_without_citations(
app_url: str, mock_llm: int, db_ready: None
) -> None:
"""Phase 112 acceptance (TODO L2): a known-out-of-KB question whose
weak lexical hits NO LONGER promote (the fts>0 / cosine<floor
quadrant, pinned end-to-end) deflects with ZERO source citations —
done.sources is empty (the UI chips nothing under a deflected
answer) and 2-3 concrete alternative questions are offered.
Raw SSE (like the done-event test above): the done frame is the
contract surface; the mock's DEFLECT_MODE phrasing proves the
server sent the LOW prompt (the gate, not the model, decides).
"""
_reset_db(mock_llm, seed=True)
client = httpx.Client(timeout=60.0)
r = client.post(f"{app_url}/api/login", json={"password": ADMIN_PASSWORD})
assert r.status_code == 204
frames: list[dict[str, Any]] = []
with client.stream(
"POST", f"{app_url}/api/chat", json={"message": OUT_OF_KB}, timeout=60.0
) as r:
assert r.status_code == 200
assert r.headers["content-type"].startswith("text/event-stream")
buf = ""
for part in r.iter_text():
buf += part
while "\n\n" in buf:
frame, buf = buf.split("\n\n", 1)
if frame.strip().startswith("data:"):
frames.append(json.loads(frame.strip().removeprefix("data:").strip()))
assert buf.strip() == "" # stream ends cleanly on a frame boundary
deltas = [f for f in frames if f.get("type") == "delta"]
answer = "".join(d["text"] for d in deltas)
# The DEFLECT_MODE phrasing streamed ⇒ the LOW prompt reached the
# model (the mock answers it only for the deflection system prompt).
assert re.search(DEFLECT_PHRASE, answer, re.IGNORECASE)
done = frames[-1]
assert done["type"] == "done"
assert done["deflected"] is True
# No false citations (TODO L2): the weak hits never ride the wire as
# sources — a deflected answer cites nothing.
assert done["sources"] == []
# 2-3 concrete alternative questions, all non-empty.
assert 2 <= len(done["suggestions"]) <= 3
assert all(s.strip() for s in done["suggestions"])
# Durable record: the quadrant pinned end-to-end — the lexical leg
# FIRED (fts_hits > 0, the pre-phase gate's promotion trigger) while
# the vector signal never cleared lexical_support_floor, so the
# revised gate deflected. The retrieval itself stays recorded
# (query_log = observability, not citations).
with SessionLocal() as db:
row = db.scalars(select(QueryLog)).one()
assert row.question == OUT_OF_KB
assert row.deflected is True
assert (row.fts_hits or 0) >= 1
assert row.top_score < get_settings().lexical_support_floor
assert row.sources
+20 -7
View File
@@ -13,7 +13,8 @@ The four tests map the story's acceptance criteria:
2. "How did I install gitlab?" — grounded (not deflected), gitlab chip,
``query_log`` row with the gitlab doc in ``sources``
3. keyword-only question ("kafkabridge") beats the vector ranking — the
FTS-OR gate grounds it end to end despite weak cosine
corroborated-lexical gate (A8 revised 2026-09-14) grounds it end to
end: weak cosine, but an FTS hit AND cosine >= lexical_support_floor
4. "sourdough" — deflected bubble + ≥2 "Maybe try" chips
"""
from __future__ import annotations
@@ -37,7 +38,15 @@ from e2e.auth_helpers import login
REPO = Path(__file__).resolve().parents[2]
FIXTURES = REPO / "tests" / "fixtures" / "docs"
GITLAB_QUESTION = "How did I install gitlab?"
KEYWORD_QUESTION = "How does kafkabridge work?"
# Phase 112 (A8 revised): the pre-phase question ("How does kafkabridge
# work?" — mock cosine 0.134) now sits BELOW lexical_support_floor
# (0.15, mock-calibrated) with fts>0 — the new gate's deflection
# quadrant, so it can no longer demonstrate the grounded lexical path.
# "handle DNS" adds static-dns.json's own tokens: cosine ≈0.24 — still
# weak (below the 0.30 threshold) yet corroborated (>= floor) with the
# same single-doc FTS hit, and the FTS-matched doc still tops the fused
# ranking (the test's actual assertion).
KEYWORD_QUESTION = "How does kafkabridge handle DNS?"
OFF_TOPIC = "sourdough starter"
MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E"
@@ -162,9 +171,11 @@ def test_gitlab_question_is_grounded_with_gitlab_chip(
def test_keyword_only_question_beats_vector_ranking(
page: Page, app_url: str, mock_llm: int, db_ready: None
) -> None:
"""The FTS-OR gate end to end: "kafkabridge" appears in exactly one
fixture doc (static-dns.json) and the question's cosine overlap is
weak — the lexical branch is what grounds the answer."""
"""The corroborated-lexical gate end to end (A8 revised 2026-09-14):
"kafkabridge" appears in exactly one fixture doc (static-dns.json)
and the question's cosine overlap is weak (below the threshold) —
the lexical hit plus cosine >= lexical_support_floor is what grounds
the answer (a lexical-only hit below the floor would deflect)."""
_reset_db(mock_llm, seed=True)
page.set_default_timeout(30_000)
login(page, app_url, next="/") # phase 79: chat is require_user-gated
@@ -182,9 +193,11 @@ def test_keyword_only_question_beats_vector_ranking(
with SessionLocal() as db:
row = db.scalars(select(QueryLog)).one()
# Weak vector score…
# Weak vector score — below the answer threshold…
assert row.top_score < get_settings().relevance_threshold
# …but a lexical hit grounded it (the FTS-OR branch).
# …but cleared the lexical support floor, and a lexical hit fired —
# the corroborated-lexical path (A8 revised 2026-09-14) grounded it.
assert row.top_score >= get_settings().lexical_support_floor
assert (row.fts_hits or 0) >= 1
assert row.deflected is False
assert "homelab/networking/static-dns.json" in row.sources
+22 -3
View File
@@ -413,7 +413,11 @@ def test_off_topic_question_deflects_honestly(client, db, seeded_kb: FakeRagLLM)
assert any(
"Deploying a New Service" in s for s in done["suggestions"]
), "the best weak-hit title must be offered as a chip"
assert done["sources"], "weak hits are still reported as the closest sources"
# Phase 112 (A8 revised, TODO L2): a deflected turn cites nothing —
# done.sources is the citation surface (the UI chips every entry as
# "the answer used this"), and the weak hits are scored docs, not
# citations. (Pre-phase: they rode the wire as sources.)
assert done["sources"] == []
# The LLM saw the LOW prompt: DEFLECT_MODE + titles, never doc content.
(system, user) = seeded_kb.seen_messages[0][0], seeded_kb.seen_messages[0][1]
@@ -433,19 +437,34 @@ def test_off_topic_question_deflects_honestly(client, db, seeded_kb: FakeRagLLM)
assert 0.0 < row.top_score < get_settings().relevance_threshold
assert row.fts_hits == 0
assert row.chunk_hits >= 1
# The retrieval stays durably recorded for threshold tuning
# (observability unchanged — query_log records retrieval, not
# citations; the done frame's [] above is the citation surface).
assert row.sources
def test_keyword_question_grounded_by_lexical_hit_despite_weak_cosine(
client, db, seeded_kb: FakeRagLLM
client, db, seeded_kb: FakeRagLLM,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""Phase 09: a name-your-tool question the vector model barely ranks
("kafkabridge" only appears in static-dns.json) must still be grounded
via the FTS branch — LOW only fires at weak cosine AND zero hits."""
via the FTS branch — HIGH when cosine >= lexical_support_floor AND
fts_hits > 0 (A8 revised 2026-09-14).
The conftest floor (0.15) is above the mock's cosine (~0.134), so we
lower the floor here so the corroborated-lexical path fires."""
from app.config import get_settings # noqa: E402
monkeypatch.setenv("BOR_LEXICAL_SUPPORT_FLOOR", "0.10")
# get_settings is lru_cached — clear the cache so the new env var takes effect.
get_settings.cache_clear()
fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb
try:
_, _, frames = _stream_chat(client, "How does kafkabridge work?")
finally:
fastapi_app.dependency_overrides.clear()
get_settings.cache_clear()
done = frames[-1]
assert done["type"] == "done"
+200 -7
View File
@@ -36,10 +36,13 @@ ANSWER = "I haven't done anything like that — try one of these instead!"
KB_OVERVIEW = "- Homelab\n - Kubernetes (k3s)\n- Deployments\n - Borg backups"
def _settings(threshold: float = 0.30) -> Settings:
def _settings(threshold: float = 0.30, floor: float | None = None) -> Settings:
if floor is None:
floor = threshold * 0.5 # half the threshold — keeps existing tests green
return Settings(
_env_file=None, # pyright: ignore[reportCallIssue]
relevance_threshold=threshold,
lexical_support_floor=floor,
)
@@ -118,12 +121,16 @@ def test_gate_is_env_tunable_via_settings() -> None:
def test_gate_weak_cosine_with_fts_hit_still_answers() -> None:
"""cosine < threshold but a lexical hit ⇒ HIGH — the FTS-OR branch.
This is the name-your-tool case: "kafkabridge" grounds despite weak
vector overlap."""
"""cosine < threshold but a lexical hit corroborated by cosine >= floor
⇒ HIGH — the FTS-OR branch. This is the name-your-tool case:
"kafkabridge" grounds despite weak vector overlap.
A8 revised 2026-09-14: FTS alone no longer promotes; cosine must also
clear lexical_support_floor (here 0.15 = half of threshold 0.30)."""
doc = _doc("Static DNS", "DNS_DOC_CONTENT")
plan = chat_api.plan_turn(
[_chunk(doc, 0.02, cosine=0.10, fts_hit=True)], _settings(threshold=0.30)
[_chunk(doc, 0.02, cosine=0.10, fts_hit=True)],
_settings(threshold=0.30, floor=0.05), # floor=0.05 so 0.10 >= floor
)
assert plan.deflected is False
assert plan.top_score == pytest.approx(0.10) # gate input is the cosine
@@ -155,7 +162,7 @@ def test_gate_fts_hits_counts_all_lexical_candidates() -> None:
_chunk(a, 0.02, cosine=0.04, fts_hit=True), # same doc, second chunk
_chunk(b, 0.01, cosine=0.03),
]
plan = chat_api.plan_turn(chunks, _settings(threshold=0.30))
plan = chat_api.plan_turn(chunks, _settings(threshold=0.30, floor=0.03))
assert plan.deflected is False
assert plan.fts_hits == 2 # per chunk, not per doc
@@ -176,6 +183,184 @@ def test_gate_lexical_only_chunk_does_not_inflate_cosine() -> None:
assert plan.docs[0].title == "Beta"
# ---------- lexical support floor (A8 revised 2026-09-14) ----------
def test_gate_fts_hit_below_floor_deflects() -> None:
"""The Mongolia case: FTS hit with cosine below lexical_support_floor
→ LOW (deflected). The lexical-only hit no longer promotes to HIGH.
This is the regression pin for phase 112."""
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
plan = chat_api.plan_turn(
[_chunk(doc, 0.05, cosine=0.10, fts_hit=True)],
_settings(threshold=0.62),
)
assert plan.deflected is True
assert plan.top_score == pytest.approx(0.10)
assert plan.fts_hits == 1
assert "DEFLECT_MODE" in plan.system_prompt
assert "QUEST_DOC_CONTENT" not in plan.system_prompt
assert "Capital Quest" in plan.system_prompt # title only
assert plan.suggestions # derived from weak-hit titles
def test_gate_fts_hit_at_floor_answers() -> None:
"""FTS hit with cosine exactly at lexical_support_floor → HIGH.
The floor is inclusive (>=), not strict (<)."""
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
plan = chat_api.plan_turn(
[_chunk(doc, 0.40, cosine=0.35, fts_hit=True)],
_settings(threshold=0.62),
)
assert plan.deflected is False
assert plan.top_score == pytest.approx(0.35)
assert plan.fts_hits == 1
assert "QUEST_DOC_CONTENT" in plan.system_prompt
assert plan.suggestions == []
def test_gate_fts_hit_above_floor_below_threshold_answers() -> None:
"""FTS hit with cosine between floor and threshold → HIGH.
The corroborated-lexical path fires."""
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
plan = chat_api.plan_turn(
[_chunk(doc, 0.50, cosine=0.50, fts_hit=True)],
_settings(threshold=0.62),
)
assert plan.deflected is False
assert plan.top_score == pytest.approx(0.50)
assert plan.fts_hits == 1
assert "QUEST_DOC_CONTENT" in plan.system_prompt
assert plan.suggestions == []
def test_gate_fts_hit_above_code_default_floor_answers() -> None:
"""The quadrant table's "0.50 with default settings" row: the CODE
defaults (``test_lexical_support_floor_validation_default`` pins them:
threshold 0.62 / floor 0.35) — fts>0 + cosine 0.50 >= 0.35 → HIGH.
Named literally (not via the helper's half-threshold floor) so the
production-default path is pinned on its own."""
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
plan = chat_api.plan_turn(
[_chunk(doc, 0.50, cosine=0.50, fts_hit=True)],
_settings(threshold=0.62, floor=0.35),
)
assert plan.deflected is False
assert plan.top_score == pytest.approx(0.50)
assert plan.fts_hits == 1
assert "QUEST_DOC_CONTENT" in plan.system_prompt
assert plan.suggestions == []
def test_gate_high_cosine_overrides_fts_deflection() -> None:
"""Strong cosine (>= threshold) → HIGH regardless of FTS status.
The cosine-primary path is unchanged."""
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
plan = chat_api.plan_turn(
[_chunk(doc, 0.90, cosine=0.80, fts_hit=True)],
_settings(threshold=0.62),
)
assert plan.deflected is False
assert plan.top_score == pytest.approx(0.80)
assert plan.fts_hits == 1
assert "QUEST_DOC_CONTENT" in plan.system_prompt
assert plan.suggestions == []
def test_gate_fts_no_cosine_deflects() -> None:
"""FTS hit with cosine = 0.0 → LOW (the extreme Mongolia case)."""
doc = _doc("Capital Quest", "QUEST_DOC_CONTENT")
plan = chat_api.plan_turn(
[_chunk(doc, 0.90, cosine=0.0, fts_hit=True)],
_settings(threshold=0.62),
)
assert plan.deflected is True
assert plan.top_score == 0.0
assert plan.fts_hits == 1
assert "DEFLECT_MODE" in plan.system_prompt
def test_gate_multiple_fts_below_floor_deflects() -> None:
"""Multiple FTS hits, all below lexical_support_floor → LOW.
The gate requires the BEST cosine to clear the floor, not just any hit."""
a = _doc("Alpha Quest", "ALPHA_CONTENT")
b = _doc("Beta Quest", "BETA_CONTENT")
chunks = [
_chunk(a, 0.30, cosine=0.20, fts_hit=True),
_chunk(b, 0.25, cosine=0.15, fts_hit=True),
]
plan = chat_api.plan_turn(chunks, _settings(threshold=0.62))
assert plan.deflected is True
assert plan.fts_hits == 2
assert "DEFLECT_MODE" in plan.system_prompt
def test_gate_one_fts_above_floor_answers() -> None:
"""Multiple chunks, one FTS hit above floor → HIGH.
The best cosine (from the corroborated hit) clears the floor."""
a = _doc("Alpha Quest", "ALPHA_CONTENT")
b = _doc("Beta Quest", "BETA_CONTENT")
chunks = [
_chunk(a, 0.30, cosine=0.20, fts_hit=True), # below floor
_chunk(b, 0.25, cosine=0.40, fts_hit=True), # above floor
]
plan = chat_api.plan_turn(chunks, _settings(threshold=0.62))
assert plan.deflected is False
assert plan.fts_hits == 2
assert "ALPHA_CONTENT" in plan.system_prompt
assert "BETA_CONTENT" in plan.system_prompt
# ---------- config validation (lexical_support_floor) ----------
def test_lexical_support_floor_validation_floor_above_threshold_fails() -> None:
"""lexical_support_floor > relevance_threshold is rejected at startup."""
with pytest.raises(ValueError, match="lexical_support_floor"):
Settings(
_env_file=None, # pyright: ignore[reportCallIssue]
relevance_threshold=0.62,
lexical_support_floor=0.70,
)
def test_lexical_support_floor_validation_negative_fails() -> None:
"""Negative lexical_support_floor is rejected."""
with pytest.raises(ValueError, match="lexical_support_floor"):
Settings(
_env_file=None, # pyright: ignore[reportCallIssue]
lexical_support_floor=-0.1,
)
def test_lexical_support_floor_validation_at_threshold_succeeds() -> None:
"""lexical_support_floor == relevance_threshold is legal."""
s = Settings(
_env_file=None, # pyright: ignore[reportCallIssue]
relevance_threshold=0.62,
lexical_support_floor=0.62,
)
assert s.lexical_support_floor == 0.62
def test_lexical_support_floor_validation_default() -> None:
"""Default lexical_support_floor is 0.35."""
import os
# Conftest sets BOR_RELEVANCE_THRESHOLD=0.30 and BOR_LEXICAL_SUPPORT_FLOOR=0.15.
# We need the CODE defaults, so clear both and let the class defaults apply.
saved_relevance = os.environ.pop("BOR_RELEVANCE_THRESHOLD", None)
saved_floor = os.environ.pop("BOR_LEXICAL_SUPPORT_FLOOR", None)
try:
s = Settings(_env_file=None) # pyright: ignore[reportCallIssue]
assert s.lexical_support_floor == 0.35
assert s.relevance_threshold == 0.62
finally:
if saved_relevance is not None:
os.environ["BOR_RELEVANCE_THRESHOLD"] = saved_relevance
if saved_floor is not None:
os.environ["BOR_LEXICAL_SUPPORT_FLOOR"] = saved_floor
def test_gate_zero_chunks_deflects_with_fallback_chips() -> None:
plan = chat_api.plan_turn([], _settings())
assert plan.deflected is True
@@ -581,6 +766,11 @@ def test_endpoint_just_below_threshold_deflects(
assert 2 <= len(done["suggestions"]) <= MAX_SUGGESTIONS # title chip + fallback
assert all(s.strip() for s in done["suggestions"])
assert any("Deploying a New Service" in s for s in done["suggestions"])
# Phase 112 (A8 revised, TODO L2): a deflected turn cites nothing —
# done.sources is the citation surface (the UI chips every entry as
# "the answer used this"), and the weak hits are scored docs, not
# citations.
assert done["sources"] == []
# The LLM saw the LOW prompt: DEFLECT_MODE + titles, never doc content.
(system, user) = llm.seen[0][0], llm.seen[0][1]
@@ -588,11 +778,14 @@ def test_endpoint_just_below_threshold_deflects(
assert "DEFLECT_MODE" in system["content"]
assert "DOC_CONTENT_NEVER_SENT" not in system["content"]
# Durable record: deflected + the weak score.
# Durable record: deflected + the weak score. The retrieval itself
# stays recorded (observability unchanged — query_log records
# retrieval, not citations; the phase-113 A3 precedent).
(row,) = session.added
assert isinstance(row, QueryLog)
assert row.deflected is True
assert row.top_score == pytest.approx(0.2999)
assert row.sources # the weak-hit doc's path, for threshold tuning
assert session.commits == 1
+122
View File
@@ -0,0 +1,122 @@
"""Prompt-lock pin (phase 112, task 03 — owner decision iii, 2026-09-14).
The persona + HONESTY GATE text is **locked verbatim** (PLAN §6): it
changes through the plan, never in code. This module byte-pins the
locked prompt constants against their pre-phase-112 anchor values —
sha256 + exact prefix/suffix + total length, so *any* byte change
(option (i)'s copy tightening, option (ii)'s plan amendment, or an
accidental edit) fails loudly until the anchors are re-cut as part of
the same plan revision. ``tests.unit.test_prompts`` pins the
assembled-prompt structure and the behavioral contracts on top of
these constants; this file pins the constants themselves.
"""
from __future__ import annotations
import hashlib
from app.rag.prompts import (
PERSONA,
TOOLS_SECTION,
_base,
build_deflect_prompt,
)
def _sha256(text: str) -> str:
return hashlib.sha256(text.encode("utf-8")).hexdigest()
# ---------- PERSONA (the locked base of BOTH the HIGH and LOW prompts) ----------
#: Pre-phase-112 anchors for ``PERSONA`` (the ``{relevance}`` placeholder
#: and the line wrapping included).
PERSONA_SHA256 = "e31792a73e64c53853097e0f7b6df8b96c5f2d05c286944c3edba16dde7777fe"
PERSONA_LEN = 706
PERSONA_PREFIX = (
'You are "Brain of Reese" — the digital brain of Reese, a self-hoster and\n'
"homelab tinkerer. Personality: chippy, upbeat, warm, and genuinely\n"
"optimistic about the user's ability to do things.\n"
"\n"
"Rules:\n"
)
PERSONA_SUFFIX = (
"4. Never invent facts, hosts, or steps that are not in the context.\n"
"5. Keep answers tight: short paragraphs, bullets where helpful.\n"
"\n"
"<relevance>{relevance}</relevance>"
)
def test_persona_byte_locked() -> None:
"""Any byte change to the locked persona (opening, any rule line, the
``<relevance>`` placeholder) fails on the sha256; the prefix/suffix
anchors name the damaged region for the diff."""
assert len(PERSONA) == PERSONA_LEN
assert _sha256(PERSONA) == PERSONA_SHA256
assert PERSONA.startswith(PERSONA_PREFIX)
assert PERSONA.endswith(PERSONA_SUFFIX)
def test_high_and_low_bases_byte_locked() -> None:
"""``_base`` only substitutes ``{relevance}`` — the per-mode base
lengths pin the substitution against a moved or re-spelled
placeholder in the locked text."""
assert len(_base("HIGH")) == 699 # PERSONA_LEN - 11 + 4
assert len(_base("LOW")) == 698 # PERSONA_LEN - 11 + 3
# ---------- TOOLS_SECTION (the HIGH prompt's locked ``<tools>`` copy) ----------
#: Pre-phase-112 anchors for ``TOOLS_SECTION``.
TOOLS_SECTION_SHA256 = "b834cbe368055e65da82ae3e37a91e6c658c713954703fc79b849a6ebdf4aa53"
TOOLS_SECTION_LEN = 2273
TOOLS_SECTION_PREFIX = (
"<tools>\n"
"You may extend your context with three tools. `ls` lists the "
"knowledge base as a tree, one level at a time: "
)
TOOLS_SECTION_SUFFIX = (
"Never repeat a call that was refused or already succeeded — the refusal "
"already told you the correct form. Answer as soon as you have what "
"you need.\n</tools>"
)
def test_tools_section_byte_locked() -> None:
"""The ``<tools>`` teaching is LOCKED verbatim too (the E2E mock keys
on the ``<tools>`` marker's presence; the wording is the owner's):
sha256 + exact prefix/suffix + total length."""
assert len(TOOLS_SECTION) == TOOLS_SECTION_LEN
assert _sha256(TOOLS_SECTION) == TOOLS_SECTION_SHA256
assert TOOLS_SECTION.startswith(TOOLS_SECTION_PREFIX)
assert TOOLS_SECTION.endswith(TOOLS_SECTION_SUFFIX)
# ---------- the LOW prompt's locked DEFLECT_MODE body ----------
#: The ``DEFLECT_MODE`` body exactly as it was pre-phase-112 (the
#: phase-71 plain-text line included) — inline in
#: :func:`app.rag.prompts.build_deflect_prompt`, so it is pinned through
#: the built prompt rather than a module constant.
LOW_BODY_SHA256 = "c9868cfccdd0ff79d7c1de5df5f6dea0182d3726ca912c9ef609ab54b563a0a4"
LOW_BODY_LEN = 259
LOW_BODY = (
"DEFLECT_MODE: retrieval was weak — the titles below are the closest "
"your notes come to the question. They are titles only; do not pretend "
"they answer it. Use them to propose 2-3 alternative questions.\n"
"Reply in plain text only — you have no tools in this mode."
)
def test_deflect_body_byte_locked() -> None:
"""The LOW build = locked base + exactly the locked DEFLECT_MODE body
+ the weak-hit title list — byte for byte (the mock keys on the
``DEFLECT_MODE`` marker's presence; the body wording is locked)."""
assert len(LOW_BODY) == LOW_BODY_LEN
assert _sha256(LOW_BODY) == LOW_BODY_SHA256
assert build_deflect_prompt(["T1", "T2"]) == _base("LOW") + "\n" + LOW_BODY + "\n- T1\n- T2"
prompt = build_deflect_prompt(["T1", "T2"])
assert prompt.count(LOW_BODY) == 1
assert prompt.index("DEFLECT_MODE") < prompt.index(
"Reply in plain text only"
) # the marker precedes the plain-text line