diff --git a/.agents/PLAN.md b/.agents/PLAN.md index 21757e3..b13619f 100644 --- a/.agents/PLAN.md +++ b/.agents/PLAN.md @@ -46,7 +46,7 @@ accounts (one admin + hand-out tokens is the model). | A5 | LLM backend | OpenAI-compatible self-hosted endpoint `https://aipi.reeseapps.com/v1` via the `openai` async client: **`turbo`** (chat, streams `reasoning_content` thinking), **`embed`** (embeddings), **`lite`** (one-shot: document summaries, KB overview, folder summaries) | | A6 | Embedding dim | **768** (verified against the live endpoint); `chunks.embedding` is fixed at table creation — a dim mismatch must **fail loudly**, never silently re-embed | | A7 | Retrieval→context | **Hybrid:** cosine top-100 ∪ Postgres FTS top-30 (OR tsquery, `ts_rank`), fused with **RRF (k=60)** → parent docs ranked by best fused chunk → **full text of top-N=2 documents, never truncated** on the retrieval path (owner 2026-08-24: "this should never happen"; the agent `read`-tool cap is a separate owner-permitted path) | -| A8 | Honesty gate | Deflect (LOW mode) **only when** best cosine < `BOR_RELEVANCE_THRESHOLD` (0.62) **and** zero FTS hits; LOW prompt carries weak-hit *titles only* + the `DEFLECT_MODE` marker (the E2E mock keys on its presence) + the plain-text no-tools line | +| A8 | Honesty gate | Deflect (LOW mode) **only when** best cosine < `BOR_RELEVANCE_THRESHOLD` (0.62) **and** (zero FTS hits **or** best cosine < `BOR_LEXICAL_SUPPORT_FLOOR` (0.35)); an FTS hit flips HIGH only when `best_cosine >= lexical_support_floor` — the vector signal must corroborate the lexical match (A8 revised 2026-09-14, owner-confirmed, TODO L2a: lexical-only hits without vector support deflect). LOW prompt carries weak-hit *titles only* + the `DEFLECT_MODE` marker (the E2E mock keys on its presence) + the plain-text no-tools line | | A9 | Content scope | Default `md, markdown, txt, yaml, yml, json, py` + Podman quadlet family + `j2`; `BOR_IMPORT_EXTENSIONS` may name **any** well-formed extension or narrow the set; hidden (dot) paths + exclusion list + per-source `ignore_paths` (raw prefixes, no globs) always apply | | A10 | State & auth | `POST /api/chat` is **stateless** (client-provided `history` only, budget-trimmed — nothing stored per conversation). Auth = single admin, signed `bor_session` cookie (Starlette `SessionMiddleware` + itsdangerous; no server-side session store); admin-issued SHA-256-hashed access tokens are the only other identity; **only** shared chats stay anonymous. Fail-loud at boot while `BOR_ADMIN_PASSWORD`/`BOR_SESSION_SECRET` are empty | | A11 | Frontend | Vanilla HTML/CSS/JS in git, **no CDN** — everything served by FastAPI `StaticFiles`; system font stack; the navbar views are views of ONE shell document (`frontend/index.html` + `router.js` deep-links) | diff --git a/.agents/phases/todo/112_honesty_gate_weak_hits/00_phase.md b/.agents/phases/complete/112_honesty_gate_weak_hits/00_phase.md similarity index 100% rename from .agents/phases/todo/112_honesty_gate_weak_hits/00_phase.md rename to .agents/phases/complete/112_honesty_gate_weak_hits/00_phase.md diff --git a/.agents/phases/todo/112_honesty_gate_weak_hits/01_gate_meaningful_fts.md b/.agents/phases/complete/112_honesty_gate_weak_hits/01_gate_meaningful_fts.md similarity index 100% rename from .agents/phases/todo/112_honesty_gate_weak_hits/01_gate_meaningful_fts.md rename to .agents/phases/complete/112_honesty_gate_weak_hits/01_gate_meaningful_fts.md diff --git a/.agents/phases/todo/112_honesty_gate_weak_hits/02_gate_tests.md b/.agents/phases/complete/112_honesty_gate_weak_hits/02_gate_tests.md similarity index 100% rename from .agents/phases/todo/112_honesty_gate_weak_hits/02_gate_tests.md rename to .agents/phases/complete/112_honesty_gate_weak_hits/02_gate_tests.md diff --git a/.agents/phases/todo/112_honesty_gate_weak_hits/03_prompt_contract_documentation.md b/.agents/phases/complete/112_honesty_gate_weak_hits/03_prompt_contract_documentation.md similarity index 100% rename from .agents/phases/todo/112_honesty_gate_weak_hits/03_prompt_contract_documentation.md rename to .agents/phases/complete/112_honesty_gate_weak_hits/03_prompt_contract_documentation.md diff --git a/.agents/phases/todo/112_honesty_gate_weak_hits/04_readme_deflection_copy.md b/.agents/phases/complete/112_honesty_gate_weak_hits/04_readme_deflection_copy.md similarity index 100% rename from .agents/phases/todo/112_honesty_gate_weak_hits/04_readme_deflection_copy.md rename to .agents/phases/complete/112_honesty_gate_weak_hits/04_readme_deflection_copy.md diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__00_phase.a1.err b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__00_phase.a1.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__00_phase.a1.md b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__00_phase.a1.md new file mode 100644 index 0000000..0332c39 --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__00_phase.a1.md @@ -0,0 +1,12 @@ +**Phase 112 — final verification pass (all 4 tasks already complete in `complete/`):** + +- Verified gate fix: `app/api/chat.py::plan_turn` — HIGH iff `best_cosine >= relevance_threshold` OR (`fts_hits > 0` AND `best_cosine >= lexical_support_floor`); `lexical_support_floor` (default 0.35, `BOR_LEXICAL_SUPPORT_FLOOR`, bounds-validated) in `app/config.py` + `.env.example`; A8 revision note (2026-09-14) in `.agents/PLAN.md`. +- Verified prompt contract: `app/rag/prompts.py` diff is docstring-only (dated owner-decision-iii entry); `tests/unit/test_prompt_lock.py` byte-pins PERSONA/TOOLS_SECTION/DEFLECT body (sha256+length). +- Verified README: L11 + L575 deflection copy refreshed; `grep "haven't done anything" README.md` → no hits; disclosed-answer behavior documented. +- Tests: `uv run pytest --cov=app --cov-report=term-missing` → **2378 passed, 99% coverage (>90%)**; includes Mongolia-quadrant unit pins (fts>0 + cosine90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a1.err b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a1.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a1.md b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a1.md new file mode 100644 index 0000000..1bf6c9b --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a1.md @@ -0,0 +1 @@ +All 2373 tests pass with 99% coverage. Now let me run the linter: diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a1.validate b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a1.validate new file mode 100644 index 0000000..323dcd8 --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a1.validate @@ -0,0 +1,101 @@ +........................................................................ [ 3%] +........................................................................ [ 6%] +........................................................................ [ 9%] +........................................................................ [ 12%] +........................................................................ [ 15%] +........................................................................ [ 18%] +........................................................................ [ 21%] +........................................................................ [ 24%] +........................................................................ [ 27%] +........................................................................ [ 30%] +........................................................................ [ 33%] +........................................................................ [ 36%] +........................................................................ [ 39%] +........................................................................ [ 42%] +........................................................................ [ 45%] +........................................................................ [ 48%] +........................................................................ [ 51%] +........................................................................ [ 54%] +........................................................................ [ 57%] +........................................................................ [ 60%] +........................................................................ [ 63%] +........................................................................ [ 66%] +........................................................................ [ 69%] +........................................................................ [ 72%] +........................................................................ [ 75%] +........................................................................ [ 78%] +........................................................................ [ 81%] +........................................................................ [ 84%] +........................................................................ [ 87%] +........................................................................ [ 91%] +........................................................................ [ 94%] +........................................................................ [ 97%] +..................................................................... [100%] +=============================== warnings summary =============================== +.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1 + /var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. + from starlette.testclient import TestClient as TestClient # noqa + +-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html +================================ tests coverage ================================ +_______________ coverage: platform linux, python 3.14.7-final-0 ________________ + +Name Stmts Miss Cover +-------------------------------------------------- +app/__init__.py 1 0 100% +app/api/__init__.py 0 0 100% +app/api/auth.py 52 0 100% +app/api/chat.py 205 1 99% +app/api/chats.py 110 0 100% +app/api/config.py 13 0 100% +app/api/doc_drafts.py 94 0 100% +app/api/docs.py 156 1 99% +app/api/git_sources.py 232 0 100% +app/api/health.py 10 0 100% +app/api/steering.py 42 0 100% +app/api/suggestions.py 33 0 100% +app/api/sync.py 139 0 100% +app/api/tokens.py 40 0 100% +app/api/ui_settings.py 55 0 100% +app/config.py 186 0 100% +app/core/__init__.py 0 0 100% +app/core/auth.py 45 0 100% +app/core/caching.py 124 0 100% +app/core/debugging.py 29 2 93% +app/core/docs_push.py 39 0 100% +app/core/errors.py 5 0 100% +app/core/logging.py 13 0 100% +app/core/rate_limit.py 44 0 100% +app/core/security_headers.py 20 0 100% +app/core/theming.py 38 0 100% +app/core/tokens.py 44 0 100% +app/db.py 22 0 100% +app/main.py 66 0 100% +app/models.py 128 0 100% +app/rag/__init__.py 0 0 100% +app/rag/agent.py 317 1 99% +app/rag/archive_upload.py 134 0 100% +app/rag/chunker.py 206 4 98% +app/rag/doc_dates.py 18 0 100% +app/rag/folder_summaries.py 123 0 100% +app/rag/git_sources.py 14 0 100% +app/rag/importer.py 215 3 99% +app/rag/llm.py 243 1 99% +app/rag/overview.py 71 0 100% +app/rag/prompts.py 88 0 100% +app/rag/retriever.py 172 3 98% +app/rag/scaffolding.py 55 0 100% +app/rag/source_removal.py 41 0 100% +app/rag/sources_meta.py 16 0 100% +app/rag/suggestions.py 27 0 100% +app/rag/summarizer.py 24 0 100% +app/schemas.py 327 0 100% +-------------------------------------------------- +TOTAL 4076 16 99% +coverage gate: app/ 99% (>90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a2.err b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a2.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a2.md b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a2.md new file mode 100644 index 0000000..1bf6c9b --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a2.md @@ -0,0 +1 @@ +All 2373 tests pass with 99% coverage. Now let me run the linter: diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a2.validate b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a2.validate new file mode 100644 index 0000000..323dcd8 --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__01_gate_meaningful_fts.a2.validate @@ -0,0 +1,101 @@ +........................................................................ [ 3%] +........................................................................ [ 6%] +........................................................................ [ 9%] +........................................................................ [ 12%] +........................................................................ [ 15%] +........................................................................ [ 18%] +........................................................................ [ 21%] +........................................................................ [ 24%] +........................................................................ [ 27%] +........................................................................ [ 30%] +........................................................................ [ 33%] +........................................................................ [ 36%] +........................................................................ [ 39%] +........................................................................ [ 42%] +........................................................................ [ 45%] +........................................................................ [ 48%] +........................................................................ [ 51%] +........................................................................ [ 54%] +........................................................................ [ 57%] +........................................................................ [ 60%] +........................................................................ [ 63%] +........................................................................ [ 66%] +........................................................................ [ 69%] +........................................................................ [ 72%] +........................................................................ [ 75%] +........................................................................ [ 78%] +........................................................................ [ 81%] +........................................................................ [ 84%] +........................................................................ [ 87%] +........................................................................ [ 91%] +........................................................................ [ 94%] +........................................................................ [ 97%] +..................................................................... [100%] +=============================== warnings summary =============================== +.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1 + /var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. + from starlette.testclient import TestClient as TestClient # noqa + +-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html +================================ tests coverage ================================ +_______________ coverage: platform linux, python 3.14.7-final-0 ________________ + +Name Stmts Miss Cover +-------------------------------------------------- +app/__init__.py 1 0 100% +app/api/__init__.py 0 0 100% +app/api/auth.py 52 0 100% +app/api/chat.py 205 1 99% +app/api/chats.py 110 0 100% +app/api/config.py 13 0 100% +app/api/doc_drafts.py 94 0 100% +app/api/docs.py 156 1 99% +app/api/git_sources.py 232 0 100% +app/api/health.py 10 0 100% +app/api/steering.py 42 0 100% +app/api/suggestions.py 33 0 100% +app/api/sync.py 139 0 100% +app/api/tokens.py 40 0 100% +app/api/ui_settings.py 55 0 100% +app/config.py 186 0 100% +app/core/__init__.py 0 0 100% +app/core/auth.py 45 0 100% +app/core/caching.py 124 0 100% +app/core/debugging.py 29 2 93% +app/core/docs_push.py 39 0 100% +app/core/errors.py 5 0 100% +app/core/logging.py 13 0 100% +app/core/rate_limit.py 44 0 100% +app/core/security_headers.py 20 0 100% +app/core/theming.py 38 0 100% +app/core/tokens.py 44 0 100% +app/db.py 22 0 100% +app/main.py 66 0 100% +app/models.py 128 0 100% +app/rag/__init__.py 0 0 100% +app/rag/agent.py 317 1 99% +app/rag/archive_upload.py 134 0 100% +app/rag/chunker.py 206 4 98% +app/rag/doc_dates.py 18 0 100% +app/rag/folder_summaries.py 123 0 100% +app/rag/git_sources.py 14 0 100% +app/rag/importer.py 215 3 99% +app/rag/llm.py 243 1 99% +app/rag/overview.py 71 0 100% +app/rag/prompts.py 88 0 100% +app/rag/retriever.py 172 3 98% +app/rag/scaffolding.py 55 0 100% +app/rag/source_removal.py 41 0 100% +app/rag/sources_meta.py 16 0 100% +app/rag/suggestions.py 27 0 100% +app/rag/summarizer.py 24 0 100% +app/schemas.py 327 0 100% +-------------------------------------------------- +TOTAL 4076 16 99% +coverage gate: app/ 99% (>90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__02_gate_tests.a1.err b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__02_gate_tests.a1.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__02_gate_tests.a1.md b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__02_gate_tests.a1.md new file mode 100644 index 0000000..dad5de2 --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__02_gate_tests.a1.md @@ -0,0 +1,19 @@ +All green. Final verification complete. + +**Task 02 (gate tests) — complete** + +Implemented: +- Unit quadrants in `tests/unit/test_chat_gate.py`: added the code-defaults quadrant (fts>0 + cosine 0.50 ≥ floor 0.35 → HIGH) and pinned the new done-frame contract on the endpoint (deflected → `done["sources"] == []`, `query_log.sources` still records retrieval) +- E2E `test_honest_deflection.py`: new `test_out_of_kb_question_deflects_without_citations` — "Who composed Rhapsody in Blue?" (LLM-known, absent from fixture KB; fts_hits≥1, cosine 0.124 < floor 0.15 — the fts>0/cosine90%) +- `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → 4 passed (in isolation); `test_chat_rag.py` → 3 passed; `test_retrieval_quality.py` → 4 passed +- At-risk audit: scanned every E2E question literal for gate flips (only the intended one + branch-independent/empty-KB suites); ran 13 gate-sensitive E2E suites (38 tests) → all passed +- `uv run ruff check .` → clean; `uv run pyright` → 0 errors, 0 warnings + +Decisions: the task's example question scores 0.303 (> E2E threshold 0.30) under the mock, so it would ground — replaced with Rhapsody in Blue, which also exercises the new quadrant; deflected-done-`sources=[]` is a deliberate contract change required by the task's acceptance (matches phase 113's assumed end state). + +Next pending task: `112_honesty_gate_weak_hits/03_prompt_contract_documentation.md` diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__02_gate_tests.a1.validate b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__02_gate_tests.a1.validate new file mode 100644 index 0000000..22336d2 --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__02_gate_tests.a1.validate @@ -0,0 +1,101 @@ +........................................................................ [ 3%] +........................................................................ [ 6%] +........................................................................ [ 9%] +........................................................................ [ 12%] +........................................................................ [ 15%] +........................................................................ [ 18%] +........................................................................ [ 21%] +........................................................................ [ 24%] +........................................................................ [ 27%] +........................................................................ [ 30%] +........................................................................ [ 33%] +........................................................................ [ 36%] +........................................................................ [ 39%] +........................................................................ [ 42%] +........................................................................ [ 45%] +........................................................................ [ 48%] +........................................................................ [ 51%] +........................................................................ [ 54%] +........................................................................ [ 57%] +........................................................................ [ 60%] +........................................................................ [ 63%] +........................................................................ [ 66%] +........................................................................ [ 69%] +........................................................................ [ 72%] +........................................................................ [ 75%] +........................................................................ [ 78%] +........................................................................ [ 81%] +........................................................................ [ 84%] +........................................................................ [ 87%] +........................................................................ [ 90%] +........................................................................ [ 94%] +........................................................................ [ 97%] +...................................................................... [100%] +=============================== warnings summary =============================== +.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1 + /var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. + from starlette.testclient import TestClient as TestClient # noqa + +-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html +================================ tests coverage ================================ +_______________ coverage: platform linux, python 3.14.7-final-0 ________________ + +Name Stmts Miss Cover +-------------------------------------------------- +app/__init__.py 1 0 100% +app/api/__init__.py 0 0 100% +app/api/auth.py 52 0 100% +app/api/chat.py 208 1 99% +app/api/chats.py 110 0 100% +app/api/config.py 13 0 100% +app/api/doc_drafts.py 94 0 100% +app/api/docs.py 156 1 99% +app/api/git_sources.py 232 0 100% +app/api/health.py 10 0 100% +app/api/steering.py 42 0 100% +app/api/suggestions.py 33 0 100% +app/api/sync.py 139 0 100% +app/api/tokens.py 40 0 100% +app/api/ui_settings.py 55 0 100% +app/config.py 186 0 100% +app/core/__init__.py 0 0 100% +app/core/auth.py 45 0 100% +app/core/caching.py 124 0 100% +app/core/debugging.py 29 2 93% +app/core/docs_push.py 39 0 100% +app/core/errors.py 5 0 100% +app/core/logging.py 13 0 100% +app/core/rate_limit.py 44 0 100% +app/core/security_headers.py 20 0 100% +app/core/theming.py 38 0 100% +app/core/tokens.py 44 0 100% +app/db.py 22 0 100% +app/main.py 66 0 100% +app/models.py 128 0 100% +app/rag/__init__.py 0 0 100% +app/rag/agent.py 317 1 99% +app/rag/archive_upload.py 134 0 100% +app/rag/chunker.py 206 4 98% +app/rag/doc_dates.py 18 0 100% +app/rag/folder_summaries.py 123 0 100% +app/rag/git_sources.py 14 0 100% +app/rag/importer.py 215 3 99% +app/rag/llm.py 243 1 99% +app/rag/overview.py 71 0 100% +app/rag/prompts.py 88 0 100% +app/rag/retriever.py 172 3 98% +app/rag/scaffolding.py 55 0 100% +app/rag/source_removal.py 41 0 100% +app/rag/sources_meta.py 16 0 100% +app/rag/suggestions.py 27 0 100% +app/rag/summarizer.py 24 0 100% +app/schemas.py 327 0 100% +-------------------------------------------------- +TOTAL 4079 16 99% +coverage gate: app/ 99% (>90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__03_prompt_contract_documentation.a1.err b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__03_prompt_contract_documentation.a1.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__03_prompt_contract_documentation.a1.md b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__03_prompt_contract_documentation.a1.md new file mode 100644 index 0000000..0710b26 --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__03_prompt_contract_documentation.a1.md @@ -0,0 +1,10 @@ +All green. Task 03 complete. + +**Report** + +- **prompts.py docstring**: dated entry recording the 2026-09-15 interactive-test finding (stochastic "Ulaanbaatar" disclosed answer vs. clean deflection on re-run, Mongolia case) and owner decision (iii, 2026-09-14): disclosed answers acceptable + documented; gate (A8 revised) is the deterministic lever; options (i)/(ii) left open. Prompt strings untouched (diff is docstring-only). +- **Prompt-lock pin**: new `tests/unit/test_prompt_lock.py` (4 tests) — sha256 + exact prefix/suffix + total length pinning `PERSONA` (706 B), `TOOLS_SECTION` (2273 B), the inline LOW `DEFLECT_MODE` body (259 B), and both `_base` lengths; any byte change fails. +- **README**: added the rare disclosed-answer sentence to the Honest-deflection bullet (work item 2). Defect fix (task 01 leftovers): refreshed the now-stale A8 gate copy in the retrieval paragraph, env table (added `BOR_LEXICAL_SUPPORT_FLOOR` row), tuning bullet, and two spots in `app/rag/retriever.py`. +- **Results**: `uv run pytest --cov=app --cov-report=term-missing` → 2378 passed, **99%** on `app/` (>90%); `uv run ruff check .` clean; `uv run pyright` 0 errors; `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → 4 passed in isolation. +- **Decisions**: kept the quoted *"I haven't done anything like that"* openings intact — task 04 rewrites those; my gate-copy fixes there will fold into it. +- **Next pending task**: `04_readme_deflection_copy.md`. diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__03_prompt_contract_documentation.a1.validate b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__03_prompt_contract_documentation.a1.validate new file mode 100644 index 0000000..6f3c2dd --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__03_prompt_contract_documentation.a1.validate @@ -0,0 +1,102 @@ +........................................................................ [ 3%] +........................................................................ [ 6%] +........................................................................ [ 9%] +........................................................................ [ 12%] +........................................................................ [ 15%] +........................................................................ [ 18%] +........................................................................ [ 21%] +........................................................................ [ 24%] +........................................................................ [ 27%] +........................................................................ [ 30%] +........................................................................ [ 33%] +........................................................................ [ 36%] +........................................................................ [ 39%] +........................................................................ [ 42%] +........................................................................ [ 45%] +........................................................................ [ 48%] +........................................................................ [ 51%] +........................................................................ [ 54%] +........................................................................ [ 57%] +........................................................................ [ 60%] +........................................................................ [ 63%] +........................................................................ [ 66%] +........................................................................ [ 69%] +........................................................................ [ 72%] +........................................................................ [ 75%] +........................................................................ [ 78%] +........................................................................ [ 81%] +........................................................................ [ 84%] +........................................................................ [ 87%] +........................................................................ [ 90%] +........................................................................ [ 93%] +........................................................................ [ 96%] +........................................................................ [ 99%] +.. [100%] +=============================== warnings summary =============================== +.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1 + /var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. + from starlette.testclient import TestClient as TestClient # noqa + +-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html +================================ tests coverage ================================ +_______________ coverage: platform linux, python 3.14.7-final-0 ________________ + +Name Stmts Miss Cover +-------------------------------------------------- +app/__init__.py 1 0 100% +app/api/__init__.py 0 0 100% +app/api/auth.py 52 0 100% +app/api/chat.py 208 1 99% +app/api/chats.py 110 0 100% +app/api/config.py 13 0 100% +app/api/doc_drafts.py 94 0 100% +app/api/docs.py 156 1 99% +app/api/git_sources.py 232 0 100% +app/api/health.py 10 0 100% +app/api/steering.py 42 0 100% +app/api/suggestions.py 33 0 100% +app/api/sync.py 139 0 100% +app/api/tokens.py 40 0 100% +app/api/ui_settings.py 55 0 100% +app/config.py 186 0 100% +app/core/__init__.py 0 0 100% +app/core/auth.py 45 0 100% +app/core/caching.py 124 0 100% +app/core/debugging.py 29 2 93% +app/core/docs_push.py 39 0 100% +app/core/errors.py 5 0 100% +app/core/logging.py 13 0 100% +app/core/rate_limit.py 44 0 100% +app/core/security_headers.py 20 0 100% +app/core/theming.py 38 0 100% +app/core/tokens.py 44 0 100% +app/db.py 22 0 100% +app/main.py 66 0 100% +app/models.py 128 0 100% +app/rag/__init__.py 0 0 100% +app/rag/agent.py 317 1 99% +app/rag/archive_upload.py 134 0 100% +app/rag/chunker.py 206 4 98% +app/rag/doc_dates.py 18 0 100% +app/rag/folder_summaries.py 123 0 100% +app/rag/git_sources.py 14 0 100% +app/rag/importer.py 215 3 99% +app/rag/llm.py 243 1 99% +app/rag/overview.py 71 0 100% +app/rag/prompts.py 88 0 100% +app/rag/retriever.py 172 3 98% +app/rag/scaffolding.py 55 0 100% +app/rag/source_removal.py 41 0 100% +app/rag/sources_meta.py 16 0 100% +app/rag/suggestions.py 27 0 100% +app/rag/summarizer.py 24 0 100% +app/schemas.py 327 0 100% +-------------------------------------------------- +TOTAL 4079 16 99% +coverage gate: app/ 99% (>90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__04_readme_deflection_copy.a1.err b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__04_readme_deflection_copy.a1.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__04_readme_deflection_copy.a1.md b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__04_readme_deflection_copy.a1.md new file mode 100644 index 0000000..864e7d5 --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__04_readme_deflection_copy.a1.md @@ -0,0 +1,12 @@ +**Task 04 complete — README deflection copy refreshed (docs-only, no code changes).** + +- Rewrote README intro (L11): now "it admits it — and offers 2-3 concrete alternative questions about things it **does** have notes on" (no exact-copy promise; the mandated opening was removed in the 2026-08-22 prompt revision). +- Rewrote Troubleshooting "Honest deflection" entry (L579): amber-bubble reference no longer quotes the removed opening; kept the accurate amber-banner + "Maybe try" chips description. +- Third stale occurrence (L588, the "reply opens with…" clause) found via the required README grep — updated to admit + 2-3 concrete alternatives. `grep -rn "haven't done anything" README.md` → no hits (exit 1). +- Tests grep: all `haven't done anything` hits in `tests/` are mock-LLM/stub-LLM *fixture* copies (the mock's own reply, which already matches the new contract; `DEFLECT_PHRASE` is documented as "the mock's deflection answer must match this") — none assert the app/prompt copy, so no test changes were needed. +- `uv run pytest` → 2378 passed. +- `uv run pytest --cov=app` → TOTAL 99% (>90% gate). +- `uv run ruff check .` → all checks passed; `uv run pyright` → 0 errors. +- `uv run pytest tests/e2e/test_honest_deflection.py -v --no-cov` → 4 passed (phase E2E green in isolation; last task of phase 112). +- No git add/commit (left in working tree for the harness phase commit). +- Next pending task: `.agents/phases/todo/113_source_chip_quality/01_usefulness_bar_sources.md` diff --git a/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__04_readme_deflection_copy.a1.validate b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__04_readme_deflection_copy.a1.validate new file mode 100644 index 0000000..6f3c2dd --- /dev/null +++ b/.agents/reports/112_honesty_gate_weak_hits/112_honesty_gate_weak_hits__04_readme_deflection_copy.a1.validate @@ -0,0 +1,102 @@ +........................................................................ [ 3%] +........................................................................ [ 6%] +........................................................................ [ 9%] +........................................................................ [ 12%] +........................................................................ [ 15%] +........................................................................ [ 18%] +........................................................................ [ 21%] +........................................................................ [ 24%] +........................................................................ [ 27%] +........................................................................ [ 30%] +........................................................................ [ 33%] +........................................................................ [ 36%] +........................................................................ [ 39%] +........................................................................ [ 42%] +........................................................................ [ 45%] +........................................................................ [ 48%] +........................................................................ [ 51%] +........................................................................ [ 54%] +........................................................................ [ 57%] +........................................................................ [ 60%] +........................................................................ [ 63%] +........................................................................ [ 66%] +........................................................................ [ 69%] +........................................................................ [ 72%] +........................................................................ [ 75%] +........................................................................ [ 78%] +........................................................................ [ 81%] +........................................................................ [ 84%] +........................................................................ [ 87%] +........................................................................ [ 90%] +........................................................................ [ 93%] +........................................................................ [ 96%] +........................................................................ [ 99%] +.. [100%] +=============================== warnings summary =============================== +.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1 + /var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. + from starlette.testclient import TestClient as TestClient # noqa + +-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html +================================ tests coverage ================================ +_______________ coverage: platform linux, python 3.14.7-final-0 ________________ + +Name Stmts Miss Cover +-------------------------------------------------- +app/__init__.py 1 0 100% +app/api/__init__.py 0 0 100% +app/api/auth.py 52 0 100% +app/api/chat.py 208 1 99% +app/api/chats.py 110 0 100% +app/api/config.py 13 0 100% +app/api/doc_drafts.py 94 0 100% +app/api/docs.py 156 1 99% +app/api/git_sources.py 232 0 100% +app/api/health.py 10 0 100% +app/api/steering.py 42 0 100% +app/api/suggestions.py 33 0 100% +app/api/sync.py 139 0 100% +app/api/tokens.py 40 0 100% +app/api/ui_settings.py 55 0 100% +app/config.py 186 0 100% +app/core/__init__.py 0 0 100% +app/core/auth.py 45 0 100% +app/core/caching.py 124 0 100% +app/core/debugging.py 29 2 93% +app/core/docs_push.py 39 0 100% +app/core/errors.py 5 0 100% +app/core/logging.py 13 0 100% +app/core/rate_limit.py 44 0 100% +app/core/security_headers.py 20 0 100% +app/core/theming.py 38 0 100% +app/core/tokens.py 44 0 100% +app/db.py 22 0 100% +app/main.py 66 0 100% +app/models.py 128 0 100% +app/rag/__init__.py 0 0 100% +app/rag/agent.py 317 1 99% +app/rag/archive_upload.py 134 0 100% +app/rag/chunker.py 206 4 98% +app/rag/doc_dates.py 18 0 100% +app/rag/folder_summaries.py 123 0 100% +app/rag/git_sources.py 14 0 100% +app/rag/importer.py 215 3 99% +app/rag/llm.py 243 1 99% +app/rag/overview.py 71 0 100% +app/rag/prompts.py 88 0 100% +app/rag/retriever.py 172 3 98% +app/rag/scaffolding.py 55 0 100% +app/rag/source_removal.py 41 0 100% +app/rag/sources_meta.py 16 0 100% +app/rag/suggestions.py 27 0 100% +app/rag/summarizer.py 24 0 100% +app/schemas.py 327 0 100% +-------------------------------------------------- +TOTAL 4079 16 99% +coverage gate: app/ 99% (>90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.env.example b/.env.example index 8793f5e..6741ed3 100644 --- a/.env.example +++ b/.env.example @@ -34,7 +34,8 @@ BOR_STREAM_THINKING=1 # stream the model's thinking as `thinking` SS # --- RAG tuning --- BOR_TOP_N_DOCS=2 -BOR_RELEVANCE_THRESHOLD=0.62 # answer when best cosine >= this OR an FTS hit; else honest deflection +BOR_RELEVANCE_THRESHOLD=0.62 # answer when best cosine >= this OR an FTS hit corroborated by cosine >= lexical_support_floor; else honest deflection +# BOR_LEXICAL_SUPPORT_FLOOR=0.35 # cosine floor for FTS hits to flip HIGH (A8 revised 2026-09-14); 0 <= floor <= relevance_threshold BOR_MAX_OUTPUT_TOKENS=32768 # max answer length in tokens (answers must not be cut off) BOR_STEERING_MAX_CHARS=8000 # char budget for the (steering notes) prompt section BOR_SUMMARY_MAX_CHARS=12000 # cap on document content sent to the lite summary model (phase 30) diff --git a/README.md b/README.md index 4ff4104..ab7fc69 100644 --- a/README.md +++ b/README.md @@ -8,8 +8,8 @@ Reciprocal Rank Fusion), feeds the **whole relevant document** to a **self-hosted LLM** (`turbo` via `https://aipi.reeseapps.com/v1`), and streams a grounded answer back. -If it doesn't have notes for your question, it admits it: *"I haven't done -anything like that"* — plus suggestions for what it **does** know. +If it doesn't have notes for your question, it admits it — and offers +2-3 concrete alternative questions about things it **does** have notes on. > **Updated your notes?** Re-run the import — it's idempotent and only > re-embeds what changed: @@ -345,9 +345,12 @@ nearly double, which lets a name-your-tool question ("gitlab") find its own document even when the question embeds close to generic templates. The **honesty gate** (A8) then answers (HIGH) when the best cosine is ≥ -`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **or** at least one chunk matched -lexically (`fts_hits > 0`) — it deflects (LOW) only when *both* signals are -absent. The top `BOR_TOP_N_DOCS` full documents are what the LLM sees. +`BOR_RELEVANCE_THRESHOLD` (default `0.62`) **or** when at least one chunk +matched lexically (`fts_hits > 0`) **and** the best cosine clears +`BOR_LEXICAL_SUPPORT_FLOOR` (default `0.35`) — the vector signal must +corroborate a lexical hit (A8 revised 2026-09-14). It deflects (LOW) +otherwise, including a weak single-token hit with vector-unsupported +docs. The top `BOR_TOP_N_DOCS` full documents are what the LLM sees. --- @@ -412,7 +415,8 @@ asset references of the known pages in flight. | `BOR_LLM_SUMMARY_MODEL` | `lite` | One-shot completions: document summaries at import, KB overview | | `BOR_EMBEDDING_DIM` | `768` | Vector dimension (fixed at table creation) | | `BOR_TOP_N_DOCS` | `2` | Full documents fed to the LLM | -| `BOR_RELEVANCE_THRESHOLD` | `0.62` | Answer when best cosine ≥ this **or** an FTS hit; below + no FTS ⇒ honest deflection | +| `BOR_RELEVANCE_THRESHOLD` | `0.62` | Answer when best cosine ≥ this **or** a lexical hit corroborated by cosine ≥ `BOR_LEXICAL_SUPPORT_FLOOR`; otherwise honest deflection | +| `BOR_LEXICAL_SUPPORT_FLOOR` | `0.35` | Best-cosine floor an FTS hit must clear to flip the gate HIGH (A8 revised 2026-09-14) | | `BOR_HYBRID_VECTOR_CANDIDATES` | `100` | Cosine list width for the RRF fusion | | `BOR_HYBRID_LEXICAL_CANDIDATES` | `30` | FTS list width for the RRF fusion | | `BOR_RRF_K` | `60` | RRF damping constant (`1/(k + rank)`) | @@ -572,22 +576,32 @@ served locally (no CDN), `BOR_ENVIRONMENT=production`. - **Embedding dimension mismatch** — aipi changed models; run `uv run python -m scripts.llm_probe`, update `BOR_EMBEDDING_DIM`, then drop + recreate the chunks table. -- **Honest deflection** (the amber *"I haven't done anything like that"* - bubble) — every question passes the honesty gate: deflection happens only +- **Honest deflection** (the amber deflection bubble) — every question + passes the honesty gate: deflection happens when the best cosine similarity is below `BOR_RELEVANCE_THRESHOLD` - (default `0.62`) **and** no chunk matched lexically (`fts_hits = 0`). - A weak cosine with a lexical hit (name-your-tool questions) still gets a - grounded answer. When it does deflect, the LLM prompt carries weak-hit - *titles only* (no document content), the reply opens with *"I haven't - done anything like that"*, the bubble renders amber with *"Maybe try"* - chips derived from the closest indexed titles, and the `query_log` row - records `deflected=true`. This is a feature, not a bug. + (default `0.62`) **and** the lexical signal is absent or not + vector-corroborated (`fts_hits = 0`, or best cosine below + `BOR_LEXICAL_SUPPORT_FLOOR` — default `0.35` — A8 revised 2026-09-14). + A weak cosine with a lexical hit (name-your-tool questions) still gets + a grounded answer, provided the best cosine clears the floor. When it + does deflect, the LLM prompt carries weak-hit *titles only* (no document + content), the reply admits it has no notes on that and offers 2-3 + concrete alternative questions about things it **does** have notes on, + the bubble renders amber with *"Maybe try"* chips derived from the + closest indexed titles, and the `query_log` row records + `deflected=true`. This is a feature, not a bug. With a small local + model, a rare turn may still answer from general knowledge with an + explicit disclosure when retrieval was borderline — the gate (A8 + revised) minimizes this by keeping vector-unsupported docs out of the + grounded prompt, and the disclosure is surfaced, never silent. - **Answers deflect too often / too rarely** — tune `BOR_RELEVANCE_THRESHOLD` (lower = answers more, higher = more honest - deflection): `0.0` ⇒ the gate leans entirely on FTS hits; `1.0` ⇒ - everything deflects unless a chunk matches lexically. The `embed` model's - cosines cluster in a ~0.6–0.85 band on the live KB, so the default is - `0.62`. Check real scores: + deflection) and `BOR_LEXICAL_SUPPORT_FLOOR` (the best-cosine floor a + lexical hit must clear to flip HIGH — lower = lexical hits ground more + easily): `0.0` ⇒ the gate leans entirely on FTS hits; `1.0` ⇒ + everything deflects unless a lexical hit is vector-corroborated. The + `embed` model's cosines cluster in a ~0.6–0.85 band on the live KB, so + the default is `0.62`. Check real scores: ```sql SELECT question, top_score, fts_hits, deflected FROM query_log ORDER BY created_at DESC LIMIT 20; diff --git a/app/api/chat.py b/app/api/chat.py index eeab02b..2eff670 100644 --- a/app/api/chat.py +++ b/app/api/chat.py @@ -15,14 +15,22 @@ turn's thinking is counted in the per-turn log line (``thinking_chars=N``); ``BOR_STREAM_THINKING=0`` suppresses the ``thinking`` frames (the pieces are still counted). -Honesty gate (A8, revised 2026-08-21): LOW — deflection — only when the -best cosine is strictly below ``BOR_RELEVANCE_THRESHOLD`` **and** no -candidate chunk FTS-matches the question (``fts_hits == 0``). A -name-your-tool question with weak vector overlap but a lexical hit still -gets a grounded answer. Deflection mode carries weak-hit *titles only* -(never document content) plus deterministic "Maybe try" chips, and the -``done`` event / ``query_log`` row record ``deflected=true``, the weak -score and the ``fts_hits`` count. +Honesty gate (A8, revised 2026-09-14): LOW — deflection — only when +``best_cosine < BOR_RELEVANCE_THRESHOLD`` **and** (``fts_hits == 0`` +or ``best_cosine < BOR_LEXICAL_SUPPORT_FLOOR``). A lexical hit alone +(cosine < lexical_support_floor) no longer promotes to HIGH — the vector +signal must corroborate the lexical match (A8 revised 2026-09-14, the +"Mongolia" fix). HIGH fires when ``best_cosine >= threshold`` OR +(fts>0 AND cosine >= lexical_support_floor). Deflection mode carries +weak-hit *titles only* (never document content) plus deterministic +"Maybe try" chips, and the ``done`` event / ``query_log`` row record +``deflected=true``, the weak score and the ``fts_hits`` count. A +deflected turn cites nothing: the ``done`` frame's ``sources`` is +``[]`` (``done.sources`` is the citation surface — the UI chips every +entry as "the answer used this" — and weak hits are scored docs, not +citations, TODO L2), while ``query_log.sources`` and the per-turn log +line keep recording the retrieval for tuning (phase 112, A8 revised). + Steering (phase 15): the owner's stored tuning notes are loaded per turn (oldest first) and injected into the system prompt as a ```` @@ -270,13 +278,15 @@ def plan_turn( """Apply the honesty gate (A8, revised) and assemble prompt + context. * **HIGH (grounded)** when ``best_cosine >= threshold`` **or** - ``fts_hits > 0``: HIGH prompt with the full top-N documents, no - suggestions. A cosine exactly at the threshold is an answer — the - gate is strict (``< threshold``). - * **LOW (deflected)** only when ``best_cosine < threshold`` **and** - ``fts_hits == 0`` (or no hits at all): LOW prompt (``DEFLECT_MODE``) - with weak-hit titles only — never document content — plus - deterministic alternative-question chips derived from those titles. + (``fts_hits > 0`` **and** ``best_cosine >= lexical_support_floor``): + HIGH prompt with the full top-N documents, no suggestions. A cosine + exactly at the threshold is an answer — the gate is strict + (``< threshold``). An FTS hit alone, without vector corroboration + (cosine < lexical_support_floor), stays LOW (A8 revised 2026-09-14). + * **LOW (deflected)** otherwise — including the fts>0 / cosine < floor + case (the "Mongolia" case): LOW prompt (``DEFLECT_MODE``) with + weak-hit titles only — never document content — plus deterministic + alternative-question chips derived from those titles. ``top_score`` (stored in ``query_log``) is the best cosine, so the gate input is always a pure vector-similarity number; the lexical @@ -305,7 +315,8 @@ def plan_turn( docs = select_documents(chunks, n=settings.top_n_docs) selected_ids = {d.id for d in docs} summary_hits = sum(1 for c in chunks if c.is_summary and c.document.id in selected_ids) - if best_cosine >= settings.relevance_threshold or fts_hits > 0: + lexical_supported = fts_hits > 0 and best_cosine >= settings.lexical_support_floor + if best_cosine >= settings.relevance_threshold or lexical_supported: return TurnPlan( best_cosine, fts_hits, @@ -736,10 +747,14 @@ async def chat( # 4. Durable record + required per-turn log line (PLAN §9). # Phase 37: the agent's read documents join the - # retrieval's — deduped by (source, path), order preserved - # — and the same combined list feeds done.sources, - # query_log.sources and the log line (empty on deflected - # turns: the agent never runs). A cancelled turn (the + # retrieval's — deduped by (source, path), order preserved. + # The combined list feeds query_log.sources and the log + # line (retrieval docs even on deflected turns — + # observability, the phase-113 A3 precedent). Phase 112: + # done.sources is the CITATION surface — it carries the + # combined list on grounded turns and [] on deflected + # ones (a deflected answer cites nothing; the weak hits + # stay in the durable record). A cancelled turn (the # generator closed by the consumer) never reaches this # step — no query_log row. cited_docs: list[Document] = [] @@ -790,15 +805,23 @@ async def chat( scaffold_stripped, ) settled = True # terminal: the done frame settles the turn + # Phase 112 (A8 revised, TODO L2): a deflected turn cites + # nothing — done.sources is [] (the UI chips every entry + # as a citation; the weak hits are scored docs, not + # citations). The retrieval stays durably recorded above + # (query_log.sources + the log line — observability + # unchanged); the phase-113 related-doc tier is the home + # for the weak hits' visibility. + cited_refs: list[SourceRef] = [] + if not plan.deflected: + cited_refs = [ + SourceRef(source=d.source, path=d.path, title=d.title) + for d in cited_docs + ] yield sse_event( ChatDoneEvent( deflected=plan.deflected, - sources=[ - SourceRef( - source=d.source, path=d.path, title=d.title - ) - for d in cited_docs - ], + sources=cited_refs, suggestions=plan.suggestions, ).model_dump() ) diff --git a/app/config.py b/app/config.py index cde93eb..c318e39 100644 --- a/app/config.py +++ b/app/config.py @@ -117,6 +117,15 @@ class Settings(BaseSettings): # default never discriminated. LOW only fires when best cosine < this # AND no candidate chunk matches the question lexically (see A8). relevance_threshold: float = 0.62 + #: Lexical support floor (A8 revised 2026-09-14): an FTS hit promotes + # a turn to HIGH only when the best cosine is >= this value — it + # requires the vector signal to corroborate the lexical match. A + # single weak token hit with vector-unsupported docs (cosine < floor) + # stays LOW (deflected). Default 0.35 ≈ half the relevance threshold; + # tunable via ``BOR_LEXICAL_SUPPORT_FLOOR``. Must be <= + # ``relevance_threshold`` (a floor above the threshold is a typo that + # would make every FTS hit require a HIGH cosine anyway). + lexical_support_floor: float = 0.35 #: Maximum output tokens a chat answer may use (owner instruction #: 2026-08-22: answers must run to their natural end — the old hard #: 700-token cap cut long answers off mid-sentence). @@ -302,6 +311,22 @@ class Settings(BaseSettings): #: separate from ``sources_dir`` (the source checkouts). docs_work_dir: str = "~/bor-docs" + @field_validator("lexical_support_floor") + @classmethod + def _lexical_support_floor_bounds(cls, v: float, info: ValidationInfo) -> float: + """The lexical support floor must be in [0, relevance_threshold]. + A value above the relevance threshold would be a typo — it would + make every FTS hit require a HIGH cosine anyway, defeating the + purpose of the floor (A8 revised 2026-09-14).""" + if v < 0: + raise ValueError("lexical_support_floor must be >= 0") + threshold = info.data.get("relevance_threshold") + if isinstance(threshold, float) and v > threshold: + raise ValueError( + f"lexical_support_floor ({v}) must be <= relevance_threshold ({threshold})" + ) + return v + @field_validator("import_extensions") @classmethod def _import_extensions_known(cls, v: str) -> str: diff --git a/app/rag/prompts.py b/app/rag/prompts.py index 9503a6b..a8bbfe4 100644 --- a/app/rag/prompts.py +++ b/app/rag/prompts.py @@ -58,6 +58,27 @@ recovery in :mod:`app.rag.scaffolding` / :mod:`app.rag.agent` is the backstop). The ``DEFLECT_MODE`` marker and everything else in the prompt stay put — the E2E mock LLM keys on the marker's *presence*, not the wording, so that contract is unchanged. + +Disclosed-answer contract (phase 112, task 03; owner decision +2026-09-14, roadmap confirmation — option (iii) of the three the +2026-09-15 interactive deflection test raised): that test found the +HONESTY GATE's compliance is **stochastic** across runs when +misleading context is injected — the "What is the capital of +Mongolia?" question had two *irrelevant* docs promoted into the HIGH +prompt by a weak single-token FTS hit (the pre-phase A8 rule: any +``fts_hits > 0`` flipped HIGH), and run 1 answered parametrically — +"Ulaanbaatar" — *with an explicit disclosure*, while the identical +one-tap re-run produced a clean, textbook deflection. The owner's +decision: the disclosed general-knowledge answer is treated as +**acceptable and documented** — a small local model cannot be relied +on to obey Rules 1/3 100% when handed misleading context, so this +prompt text stays byte-identical (LOCKED verbatim; options (i) tighten +the copy / (ii) amend this locked prompt via the plan to explicitly +permit disclosed general-knowledge answers remain open to a future +owner decision). The deterministic lever is the honesty gate itself +(A8 revised 2026-09-14: an FTS hit flips HIGH only when +``best_cosine >= lexical_support_floor``, so the misleading docs are +no longer injected — see :mod:`app.api.chat`). """ from __future__ import annotations diff --git a/app/rag/retriever.py b/app/rag/retriever.py index 1cb3a85..cbb339d 100644 --- a/app/rag/retriever.py +++ b/app/rag/retriever.py @@ -21,8 +21,9 @@ match a document's ``qwen3``/``8``/``27b`` tokens, while every unrelated llama.cpp quadlet out-ranks the target on the shared ``llama``/``cpp`` tokens). The name-hit rows carry ``fts_hit=True`` - (they ARE the lexical signal — the A8 honesty gate then answers - instead of deflecting) and ``cosine=0.0``; the RRF fusion is + (they ARE the lexical signal — the A8 honesty gate answers on them + only when the best cosine clears ``BOR_LEXICAL_SUPPORT_FLOOR``, + A8 revised 2026-09-14) and ``cosine=0.0``; the RRF fusion is unchanged (same lists, same ``1/(k+rank)`` terms). * **Fusion** — Reciprocal Rank Fusion (``score = Σ 1/(k + rank)`` over the lists a chunk appears in; chunks hit by both lists get both terms). The @@ -445,7 +446,7 @@ def _name_hit_chunks(db: Session, question: str) -> list[RetrievedChunk]: score=0.0, # filled in by :func:`fuse` document=doc, cosine=0.0, # no vector rank — name-only hit - fts_hit=True, # the lexical signal — the A8 gate answers + fts_hit=True, # lexical signal — A8 answers if cosine corroborates is_summary=bool(row.is_summary), ) ) diff --git a/tests/conftest.py b/tests/conftest.py index 6f0eb60..50f7c5c 100644 --- a/tests/conftest.py +++ b/tests/conftest.py @@ -15,6 +15,8 @@ from sqlalchemy.orm import Session # conftest.py. Must be set before ``app.main`` (below) caches settings. # The production default stays 0.62 (app/config.py, A8 revised). os.environ.setdefault("BOR_RELEVANCE_THRESHOLD", "0.30") +# A8 revised 2026-09-14: lexical_support_floor must be <= relevance_threshold. +os.environ.setdefault("BOR_LEXICAL_SUPPORT_FLOOR", "0.15") # Phase 16: single-admin auth is fail-loud — create_app() refuses to boot # without both vars, and app.main (imported below) builds the app at diff --git a/tests/e2e/test_honest_deflection.py b/tests/e2e/test_honest_deflection.py index e7e6140..76fea4c 100644 --- a/tests/e2e/test_honest_deflection.py +++ b/tests/e2e/test_honest_deflection.py @@ -34,6 +34,17 @@ from e2e.auth_helpers import ADMIN_PASSWORD, login REPO = Path(__file__).resolve().parents[2] FIXTURES = REPO / "tests" / "fixtures" / "docs" OFF_TOPIC = "How do I bake sourdough bread?" +# Phase 112 (A8 revised 2026-09-14, TODO L2a — the "Mongolia case"): a +# question the LLM knows (Gershwin) but the fixture KB does not cover. +# Unlike the plain no-hit deflection above, it carries WEAK FTS hits +# (the "compos" stem matches the compose fixture docs — fts_hits >= 1) +# while its best mock token-overlap cosine (~0.12) sits BELOW +# lexical_support_floor (0.15, the mock-calibrated conftest value). The +# pre-phase gate (fts>0 → HIGH) grounded it and injected irrelevant docs +# into the prompt; the revised gate (cosine corroboration) must keep it +# LOW. The mock keys on DEFLECT_MODE, so the test pins the gate, not +# model compliance. +OUT_OF_KB = "Who composed Rhapsody in Blue?" # The mock's deflection answer (tests/e2e/mock_llm.py) must match this. DEFLECT_PHRASE = r"haven't done anything like that" MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E" @@ -201,3 +212,67 @@ def test_deflected_done_event_and_query_log(app_url: str, mock_llm: int, db_read assert row.deflected is True assert 0.0 < row.top_score < get_settings().relevance_threshold assert row.chunk_hits >= 1 + + +def test_out_of_kb_question_deflects_without_citations( + app_url: str, mock_llm: int, db_ready: None +) -> None: + """Phase 112 acceptance (TODO L2): a known-out-of-KB question whose + weak lexical hits NO LONGER promote (the fts>0 / cosine 0, the pre-phase gate's promotion trigger) while + # the vector signal never cleared lexical_support_floor, so the + # revised gate deflected. The retrieval itself stays recorded + # (query_log = observability, not citations). + with SessionLocal() as db: + row = db.scalars(select(QueryLog)).one() + assert row.question == OUT_OF_KB + assert row.deflected is True + assert (row.fts_hits or 0) >= 1 + assert row.top_score < get_settings().lexical_support_floor + assert row.sources diff --git a/tests/e2e/test_retrieval_quality.py b/tests/e2e/test_retrieval_quality.py index e982f88..7ef9c08 100644 --- a/tests/e2e/test_retrieval_quality.py +++ b/tests/e2e/test_retrieval_quality.py @@ -13,7 +13,8 @@ The four tests map the story's acceptance criteria: 2. "How did I install gitlab?" — grounded (not deflected), gitlab chip, ``query_log`` row with the gitlab doc in ``sources`` 3. keyword-only question ("kafkabridge") beats the vector ranking — the - FTS-OR gate grounds it end to end despite weak cosine + corroborated-lexical gate (A8 revised 2026-09-14) grounds it end to + end: weak cosine, but an FTS hit AND cosine >= lexical_support_floor 4. "sourdough" — deflected bubble + ≥2 "Maybe try" chips """ from __future__ import annotations @@ -37,7 +38,15 @@ from e2e.auth_helpers import login REPO = Path(__file__).resolve().parents[2] FIXTURES = REPO / "tests" / "fixtures" / "docs" GITLAB_QUESTION = "How did I install gitlab?" -KEYWORD_QUESTION = "How does kafkabridge work?" +# Phase 112 (A8 revised): the pre-phase question ("How does kafkabridge +# work?" — mock cosine 0.134) now sits BELOW lexical_support_floor +# (0.15, mock-calibrated) with fts>0 — the new gate's deflection +# quadrant, so it can no longer demonstrate the grounded lexical path. +# "handle DNS" adds static-dns.json's own tokens: cosine ≈0.24 — still +# weak (below the 0.30 threshold) yet corroborated (>= floor) with the +# same single-doc FTS hit, and the FTS-matched doc still tops the fused +# ranking (the test's actual assertion). +KEYWORD_QUESTION = "How does kafkabridge handle DNS?" OFF_TOPIC = "sourdough starter" MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E" @@ -162,9 +171,11 @@ def test_gitlab_question_is_grounded_with_gitlab_chip( def test_keyword_only_question_beats_vector_ranking( page: Page, app_url: str, mock_llm: int, db_ready: None ) -> None: - """The FTS-OR gate end to end: "kafkabridge" appears in exactly one - fixture doc (static-dns.json) and the question's cosine overlap is - weak — the lexical branch is what grounds the answer.""" + """The corroborated-lexical gate end to end (A8 revised 2026-09-14): + "kafkabridge" appears in exactly one fixture doc (static-dns.json) + and the question's cosine overlap is weak (below the threshold) — + the lexical hit plus cosine >= lexical_support_floor is what grounds + the answer (a lexical-only hit below the floor would deflect).""" _reset_db(mock_llm, seed=True) page.set_default_timeout(30_000) login(page, app_url, next="/") # phase 79: chat is require_user-gated @@ -182,9 +193,11 @@ def test_keyword_only_question_beats_vector_ranking( with SessionLocal() as db: row = db.scalars(select(QueryLog)).one() - # Weak vector score… + # Weak vector score — below the answer threshold… assert row.top_score < get_settings().relevance_threshold - # …but a lexical hit grounded it (the FTS-OR branch). + # …but cleared the lexical support floor, and a lexical hit fired — + # the corroborated-lexical path (A8 revised 2026-09-14) grounded it. + assert row.top_score >= get_settings().lexical_support_floor assert (row.fts_hits or 0) >= 1 assert row.deflected is False assert "homelab/networking/static-dns.json" in row.sources diff --git a/tests/integration/test_chat_api.py b/tests/integration/test_chat_api.py index 1c1c064..d59067c 100644 --- a/tests/integration/test_chat_api.py +++ b/tests/integration/test_chat_api.py @@ -413,7 +413,11 @@ def test_off_topic_question_deflects_honestly(client, db, seeded_kb: FakeRagLLM) assert any( "Deploying a New Service" in s for s in done["suggestions"] ), "the best weak-hit title must be offered as a chip" - assert done["sources"], "weak hits are still reported as the closest sources" + # Phase 112 (A8 revised, TODO L2): a deflected turn cites nothing — + # done.sources is the citation surface (the UI chips every entry as + # "the answer used this"), and the weak hits are scored docs, not + # citations. (Pre-phase: they rode the wire as sources.) + assert done["sources"] == [] # The LLM saw the LOW prompt: DEFLECT_MODE + titles, never doc content. (system, user) = seeded_kb.seen_messages[0][0], seeded_kb.seen_messages[0][1] @@ -433,19 +437,34 @@ def test_off_topic_question_deflects_honestly(client, db, seeded_kb: FakeRagLLM) assert 0.0 < row.top_score < get_settings().relevance_threshold assert row.fts_hits == 0 assert row.chunk_hits >= 1 + # The retrieval stays durably recorded for threshold tuning + # (observability unchanged — query_log records retrieval, not + # citations; the done frame's [] above is the citation surface). + assert row.sources def test_keyword_question_grounded_by_lexical_hit_despite_weak_cosine( - client, db, seeded_kb: FakeRagLLM + client, db, seeded_kb: FakeRagLLM, + monkeypatch: pytest.MonkeyPatch, ) -> None: """Phase 09: a name-your-tool question the vector model barely ranks ("kafkabridge" only appears in static-dns.json) must still be grounded - via the FTS branch — LOW only fires at weak cosine AND zero hits.""" + via the FTS branch — HIGH when cosine >= lexical_support_floor AND + fts_hits > 0 (A8 revised 2026-09-14). + + The conftest floor (0.15) is above the mock's cosine (~0.134), so we + lower the floor here so the corroborated-lexical path fires.""" + from app.config import get_settings # noqa: E402 + + monkeypatch.setenv("BOR_LEXICAL_SUPPORT_FLOOR", "0.10") + # get_settings is lru_cached — clear the cache so the new env var takes effect. + get_settings.cache_clear() fastapi_app.dependency_overrides[chat_api.get_llm] = lambda: seeded_kb try: _, _, frames = _stream_chat(client, "How does kafkabridge work?") finally: fastapi_app.dependency_overrides.clear() + get_settings.cache_clear() done = frames[-1] assert done["type"] == "done" diff --git a/tests/unit/test_chat_gate.py b/tests/unit/test_chat_gate.py index f3a5fad..6b898f3 100644 --- a/tests/unit/test_chat_gate.py +++ b/tests/unit/test_chat_gate.py @@ -36,10 +36,13 @@ ANSWER = "I haven't done anything like that — try one of these instead!" KB_OVERVIEW = "- Homelab\n - Kubernetes (k3s)\n- Deployments\n - Borg backups" -def _settings(threshold: float = 0.30) -> Settings: +def _settings(threshold: float = 0.30, floor: float | None = None) -> Settings: + if floor is None: + floor = threshold * 0.5 # half the threshold — keeps existing tests green return Settings( _env_file=None, # pyright: ignore[reportCallIssue] relevance_threshold=threshold, + lexical_support_floor=floor, ) @@ -118,12 +121,16 @@ def test_gate_is_env_tunable_via_settings() -> None: def test_gate_weak_cosine_with_fts_hit_still_answers() -> None: - """cosine < threshold but a lexical hit ⇒ HIGH — the FTS-OR branch. - This is the name-your-tool case: "kafkabridge" grounds despite weak - vector overlap.""" + """cosine < threshold but a lexical hit corroborated by cosine >= floor + ⇒ HIGH — the FTS-OR branch. This is the name-your-tool case: + "kafkabridge" grounds despite weak vector overlap. + + A8 revised 2026-09-14: FTS alone no longer promotes; cosine must also + clear lexical_support_floor (here 0.15 = half of threshold 0.30).""" doc = _doc("Static DNS", "DNS_DOC_CONTENT") plan = chat_api.plan_turn( - [_chunk(doc, 0.02, cosine=0.10, fts_hit=True)], _settings(threshold=0.30) + [_chunk(doc, 0.02, cosine=0.10, fts_hit=True)], + _settings(threshold=0.30, floor=0.05), # floor=0.05 so 0.10 >= floor ) assert plan.deflected is False assert plan.top_score == pytest.approx(0.10) # gate input is the cosine @@ -155,7 +162,7 @@ def test_gate_fts_hits_counts_all_lexical_candidates() -> None: _chunk(a, 0.02, cosine=0.04, fts_hit=True), # same doc, second chunk _chunk(b, 0.01, cosine=0.03), ] - plan = chat_api.plan_turn(chunks, _settings(threshold=0.30)) + plan = chat_api.plan_turn(chunks, _settings(threshold=0.30, floor=0.03)) assert plan.deflected is False assert plan.fts_hits == 2 # per chunk, not per doc @@ -176,6 +183,184 @@ def test_gate_lexical_only_chunk_does_not_inflate_cosine() -> None: assert plan.docs[0].title == "Beta" +# ---------- lexical support floor (A8 revised 2026-09-14) ---------- + + +def test_gate_fts_hit_below_floor_deflects() -> None: + """The Mongolia case: FTS hit with cosine below lexical_support_floor + → LOW (deflected). The lexical-only hit no longer promotes to HIGH. + This is the regression pin for phase 112.""" + doc = _doc("Capital Quest", "QUEST_DOC_CONTENT") + plan = chat_api.plan_turn( + [_chunk(doc, 0.05, cosine=0.10, fts_hit=True)], + _settings(threshold=0.62), + ) + assert plan.deflected is True + assert plan.top_score == pytest.approx(0.10) + assert plan.fts_hits == 1 + assert "DEFLECT_MODE" in plan.system_prompt + assert "QUEST_DOC_CONTENT" not in plan.system_prompt + assert "Capital Quest" in plan.system_prompt # title only + assert plan.suggestions # derived from weak-hit titles + + +def test_gate_fts_hit_at_floor_answers() -> None: + """FTS hit with cosine exactly at lexical_support_floor → HIGH. + The floor is inclusive (>=), not strict (<).""" + doc = _doc("Capital Quest", "QUEST_DOC_CONTENT") + plan = chat_api.plan_turn( + [_chunk(doc, 0.40, cosine=0.35, fts_hit=True)], + _settings(threshold=0.62), + ) + assert plan.deflected is False + assert plan.top_score == pytest.approx(0.35) + assert plan.fts_hits == 1 + assert "QUEST_DOC_CONTENT" in plan.system_prompt + assert plan.suggestions == [] + + +def test_gate_fts_hit_above_floor_below_threshold_answers() -> None: + """FTS hit with cosine between floor and threshold → HIGH. + The corroborated-lexical path fires.""" + doc = _doc("Capital Quest", "QUEST_DOC_CONTENT") + plan = chat_api.plan_turn( + [_chunk(doc, 0.50, cosine=0.50, fts_hit=True)], + _settings(threshold=0.62), + ) + assert plan.deflected is False + assert plan.top_score == pytest.approx(0.50) + assert plan.fts_hits == 1 + assert "QUEST_DOC_CONTENT" in plan.system_prompt + assert plan.suggestions == [] + + +def test_gate_fts_hit_above_code_default_floor_answers() -> None: + """The quadrant table's "0.50 with default settings" row: the CODE + defaults (``test_lexical_support_floor_validation_default`` pins them: + threshold 0.62 / floor 0.35) — fts>0 + cosine 0.50 >= 0.35 → HIGH. + Named literally (not via the helper's half-threshold floor) so the + production-default path is pinned on its own.""" + doc = _doc("Capital Quest", "QUEST_DOC_CONTENT") + plan = chat_api.plan_turn( + [_chunk(doc, 0.50, cosine=0.50, fts_hit=True)], + _settings(threshold=0.62, floor=0.35), + ) + assert plan.deflected is False + assert plan.top_score == pytest.approx(0.50) + assert plan.fts_hits == 1 + assert "QUEST_DOC_CONTENT" in plan.system_prompt + assert plan.suggestions == [] + + +def test_gate_high_cosine_overrides_fts_deflection() -> None: + """Strong cosine (>= threshold) → HIGH regardless of FTS status. + The cosine-primary path is unchanged.""" + doc = _doc("Capital Quest", "QUEST_DOC_CONTENT") + plan = chat_api.plan_turn( + [_chunk(doc, 0.90, cosine=0.80, fts_hit=True)], + _settings(threshold=0.62), + ) + assert plan.deflected is False + assert plan.top_score == pytest.approx(0.80) + assert plan.fts_hits == 1 + assert "QUEST_DOC_CONTENT" in plan.system_prompt + assert plan.suggestions == [] + + +def test_gate_fts_no_cosine_deflects() -> None: + """FTS hit with cosine = 0.0 → LOW (the extreme Mongolia case).""" + doc = _doc("Capital Quest", "QUEST_DOC_CONTENT") + plan = chat_api.plan_turn( + [_chunk(doc, 0.90, cosine=0.0, fts_hit=True)], + _settings(threshold=0.62), + ) + assert plan.deflected is True + assert plan.top_score == 0.0 + assert plan.fts_hits == 1 + assert "DEFLECT_MODE" in plan.system_prompt + + +def test_gate_multiple_fts_below_floor_deflects() -> None: + """Multiple FTS hits, all below lexical_support_floor → LOW. + The gate requires the BEST cosine to clear the floor, not just any hit.""" + a = _doc("Alpha Quest", "ALPHA_CONTENT") + b = _doc("Beta Quest", "BETA_CONTENT") + chunks = [ + _chunk(a, 0.30, cosine=0.20, fts_hit=True), + _chunk(b, 0.25, cosine=0.15, fts_hit=True), + ] + plan = chat_api.plan_turn(chunks, _settings(threshold=0.62)) + assert plan.deflected is True + assert plan.fts_hits == 2 + assert "DEFLECT_MODE" in plan.system_prompt + + +def test_gate_one_fts_above_floor_answers() -> None: + """Multiple chunks, one FTS hit above floor → HIGH. + The best cosine (from the corroborated hit) clears the floor.""" + a = _doc("Alpha Quest", "ALPHA_CONTENT") + b = _doc("Beta Quest", "BETA_CONTENT") + chunks = [ + _chunk(a, 0.30, cosine=0.20, fts_hit=True), # below floor + _chunk(b, 0.25, cosine=0.40, fts_hit=True), # above floor + ] + plan = chat_api.plan_turn(chunks, _settings(threshold=0.62)) + assert plan.deflected is False + assert plan.fts_hits == 2 + assert "ALPHA_CONTENT" in plan.system_prompt + assert "BETA_CONTENT" in plan.system_prompt + + +# ---------- config validation (lexical_support_floor) ---------- + + +def test_lexical_support_floor_validation_floor_above_threshold_fails() -> None: + """lexical_support_floor > relevance_threshold is rejected at startup.""" + with pytest.raises(ValueError, match="lexical_support_floor"): + Settings( + _env_file=None, # pyright: ignore[reportCallIssue] + relevance_threshold=0.62, + lexical_support_floor=0.70, + ) + + +def test_lexical_support_floor_validation_negative_fails() -> None: + """Negative lexical_support_floor is rejected.""" + with pytest.raises(ValueError, match="lexical_support_floor"): + Settings( + _env_file=None, # pyright: ignore[reportCallIssue] + lexical_support_floor=-0.1, + ) + + +def test_lexical_support_floor_validation_at_threshold_succeeds() -> None: + """lexical_support_floor == relevance_threshold is legal.""" + s = Settings( + _env_file=None, # pyright: ignore[reportCallIssue] + relevance_threshold=0.62, + lexical_support_floor=0.62, + ) + assert s.lexical_support_floor == 0.62 + + +def test_lexical_support_floor_validation_default() -> None: + """Default lexical_support_floor is 0.35.""" + import os + # Conftest sets BOR_RELEVANCE_THRESHOLD=0.30 and BOR_LEXICAL_SUPPORT_FLOOR=0.15. + # We need the CODE defaults, so clear both and let the class defaults apply. + saved_relevance = os.environ.pop("BOR_RELEVANCE_THRESHOLD", None) + saved_floor = os.environ.pop("BOR_LEXICAL_SUPPORT_FLOOR", None) + try: + s = Settings(_env_file=None) # pyright: ignore[reportCallIssue] + assert s.lexical_support_floor == 0.35 + assert s.relevance_threshold == 0.62 + finally: + if saved_relevance is not None: + os.environ["BOR_RELEVANCE_THRESHOLD"] = saved_relevance + if saved_floor is not None: + os.environ["BOR_LEXICAL_SUPPORT_FLOOR"] = saved_floor + + def test_gate_zero_chunks_deflects_with_fallback_chips() -> None: plan = chat_api.plan_turn([], _settings()) assert plan.deflected is True @@ -581,6 +766,11 @@ def test_endpoint_just_below_threshold_deflects( assert 2 <= len(done["suggestions"]) <= MAX_SUGGESTIONS # title chip + fallback assert all(s.strip() for s in done["suggestions"]) assert any("Deploying a New Service" in s for s in done["suggestions"]) + # Phase 112 (A8 revised, TODO L2): a deflected turn cites nothing — + # done.sources is the citation surface (the UI chips every entry as + # "the answer used this"), and the weak hits are scored docs, not + # citations. + assert done["sources"] == [] # The LLM saw the LOW prompt: DEFLECT_MODE + titles, never doc content. (system, user) = llm.seen[0][0], llm.seen[0][1] @@ -588,11 +778,14 @@ def test_endpoint_just_below_threshold_deflects( assert "DEFLECT_MODE" in system["content"] assert "DOC_CONTENT_NEVER_SENT" not in system["content"] - # Durable record: deflected + the weak score. + # Durable record: deflected + the weak score. The retrieval itself + # stays recorded (observability unchanged — query_log records + # retrieval, not citations; the phase-113 A3 precedent). (row,) = session.added assert isinstance(row, QueryLog) assert row.deflected is True assert row.top_score == pytest.approx(0.2999) + assert row.sources # the weak-hit doc's path, for threshold tuning assert session.commits == 1 diff --git a/tests/unit/test_prompt_lock.py b/tests/unit/test_prompt_lock.py new file mode 100644 index 0000000..64e239b --- /dev/null +++ b/tests/unit/test_prompt_lock.py @@ -0,0 +1,122 @@ +"""Prompt-lock pin (phase 112, task 03 — owner decision iii, 2026-09-14). + +The persona + HONESTY GATE text is **locked verbatim** (PLAN §6): it +changes through the plan, never in code. This module byte-pins the +locked prompt constants against their pre-phase-112 anchor values — +sha256 + exact prefix/suffix + total length, so *any* byte change +(option (i)'s copy tightening, option (ii)'s plan amendment, or an +accidental edit) fails loudly until the anchors are re-cut as part of +the same plan revision. ``tests.unit.test_prompts`` pins the +assembled-prompt structure and the behavioral contracts on top of +these constants; this file pins the constants themselves. +""" +from __future__ import annotations + +import hashlib + +from app.rag.prompts import ( + PERSONA, + TOOLS_SECTION, + _base, + build_deflect_prompt, +) + + +def _sha256(text: str) -> str: + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + +# ---------- PERSONA (the locked base of BOTH the HIGH and LOW prompts) ---------- + +#: Pre-phase-112 anchors for ``PERSONA`` (the ``{relevance}`` placeholder +#: and the line wrapping included). +PERSONA_SHA256 = "e31792a73e64c53853097e0f7b6df8b96c5f2d05c286944c3edba16dde7777fe" +PERSONA_LEN = 706 +PERSONA_PREFIX = ( + 'You are "Brain of Reese" — the digital brain of Reese, a self-hoster and\n' + "homelab tinkerer. Personality: chippy, upbeat, warm, and genuinely\n" + "optimistic about the user's ability to do things.\n" + "\n" + "Rules:\n" +) +PERSONA_SUFFIX = ( + "4. Never invent facts, hosts, or steps that are not in the context.\n" + "5. Keep answers tight: short paragraphs, bullets where helpful.\n" + "\n" + "{relevance}" +) + + +def test_persona_byte_locked() -> None: + """Any byte change to the locked persona (opening, any rule line, the + ```` placeholder) fails on the sha256; the prefix/suffix + anchors name the damaged region for the diff.""" + assert len(PERSONA) == PERSONA_LEN + assert _sha256(PERSONA) == PERSONA_SHA256 + assert PERSONA.startswith(PERSONA_PREFIX) + assert PERSONA.endswith(PERSONA_SUFFIX) + + +def test_high_and_low_bases_byte_locked() -> None: + """``_base`` only substitutes ``{relevance}`` — the per-mode base + lengths pin the substitution against a moved or re-spelled + placeholder in the locked text.""" + assert len(_base("HIGH")) == 699 # PERSONA_LEN - 11 + 4 + assert len(_base("LOW")) == 698 # PERSONA_LEN - 11 + 3 + + +# ---------- TOOLS_SECTION (the HIGH prompt's locked ```` copy) ---------- + +#: Pre-phase-112 anchors for ``TOOLS_SECTION``. +TOOLS_SECTION_SHA256 = "b834cbe368055e65da82ae3e37a91e6c658c713954703fc79b849a6ebdf4aa53" +TOOLS_SECTION_LEN = 2273 +TOOLS_SECTION_PREFIX = ( + "\n" + "You may extend your context with three tools. `ls` lists the " + "knowledge base as a tree, one level at a time: " +) +TOOLS_SECTION_SUFFIX = ( + "Never repeat a call that was refused or already succeeded — the refusal " + "already told you the correct form. Answer as soon as you have what " + "you need.\n" +) + + +def test_tools_section_byte_locked() -> None: + """The ```` teaching is LOCKED verbatim too (the E2E mock keys + on the ```` marker's presence; the wording is the owner's): + sha256 + exact prefix/suffix + total length.""" + assert len(TOOLS_SECTION) == TOOLS_SECTION_LEN + assert _sha256(TOOLS_SECTION) == TOOLS_SECTION_SHA256 + assert TOOLS_SECTION.startswith(TOOLS_SECTION_PREFIX) + assert TOOLS_SECTION.endswith(TOOLS_SECTION_SUFFIX) + + +# ---------- the LOW prompt's locked DEFLECT_MODE body ---------- + +#: The ``DEFLECT_MODE`` body exactly as it was pre-phase-112 (the +#: phase-71 plain-text line included) — inline in +#: :func:`app.rag.prompts.build_deflect_prompt`, so it is pinned through +#: the built prompt rather than a module constant. +LOW_BODY_SHA256 = "c9868cfccdd0ff79d7c1de5df5f6dea0182d3726ca912c9ef609ab54b563a0a4" +LOW_BODY_LEN = 259 +LOW_BODY = ( + "DEFLECT_MODE: retrieval was weak — the titles below are the closest " + "your notes come to the question. They are titles only; do not pretend " + "they answer it. Use them to propose 2-3 alternative questions.\n" + "Reply in plain text only — you have no tools in this mode." +) + + +def test_deflect_body_byte_locked() -> None: + """The LOW build = locked base + exactly the locked DEFLECT_MODE body + + the weak-hit title list — byte for byte (the mock keys on the + ``DEFLECT_MODE`` marker's presence; the body wording is locked).""" + assert len(LOW_BODY) == LOW_BODY_LEN + assert _sha256(LOW_BODY) == LOW_BODY_SHA256 + assert build_deflect_prompt(["T1", "T2"]) == _base("LOW") + "\n" + LOW_BODY + "\n- T1\n- T2" + prompt = build_deflect_prompt(["T1", "T2"]) + assert prompt.count(LOW_BODY) == 1 + assert prompt.index("DEFLECT_MODE") < prompt.index( + "Reply in plain text only" + ) # the marker precedes the plain-text line