diff --git a/.agents/phases/todo/114_embed_question_length/00_phase.md b/.agents/phases/complete/114_embed_question_length/00_phase.md similarity index 100% rename from .agents/phases/todo/114_embed_question_length/00_phase.md rename to .agents/phases/complete/114_embed_question_length/00_phase.md diff --git a/.agents/phases/todo/114_embed_question_length/01_embed_truncation.md b/.agents/phases/complete/114_embed_question_length/01_embed_truncation.md similarity index 100% rename from .agents/phases/todo/114_embed_question_length/01_embed_truncation.md rename to .agents/phases/complete/114_embed_question_length/01_embed_truncation.md diff --git a/.agents/phases/todo/114_embed_question_length/02_too_long_error_mapping.md b/.agents/phases/complete/114_embed_question_length/02_too_long_error_mapping.md similarity index 100% rename from .agents/phases/todo/114_embed_question_length/02_too_long_error_mapping.md rename to .agents/phases/complete/114_embed_question_length/02_too_long_error_mapping.md diff --git a/.agents/phases/todo/114_embed_question_length/03_embed_length_tests.md b/.agents/phases/complete/114_embed_question_length/03_embed_length_tests.md similarity index 100% rename from .agents/phases/todo/114_embed_question_length/03_embed_length_tests.md rename to .agents/phases/complete/114_embed_question_length/03_embed_length_tests.md diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__00_phase.a1.err b/.agents/reports/114_embed_question_length/114_embed_question_length__00_phase.a1.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__00_phase.a1.md b/.agents/reports/114_embed_question_length/114_embed_question_length__00_phase.a1.md new file mode 100644 index 0000000..fbb6f60 --- /dev/null +++ b/.agents/reports/114_embed_question_length/114_embed_question_length__00_phase.a1.md @@ -0,0 +1,14 @@ +All verification passes complete — the phase was already fully implemented in the working tree, and every gate is green. No defects found; no code changes were needed. + +**Final verification pass — Phase 114 (embed question length):** +- Verified truncation: `chat.py:459` embeds `request.message[:settings.embed_question_max_chars]` (default 1200, `BOR_EMBED_QUESTION_MAX_CHARS`, `>0` validator); full question still reaches the LLM prompt/log. +- Verified error mapping: `EmbeddingInputTooLargeError(EmbeddingError)` (byte-identical message) caught **before** `EmbeddingError` → no retry, terminal frame `detail="Question too long — trim it and re-ask."` + reachability-fine hint; `ChatErrorEvent.hint` additive. +- Verified frontend chain: frame `hint` → `err.hint` → `setUiState(error, …, {hint})` → `showErrorBanner(…, {retryable: true})` — hint replaces default `ERROR_HINT`, phase-111 `#banner-retry` button revealed. 4,000-char clamp untouched (A2). +- `uv run pytest tests/unit/test_embed_question_length.py -v --no-cov` → 21 passed +- `uv run pytest tests/e2e/test_embed_question_length.py -v --no-cov` (isolation, DB up) → 1 passed (4,000-char question → done, no banner) +- Regression: `test_llm_retry.py` 4 passed · `test_oneshot_llm_retry.py` 2 passed · `test_chip_sizing_question_cap.py` 6 passed +- `uv run pytest --cov=app --cov-report=term-missing` → 2444 passed, TOTAL **99%** (>90% gate) +- `uv run ruff check .` → All checks passed; `uv run pyright` → 0 errors, 0 warnings + +**Completion criteria:** (1) 4,000-char question embeds prefix + full prompt ✅ · (2) too-large → accurate frame + hint + Retry button ✅ · (3) reachability failure byte-identical (retries + old copy) ✅ · (4) all gates green ✅ · (5) commit/phase-move → left to the harness per instructions (no `git add`/`commit` run). +**Deviations:** none. **Next pending phase:** `115_doc_draft_discard`. diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__00_phase.a1.validate b/.agents/reports/114_embed_question_length/114_embed_question_length__00_phase.a1.validate new file mode 100644 index 0000000..14beaee --- /dev/null +++ b/.agents/reports/114_embed_question_length/114_embed_question_length__00_phase.a1.validate @@ -0,0 +1,102 @@ +........................................................................ [ 2%] +........................................................................ [ 5%] +........................................................................ [ 8%] +........................................................................ [ 11%] +........................................................................ [ 14%] +........................................................................ [ 17%] +........................................................................ [ 20%] +........................................................................ [ 23%] +........................................................................ [ 26%] +........................................................................ [ 29%] +........................................................................ [ 32%] +........................................................................ [ 35%] +........................................................................ [ 38%] +........................................................................ [ 41%] +........................................................................ [ 44%] +........................................................................ [ 47%] +........................................................................ [ 50%] +........................................................................ [ 53%] +........................................................................ [ 55%] +........................................................................ [ 58%] +........................................................................ [ 61%] +........................................................................ [ 64%] +........................................................................ [ 67%] +........................................................................ [ 70%] +........................................................................ [ 73%] +........................................................................ [ 76%] +........................................................................ [ 79%] +........................................................................ [ 82%] +........................................................................ [ 85%] +........................................................................ [ 88%] +........................................................................ [ 91%] +........................................................................ [ 94%] +........................................................................ [ 97%] +.................................................................... [100%] +=============================== warnings summary =============================== +.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1 + /var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. + from starlette.testclient import TestClient as TestClient # noqa + +-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html +================================ tests coverage ================================ +_______________ coverage: platform linux, python 3.14.7-final-0 ________________ + +Name Stmts Miss Cover +-------------------------------------------------- +app/__init__.py 1 0 100% +app/api/__init__.py 0 0 100% +app/api/auth.py 52 0 100% +app/api/chat.py 224 1 99% +app/api/chats.py 110 0 100% +app/api/config.py 13 0 100% +app/api/doc_drafts.py 94 0 100% +app/api/docs.py 156 1 99% +app/api/git_sources.py 232 0 100% +app/api/health.py 10 0 100% +app/api/steering.py 42 0 100% +app/api/suggestions.py 33 0 100% +app/api/sync.py 139 0 100% +app/api/tokens.py 40 0 100% +app/api/ui_settings.py 55 0 100% +app/config.py 210 0 100% +app/core/__init__.py 0 0 100% +app/core/auth.py 45 0 100% +app/core/caching.py 124 0 100% +app/core/debugging.py 29 2 93% +app/core/docs_push.py 39 0 100% +app/core/errors.py 5 0 100% +app/core/logging.py 13 0 100% +app/core/rate_limit.py 44 0 100% +app/core/security_headers.py 20 0 100% +app/core/theming.py 38 0 100% +app/core/tokens.py 44 0 100% +app/db.py 22 0 100% +app/main.py 66 0 100% +app/models.py 128 0 100% +app/rag/__init__.py 0 0 100% +app/rag/agent.py 317 1 99% +app/rag/archive_upload.py 134 0 100% +app/rag/chunker.py 206 4 98% +app/rag/doc_dates.py 18 0 100% +app/rag/folder_summaries.py 123 0 100% +app/rag/git_sources.py 14 0 100% +app/rag/importer.py 215 3 99% +app/rag/llm.py 244 1 99% +app/rag/overview.py 71 0 100% +app/rag/prompts.py 88 0 100% +app/rag/retriever.py 192 3 98% +app/rag/scaffolding.py 55 0 100% +app/rag/source_removal.py 41 0 100% +app/rag/sources_meta.py 16 0 100% +app/rag/suggestions.py 27 0 100% +app/rag/summarizer.py 24 0 100% +app/schemas.py 329 0 100% +-------------------------------------------------- +TOTAL 4142 16 99% +coverage gate: app/ 99% (>90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__01_embed_truncation.a1.err b/.agents/reports/114_embed_question_length/114_embed_question_length__01_embed_truncation.a1.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__01_embed_truncation.a1.md b/.agents/reports/114_embed_question_length/114_embed_question_length__01_embed_truncation.a1.md new file mode 100644 index 0000000..83facb7 --- /dev/null +++ b/.agents/reports/114_embed_question_length/114_embed_question_length__01_embed_truncation.a1.md @@ -0,0 +1,19 @@ +Task 01 is complete. All gates green. + +**Implemented** +- `app/config.py`: new `embed_question_max_chars: int = 1200` (env `BOR_EMBED_QUESTION_MAX_CHARS`, RAG-tuning section) + `> 0` startup validator (`agent_max_rounds` pattern) +- `app/api/chat.py`: embed step now embeds `request.message[: settings.embed_question_max_chars]` with a call-site comment; LLM prompt build, retrieval, and per-turn log line untouched +- `.env.example`: commented entry citing the chunker rationale (`chunker.py:30–51`, ~1.4 chars/token, ~1024-token cap) +- `tests/unit/test_embed_question_length.py` (new, 13 tests): 4,000-char question → `embed_one` gets exactly the 1200-char prefix (one call) and the turn completes; full 4,000-char question reaches the LLM request on both the deflected and grounded (agent) branches; short question → byte-identical embed call; exactly-at-budget boundary; env-tunable budget (500); default == chunker `HARD_MAX_CHARS`; validator rejects 0/negative + +**Results** +- `uv run pytest --cov=app --cov-report=term-missing` → 2435 passed, TOTAL **99%** (>90% ✓; `config.py` 100%, `api/chat.py` 99%) +- `uv run pytest tests/e2e/test_llm_retry.py tests/e2e/test_oneshot_llm_retry.py tests/e2e/test_chip_sizing_question_cap.py --no-cov` → 12 passed (phase regression gates) +- `uv run ruff check .` → All checks passed; `uv run pyright` → 0 errors, 0 warnings + +**Decisions** +- Plain `= 1200` default (not literal `Field(default=1200)`): behaviorally identical in pydantic-settings and matches every sibling setting's style +- Call-site comment wraps to 4 lines (100-char ruff limit) +- Test file created with task-01's truncation pins only; task 03 extends it with error-mapping pins + E2E + +**Next pending task:** `114_embed_question_length/02_too_long_error_mapping.md` diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__01_embed_truncation.a1.validate b/.agents/reports/114_embed_question_length/114_embed_question_length__01_embed_truncation.a1.validate new file mode 100644 index 0000000..3996744 --- /dev/null +++ b/.agents/reports/114_embed_question_length/114_embed_question_length__01_embed_truncation.a1.validate @@ -0,0 +1,102 @@ +........................................................................ [ 2%] +........................................................................ [ 5%] +........................................................................ [ 8%] +........................................................................ [ 11%] +........................................................................ [ 14%] +........................................................................ [ 17%] +........................................................................ [ 20%] +........................................................................ [ 23%] +........................................................................ [ 26%] +........................................................................ [ 29%] +........................................................................ [ 32%] +........................................................................ [ 35%] +........................................................................ [ 38%] +........................................................................ [ 41%] +........................................................................ [ 44%] +........................................................................ [ 47%] +........................................................................ [ 50%] +........................................................................ [ 53%] +........................................................................ [ 56%] +........................................................................ [ 59%] +........................................................................ [ 62%] +........................................................................ [ 65%] +........................................................................ [ 68%] +........................................................................ [ 70%] +........................................................................ [ 73%] +........................................................................ [ 76%] +........................................................................ [ 79%] +........................................................................ [ 82%] +........................................................................ [ 85%] +........................................................................ [ 88%] +........................................................................ [ 91%] +........................................................................ [ 94%] +........................................................................ [ 97%] +........................................................... [100%] +=============================== warnings summary =============================== +.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1 + /var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. + from starlette.testclient import TestClient as TestClient # noqa + +-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html +================================ tests coverage ================================ +_______________ coverage: platform linux, python 3.14.7-final-0 ________________ + +Name Stmts Miss Cover +-------------------------------------------------- +app/__init__.py 1 0 100% +app/api/__init__.py 0 0 100% +app/api/auth.py 52 0 100% +app/api/chat.py 217 1 99% +app/api/chats.py 110 0 100% +app/api/config.py 13 0 100% +app/api/doc_drafts.py 94 0 100% +app/api/docs.py 156 1 99% +app/api/git_sources.py 232 0 100% +app/api/health.py 10 0 100% +app/api/steering.py 42 0 100% +app/api/suggestions.py 33 0 100% +app/api/sync.py 139 0 100% +app/api/tokens.py 40 0 100% +app/api/ui_settings.py 55 0 100% +app/config.py 210 0 100% +app/core/__init__.py 0 0 100% +app/core/auth.py 45 0 100% +app/core/caching.py 124 0 100% +app/core/debugging.py 29 2 93% +app/core/docs_push.py 39 0 100% +app/core/errors.py 5 0 100% +app/core/logging.py 13 0 100% +app/core/rate_limit.py 44 0 100% +app/core/security_headers.py 20 0 100% +app/core/theming.py 38 0 100% +app/core/tokens.py 44 0 100% +app/db.py 22 0 100% +app/main.py 66 0 100% +app/models.py 128 0 100% +app/rag/__init__.py 0 0 100% +app/rag/agent.py 317 1 99% +app/rag/archive_upload.py 134 0 100% +app/rag/chunker.py 206 4 98% +app/rag/doc_dates.py 18 0 100% +app/rag/folder_summaries.py 123 0 100% +app/rag/git_sources.py 14 0 100% +app/rag/importer.py 215 3 99% +app/rag/llm.py 243 1 99% +app/rag/overview.py 71 0 100% +app/rag/prompts.py 88 0 100% +app/rag/retriever.py 192 3 98% +app/rag/scaffolding.py 55 0 100% +app/rag/source_removal.py 41 0 100% +app/rag/sources_meta.py 16 0 100% +app/rag/suggestions.py 27 0 100% +app/rag/summarizer.py 24 0 100% +app/schemas.py 328 0 100% +-------------------------------------------------- +TOTAL 4133 16 99% +coverage gate: app/ 99% (>90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__02_too_long_error_mapping.a1.err b/.agents/reports/114_embed_question_length/114_embed_question_length__02_too_long_error_mapping.a1.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__02_too_long_error_mapping.a1.md b/.agents/reports/114_embed_question_length/114_embed_question_length__02_too_long_error_mapping.a1.md new file mode 100644 index 0000000..1a69af2 --- /dev/null +++ b/.agents/reports/114_embed_question_length/114_embed_question_length__02_too_long_error_mapping.a1.md @@ -0,0 +1,20 @@ +All green. Final state verified: full suite 2446/2446 passing, coverage 99%, ruff + pyright clean, regression E2E suites passing in isolation. + +## Report — Task 02: too-long embed error mapping (phase 114) + +**Implemented** +- `app/rag/llm.py`: new `EmbeddingInputTooLargeError(EmbeddingError)`; single-text `_TooLarge` branch raises it with byte-identical message (importer drop-in, still catches `EmbeddingError`) +- `app/schemas.py`: `ChatErrorEvent` gains additive `hint: str | None = None` (PLAN §4 old-client-ignore contract) +- `app/api/chat.py`: catches the subclass **before** `EmbeddingError` in the phase-67 loop → no retry (locked A3), error log with `too-large` marker, terminal frame `{detail: "Question too long — trim it and re-ask.", hint: "The app reached the embedding model fine — …"}` +- `frontend/assets/app.js`: `showErrorBanner` honors `opts.hint ?? ERROR_HINT`; SSE error branch threads `ev.hint` through the throw → `setUiState(state, detail, opts)` → `{ retryable: true, ...opts }` (phase-111 Retry button rides the same path) +- Tests: 3 error-mapping unit tests (too-large → 1 terminal frame, no retry, 1 attempt; 502 → legacy retry path + old copy + `hint: null`; subclass message pin) + 4 frontend source-assertion pins in `tests/unit/test_embed_question_length.py`; updated stale `{type, detail}`-only shape pins in `test_sse_events.py` and `test_chat_api.py` (2 tests) + +**Verification** +- `uv run pytest --cov=app` → 2446 passed, 0 failed; app/ coverage **99%** (>90%) +- `uv run ruff check .` → clean; `uv run pyright` → 0 errors +- `uv run pytest tests/e2e/test_llm_retry.py --no-cov` → 4 passed; `test_oneshot_llm_retry.py` → 2 passed; `test_chip_sizing_question_cap.py` → 6 passed +- `node --check frontend/assets/app.js` → clean + +**Decisions**: bare `model_dump()` (hintless frames serialize `"hint": null`) — matches the `ChatDoneEvent.related` additive-field house convention; E2E file `tests/e2e/test_embed_question_length.py` is task 03's scope (mock LLM verified to have no embed size cap, so it will pass). + +**Next pending task**: `.agents/phases/todo/114_embed_question_length/03_embed_length_tests.md` diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__02_too_long_error_mapping.a1.validate b/.agents/reports/114_embed_question_length/114_embed_question_length__02_too_long_error_mapping.a1.validate new file mode 100644 index 0000000..cfa7bce --- /dev/null +++ b/.agents/reports/114_embed_question_length/114_embed_question_length__02_too_long_error_mapping.a1.validate @@ -0,0 +1,102 @@ +........................................................................ [ 2%] +........................................................................ [ 5%] +........................................................................ [ 8%] +........................................................................ [ 11%] +........................................................................ [ 14%] +........................................................................ [ 17%] +........................................................................ [ 20%] +........................................................................ [ 23%] +........................................................................ [ 26%] +........................................................................ [ 29%] +........................................................................ [ 32%] +........................................................................ [ 35%] +........................................................................ [ 38%] +........................................................................ [ 41%] +........................................................................ [ 44%] +........................................................................ [ 47%] +........................................................................ [ 50%] +........................................................................ [ 53%] +........................................................................ [ 55%] +........................................................................ [ 58%] +........................................................................ [ 61%] +........................................................................ [ 64%] +........................................................................ [ 67%] +........................................................................ [ 70%] +........................................................................ [ 73%] +........................................................................ [ 76%] +........................................................................ [ 79%] +........................................................................ [ 82%] +........................................................................ [ 85%] +........................................................................ [ 88%] +........................................................................ [ 91%] +........................................................................ [ 94%] +........................................................................ [ 97%] +................................................................... [100%] +=============================== warnings summary =============================== +.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1 + /var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. + from starlette.testclient import TestClient as TestClient # noqa + +-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html +================================ tests coverage ================================ +_______________ coverage: platform linux, python 3.14.7-final-0 ________________ + +Name Stmts Miss Cover +-------------------------------------------------- +app/__init__.py 1 0 100% +app/api/__init__.py 0 0 100% +app/api/auth.py 52 0 100% +app/api/chat.py 224 1 99% +app/api/chats.py 110 0 100% +app/api/config.py 13 0 100% +app/api/doc_drafts.py 94 0 100% +app/api/docs.py 156 1 99% +app/api/git_sources.py 232 0 100% +app/api/health.py 10 0 100% +app/api/steering.py 42 0 100% +app/api/suggestions.py 33 0 100% +app/api/sync.py 139 0 100% +app/api/tokens.py 40 0 100% +app/api/ui_settings.py 55 0 100% +app/config.py 210 0 100% +app/core/__init__.py 0 0 100% +app/core/auth.py 45 0 100% +app/core/caching.py 124 0 100% +app/core/debugging.py 29 2 93% +app/core/docs_push.py 39 0 100% +app/core/errors.py 5 0 100% +app/core/logging.py 13 0 100% +app/core/rate_limit.py 44 0 100% +app/core/security_headers.py 20 0 100% +app/core/theming.py 38 0 100% +app/core/tokens.py 44 0 100% +app/db.py 22 0 100% +app/main.py 66 0 100% +app/models.py 128 0 100% +app/rag/__init__.py 0 0 100% +app/rag/agent.py 317 1 99% +app/rag/archive_upload.py 134 0 100% +app/rag/chunker.py 206 4 98% +app/rag/doc_dates.py 18 0 100% +app/rag/folder_summaries.py 123 0 100% +app/rag/git_sources.py 14 0 100% +app/rag/importer.py 215 3 99% +app/rag/llm.py 244 1 99% +app/rag/overview.py 71 0 100% +app/rag/prompts.py 88 0 100% +app/rag/retriever.py 192 3 98% +app/rag/scaffolding.py 55 0 100% +app/rag/source_removal.py 41 0 100% +app/rag/sources_meta.py 16 0 100% +app/rag/suggestions.py 27 0 100% +app/rag/summarizer.py 24 0 100% +app/schemas.py 329 0 100% +-------------------------------------------------- +TOTAL 4142 16 99% +coverage gate: app/ 99% (>90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__03_embed_length_tests.a1.err b/.agents/reports/114_embed_question_length/114_embed_question_length__03_embed_length_tests.a1.err new file mode 100644 index 0000000..e69de29 diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__03_embed_length_tests.a1.md b/.agents/reports/114_embed_question_length/114_embed_question_length__03_embed_length_tests.a1.md new file mode 100644 index 0000000..6ba8b82 --- /dev/null +++ b/.agents/reports/114_embed_question_length/114_embed_question_length__03_embed_length_tests.a1.md @@ -0,0 +1,16 @@ +Task 03 complete — all green. + +**Implemented** +- `tests/e2e/test_embed_question_length.py` (new): 4,000-char (composer-clamp) question typed to the counter's "4000/4000 — character limit" + `.is-max` → send → grounded mock answer streams to done, no error banner, input/counter cleared (the L6 acceptance pin — pre-phase this died with "couldn't reach the embedding model"). +- `tests/unit/test_embed_question_length.py` (existed from tasks 01/02; refined to task text): short-question byte-identical pin parametrized [100, 900]; legacy regression pin now a 500 *without* the too-large signature (retry frames + old copy + `hint: null`). +- Verified tasks 01/02 code in-tree (1200-char prefix embed, `EmbeddingInputTooLargeError` → accurate "Question too long" frame, no retry, hint; `ChatErrorEvent.hint`; frontend `showErrorBanner` opts.hint) — no app changes needed. + +**Results** +- `uv run pytest tests/e2e/test_embed_question_length.py -v --no-cov` → 1 passed (DB up) +- Regression E2E `test_llm_retry.py` + `test_oneshot_llm_retry.py` + `test_chip_sizing_question_cap.py` → 12 passed +- `uv run pytest --cov=app --cov-report=term-missing` → 2444 passed, app/ **99%** (>90%) +- `uv run ruff check .` → clean; `uv run pyright` → 0 errors, 0 warnings + +**Decisions** — E2E question repeats the fixture-KB "kubernetes" sentence cut to exactly 4,000 chars so retrieval grounds (strongest "turn completed" reading); mock LLM has no token cap, so the prefix itself is unit-pinned (phase split). No defects found in prior work. + +**Next pending task** — phase 114 is now complete (harness commits/moves it); next: `115_doc_draft_discard` (first task in `.agents/phases/todo/`). diff --git a/.agents/reports/114_embed_question_length/114_embed_question_length__03_embed_length_tests.a1.validate b/.agents/reports/114_embed_question_length/114_embed_question_length__03_embed_length_tests.a1.validate new file mode 100644 index 0000000..14beaee --- /dev/null +++ b/.agents/reports/114_embed_question_length/114_embed_question_length__03_embed_length_tests.a1.validate @@ -0,0 +1,102 @@ +........................................................................ [ 2%] +........................................................................ [ 5%] +........................................................................ [ 8%] +........................................................................ [ 11%] +........................................................................ [ 14%] +........................................................................ [ 17%] +........................................................................ [ 20%] +........................................................................ [ 23%] +........................................................................ [ 26%] +........................................................................ [ 29%] +........................................................................ [ 32%] +........................................................................ [ 35%] +........................................................................ [ 38%] +........................................................................ [ 41%] +........................................................................ [ 44%] +........................................................................ [ 47%] +........................................................................ [ 50%] +........................................................................ [ 53%] +........................................................................ [ 55%] +........................................................................ [ 58%] +........................................................................ [ 61%] +........................................................................ [ 64%] +........................................................................ [ 67%] +........................................................................ [ 70%] +........................................................................ [ 73%] +........................................................................ [ 76%] +........................................................................ [ 79%] +........................................................................ [ 82%] +........................................................................ [ 85%] +........................................................................ [ 88%] +........................................................................ [ 91%] +........................................................................ [ 94%] +........................................................................ [ 97%] +.................................................................... [100%] +=============================== warnings summary =============================== +.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1 + /var/home/ducoterra/Projects/Personal/brain-of-reese/.venv/lib64/python3.14/site-packages/fastapi/testclient.py:1: StarletteDeprecationWarning: Using `httpx` with `starlette.testclient` is deprecated; install `httpx2` instead. + from starlette.testclient import TestClient as TestClient # noqa + +-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html +================================ tests coverage ================================ +_______________ coverage: platform linux, python 3.14.7-final-0 ________________ + +Name Stmts Miss Cover +-------------------------------------------------- +app/__init__.py 1 0 100% +app/api/__init__.py 0 0 100% +app/api/auth.py 52 0 100% +app/api/chat.py 224 1 99% +app/api/chats.py 110 0 100% +app/api/config.py 13 0 100% +app/api/doc_drafts.py 94 0 100% +app/api/docs.py 156 1 99% +app/api/git_sources.py 232 0 100% +app/api/health.py 10 0 100% +app/api/steering.py 42 0 100% +app/api/suggestions.py 33 0 100% +app/api/sync.py 139 0 100% +app/api/tokens.py 40 0 100% +app/api/ui_settings.py 55 0 100% +app/config.py 210 0 100% +app/core/__init__.py 0 0 100% +app/core/auth.py 45 0 100% +app/core/caching.py 124 0 100% +app/core/debugging.py 29 2 93% +app/core/docs_push.py 39 0 100% +app/core/errors.py 5 0 100% +app/core/logging.py 13 0 100% +app/core/rate_limit.py 44 0 100% +app/core/security_headers.py 20 0 100% +app/core/theming.py 38 0 100% +app/core/tokens.py 44 0 100% +app/db.py 22 0 100% +app/main.py 66 0 100% +app/models.py 128 0 100% +app/rag/__init__.py 0 0 100% +app/rag/agent.py 317 1 99% +app/rag/archive_upload.py 134 0 100% +app/rag/chunker.py 206 4 98% +app/rag/doc_dates.py 18 0 100% +app/rag/folder_summaries.py 123 0 100% +app/rag/git_sources.py 14 0 100% +app/rag/importer.py 215 3 99% +app/rag/llm.py 244 1 99% +app/rag/overview.py 71 0 100% +app/rag/prompts.py 88 0 100% +app/rag/retriever.py 192 3 98% +app/rag/scaffolding.py 55 0 100% +app/rag/source_removal.py 41 0 100% +app/rag/sources_meta.py 16 0 100% +app/rag/suggestions.py 27 0 100% +app/rag/summarizer.py 24 0 100% +app/schemas.py 329 0 100% +-------------------------------------------------- +TOTAL 4142 16 99% +coverage gate: app/ 99% (>90%) OK +All checks passed! +0 errors, 0 warnings, 0 informations +WARNING: there is a new pyright version available (v1.1.411 -> v1.1.414). +Please install the new version or set PYRIGHT_PYTHON_FORCE_VERSION to `latest` + +validation OK diff --git a/.env.example b/.env.example index 6f2a195..d6ebd65 100644 --- a/.env.example +++ b/.env.example @@ -47,6 +47,11 @@ BOR_FOLDER_SUMMARY_INPUT_MAX_CHARS=8000 # cap on a folder's document list sent t BOR_CHUNK_TARGET_CHARS=2000 BOR_CHUNK_OVERLAP_CHARS=200 BOR_EMBED_BATCH_SIZE=16 +# BOR_EMBED_QUESTION_MAX_CHARS=1200 # chat question-embed prefix cap (phase 114): the embed step embeds at most this +# # many chars of the question. Default = the chunker's HARD_MAX_CHARS budget +# # (app/rag/chunker.py:30-51 — worst-case ~1.4 chars/token stays under the +# # endpoint's ~1024-token per-request input cap). The FULL question still +# # reaches the LLM prompt — only the embedding is bounded (TODO L6). # --- Hybrid retrieval (vector + Postgres FTS, RRF-fused) --- BOR_HYBRID_VECTOR_CANDIDATES=100 # cosine list width for the fusion diff --git a/app/api/chat.py b/app/api/chat.py index 39c502a..84ced78 100644 --- a/app/api/chat.py +++ b/app/api/chat.py @@ -184,6 +184,7 @@ from app.rag.agent import ( ) from app.rag.llm import ( EmbeddingError, + EmbeddingInputTooLargeError, # phase 114: the deterministic too-large failure LLMClient, LLMError, RetryPiece, # phase 67: one LLM request restart (an SSE retry frame) @@ -450,8 +451,44 @@ async def chat( attempt = 1 while True: try: - question_vec = await llm.embed_one(request.message) + # Phase 114 (TODO L6): the embed input is bounded to the + # model's per-request input cap (the chunker's 1200-char + # budget, env-tunable) — the FULL question still reaches + # the LLM prompt (prompt build + log line untouched). + question_vec = await llm.embed_one( + request.message[: settings.embed_question_max_chars] + ) break + except EmbeddingInputTooLargeError as e: + # Phase 114 (TODO L6, locked A3): a too-large input is + # DETERMINISTIC — retrying the same size is guaranteed + # to repeat — so this short-circuits the phase-67 retry + # loop: no ``retry`` frame, no restart, one terminal + # error frame with the accurate "question too long" + # detail + the reachability-fine hint (the + # "couldn't reach" copy and the retry budget stay for + # reachability failures only — the branch below). + embed_ms = int((time.monotonic() - t0) * 1000) + total_ms = int((time.monotonic() - started) * 1000) + logger.error( + "chat: question=%r embed_ms=%d total_ms=%d — " + "embedding failed (too-large): %s", + request.message, + embed_ms, + total_ms, + e, + ) + settled = True # terminal: the error frame settles the turn + yield sse_event( + ChatErrorEvent( + detail="Question too long — trim it and re-ask.", + hint=( + "The app reached the embedding model fine — " + "only the question length is the problem." + ), + ).model_dump() + ) + return except EmbeddingError as e: embed_ms = int((time.monotonic() - t0) * 1000) if attempt >= max_attempts: diff --git a/app/config.py b/app/config.py index c0d58cb..afa3b8e 100644 --- a/app/config.py +++ b/app/config.py @@ -156,6 +156,18 @@ class Settings(BaseSettings): chunk_target_chars: int = 2_000 chunk_overlap_chars: int = 200 embed_batch_size: int = 16 + #: Char budget for the chat turn's question embed (phase 114, TODO L6; + #: LOCKED A1): the embed step embeds at most this many chars of the + #: question — the default 1200 is the chunker's ``HARD_MAX_CHARS`` + #: budget (``app/rag/chunker.py``: worst-case ~1.4 chars/token, so it + #: stays under the endpoint's ~1024-token per-request input cap). Only + #: the embedding is bounded: the FULL question still reaches the LLM + #: prompt (prompt build untouched), and a question at or under the + #: budget embeds byte-identically to the pre-phase path. A model with a + #: larger/smaller cap is accommodated by env, no code change (A1). + #: ``0``/negative is a typo — the validator fails loudly at startup + #: (the ``agent_max_rounds`` pattern). + embed_question_max_chars: int = 1200 #: Total char budget for the ```` section of the system prompt #: (phase 15, steering notes). The newest-fitting notes are kept and the #: overflow is replaced by the ``[…truncated…]`` marker. @@ -415,6 +427,15 @@ class Settings(BaseSettings): raise ValueError("read_max_chars must be >= 0 (chars)") return v + @field_validator("embed_question_max_chars") + @classmethod + def _embed_question_max_chars_positive(cls, v: int) -> int: + """``0``/negative would embed an empty/absent prefix — fail loud at + startup (the ``agent_max_rounds`` pattern, phase 114).""" + if v <= 0: + raise ValueError("embed_question_max_chars must be > 0 (chars)") + return v + @field_validator("llm_retries") @classmethod def _llm_retries_non_negative(cls, v: int) -> int: diff --git a/app/rag/llm.py b/app/rag/llm.py index aced9cb..3a975b5 100644 --- a/app/rag/llm.py +++ b/app/rag/llm.py @@ -43,6 +43,21 @@ class EmbeddingError(RuntimeError): """The embeddings endpoint failed (network, HTTP, or malformed reply).""" +class EmbeddingInputTooLargeError(EmbeddingError): + """A single text exceeded the endpoint's per-request input token cap + (phase 114, TODO L6). + + The endpoint REACHED and answered — this is a deterministic input- + SIZE failure, not a reachability problem: retrying the same input is + guaranteed to repeat (locked A3), so the chat endpoint catches this + BEFORE :class:`EmbeddingError` and settles with the accurate + "question too long" terminal error (no phase-67 retry frames). The + importer keeps catching the parent :class:`EmbeddingError` — this + subclass is a drop-in there and its message text is byte-identical + to the pre-phase import-oriented copy. + """ + + class EmbeddingDimensionError(EmbeddingError): """Embedding dimension != BOR_EMBEDDING_DIM — import must fail loudly.""" @@ -278,7 +293,13 @@ class LLMClient: return await self._post_embeddings(chunk) except _TooLarge: if len(chunk) == 1: - raise EmbeddingError( + # Phase 114 (TODO L6): a single text over the cap is a + # deterministic input-size failure (never reachability) — + # the subclass lets the chat endpoint map it to the + # accurate "question too long" error and skip the retry + # loop (locked A3). The message text stays byte-identical: + # the importer catches the parent EmbeddingError. + raise EmbeddingInputTooLargeError( f"a single {len(chunk[0])}-char chunk exceeded the endpoint's " "per-request input token cap — lower BOR_CHUNK_TARGET_CHARS " "and re-import" diff --git a/app/schemas.py b/app/schemas.py index 2bbb5ca..adbbf95 100644 --- a/app/schemas.py +++ b/app/schemas.py @@ -206,10 +206,18 @@ class ChatErrorEvent(BaseModel): The client's loading-feedback state machine (phase 06) keys off this exact shape — ``{type: "error", detail: str}`` — to flip to the error state and re-enable the send button. + + ``hint`` (phase 114, TODO L6): an optional one-line clarification the + client shows IN PLACE of its default reachability hint when present + (the "question too long" frame: the embedding model WAS reached — + only the question's length is the problem). Additive: frames without + it serialize ``"hint": null``, and old clients ignore the field + (PLAN §4 house contract — the ``ChatDoneEvent.related`` pattern). """ type: str = "error" detail: str + hint: str | None = None class ChatRetryEvent(BaseModel): diff --git a/frontend/assets/app.js b/frontend/assets/app.js index 58882b2..1809084 100644 --- a/frontend/assets/app.js +++ b/frontend/assets/app.js @@ -1251,7 +1251,7 @@ document.addEventListener("visibilitychange", () => { * turn is in flight (click or Enter aborts it). It is never disabled * anymore, and the spinner never shows: the "Stop" label + the .is-stop * treatment carry the in-flight state. */ -export function setUiState(state, errorDetail = "") { +export function setUiState(state, errorDetail = "", opts = {}) { uiState = state; stopThinkingClock(); clearTurnTimeout(); // the guard only owns the pre-token window @@ -1278,7 +1278,11 @@ export function setUiState(state, errorDetail = "") { } else { removeTyping(); } - if (state === UI_STATE.error) showErrorBanner(errorDetail, { retryable: true }); + // Phase 114 (TODO L6): turn-error opts (the SSE error frame's optional + // hint) merge into the banner call — the frame's hint replaces the + // default reachability hint when present (showErrorBanner's opts.hint). + if (state === UI_STATE.error) + showErrorBanner(errorDetail, { retryable: true, ...opts }); } /* Phase 104 (owner 2026-09-12, A3/A4): the visible question-length cap. @@ -2135,7 +2139,12 @@ function showErrorBanner(detail, opts = {}) { banner.hidden = false; banner.classList.add("is-error"); banner.setAttribute("role", "alert"); - bannerText.textContent = detail ? `${detail} ${ERROR_HINT}` : ERROR_HINT; + // Phase 114 (TODO L6): a frame-carried hint (the "question too long" + // case — reachability is fine, only the length is the problem) replaces + // the default reachability hint when present. + bannerText.textContent = detail + ? `${detail} ${opts.hint ?? ERROR_HINT}` + : (opts.hint ?? ERROR_HINT); // Phase 111 (task 01): reveal the banner Retry button only for failed // chat turns (opts.retryable) AND when a retryable bubble exists. if (opts.retryable) { @@ -2558,7 +2567,12 @@ async function runTurn(text, { reask = false } = {}) { } } } else if (ev.type === "error") { - throw new Error(ev.detail || "Something went wrong on my side."); + // Phase 114 (TODO L6): carry the frame's optional hint (the + // "question too long" frame has one) through the throw — the + // banner shows it in place of the default reachability hint. + const err = new Error(ev.detail || "Something went wrong on my side."); + err.hint = ev.hint; + throw err; } }); // Stream-drop guard (phase 17): frames arrived but no `done` event — @@ -2619,7 +2633,10 @@ async function runTurn(text, { reask = false } = {}) { err instanceof Error && err.message ? err.message : "Something went wrong on my side."; - setUiState(UI_STATE.error, detail); + // Phase 114 (TODO L6): the SSE error frame's optional hint flows to + // the banner (setUiState → showErrorBanner's opts.hint); the + // phase-111 Retry button rides along on the same turn-error path. + setUiState(UI_STATE.error, detail, err instanceof Error ? { hint: err.hint } : {}); } } finally { // done | error | stop → idle: always settle, always focus back. diff --git a/tests/e2e/test_embed_question_length.py b/tests/e2e/test_embed_question_length.py new file mode 100644 index 0000000..9c81c83 --- /dev/null +++ b/tests/e2e/test_embed_question_length.py @@ -0,0 +1,195 @@ +"""Phase 114 E2E (Playwright): the 4,000-char (composer-clamp) question +sends a CLEAN turn — the L6 acceptance pin (TODO.md L179–181). + +Source: TODO.md L149–181 — "L6 — 4,000-char question clamp exceeds the +embed model's input cap → misleading 'couldn't reach the embedding model' +error (2026-09-15, brain-of-reese interactive test)". The repro was 100% +reliable: a question at the composer's 4,000-char clamp (~903 tokens) +made the REAL aipi endpoint's litellm reject the embedding with +``input (903 tokens) is too large to process`` (HTTP 500) — and the +turn died pre-token with the banner "I couldn't reach the embedding +model — please try again." + +The fix under test (LOCKED A1 + A2, 00_phase.md): the embed step now +embeds at most ``embed_question_max_chars`` (default 1200 — the +chunker's ``HARD_MAX_CHARS`` budget, env-tunable) of the question — +the unit suite (``tests/unit/test_embed_question_length.py``) pins that +``embed_one`` receives EXACTLY the 1200-char prefix while the FULL +question still reaches the LLM prompt — while the 4,000-char composer +clamp stays (locked A2: truncation, not a lower clamp). + +Pinned here (the truncated-embed success path — the unit suite pins the +prefix itself and the too-long error mapping): + +* a question typed to the FULL clamp (EXACTLY 4,000 chars — the counter + reads ``4000/4000 — character limit`` + ``.is-max``) sends: the turn + streams to ``done`` on the mock LLM with the grounded answer marker, + NO error banner (the pre-phase "couldn't reach the embedding model" + death is gone), the user bubble carries the FULL 4,000-char question, + and the input + counter clear (never stale, PLAN §7.4). + +The question repeats a fixture-KB sentence (``kubernetes`` — indexed by +``tests/fixtures/docs/homelab/kubernetes.md``) so the hybrid retrieval +grounds (the mock-calibrated 0.30 threshold in the e2e conftest) and +the brain bubble carries ``MOCK_ANSWER_MARKER`` — the strongest +"the turn completed" reading. + +The endpoint and the chat are authed (phase 79, ``require_user``), so +the test signs in as admin first (``auth_helpers.login``). The send +auto-saves a ``saved_chats`` row, so the autouse fixture truncates that +table before and after the test (the phase-80/103 isolation pattern). + +Run in isolation (DB must be up: ``podman compose up -d db``): + + uv run pytest tests/e2e/test_embed_question_length.py -v --no-cov +""" +from __future__ import annotations + +import asyncio +import re +from collections.abc import Iterator +from pathlib import Path +from threading import Thread +from typing import Any + +import pytest +from playwright.sync_api import Page, expect +from sqlalchemy import text + +from app.config import Settings +from app.db import SessionLocal +from app.rag.importer import ImportSummary, import_sources +from app.rag.llm import LLMClient +from e2e.auth_helpers import login + +REPO = Path(__file__).resolve().parents[2] +FIXTURES = REPO / "tests" / "fixtures" / "docs" +MOCK_ANSWER_MARKER = "Deterministic mock answer for E2E" + +#: The composer's hard cap (``maxlength="4000"`` mirroring the server +#: ``ChatRequest.message max_length=4000`` — the phase-104 A3 clamp). +_CAP = 4_000 + +#: A sentence the fixture KB indexes (``kubernetes`` — the FTS leg of +#: the hybrid retrieval hits, so the turn GROUNDS) repeated to the +#: clamp: the L6 repro shape — a legal 4,000-char question whose pre- +#: phase embedding input (~903 tokens) exceeded the real endpoint's +#: per-request input cap. +_QUESTION_SENTENCE = ( + "How is my homelab kubernetes cluster configured for long-running batch jobs? " +) +QUESTION = (_QUESTION_SENTENCE * 52)[:_CAP] +assert len(QUESTION) == _CAP, "the question must land EXACTLY at the clamp" + + +@pytest.fixture(autouse=True) +def clean_chats(db_ready: None) -> Iterator[None]: + """The send auto-saves a row per turn — truncate ``saved_chats`` + before and after the test so it starts from (and leaves) an empty + deployment (the phase-80/103 autouse pattern).""" + with SessionLocal() as db: + db.execute(text("TRUNCATE saved_chats")) + db.commit() + yield + with SessionLocal() as db: + db.execute(text("TRUNCATE saved_chats")) + db.commit() + + +async def _import_fixtures(mock_port: int) -> ImportSummary: + kwargs: dict[str, Any] = {"_env_file": None, "llm_base_url": f"http://127.0.0.1:{mock_port}/v1"} + settings = Settings(**kwargs) # pyright: ignore[reportCallIssue] + return await import_sources([FIXTURES], LLMClient(settings)) + + +def _run_in_thread(coro: Any) -> Any: + """Run a coroutine on a worker thread (Playwright owns the test loop).""" + box: dict[str, Any] = {} + + def runner() -> None: + try: + box["value"] = asyncio.run(coro) + except BaseException as e: # noqa: BLE001 — re-raised on the test thread + box["error"] = e + + t = Thread(target=runner) + t.start() + t.join() + if "error" in box: + raise box["error"] + return box["value"] + + +def _seed_kb(mock_port: int) -> ImportSummary: + """Deterministic KB: truncate the KB tables, import the fixture + docs (needed for the grounded answer).""" + with SessionLocal() as db: + db.execute(text("TRUNCATE chunks, documents, query_log")) + db.commit() + summary = _run_in_thread(_import_fixtures(mock_port)) + assert summary is not None and summary.added == 13 # A9 formats (phase 47 added quadlet+j2) + return summary + + +def _wait_chat_booted(page: Page) -> None: + """Wait until app.js has FINISHED booting the chat page. The login + helper returns on the URL change (navigation commit) — the page's + module script may still be executing, and an ``input`` event + dispatched before its top-level listener registrations land on a + page whose listeners do not exist yet (the event is simply lost). + ``#view-chat.chat-booted`` is added two frames after the boot + settles (AFTER every top-level listener), so it is the "the app's + JS is live" sentinel.""" + page.wait_for_function( + "() => document.getElementById('view-chat')?." + "classList.contains('chat-booted')", + timeout=15_000, + ) + expect(page.locator("#view-chat")).to_have_class(re.compile(r"\bchat-booted\b")) + + +def test_question_at_the_composer_clamp_sends_a_clean_turn( + page: Page, app_url: str, mock_llm: int, db_ready: None +) -> None: + """L6 acceptance: a question typed to the FULL 4,000-char clamp + (counter ``4000/4000 — character limit`` + ``.is-max``) sends — the + embed step sends only the bounded 1200-char prefix to the model + (unit-pinned), so the turn streams to ``done`` on the mock LLM: + the user bubble carries the FULL 4,000-char question, the brain + bubble carries the grounded mock marker, NO error banner (the + pre-phase "couldn't reach the embedding model" death), and the + input + counter clear.""" + _seed_kb(mock_llm) + page.set_default_timeout(30_000) + login(page, app_url, next="/") + _wait_chat_booted(page) + + counter = page.locator("#char-count") + input_el = page.locator("#message-input") + banner = page.locator("#kb-banner") + expect(counter).to_be_hidden() + expect(banner).to_be_hidden() + + # Type the full-clamp question: fill sets the value + dispatches + # the input event (the counter path) — EXACTLY 4,000 chars. + page.fill("#message-input", QUESTION) + expect(input_el).to_have_value(QUESTION) + expect(counter).to_be_visible(timeout=5_000) + expect(counter).to_have_text("4000/4000 — character limit") + expect(counter).to_have_class(re.compile("is-max")) + + # Send at the clamp: the bounded-prefix embed succeeds and the turn + # streams to done — no error frame of any kind. + page.click("#send-btn") + expect(page.locator(".msg.user .bubble")).to_have_count(1, timeout=30_000) + expect(page.locator(".msg.user .bubble")).to_have_text(QUESTION) + brain = page.locator(".msg.brain .bubble").first + expect(brain).to_contain_text(MOCK_ANSWER_MARKER, timeout=30_000) + expect(page.locator(".msg.brain.is-deflected")).to_have_count(0) + expect(banner).to_be_hidden() # NO "couldn't reach the embedding model" death + + # Never stale: the turn cleared the input AND the counter, and the + # send button recovered. + expect(input_el).to_have_value("") + expect(counter).to_be_hidden() + expect(page.locator("#send-btn")).to_be_enabled() diff --git a/tests/integration/test_chat_api.py b/tests/integration/test_chat_api.py index 693004b..681a16b 100644 --- a/tests/integration/test_chat_api.py +++ b/tests/integration/test_chat_api.py @@ -668,9 +668,11 @@ def test_chat_mid_stream_failure_yields_error_after_partial_deltas(client, db) - def test_error_event_matches_contract_shape( client, db, seeded_kb, monkeypatch: pytest.MonkeyPatch ) -> None: - """The SSE error event (PLAN §4) is exactly ``{type, detail}`` — the - client's loading-feedback state machine (phase 06) keys off this shape - to flip to the error state and re-enable the send button. + """The SSE error event (PLAN §4) is exactly ``{type, detail, hint}`` — + the client's loading-feedback state machine (phase 06) keys off the + ``type``/``detail`` shape to flip to the error state and re-enable + the send button; ``hint`` (phase 114, TODO L6) is additive — present + as ``null`` on reachability frames, old clients ignore it. ``llm_retries=0`` keeps this a single-attempt turn: the contract under test is the error frame itself, not the phase-67 retry loop.""" broken = FakeRagLLM(embed_error=EmbeddingError("embeddings endpoint down")) @@ -686,9 +688,10 @@ def test_error_event_matches_contract_shape( assert len(frames) == 1 event = frames[0] - assert set(event.keys()) == {"type", "detail"} + assert set(event.keys()) == {"type", "detail", "hint"} assert event["type"] == "error" assert isinstance(event["detail"], str) and event["detail"] + assert event["hint"] is None # reachability frame — no too-long hint def test_chat_db_down_returns_503_json(client, monkeypatch) -> None: @@ -1704,9 +1707,9 @@ def test_deflected_scaffolding_twice_settles_malformed( ) -> None: """(b) The recovery answer is scaffolding again — a second empty reply is terminal: the DEDICATED error frame (the exact copy), no - ``done``, no query_log row — byte-for-byte today's ``LLMError`` - terminal shape — and no third request (at most one recovery per - turn).""" + ``done``, no query_log row — the standard ``LLMError`` terminal + shape (phase 114: the additive ``hint`` field is ``null`` here) — + and no third request (at most one recovery per turn).""" span = _scaffold_span() dead = FakeRagLLM(answer_sequence=[span, span]) live = get_settings() @@ -1721,7 +1724,12 @@ def test_deflected_scaffolding_twice_settles_malformed( assert frames[0]["detail"] == ( "The model returned a malformed reply — please try again." ) - assert set(frames[0].keys()) == {"type", "detail"} # the contract shape + assert set(frames[0].keys()) == { + "type", + "detail", + "hint", + } # the contract shape (phase 114: additive hint — null here) + assert frames[0]["hint"] is None assert span not in json.dumps(frames) assert not any(f["type"] == "done" for f in frames) assert db.scalars(select(QueryLog)).all() == [] diff --git a/tests/unit/test_embed_question_length.py b/tests/unit/test_embed_question_length.py new file mode 100644 index 0000000..c1e5249 --- /dev/null +++ b/tests/unit/test_embed_question_length.py @@ -0,0 +1,569 @@ +"""Unit: the chat question-embed prefix budget (phase 114, TODO L6; LOCKED A1) +and the too-large embed error mapping (task 02; LOCKED A3). + +The embed step of ``POST /api/chat`` embeds at most +``settings.embed_question_max_chars`` (default 1200 — the chunker's +``HARD_MAX_CHARS`` budget: worst-case ~1.4 chars/token, so it stays +under the endpoint's ~1024-token per-request input cap) of the +question; the FULL question still reaches the LLM prompt. A question +at or under the budget embeds byte-identically to the pre-phase path. + +Error mapping (task 02): a single text over the endpoint's input cap +is a DETERMINISTIC size failure (``EmbeddingInputTooLargeError``, the +real ``_post_embeddings`` → ``_TooLarge`` branch) — the chat endpoint +settles it with the accurate "question too long" terminal frame + the +reachability-fine hint, ONE attempt, NO retry frame (locked A3). An +embed failure without the too-large signature keeps the phase-67 +reachability path byte-identically (retry frames + the old copy). + +The endpoint-level tests drive ``POST /api/chat`` with the LLM (a +recording fake that captures every ``embed_one`` input and the +messages of each request, or a real ``LLMClient`` on a canned-failure +transport for the error-mapping tests), the retriever, and the DB +session all faked (the ``test_chat_gate.py`` wiring), so the whole +embed → retrieve → prompt contract runs without a stack. + +Frontend pins (task 02, source-assertion house style): the SSE +error frame's optional ``hint`` threads through the stream state +machine to ``showErrorBanner`` (shown in place of the default +reachability hint); the phase-111 Retry button rides the same +turn-error path. +""" +from __future__ import annotations + +import asyncio +import json +import uuid +from collections.abc import Iterator +from datetime import UTC, datetime +from pathlib import Path +from types import SimpleNamespace +from typing import TYPE_CHECKING, Any + +import pytest +from fastapi.testclient import TestClient +from pydantic import ValidationError + +from app.api import chat as chat_api +from app.config import Settings +from app.main import app as fastapi_app +from app.models import Document, KbOverview +from app.rag.chunker import HARD_MAX_CHARS +from app.rag.llm import ( + EmbeddingError, + EmbeddingInputTooLargeError, + LLMClient, + StreamPiece, +) +from app.rag.retriever import RetrievedChunk +from tests.conftest import ADMIN_PASSWORD + +if TYPE_CHECKING: + from app.rag.scaffolding import ScaffoldingFilter + +ANSWER = "Here is what your notes say about that." + + +def _question(n: int) -> str: + """A deterministic *n*-char question with a distinct head and tail. + + Repeated, index-marked sentences cut at exactly *n* chars — the + 4,000-char case is the composer's schema clamp + (``ChatRequest.message`` ``max_length=4000``), the L6 repro. + """ + sentence = "How is my homelab kubernetes cluster configured for long-running batch jobs? " + parts: list[str] = [] + total = 0 + i = 0 + while total < n: + part = f"[{i}] " + sentence + parts.append(part) + total += len(part) + i += 1 + return "".join(parts)[:n] + + +# ---------- the setting (default + validator) ---------- + + +def test_default_budget_matches_the_chunker_hard_cap() -> None: + """LOCKED A1: the default is the chunker's ``HARD_MAX_CHARS`` budget.""" + assert Settings(_env_file=None).embed_question_max_chars == 1200 # pyright: ignore[reportCallIssue] + assert Settings.model_fields["embed_question_max_chars"].default == HARD_MAX_CHARS + + +@pytest.mark.parametrize("bad", [0, -1, -1200]) +def test_budget_rejects_zero_and_negative(bad: int) -> None: + """``0``/negative would embed an empty/absent prefix — a typo that + must fail loudly at startup (the ``agent_max_rounds`` pattern).""" + with pytest.raises(ValidationError, match="embed_question_max_chars must be > 0"): + Settings(_env_file=None, embed_question_max_chars=bad) # pyright: ignore[reportCallIssue] + + +@pytest.mark.parametrize("good", [1, 500, 10_000]) +def test_budget_accepts_positive_values(good: int) -> None: + """A model with a smaller/larger cap is env-tunable, no code change.""" + settings = Settings(_env_file=None, embed_question_max_chars=good) # pyright: ignore[reportCallIssue] + assert settings.embed_question_max_chars == good + + +# ---------- endpoint-level (fake LLM + fake retriever + fake session) ---------- + + +class _RecordingLLM: + """Records every ``embed_one`` input and the messages of each request. + + Streams a canned answer and never emits tool calls, so a grounded + turn through the agent loop ends after the single (tools-offered) + request. Mirrors the ``test_chat_gate.py`` fake LLM. + """ + + def __init__(self, answer: str = ANSWER) -> None: + self.settings = Settings(_env_file=None) # pyright: ignore[reportCallIssue] + self.embedded: list[str] = [] + self.answer = answer + self.seen: list[list[dict[str, str]]] = [] + + async def embed_one(self, text: str) -> list[float]: + self.embedded.append(text) + return [0.0] * 768 + + async def chat_stream( + self, + messages: list[dict[str, str]], + tools: list[dict[str, Any]] | None = None, + scaffolding: ScaffoldingFilter | None = None, # phase 71 pass-through + ): + self.seen.append(messages) + for i in range(0, len(self.answer), 12): + yield StreamPiece("content", self.answer[i : i + 12]) + + +class _FakeSteeringResult: + """Empty steering-note result (no stored notes in these tests).""" + + def all(self) -> list[Any]: + return [] + + +class _FakeSession: + """Stands in for the DB session (the ``test_chat_gate.py`` fake).""" + + def __init__(self) -> None: + self.added: list[Any] = [] + self.commits = 0 + + def __enter__(self) -> _FakeSession: + return self + + def __exit__(self, *args: Any) -> None: + pass + + def add(self, obj: Any) -> None: + self.added.append(obj) + + def commit(self) -> None: + self.commits += 1 + + def scalars(self, _stmt: Any) -> _FakeSteeringResult: + return _FakeSteeringResult() + + def get(self, model: Any, pk: Any) -> Any: + if model is KbOverview: + return None + return None + + +def _doc(title: str, content: str) -> Document: + return Document( + id=uuid.uuid4(), + source="Homelab", + path=f"{title.lower().replace(' ', '-')}.md", + full_path="/tmp/doc.md", + title=title, + content=content, + content_hash="0" * 64, + created_at=datetime(2024, 6, 15, 12, 0, 0, tzinfo=UTC), + ) + + +def _chunk(doc: Document, score: float) -> RetrievedChunk: + return RetrievedChunk( + chunk_id=uuid.uuid4(), + position=0, + content=doc.content[:32], + score=score, + document=doc, + cosine=score, + fts_hit=False, + is_summary=False, + ) + + +def _fake_retriever(chunks: list[RetrievedChunk]) -> Any: + def retrieve(_db: Any, _question: str, _vec: list[float]) -> list[RetrievedChunk]: + return chunks + + return retrieve + + +@pytest.fixture(autouse=True) +def _admin_signed_in(client: TestClient) -> None: + """``POST /api/chat`` is user-gated — the endpoint-level tests run + as the signed-in ADMIN (the ``test_chat_gate.py`` pattern).""" + r = client.post("/api/login", json={"password": ADMIN_PASSWORD}) + assert r.status_code == 204, f"admin login failed: {r.status_code} {r.text}" + + +@pytest.fixture() +def embed_env(monkeypatch: pytest.MonkeyPatch) -> Iterator[tuple[_FakeSession, _RecordingLLM]]: + """``POST /api/chat`` with retriever, session, and LLM all faked. + + The code defaults apply (``embed_question_max_chars=1200``); the + gate threshold is pinned low so a 0.9-chunk goes grounded and a + 0.29-chunk deflects, regardless of any local ``.env``. + """ + monkeypatch.setattr(chat_api, "db_available", lambda: True) + session = _FakeSession() + llm = _RecordingLLM() + monkeypatch.setattr(chat_api, "SessionLocal", lambda: session) + monkeypatch.setitem(fastapi_app.dependency_overrides, chat_api.get_llm, lambda: llm) + monkeypatch.setattr( + chat_api, + "get_settings", + lambda: Settings(_env_file=None, relevance_threshold=0.30), # pyright: ignore[reportCallIssue] + ) + yield session, llm + fastapi_app.dependency_overrides.clear() + + +def _ask(client: TestClient, message: str) -> list[dict[str, Any]]: + with client.stream("POST", "/api/chat", json={"message": message}) as r: + assert r.status_code == 200 + frames: list[dict[str, Any]] = [] + buf = "" + for part in r.iter_text(): + buf += part + while "\n\n" in buf: + frame, buf = buf.split("\n\n", 1) + frame = frame.strip() + if frame.startswith("data:"): + frames.append(json.loads(frame.removeprefix("data:").strip())) + assert buf.strip() == "" + return frames + + +def test_long_question_embeds_exactly_the_bounded_prefix( + client: TestClient, + embed_env: tuple[_FakeSession, _RecordingLLM], + monkeypatch: pytest.MonkeyPatch, +) -> None: + """L6 repro: the 4,000-char (composer-clamp) question embeds ONLY the + 1200-char prefix — one embed call, exactly the head, and the turn + completes (no error frame).""" + _session, llm = embed_env + question = _question(4_000) + monkeypatch.setattr(chat_api, "retrieve", _fake_retriever([_chunk(_doc("T", "C"), 0.29)])) + + frames = _ask(client, question) + + # The code default (1200), derived from the field so this never drifts. + budget = Settings.model_fields["embed_question_max_chars"].default + assert llm.embedded == [question[:budget]] + assert llm.embedded[0] != question # it really was cut + assert all(f["type"] != "error" for f in frames) + assert frames[-1]["type"] == "done" + + +def test_long_question_full_text_reaches_llm_prompt( + client: TestClient, + embed_env: tuple[_FakeSession, _RecordingLLM], + monkeypatch: pytest.MonkeyPatch, +) -> None: + """Truncation is the embed step ONLY: the deflected turn's request + carries the FULL 4,000-char question as the user message.""" + _session, llm = embed_env + question = _question(4_000) + monkeypatch.setattr(chat_api, "retrieve", _fake_retriever([_chunk(_doc("T", "C"), 0.29)])) + + _frames = _ask(client, question) + + assert len(llm.seen) == 1 + assert llm.seen[0][-1] == {"role": "user", "content": question} + assert len(llm.seen[0][-1]["content"]) == 4_000 + + +def test_grounded_turn_agent_request_carries_full_question( + client: TestClient, + embed_env: tuple[_FakeSession, _RecordingLLM], + monkeypatch: pytest.MonkeyPatch, +) -> None: + """Same contract on the grounded (agent-loop) branch: the single + tools-offered request carries the FULL question.""" + _session, llm = embed_env + question = _question(4_000) + monkeypatch.setattr(chat_api, "retrieve", _fake_retriever([_chunk(_doc("T", "C"), 0.90)])) + + frames = _ask(client, question) + + assert frames[-1]["deflected"] is False + assert len(llm.seen) == 1 + assert llm.seen[0][-1] == {"role": "user", "content": question} + assert llm.embedded == [question[:1200]] # the prefix, not the full text + + +@pytest.mark.parametrize("n", [100, 900]) +def test_short_question_embeds_byte_identically( + client: TestClient, + embed_env: tuple[_FakeSession, _RecordingLLM], + monkeypatch: pytest.MonkeyPatch, + n: int, +) -> None: + """A question under the budget embeds the WHOLE question — the + pre-phase call, byte for byte (one call, the exact string).""" + _session, llm = embed_env + question = _question(n) + monkeypatch.setattr(chat_api, "retrieve", _fake_retriever([_chunk(_doc("T", "C"), 0.29)])) + + _frames = _ask(client, question) + + assert llm.embedded == [question] + + +def test_question_at_exactly_the_budget_embeds_whole( + client: TestClient, + embed_env: tuple[_FakeSession, _RecordingLLM], + monkeypatch: pytest.MonkeyPatch, +) -> None: + """The budget is an INCLUSIVE cap (``[:budget]``): a question exactly + 1200 chars long embeds in full — no char lost at the boundary.""" + _session, llm = embed_env + question = _question(1_200) + monkeypatch.setattr(chat_api, "retrieve", _fake_retriever([_chunk(_doc("T", "C"), 0.29)])) + + _frames = _ask(client, question) + + assert llm.embedded == [question] + + +def test_budget_is_env_tunable_via_settings( + client: TestClient, + embed_env: tuple[_FakeSession, _RecordingLLM], + monkeypatch: pytest.MonkeyPatch, +) -> None: + """LOCKED A1: the budget is a setting — a deployment with a smaller-cap + model lowers it via ``BOR_EMBED_QUESTION_MAX_CHARS`` (here: 500) and + the prefix follows, no code change.""" + _session, llm = embed_env + monkeypatch.setattr( + chat_api, + "get_settings", + lambda: Settings( + _env_file=None, # pyright: ignore[reportCallIssue] + relevance_threshold=0.30, + embed_question_max_chars=500, + ), + ) + question = _question(4_000) + monkeypatch.setattr(chat_api, "retrieve", _fake_retriever([_chunk(_doc("T", "C"), 0.29)])) + + _frames = _ask(client, question) + + assert llm.embedded == [question[:500]] + assert llm.seen[0][-1] == {"role": "user", "content": question} + + +# ---------- error mapping (task 02, LOCKED A3) ---------- + +#: The aipi/litellm signature of a too-large input (the L6 repro body, +#: truncated the way the client sees it — the client keys off "too +#: large" in the body). +_TOO_LARGE_BODY = ( + "input (903 tokens) is too large to process. increase the physical " + "batch size (current batch size: 512)" +) + + +class _CannedHttpResponse: + """One canned transport reply (status + text body, no JSON).""" + + def __init__(self, status_code: int, text: str) -> None: + self.status_code = status_code + self.text = text + + +class _CannedHttp: + """Stands in for the httpx transport the openai client owns. + + Every POST returns the same canned failure and records the request + body — the embed attempt counter. + """ + + def __init__(self, status_code: int, text: str) -> None: + self.status_code = status_code + self.text = text + self.posts: list[dict[str, Any]] = [] + + async def post( + self, url: str, *, json: dict[str, Any], headers: dict[str, str] | None = None + ) -> _CannedHttpResponse: + self.posts.append(json) + return _CannedHttpResponse(self.status_code, self.text) + + +def _canned_embed_llm( + settings: Settings, status_code: int, text: str +) -> tuple[LLMClient, _CannedHttp]: + """A REAL ``LLMClient`` whose transport is a canned failure. + + ``embed_one`` runs the real ``_post_embeddings`` → ``_embed_batch`` + path — the ``_TooLarge`` branch fires for real. The chat stream is + never reached: the embed step settles the turn first. + """ + llm = LLMClient(settings) + http = _CannedHttp(status_code, text) + llm._client = SimpleNamespace(_client=http) # pyright: ignore[reportAttributeAccessIssue] + return llm, http + + +def test_single_oversized_text_raises_the_too_large_subclass() -> None: + """The single-text ``_TooLarge`` branch of ``LLMClient._embed_batch`` + raises ``EmbeddingInputTooLargeError`` — still an + ``EmbeddingError`` (the importer path is a drop-in) with the + byte-identical import-oriented message.""" + settings = Settings(_env_file=None) # pyright: ignore[reportCallIssue] + llm, http = _canned_embed_llm(settings, 500, _TOO_LARGE_BODY) + with pytest.raises(EmbeddingInputTooLargeError, match="token cap") as exc: + asyncio.run(llm.embed_one("x" * 3000)) + assert isinstance(exc.value, EmbeddingError) + assert str(exc.value) == ( + "a single 3000-char chunk exceeded the endpoint's per-request input " + "token cap — lower BOR_CHUNK_TARGET_CHARS and re-import" + ) + assert len(http.posts) == 1 + assert llm.embed_batches == 0 + + +def test_too_large_embed_maps_to_terminal_too_long_frame( + client: TestClient, + embed_env: tuple[_FakeSession, _RecordingLLM], + monkeypatch: pytest.MonkeyPatch, +) -> None: + """L6 acceptance: the litellm "too large to process" 500 maps to the + ACCURATE terminal error — exactly ONE error frame with the precise + detail + the reachability-fine hint, NO retry frame, ONE embed + attempt (locked A3: deterministic — never retried), no "couldn't + reach" copy.""" + _session, _recording = embed_env + settings = Settings( + _env_file=None, # pyright: ignore[reportCallIssue] + relevance_threshold=0.30, + llm_retry_delay=0.01, # keep the (unused here) budget cheap + ) + monkeypatch.setattr(chat_api, "get_settings", lambda: settings) + llm, http = _canned_embed_llm(settings, 500, _TOO_LARGE_BODY) + monkeypatch.setitem(fastapi_app.dependency_overrides, chat_api.get_llm, lambda: llm) + + frames = _ask(client, "How is my homelab kubernetes cluster configured?") + + assert len(http.posts) == 1 # ONE attempt — no restart (locked A3) + assert [f["type"] for f in frames] == ["error"] # terminal: no retry, no done + frame = frames[0] + assert frame["detail"] == "Question too long — trim it and re-ask." + assert frame["hint"] == ( + "The app reached the embedding model fine — only the question length is the problem." + ) + + +def test_transport_embed_failure_keeps_the_legacy_retry_path( + client: TestClient, + embed_env: tuple[_FakeSession, _RecordingLLM], + monkeypatch: pytest.MonkeyPatch, +) -> None: + """Regression pin: an embed 500 WITHOUT the too-large signature + keeps the phase-67 reachability behavior byte-identical — one + retry frame per restart (attempts 2–4 of 4), then the OLD + "couldn't reach" copy (``hint`` null) after the full attempt + budget.""" + _session, _recording = embed_env + settings = Settings( + _env_file=None, # pyright: ignore[reportCallIssue] + relevance_threshold=0.30, + llm_retry_delay=0.01, # 3 restarts × 0.01 s — the shape is what is pinned + ) + monkeypatch.setattr(chat_api, "get_settings", lambda: settings) + llm, http = _canned_embed_llm(settings, 500, "internal server error") + monkeypatch.setitem(fastapi_app.dependency_overrides, chat_api.get_llm, lambda: llm) + + frames = _ask(client, "How is my homelab kubernetes cluster configured?") + + assert len(http.posts) == 4 # the full budget was spent (it retried) + assert [f["type"] for f in frames] == ["retry", "retry", "retry", "error"] + for i, frame in enumerate(frames[:3], start=2): + assert frame == {"type": "retry", "attempt": i, "max_attempts": 4} + error = frames[-1] + assert error["detail"] == ( + "I couldn't reach the embedding model — please try again." + ) + assert error["hint"] is None # the additive field is null, never a too-long hint + + +# ---------- frontend hint support (task 02, source-assertion house style) ---------- + +APP_JS = Path(__file__).resolve().parents[2] / "frontend" / "assets" / "app.js" + + +def _js() -> str: + return APP_JS.read_text(encoding="utf-8") + + +def test_sse_error_branch_threads_the_frame_hint() -> None: + """The stream state machine's error branch carries the frame's + optional ``hint`` through the throw (``err.hint``) — the banner + shows it in place of the default reachability hint.""" + js = _js() + idx = js.find('ev.type === "error"') + assert idx != -1, "the readSSE handler must branch on error frames" + end = js.find("throw err;", idx) + assert end != -1, "the error branch throws the detail as the turn error" + branch = js[idx:end] + assert "err.hint = ev.hint;" in branch, ( + "the frame's optional hint must ride the thrown error" + ) + + +def test_catch_passes_the_thrown_hint_to_set_ui_state() -> None: + """The turn catch passes the thrown error's hint to ``setUiState`` + (the third argument) — the banner call gets it via the opts merge. + Non-Error throws pass no opts (the default hint applies).""" + js = _js() + idx = js.find("setUiState(UI_STATE.error, detail,") + assert idx != -1, "the turn-error landing must flow through setUiState" + call = js[idx : js.find(");", idx)] + assert "{ hint: err.hint }" in call, "the thrown hint must reach setUiState" + + +def test_set_ui_state_merges_opts_into_the_banner_call() -> None: + """``setUiState``'s error transition keeps the phase-111 + ``{ retryable: true }`` (the banner Retry button) AND merges the + turn opts (the hint) into the ``showErrorBanner`` call.""" + js = _js() + idx = js.find("export function setUiState") + assert idx != -1 + body = js[idx : js.find("\n}\n", idx)] + assert "opts = {}" in body, "the opts parameter carries the turn hint" + assert "showErrorBanner(errorDetail, { retryable: true, ...opts });" in body + + +def test_show_error_banner_honors_opts_hint() -> None: + """``showErrorBanner`` shows ``opts.hint`` in place of the default + ``ERROR_HINT`` — in BOTH the with-detail and the detail-less forms + (``??`` falls back on null/undefined, so hint-less frames keep the + old copy byte-identically).""" + js = _js() + idx = js.find("function showErrorBanner") + assert idx != -1 + body = js[idx : js.find("\n}\n", idx)] + assert body.count("opts.hint ?? ERROR_HINT") == 2, ( + "the hint fallback must cover the detail and no-detail forms" + ) diff --git a/tests/unit/test_sse_events.py b/tests/unit/test_sse_events.py index 63ec1f8..c339c33 100644 --- a/tests/unit/test_sse_events.py +++ b/tests/unit/test_sse_events.py @@ -56,16 +56,35 @@ def test_multi_line_text_stays_one_frame() -> None: def test_error_event_model_serializes_exact_frame() -> None: """The ``ChatErrorEvent`` model is the wire shape of every server-side - failure the UI's state machine (phase 06) must recover from.""" + failure the UI's state machine (phase 06) must recover from. + + Phase 114: the additive ``hint`` field serializes ``null`` when + absent (old clients ignore the field — PLAN §4; the + ``ChatDoneEvent.related`` pattern).""" frame = sse_event(ChatErrorEvent(detail="boom").model_dump()) - assert frame == 'data: {"type": "error", "detail": "boom"}\n\n' - assert _payload(frame) == {"type": "error", "detail": "boom"} + assert frame == 'data: {"type": "error", "detail": "boom", "hint": null}\n\n' + assert _payload(frame) == {"type": "error", "detail": "boom", "hint": None} -def test_error_event_shape_is_type_and_detail_only() -> None: +def test_error_event_shape_is_type_detail_and_optional_hint() -> None: dumped = ChatErrorEvent(detail="The chat model dropped the connection").model_dump() - assert set(dumped.keys()) == {"type", "detail"} + assert set(dumped.keys()) == {"type", "detail", "hint"} assert dumped["type"] == "error" # default — call sites never spell it out + assert dumped["hint"] is None # absent hint serializes null, not dropped + + +def test_error_event_hint_serializes_verbatim_when_set() -> None: + """Phase 114 (TODO L6): the too-long frame carries the reachability- + fine hint — it survives the roundtrip byte-identically (em-dash and + all) for the banner to show in place of the default hint.""" + hint = "The app reached the embedding model fine — only the question length is the problem." + dumped = ChatErrorEvent( + detail="Question too long — trim it and re-ask.", hint=hint + ).model_dump() + assert dumped["hint"] == hint + payload = _payload(sse_event(dumped)) + assert payload["hint"] == hint + assert payload["detail"] == "Question too long — trim it and re-ask." def test_thinking_frame_serializes_exactly() -> None: