All verification passes complete — the phase was already fully implemented in the working tree, and every gate is green. No defects found; no code changes were needed.
**Final verification pass — Phase 114 (embed question length):**
- Verified truncation: `chat.py:459` embeds `request.message[:settings.embed_question_max_chars]` (default 1200, `BOR_EMBED_QUESTION_MAX_CHARS`, `>0` validator); full question still reaches the LLM prompt/log.
- Verified error mapping: `EmbeddingInputTooLargeError(EmbeddingError)` (byte-identical message) caught **before** `EmbeddingError` → no retry, terminal frame `detail="Question too long — trim it and re-ask."` + reachability-fine hint; `ChatErrorEvent.hint` additive.
- Verified frontend chain: frame `hint` → `err.hint` → `setUiState(error, …, {hint})` → `showErrorBanner(…, {retryable: true})` — hint replaces default `ERROR_HINT`, phase-111 `#banner-retry` button revealed. 4,000-char clamp untouched (A2).
- `uv run pytest tests/unit/test_embed_question_length.py -v --no-cov` → 21 passed
- `uv run pytest tests/e2e/test_embed_question_length.py -v --no-cov` (isolation, DB up) → 1 passed (4,000-char question → done, no banner)
- Regression: `test_llm_retry.py` 4 passed · `test_oneshot_llm_retry.py` 2 passed · `test_chip_sizing_question_cap.py` 6 passed
- `uv run pytest --cov=app --cov-report=term-missing` → 2444 passed, TOTAL **99%** (>90% gate)
- `uv run ruff check .` → All checks passed; `uv run pyright` → 0 errors, 0 warnings
**Completion criteria:** (1) 4,000-char question embeds prefix + full prompt ✅ · (2) too-large → accurate frame + hint + Retry button ✅ · (3) reachability failure byte-identical (retries + old copy) ✅ · (4) all gates green ✅ · (5) commit/phase-move → left to the harness per instructions (no `git add`/`commit` run).
**Deviations:** none. **Next pending phase:** `115_doc_draft_discard`.
6.3 KiB
Phase 114 — Embed question length: truncation + accurate error (TODO L6)
Source: TODO.md L149–181 — "L6 — 4,000-char question clamp exceeds the embed model's input cap → misleading 'couldn't reach the embedding model' error (2026-09-15, brain-of-reese interactive test)"
Story: n/a (interactive-test follow-up fix; extends the phase-67 LLM-retry and phase-06 loading-feedback assets).
Context: The composer clamps at 4,000 chars; app/api/chat.py:423 embeds the full question via llm.embed_one; aipi's litellm rejects the ~903-token input with HTTP 500: "input (903 tokens) is too large to process. increase the physical batch size (current batch size: 512)". The single-text path in app/rag/llm.py (_embed_batch → _TooLarge, ~L279–284) turns that into EmbeddingError("a single …-char chunk exceeded the endpoint's per-request input token cap — lower BOR_CHUNK_TARGET_CHARS and re-import") — an import-oriented message — and the chat endpoint's catch-all (~L428–444) maps EVERY EmbeddingError to "I couldn't reach the embedding model — please try again." Both diagnoses are wrong (reachability is fine; the chunker constant is irrelevant to a question). The chunker's own HARD_MAX_CHARS = 1200 (app/rag/chunker.py:51, ~1024 tokens at ~1.4 chars/token) shows the question path never got the same treatment.
Objective
Every legal question (≤ the UI clamp) is embeddable: the embed step gets a bounded prefix of the question (the chunker's 1200-char budget) while the full question still reaches the LLM prompt; and if the input is still too large (a smaller-cap model, a misconfiguration), the turn fails with an accurate "question too long" error — no false reachability diagnosis, no wasted retries — and the banner carries the phase-111 Retry button.
Dependencies
111_chat_banner_retry(todo) — L6's acceptance: "the L1 'Try again' button fix should also apply to this banner" — the too-long error flows through the same turn-error state machine, so the phase-111 Retry button is offered on it.
Design (shared by all tasks — the executor reads this, not the chat)
- Truncation (task 01): new setting
embed_question_max_chars: int = 1200(envBOR_EMBED_QUESTION_MAX_CHARS, default = the chunker'sHARD_MAX_CHARSbudget, validated> 0). The chat embed step (chat.py:423) embedsrequest.message[:settings.embed_question_max_chars]; the LLM prompt build is unchanged (the full question still reaches the model). Questions shorter than the budget are byte-identical to today. - Error mapping (task 02):
app/rag/llm.py— newEmbeddingInputTooLargeError(EmbeddingError)subclass; the single-text_TooLargebranch of_embed_batchraises it (same message text — the importer path is byte-identical, it still catchesEmbeddingError). The chat endpoint catchesEmbeddingInputTooLargeErrorbeforeEmbeddingErrorinside the phase-67 retry loop → no retry (a deterministic failure — locked A3) → terminalChatErrorEventwithdetail="Question too long — trim it and re-ask."and a new optionalhintfield:hint="The app reached the embedding model fine — only the question length is the problem."ChatErrorEventgainshint: str | None = None(additive; PLAN §4 old-client ignore contract). The frontend's phase-111 reworkedshowErrorBanner(detail, opts)showsopts.hintwhen the frame carries one, else the defaultERROR_HINT. - Retry: the too-long frame flows through the turn-error state machine → the phase-111 banner Retry button is offered (re-asking is the user's call after trimming; the composer clamp still applies).
- NOT touched: the importer's embed path and its batch-halving
_TooLargebehavior/error copy, the 4,000-char composer clamp (locked A2 — truncation, not a lower clamp), the reachability-failure retry semantics (phase 67 — byte-identical).
Tasks
01_embed_truncation.md— the bounded-prefix embed + the setting.02_too_long_error_mapping.md—EmbeddingInputTooLargeError, the chat-path mapping,ChatErrorEvent.hint, the frontend hint support.03_embed_length_tests.md— the unit pins + the 4,000-char E2E.
Testing & Quality
- Unit:
tests/unit/test_embed_question_length.py(new, task 03) — truncation (long → prefix embedded, LLM prompt carries the full text; short → byte-identical), error mapping (too-large failure → exact detail + hint, no retry frame, one attempt; transport failure → legacy reachability path with retries — the regression pin), the config validator. - E2E:
tests/e2e/test_embed_question_length.py(new, task 03; run in isolation:uv run pytest tests/e2e/test_embed_question_length.py -v --no-cov) — a 4,000-char question (the composer clamp) streams to done (mock LLM), no error banner. - Regression:
tests/e2e/test_llm_retry.py,test_oneshot_llm_retry.py,test_chip_sizing_question_cap.py(the 4,000-char counter) stay green. - Coverage: >90% on
app/(validate.sh gate).
Completion Criteria
- A 4,000-char question embeds (bounded prefix) and the turn succeeds; the LLM prompt carries the full question.
- A too-large embed failure (forced in a unit test) → the accurate "Question too long" frame + the reachability-fine hint; the banner offers the phase-111 Retry button.
- A reachability embed failure behaves byte-identically to pre-phase (retries + old copy).
uv run pytestgreen; coverage >90%; the e2e file green in isolation;uv run ruff check . && uv run pyrightclean.- One
--no-gpg-signcommit; phase dir moved tocomplete/by the pipeline gate.
Locked decisions
- A1 — both fixes combined: 1200-char embed truncation (default = the chunker budget, env-tunable) + the precise too-long error mapping (owner-confirmed 2026-09-14, roadmap confirmation).
- A2 — the 4,000-char composer clamp stays (owner-confirmed 2026-09-14) — truncation, not a lower clamp.
- A3 — a too-large embed failure is NOT retried (deterministic failure) — it short-circuits the phase-67 retry loop (owner-confirmed 2026-09-14).
Commit
git add app/ tests/ frontend/ .agents/phases/ && git commit --no-gpg-sign -m "fix(rag): embed a bounded question prefix (1200-char budget) and map the embed too-large failure to an accurate too-long error with a reachability-fine hint"