Files
brain-of-reese/.agents/phases/complete/114_embed_question_length/00_phase.md
T
ducoterra 3846f26a58
Build and Push Containers / build-and-push-app (push) Successful in 2m6s
Build and Push Containers / build-and-push-db (push) Successful in 13s
phase: 114_embed_question_length
All verification passes complete — the phase was already fully implemented in the working tree, and every gate is green. No defects found; no code changes were needed.

**Final verification pass — Phase 114 (embed question length):**
- Verified truncation: `chat.py:459` embeds `request.message[:settings.embed_question_max_chars]` (default 1200, `BOR_EMBED_QUESTION_MAX_CHARS`, `>0` validator); full question still reaches the LLM prompt/log.
- Verified error mapping: `EmbeddingInputTooLargeError(EmbeddingError)` (byte-identical message) caught **before** `EmbeddingError` → no retry, terminal frame `detail="Question too long — trim it and re-ask."` + reachability-fine hint; `ChatErrorEvent.hint` additive.
- Verified frontend chain: frame `hint` → `err.hint` → `setUiState(error, …, {hint})` → `showErrorBanner(…, {retryable: true})` — hint replaces default `ERROR_HINT`, phase-111 `#banner-retry` button revealed. 4,000-char clamp untouched (A2).
- `uv run pytest tests/unit/test_embed_question_length.py -v --no-cov` → 21 passed
- `uv run pytest tests/e2e/test_embed_question_length.py -v --no-cov` (isolation, DB up) → 1 passed (4,000-char question → done, no banner)
- Regression: `test_llm_retry.py` 4 passed · `test_oneshot_llm_retry.py` 2 passed · `test_chip_sizing_question_cap.py` 6 passed
- `uv run pytest --cov=app --cov-report=term-missing` → 2444 passed, TOTAL **99%** (>90% gate)
- `uv run ruff check .` → All checks passed; `uv run pyright` → 0 errors, 0 warnings

**Completion criteria:** (1) 4,000-char question embeds prefix + full prompt ✅ · (2) too-large → accurate frame + hint + Retry button ✅ · (3) reachability failure byte-identical (retries + old copy) ✅ · (4) all gates green ✅ · (5) commit/phase-move → left to the harness per instructions (no `git add`/`commit` run).
**Deviations:** none. **Next pending phase:** `115_doc_draft_discard`.
2026-09-15 04:16:55 -04:00

6.3 KiB
Raw Blame History

Phase 114 — Embed question length: truncation + accurate error (TODO L6)

Source: TODO.md L149–181 — "L6 — 4,000-char question clamp exceeds the embed model's input cap → misleading 'couldn't reach the embedding model' error (2026-09-15, brain-of-reese interactive test)" Story: n/a (interactive-test follow-up fix; extends the phase-67 LLM-retry and phase-06 loading-feedback assets). Context: The composer clamps at 4,000 chars; app/api/chat.py:423 embeds the full question via llm.embed_one; aipi's litellm rejects the ~903-token input with HTTP 500: "input (903 tokens) is too large to process. increase the physical batch size (current batch size: 512)". The single-text path in app/rag/llm.py (_embed_batch → _TooLarge, ~L279–284) turns that into EmbeddingError("a single …-char chunk exceeded the endpoint's per-request input token cap — lower BOR_CHUNK_TARGET_CHARS and re-import") — an import-oriented message — and the chat endpoint's catch-all (~L428–444) maps EVERY EmbeddingError to "I couldn't reach the embedding model — please try again." Both diagnoses are wrong (reachability is fine; the chunker constant is irrelevant to a question). The chunker's own HARD_MAX_CHARS = 1200 (app/rag/chunker.py:51, ~1024 tokens at ~1.4 chars/token) shows the question path never got the same treatment.

Objective

Every legal question (≤ the UI clamp) is embeddable: the embed step gets a bounded prefix of the question (the chunker's 1200-char budget) while the full question still reaches the LLM prompt; and if the input is still too large (a smaller-cap model, a misconfiguration), the turn fails with an accurate "question too long" error — no false reachability diagnosis, no wasted retries — and the banner carries the phase-111 Retry button.

Dependencies

  • 111_chat_banner_retry (todo) — L6's acceptance: "the L1 'Try again' button fix should also apply to this banner" — the too-long error flows through the same turn-error state machine, so the phase-111 Retry button is offered on it.

Design (shared by all tasks — the executor reads this, not the chat)

  • Truncation (task 01): new setting embed_question_max_chars: int = 1200 (env BOR_EMBED_QUESTION_MAX_CHARS, default = the chunker's HARD_MAX_CHARS budget, validated > 0). The chat embed step (chat.py:423) embeds request.message[:settings.embed_question_max_chars]; the LLM prompt build is unchanged (the full question still reaches the model). Questions shorter than the budget are byte-identical to today.
  • Error mapping (task 02): app/rag/llm.py — new EmbeddingInputTooLargeError(EmbeddingError) subclass; the single-text _TooLarge branch of _embed_batch raises it (same message text — the importer path is byte-identical, it still catches EmbeddingError). The chat endpoint catches EmbeddingInputTooLargeError before EmbeddingError inside the phase-67 retry loop → no retry (a deterministic failure — locked A3) → terminal ChatErrorEvent with detail="Question too long — trim it and re-ask." and a new optional hint field: hint="The app reached the embedding model fine — only the question length is the problem." ChatErrorEvent gains hint: str | None = None (additive; PLAN §4 old-client ignore contract). The frontend's phase-111 reworked showErrorBanner(detail, opts) shows opts.hint when the frame carries one, else the default ERROR_HINT.
  • Retry: the too-long frame flows through the turn-error state machine → the phase-111 banner Retry button is offered (re-asking is the user's call after trimming; the composer clamp still applies).
  • NOT touched: the importer's embed path and its batch-halving _TooLarge behavior/error copy, the 4,000-char composer clamp (locked A2 — truncation, not a lower clamp), the reachability-failure retry semantics (phase 67 — byte-identical).

Tasks

  1. 01_embed_truncation.md — the bounded-prefix embed + the setting.
  2. 02_too_long_error_mapping.md — EmbeddingInputTooLargeError, the chat-path mapping, ChatErrorEvent.hint, the frontend hint support.
  3. 03_embed_length_tests.md — the unit pins + the 4,000-char E2E.

Testing & Quality

  • Unit: tests/unit/test_embed_question_length.py (new, task 03) — truncation (long → prefix embedded, LLM prompt carries the full text; short → byte-identical), error mapping (too-large failure → exact detail + hint, no retry frame, one attempt; transport failure → legacy reachability path with retries — the regression pin), the config validator.
  • E2E: tests/e2e/test_embed_question_length.py (new, task 03; run in isolation: uv run pytest tests/e2e/test_embed_question_length.py -v --no-cov) — a 4,000-char question (the composer clamp) streams to done (mock LLM), no error banner.
  • Regression: tests/e2e/test_llm_retry.py, test_oneshot_llm_retry.py, test_chip_sizing_question_cap.py (the 4,000-char counter) stay green.
  • Coverage: >90% on app/ (validate.sh gate).

Completion Criteria

  • A 4,000-char question embeds (bounded prefix) and the turn succeeds; the LLM prompt carries the full question.
  • A too-large embed failure (forced in a unit test) → the accurate "Question too long" frame + the reachability-fine hint; the banner offers the phase-111 Retry button.
  • A reachability embed failure behaves byte-identically to pre-phase (retries + old copy).
  • uv run pytest green; coverage >90%; the e2e file green in isolation; uv run ruff check . && uv run pyright clean.
  • One --no-gpg-sign commit; phase dir moved to complete/ by the pipeline gate.

Locked decisions

  • A1 — both fixes combined: 1200-char embed truncation (default = the chunker budget, env-tunable) + the precise too-long error mapping (owner-confirmed 2026-09-14, roadmap confirmation).
  • A2 — the 4,000-char composer clamp stays (owner-confirmed 2026-09-14) — truncation, not a lower clamp.
  • A3 — a too-large embed failure is NOT retried (deterministic failure) — it short-circuits the phase-67 retry loop (owner-confirmed 2026-09-14).

Commit

git add app/ tests/ frontend/ .agents/phases/ && git commit --no-gpg-sign -m "fix(rag): embed a bounded question prefix (1200-char budget) and map the embed too-large failure to an accurate too-long error with a reachability-fine hint"