Files
brain-of-reese/.agents/phases/todo/114_embed_question_length/00_phase.md
T
ducoterra f37c517590
Build and Push Containers / build-and-push-app (push) Successful in 16s
Build and Push Containers / build-and-push-db (push) Successful in 13s
chore(agent): phase roadmap from TODO.md — 6 phases (111–116): banner retry, honesty gate, chip quality, embed length, draft discard, modal scrollbar
2026-09-14 22:07:32 -04:00

6.3 KiB
Raw Blame History

Phase 114 — Embed question length: truncation + accurate error (TODO L6)

Source: TODO.md L149–181 — "L6 — 4,000-char question clamp exceeds the embed model's input cap → misleading 'couldn't reach the embedding model' error (2026-09-15, brain-of-reese interactive test)" Story: n/a (interactive-test follow-up fix; extends the phase-67 LLM-retry and phase-06 loading-feedback assets). Context: The composer clamps at 4,000 chars; app/api/chat.py:423 embeds the full question via llm.embed_one; aipi's litellm rejects the ~903-token input with HTTP 500: "input (903 tokens) is too large to process. increase the physical batch size (current batch size: 512)". The single-text path in app/rag/llm.py (_embed_batch → _TooLarge, ~L279–284) turns that into EmbeddingError("a single …-char chunk exceeded the endpoint's per-request input token cap — lower BOR_CHUNK_TARGET_CHARS and re-import") — an import-oriented message — and the chat endpoint's catch-all (~L428–444) maps EVERY EmbeddingError to "I couldn't reach the embedding model — please try again." Both diagnoses are wrong (reachability is fine; the chunker constant is irrelevant to a question). The chunker's own HARD_MAX_CHARS = 1200 (app/rag/chunker.py:51, ~1024 tokens at ~1.4 chars/token) shows the question path never got the same treatment.

Objective

Every legal question (≤ the UI clamp) is embeddable: the embed step gets a bounded prefix of the question (the chunker's 1200-char budget) while the full question still reaches the LLM prompt; and if the input is still too large (a smaller-cap model, a misconfiguration), the turn fails with an accurate "question too long" error — no false reachability diagnosis, no wasted retries — and the banner carries the phase-111 Retry button.

Dependencies

  • 111_chat_banner_retry (todo) — L6's acceptance: "the L1 'Try again' button fix should also apply to this banner" — the too-long error flows through the same turn-error state machine, so the phase-111 Retry button is offered on it.

Design (shared by all tasks — the executor reads this, not the chat)

  • Truncation (task 01): new setting embed_question_max_chars: int = 1200 (env BOR_EMBED_QUESTION_MAX_CHARS, default = the chunker's HARD_MAX_CHARS budget, validated > 0). The chat embed step (chat.py:423) embeds request.message[:settings.embed_question_max_chars]; the LLM prompt build is unchanged (the full question still reaches the model). Questions shorter than the budget are byte-identical to today.
  • Error mapping (task 02): app/rag/llm.py — new EmbeddingInputTooLargeError(EmbeddingError) subclass; the single-text _TooLarge branch of _embed_batch raises it (same message text — the importer path is byte-identical, it still catches EmbeddingError). The chat endpoint catches EmbeddingInputTooLargeError before EmbeddingError inside the phase-67 retry loop → no retry (a deterministic failure — locked A3) → terminal ChatErrorEvent with detail="Question too long — trim it and re-ask." and a new optional hint field: hint="The app reached the embedding model fine — only the question length is the problem." ChatErrorEvent gains hint: str | None = None (additive; PLAN §4 old-client ignore contract). The frontend's phase-111 reworked showErrorBanner(detail, opts) shows opts.hint when the frame carries one, else the default ERROR_HINT.
  • Retry: the too-long frame flows through the turn-error state machine → the phase-111 banner Retry button is offered (re-asking is the user's call after trimming; the composer clamp still applies).
  • NOT touched: the importer's embed path and its batch-halving _TooLarge behavior/error copy, the 4,000-char composer clamp (locked A2 — truncation, not a lower clamp), the reachability-failure retry semantics (phase 67 — byte-identical).

Tasks

  1. 01_embed_truncation.md — the bounded-prefix embed + the setting.
  2. 02_too_long_error_mapping.md — EmbeddingInputTooLargeError, the chat-path mapping, ChatErrorEvent.hint, the frontend hint support.
  3. 03_embed_length_tests.md — the unit pins + the 4,000-char E2E.

Testing & Quality

  • Unit: tests/unit/test_embed_question_length.py (new, task 03) — truncation (long → prefix embedded, LLM prompt carries the full text; short → byte-identical), error mapping (too-large failure → exact detail + hint, no retry frame, one attempt; transport failure → legacy reachability path with retries — the regression pin), the config validator.
  • E2E: tests/e2e/test_embed_question_length.py (new, task 03; run in isolation: uv run pytest tests/e2e/test_embed_question_length.py -v --no-cov) — a 4,000-char question (the composer clamp) streams to done (mock LLM), no error banner.
  • Regression: tests/e2e/test_llm_retry.py, test_oneshot_llm_retry.py, test_chip_sizing_question_cap.py (the 4,000-char counter) stay green.
  • Coverage: >90% on app/ (validate.sh gate).

Completion Criteria

  • A 4,000-char question embeds (bounded prefix) and the turn succeeds; the LLM prompt carries the full question.
  • A too-large embed failure (forced in a unit test) → the accurate "Question too long" frame + the reachability-fine hint; the banner offers the phase-111 Retry button.
  • A reachability embed failure behaves byte-identically to pre-phase (retries + old copy).
  • uv run pytest green; coverage >90%; the e2e file green in isolation; uv run ruff check . && uv run pyright clean.
  • One --no-gpg-sign commit; phase dir moved to complete/ by the pipeline gate.

Locked decisions

  • A1 — both fixes combined: 1200-char embed truncation (default = the chunker budget, env-tunable) + the precise too-long error mapping (owner-confirmed 2026-09-14, roadmap confirmation).
  • A2 — the 4,000-char composer clamp stays (owner-confirmed 2026-09-14) — truncation, not a lower clamp.
  • A3 — a too-large embed failure is NOT retried (deterministic failure) — it short-circuits the phase-67 retry loop (owner-confirmed 2026-09-14).

Commit

git add app/ tests/ frontend/ .agents/phases/ && git commit --no-gpg-sign -m "fix(rag): embed a bounded question prefix (1200-char budget) and map the embed too-large failure to an accurate too-long error with a reachability-fine hint"