phase: 114_embed_question_length
All verification passes complete — the phase was already fully implemented in the working tree, and every gate is green. No defects found; no code changes were needed.
**Final verification pass — Phase 114 (embed question length):**
- Verified truncation: `chat.py:459` embeds `request.message[:settings.embed_question_max_chars]` (default 1200, `BOR_EMBED_QUESTION_MAX_CHARS`, `>0` validator); full question still reaches the LLM prompt/log.
- Verified error mapping: `EmbeddingInputTooLargeError(EmbeddingError)` (byte-identical message) caught **before** `EmbeddingError` → no retry, terminal frame `detail="Question too long — trim it and re-ask."` + reachability-fine hint; `ChatErrorEvent.hint` additive.
- Verified frontend chain: frame `hint` → `err.hint` → `setUiState(error, …, {hint})` → `showErrorBanner(…, {retryable: true})` — hint replaces default `ERROR_HINT`, phase-111 `#banner-retry` button revealed. 4,000-char clamp untouched (A2).
- `uv run pytest tests/unit/test_embed_question_length.py -v --no-cov` → 21 passed
- `uv run pytest tests/e2e/test_embed_question_length.py -v --no-cov` (isolation, DB up) → 1 passed (4,000-char question → done, no banner)
- Regression: `test_llm_retry.py` 4 passed · `test_oneshot_llm_retry.py` 2 passed · `test_chip_sizing_question_cap.py` 6 passed
- `uv run pytest --cov=app --cov-report=term-missing` → 2444 passed, TOTAL **99%** (>90% gate)
- `uv run ruff check .` → All checks passed; `uv run pyright` → 0 errors, 0 warnings
**Completion criteria:** (1) 4,000-char question embeds prefix + full prompt ✅ · (2) too-large → accurate frame + hint + Retry button ✅ · (3) reachability failure byte-identical (retries + old copy) ✅ · (4) all gates green ✅ · (5) commit/phase-move → left to the harness per instructions (no `git add`/`commit` run).
**Deviations:** none. **Next pending phase:** `115_doc_draft_discard`.
This commit is contained in:
+38
-1
@@ -184,6 +184,7 @@ from app.rag.agent import (
|
||||
)
|
||||
from app.rag.llm import (
|
||||
EmbeddingError,
|
||||
EmbeddingInputTooLargeError, # phase 114: the deterministic too-large failure
|
||||
LLMClient,
|
||||
LLMError,
|
||||
RetryPiece, # phase 67: one LLM request restart (an SSE retry frame)
|
||||
@@ -450,8 +451,44 @@ async def chat(
|
||||
attempt = 1
|
||||
while True:
|
||||
try:
|
||||
question_vec = await llm.embed_one(request.message)
|
||||
# Phase 114 (TODO L6): the embed input is bounded to the
|
||||
# model's per-request input cap (the chunker's 1200-char
|
||||
# budget, env-tunable) — the FULL question still reaches
|
||||
# the LLM prompt (prompt build + log line untouched).
|
||||
question_vec = await llm.embed_one(
|
||||
request.message[: settings.embed_question_max_chars]
|
||||
)
|
||||
break
|
||||
except EmbeddingInputTooLargeError as e:
|
||||
# Phase 114 (TODO L6, locked A3): a too-large input is
|
||||
# DETERMINISTIC — retrying the same size is guaranteed
|
||||
# to repeat — so this short-circuits the phase-67 retry
|
||||
# loop: no ``retry`` frame, no restart, one terminal
|
||||
# error frame with the accurate "question too long"
|
||||
# detail + the reachability-fine hint (the
|
||||
# "couldn't reach" copy and the retry budget stay for
|
||||
# reachability failures only — the branch below).
|
||||
embed_ms = int((time.monotonic() - t0) * 1000)
|
||||
total_ms = int((time.monotonic() - started) * 1000)
|
||||
logger.error(
|
||||
"chat: question=%r embed_ms=%d total_ms=%d — "
|
||||
"embedding failed (too-large): %s",
|
||||
request.message,
|
||||
embed_ms,
|
||||
total_ms,
|
||||
e,
|
||||
)
|
||||
settled = True # terminal: the error frame settles the turn
|
||||
yield sse_event(
|
||||
ChatErrorEvent(
|
||||
detail="Question too long — trim it and re-ask.",
|
||||
hint=(
|
||||
"The app reached the embedding model fine — "
|
||||
"only the question length is the problem."
|
||||
),
|
||||
).model_dump()
|
||||
)
|
||||
return
|
||||
except EmbeddingError as e:
|
||||
embed_ms = int((time.monotonic() - t0) * 1000)
|
||||
if attempt >= max_attempts:
|
||||
|
||||
@@ -156,6 +156,18 @@ class Settings(BaseSettings):
|
||||
chunk_target_chars: int = 2_000
|
||||
chunk_overlap_chars: int = 200
|
||||
embed_batch_size: int = 16
|
||||
#: Char budget for the chat turn's question embed (phase 114, TODO L6;
|
||||
#: LOCKED A1): the embed step embeds at most this many chars of the
|
||||
#: question — the default 1200 is the chunker's ``HARD_MAX_CHARS``
|
||||
#: budget (``app/rag/chunker.py``: worst-case ~1.4 chars/token, so it
|
||||
#: stays under the endpoint's ~1024-token per-request input cap). Only
|
||||
#: the embedding is bounded: the FULL question still reaches the LLM
|
||||
#: prompt (prompt build untouched), and a question at or under the
|
||||
#: budget embeds byte-identically to the pre-phase path. A model with a
|
||||
#: larger/smaller cap is accommodated by env, no code change (A1).
|
||||
#: ``0``/negative is a typo — the validator fails loudly at startup
|
||||
#: (the ``agent_max_rounds`` pattern).
|
||||
embed_question_max_chars: int = 1200
|
||||
#: Total char budget for the ``<tuning>`` section of the system prompt
|
||||
#: (phase 15, steering notes). The newest-fitting notes are kept and the
|
||||
#: overflow is replaced by the ``[…truncated…]`` marker.
|
||||
@@ -415,6 +427,15 @@ class Settings(BaseSettings):
|
||||
raise ValueError("read_max_chars must be >= 0 (chars)")
|
||||
return v
|
||||
|
||||
@field_validator("embed_question_max_chars")
|
||||
@classmethod
|
||||
def _embed_question_max_chars_positive(cls, v: int) -> int:
|
||||
"""``0``/negative would embed an empty/absent prefix — fail loud at
|
||||
startup (the ``agent_max_rounds`` pattern, phase 114)."""
|
||||
if v <= 0:
|
||||
raise ValueError("embed_question_max_chars must be > 0 (chars)")
|
||||
return v
|
||||
|
||||
@field_validator("llm_retries")
|
||||
@classmethod
|
||||
def _llm_retries_non_negative(cls, v: int) -> int:
|
||||
|
||||
+22
-1
@@ -43,6 +43,21 @@ class EmbeddingError(RuntimeError):
|
||||
"""The embeddings endpoint failed (network, HTTP, or malformed reply)."""
|
||||
|
||||
|
||||
class EmbeddingInputTooLargeError(EmbeddingError):
|
||||
"""A single text exceeded the endpoint's per-request input token cap
|
||||
(phase 114, TODO L6).
|
||||
|
||||
The endpoint REACHED and answered — this is a deterministic input-
|
||||
SIZE failure, not a reachability problem: retrying the same input is
|
||||
guaranteed to repeat (locked A3), so the chat endpoint catches this
|
||||
BEFORE :class:`EmbeddingError` and settles with the accurate
|
||||
"question too long" terminal error (no phase-67 retry frames). The
|
||||
importer keeps catching the parent :class:`EmbeddingError` — this
|
||||
subclass is a drop-in there and its message text is byte-identical
|
||||
to the pre-phase import-oriented copy.
|
||||
"""
|
||||
|
||||
|
||||
class EmbeddingDimensionError(EmbeddingError):
|
||||
"""Embedding dimension != BOR_EMBEDDING_DIM — import must fail loudly."""
|
||||
|
||||
@@ -278,7 +293,13 @@ class LLMClient:
|
||||
return await self._post_embeddings(chunk)
|
||||
except _TooLarge:
|
||||
if len(chunk) == 1:
|
||||
raise EmbeddingError(
|
||||
# Phase 114 (TODO L6): a single text over the cap is a
|
||||
# deterministic input-size failure (never reachability) —
|
||||
# the subclass lets the chat endpoint map it to the
|
||||
# accurate "question too long" error and skip the retry
|
||||
# loop (locked A3). The message text stays byte-identical:
|
||||
# the importer catches the parent EmbeddingError.
|
||||
raise EmbeddingInputTooLargeError(
|
||||
f"a single {len(chunk[0])}-char chunk exceeded the endpoint's "
|
||||
"per-request input token cap — lower BOR_CHUNK_TARGET_CHARS "
|
||||
"and re-import"
|
||||
|
||||
@@ -206,10 +206,18 @@ class ChatErrorEvent(BaseModel):
|
||||
The client's loading-feedback state machine (phase 06) keys off this
|
||||
exact shape — ``{type: "error", detail: str}`` — to flip to the error
|
||||
state and re-enable the send button.
|
||||
|
||||
``hint`` (phase 114, TODO L6): an optional one-line clarification the
|
||||
client shows IN PLACE of its default reachability hint when present
|
||||
(the "question too long" frame: the embedding model WAS reached —
|
||||
only the question's length is the problem). Additive: frames without
|
||||
it serialize ``"hint": null``, and old clients ignore the field
|
||||
(PLAN §4 house contract — the ``ChatDoneEvent.related`` pattern).
|
||||
"""
|
||||
|
||||
type: str = "error"
|
||||
detail: str
|
||||
hint: str | None = None
|
||||
|
||||
|
||||
class ChatRetryEvent(BaseModel):
|
||||
|
||||
Reference in New Issue
Block a user