2026-09-06 fixture runs: contract 100 %, executed 100 %, wall ~113 s (2 runs). Derived battery: FAIL only on usage floor (5/10 tool-turns) — answers seeded questions from context, which is ideal grounded behavior. Wall time ~2.8× lite (113 s vs 40 s). Model is clean.
2026-09-06 fixture runs: contract 92–93 %, executed 64–75 %, wall ~40.5 s (2 runs). Derived battery: FAIL, 36 % executed (38.3 s). Same pattern — copy-invariant re-read habit blocks the ≥90 % executed bar under current ALREADY_IN_CONTEXT refusal semantics. Model is working correctly; the bottleneck is the app's dedupe refusal, not the model.
Root cause (owner repro, verified in a real browser 2026-09-06): the
five navbar views (Chat, RAG, Sources, Tuning, History) were separate
HTML documents, so a navbar click was a REAL cross-document navigation
— the chat page unloaded, the in-flight SSE fetch was aborted, and the
phase-48 teardown (app/api/chat.py `finally`, "chat: turn cancelled")
stopped the model. Observed: send question -> click RAG mid-stream ->
click Chat -> the answer never finished: no `query_log` row, and on
return a dangling question with no brain record (the pre-token pagehide
partial persist skips because `acc` is empty).
Phase-48 LOCKED-DECISION REFINEMENT (owner-confirmed 2026-09-06,
flagged per AGENTS.md rule 3, not silently deviated): "real navigation
cancels the fetch" now means LEAVING THE APP — tab close,
external/other-document navigation, the Stop button. In-app navbar
switches are client-side view switches and no longer cancel.
Fix — Option A (SPA shell), chosen over B (Service Worker owns the
stream) and C (server-side turn registry + resume):
- frontend/index.html is the shell: ONE `<main id="main">` holds the
five `<section class="view">` blocks; hidden views carry BOTH
`hidden` and `inert` (WCAG — no focus/keyboard traversal). The
shared header, the single `doc-modal-*` skeleton, and the
`#app-version` footer each exist exactly once; the per-view copies
from the four folded pages are dropped.
- New frontend/assets/router.js (vanilla module — no framework, no
bundler, No-CDN rule intact): lazy-imports a view module on FIRST
show only (mount-once, hide-forever — the chat view's in-flight SSE
reader persists across switches; that persistence IS the fix);
intercepts same-shell navbar links with preventDefault +
history.pushState (never a document load); handles popstate; single
writer of `.nav-link` active state (is-active + aria-current),
document.title, and the per-view meta description (values carried
over from the old pages' heads, brand-resolved at write time).
- Each folded page's JS becomes `export async function mount(root)` —
root-scoped queries; `initSharedHeader()` dropped (the header boots
once in the shell via the chat module; the admin flag comes from the
same cached `fetchIsAdmin()` promise — zero extra requests).
- app/main.py: a small list-driven route factory serves the shell for
/tuning.html, /sources.html, /git-sources.html, /history.html —
registered AFTER the API routers and BEFORE the static catch-all
(routes-first). The phase-33 caching middleware applies no-cache +
`?v=` rewriting unchanged; app/core/caching.py needed NO change
(the view paths did not change — pinned by the integration tests).
- The four old view .html files are DELETED (one source of truth);
deep links to the old URLs keep working (the router picks the view
from the pathname); `/?chat=<id>` is unaffected; the Containerfile
bundles router.js (inlining the lazy view modules) and drops the
folded page files.
- app/schemas.py: HistoryTurn.text cap 4000 -> 32000 — the shell
keeps long saved answers in the chat, and the old cap (stricter than
the 24_000-char total history budget) 422-rejected any second turn
in such a chat (found by the phase-42 E2E suite on the shell).
Boundaries: login.html, shared.html, doc-edit.html, document.html
REMAIN separate documents (flow pages, not navbar tabs); a mid-stream
navigation to doc-edit/document.html still cancels per phase 48
(follow-up candidate, out of scope). The SSE API is unchanged. Real
departures still cancel the turn — phase 48 intact (pinned by
tests/e2e/test_stop_generation.py, unchanged, and by the new suite's
real-departure control).
Tests:
- Phase-20 suite REWRITTEN to the new semantics
(tests/e2e/test_sources_midstream_bug.py): a navbar switch no longer
cancels — the stream survives the switch and the FULL answer
settles; the pagehide partial persist REMAINS for real departures
(the partial's exact shape — first streamed chunk prefix, no done
metadata — is still pinned there).
- NEW story suite tests/e2e/test_nav_switch_keeps_stream.py (mock
LLM): the owner repro (send -> RAG mid-stream -> Chat: window
sentinel survives = same document, FULL answer, exactly one brain
turn in bor.chat.v1, exactly one settled query_log row, auto-saved
row matches) + the same mid-stream switch against the other three
views + the real-departure-still-cancels control + the no-switch
baseline.
- tests/unit/test_frontend_router.py: source-level pins of the router
invariants (click interceptor targets ONLY same-shell view paths,
pushState-only switches, mount-once guard, hidden+inert pair,
single-writer active state/title); shell-route integration tests
(each folded path serves the shell with no-cache + `?v=` body; a
non-view path still 404s); the file-reading unit pins re-pointed at
the shell (the four view files are gone — the shell is the source
of truth).
Verification (this commit): full suite green — 1565 unit+integration
tests, app/ coverage 99% (>90% floor); ruff + pyright clean; the
phase's E2E suites green in isolation (house protocol, AGENTS.md rule
9). Owner repro verified in a real browser against the real LLM
(dev server :8010, headful Chromium): "tell me about everquest" ->
RAG mid-stream -> Chat — the answer completed with one brain bubble
and no error banner, `query_log` gained exactly one settled row
(deflected=True: the dev KB holds no EverQuest docs — the settle, not
the topic, is the proof), zero "chat: turn cancelled" lines for that
turn; the control (real navigation to /shared.html mid-stream) still
cancelled (no settled row, the cancel line logged, the partial
persisted on return). Screenshots: .agents/screenshots/76_manual_*.
Phase 76 (76_spa_nav_shell) complete — moved to
.agents/phases/complete/.
Phase 75 (TODO.md L4): "Save as doc" now drafts a document from the
ENTIRE chat session — every question and answer up to the click, in
order — instead of only the clicked bubble's answer; the existing
doc-edit screen's free-form body editing is how the user edits out
anything they don't want to keep from previous replies (no new UI
surface).
Task 01 (frontend):
- app.js buildSessionTranscript(): walks the bor.chat.v1 conversation
record in order — a numbered section per user turn ("## N.
<question, raw>" + blank line + the raw answer text; more answers
join under the same heading), sections blank-line separated, all
trailing whitespace collapsed to one final newline. Only the raw
persisted text travels (m.who + m.text — no thinking blocks, no
source chips, no tune metadata); a brain record before the first
user record is skipped; a heading-only section marks a user turn
whose answer never landed (A6, owner-confirmed 2026-09-08).
- saveAsDoc(btn): the draft body is buildSessionTranscript(); the
dead single-bubble markdown parameter is dropped (the button's
appendSaveAsDocButton signature is unchanged — one button per
bubble). Title/path/double-click guard/hand-off are unchanged
(defaultDocTitle: the last question, whitespace-collapsed,
<=120 chars; docs/<slug>.md).
- Unit: the app.js source pins move to the transcript shape (whole
session, no thinking, no dead parameter).
Task 02 (E2E):
- tests/e2e/test_save_doc_session.py (bare-repo fixture, the
phase-59 convention — git as source of truth): three DISTINCT
on-topic turns in one session (turn 1 carries the phase-17
"think out loud" trigger so its record has a thinking block the
transcript must exclude) -> save on the LAST bubble -> the
prefilled body is ## 1./## 2./## 3. in order, byte-exact against
the deterministic mock, thinking-free -> edit the whole
section-2 block out of the body -> push -> git show
bor-docs:<path> equals the EDITED body byte-for-byte (section 2's
question and answer provably absent; sections 1 and 3 byte-exact;
the UI's sha prefix is git rev-parse bor-docs). Second test:
the button on the FIRST bubble still drafts the whole session
(A6 — the transcript is the session at click time, title stays
the last question); canceling leaves the branch tip untouched.
- tests/e2e/test_response_to_docs.py: the phase-59 single-turn body
expectation moves to the transcript shape ("## 1. <question>" +
the answer's markdown) — the rest of the suite unchanged.
Also lands the phase-74 file moves (00_phase.md /
03_mock_marker_e2e.md -> complete/) and the phase reports — the
house convention of committing .agents/ with the phase.
Phase 74 (TODO.md L4): a follow-up question now reaches the model WITH
the conversation so far — every prior user/brain turn and the prior
thinking blocks on brain turns (preserve-thinking) — while
POST /api/chat stays stateless (A10): the client provides the history
in the request body and the server stores nothing new.
Server (task 01):
- ChatRequest.history: optional list[HistoryTurn] (who: user|brain,
text, optional thinking) — absent/empty keeps the request
byte-identical to pre-phase-74 (the two-message [system, user]
request; the kill-switch semantics are pinned in the integration
suite).
- app.rag.prompts.history_to_messages: pure mapper — walks the turns
newest-first against the settings budgets (history_max_turns=40 /
history_max_chars=24000, BOR_HISTORY_MAX_TURNS /
BOR_HISTORY_MAX_CHARS); a capped turn is dropped WHOLE (never cut
mid-answer); the kept window is returned oldest-first; brain turns
carry their thinking as reasoning_content (A4) only when
non-empty.
- Both branches feed it: the deflected path splices it between the
system prompt and the current user message (the phase-71 recovery
still rebuilds from messages[1:]), the grounded agent receives
run_agent(..., history=hist); llm.py's message params widen to
list[dict[str, Any]] (string-only messages stay byte-identical on
the wire — the SDK passes message dicts through verbatim).
- The per-turn log line (PLAN §9) gains history_msgs=N after
kb_chars=N.
- Pins: tests/unit/test_history.py (mapper: mapping, reasoning
gating, both budgets, drop-whole, ordering, empty default),
tests/unit/test_config.py (the two settings + env overrides),
tests/unit/test_agent.py (the history splice + the default),
tests/integration/test_chat_api.py (deflected AND grounded forward
the history incl. reasoning_content, no-history byte-identity, 422
pins, the log field).
Client (task 02):
- runTurn — the single funnel for fresh send / phase-49 retry /
phase-53 stale-regen — sends history = the conversation record
minus the current question, with thinking only on brain records
that streamed one (undefined drops the key from the JSON, the
record's convention); the question is never duplicated into the
history.
Wire proof (task 03):
- The mock's echo my history marker (HISTORY_TRIGGER) answers with
the deterministic history echo — history: N prior messages; last
answer tail: <last 24 chars>; thinking: yes|no — checked BEFORE
the DEFLECT_MODE branch (like TABLE_TRIGGER), so it fires on both
turn branches whatever the gate says; the module docstring records
the user/assistant-only history invariant that keeps every
existing (tool-result-classified) marker flow unaffected.
- tests/e2e/test_llm_history.py (isolated): a grounded follow-up and
a deflected follow-up both receive history: 2 prior messages +
thinking: yes + the byte-exact tail of turn 1's answer (derived
from the persisted bor.chat.v1 record — the same array the client
maps into the body); a cold start receives history: 0 prior
messages / last answer tail: none / thinking: no.
- Regressions green in isolation: chat_rag, chat_history (phase 50),
agent_document_tools, harness_aligned_tools, stop_generation,
retry_answer, response_to_docs.
Root cause (task 01): none of C1-C3 - in Chromium 151 (real mode) a
merely-hidden tab neither stops the stream (frames arrive at full rate;
turn completes) nor fires pagehide on tab switch; C1's double-record
path was proven latent via a synthetic pagehide (trigger is
browser-dependent, e.g. Safari) and C2 (the 120s pre-token guard) was
confirmed to fire while hidden.
- C1: the pagehide partial-persist is correlated with the turn's settle
(leavePartialIndex) - the done/stop settle REPLACES it in place
(identity-guarded rememberBrainTurn in-place mode), so bor.chat.v1
and the auto-saved saved_chats row keep exactly ONE brain turn per
question; a real navigation never runs a settle, so the leave-save
is unchanged.
- C2: the visibility re-arm gives the still-armed pre-token guard a
fresh TURN_TIMEOUT_MS when the tab returns to visible - hidden time
no longer counts toward the 120s guard.
- Phase-48 teardown contract untouched: Stop / tab close / real
navigation still cancel the fetch and stop the model.
- Unit pins: tests/unit/test_frontend_hidden_tab.py (the app.js
mechanisms without a browser).
- E2E pins: tests/e2e/test_hidden_tab_stream.py - synthetic pagehide
mid-stream completes exactly once with one brain turn (localStorage
+ auto-saved row), reload restores one bubble, no-event baseline,
and the fake-clock pre-token guard re-arm (discriminating: fails
with the re-arm disabled).
Convert the two unchecked TODO items into executable phases (Protocol B,
appended after the 72 completed phases):
- 73_hidden_tab_stream (TODO L3): a merely-hidden browser tab must never
stop a generating answer; repro/root-cause decision tree + the pagehide
partial-correlation fix + the hidden-tab E2E pin.
- 74_llm_chat_history (TODO L4, history): client-provided history in
POST /api/chat (stateless, A10) mapped through both the deflected and
grounded agent paths, prior thinking blocks preserved via
reasoning_content, capped oldest-first; mock echo marker + E2E.
- 75_save_doc_full_session (TODO L4, save-as-doc): the Save-as-doc draft
body becomes the full session transcript; edit-out happens in the
existing doc-edit body; multi-turn git-verified E2E.
Owner-confirmed assumptions A1-A7 are recorded as ASSUMPTION lines in the
task files. TODO.md is cleared (items now live in .agents/phases/todo/).
Standardize on the .agents/ directory (shared with project skills):
phases/, user_stories/, reports/, screenshots/, validate.sh, and
phase-sessions/ + pipeline.log all move to .agents/ (git mv preserves
history; runtime artifacts move alongside).
Updates every reference in AGENTS.md, README.md, .gitignore, app
docstrings, and test story headers. Historical KB content in data/
and the runtime pipeline.log transcript are left untouched.
Codifies the 2026-09-05 turbo comparison workflow as a project skill under
.agents/skills/: switch BOR_LLM_CHAT_MODEL in .env, run the fixture gate
(twice, for variance) + the locked derived gate with per-turn wall timing,
interpret the two metrics against the reference model rates (re-read habit:
lite ~100%, turbo ~12%; usage-floor MISS as test artifact; caps as real
regression), record the verdicts byte-exact in TOOL_CALLING_TESTING.md, and
commit the doc. Rules baked in: never touch the battery/thresholds/fixtures,
never edit app code, never commit .env.
turbo (2026-09-05, same fixture KB): fixture gate PASS 100%/100% on both
metrics, two runs (wall 105-135s vs lite 43-55s); the redundant re-read
of seeded documents that capped lite's executed ratio at 58-73% is
model-specific (turbo re-read rate ~12% vs ~100% in-sample), corroborating
section 7's framing. Locked derived battery: turbo fails only the >=6/10
tool-turn usage floor (it answers seeded read-target questions from
context instead of making the refusable read call) - accuracy on all
emitted calls still 100%/100%.
Phase 72 (72_teaching_refusals) — completed under the 2026-09-04 controlled
methodology (owner directive: stop clearing/re-importing the homelab KB per
iteration; measure tool-calling accuracy on a controlled fixture KB, target
>90%).
Real-model gate verdicts (live, configured chat model 'lite', fixture KB):
- Controlled fixture battery (the new methodology's pass condition —
contract accuracy >= 90%): PASS, 4 consecutive runs:
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 8/11 executed (73%) contract 11/11 (100%) 2026-09-04 (wall 43.4s)
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 8/13 executed (62%) contract 12/13 (92%) 2026-09-04 (wall 50.6s)
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 7/11 executed (64%) contract 11/11 (100%) 2026-09-04 (wall 46.8s)
gate: lite PASS turns=10 answered=10 caps=0 tool-turns=10 calls 9/15 executed (60%) contract 14/15 (93%) 2026-09-04 (wall 54.8s)
- Locked derived battery (phase-72 task 05, executed >= 90% bar, run
unchanged on the same fixture KB):
gate: lite FAIL turns=10 answered=10 caps=0 tool-turns=10 calls 5/15 executed (33%) contract 12/15 (80%) 2026-09-04 (wall 47.7s)
The teaching works — every bare-path trap self-corrects in exactly one
round, zero cap hits, zero repeat loops, 10/10 answered. The locked
executed bar is blocked by ALREADY_IN_CONTEXT dedupe refusals on the
corrected re-reads (the trap question seeds its target, so the correct
combined-form read is refused for redundancy) — a copy-invariant model
behavior (five copy variants, 0/15 re-reads flipped, 2026-09-03 -> 04)
and an app-semantics decision for the owner (TOOL_CALLING_TESTING.md
sections 5 and 7), not a copy lever.
Copy changes this phase owns (unit pins updated to follow):
- app/rag/agent.py: ls teaching refusals (path-like scope -> document-path
line; unknown source -> no-source line with the source-name
parenthetical), read/grep 'did you mean source/path?' teaching
(find_path_candidates: exact or suffix path match, catalog order, cap 3),
ALREADY_IN_CONTEXT naming the correct action (answer from the text
already in the prompt), read tool description front-loaded with the
do-not-read rule (the 2026-09-04 controlled telemetry: the re-read is
the only remaining refusal class; contract accuracy 92-100% across runs)
- app/rag/prompts.py: TOOLS_SECTION states the document-identity contract
up front (ls path = source name; read/grep = combined source/path
including the source name; do-not-read for <documents> documents placed
next to the read teaching; one-call-per-reply and never-repeat rules)
- tests: refusal pins (unit + integration), new dedicated E2E suite
tests/e2e/test_tool_path_teaching.py (mock misuse flow, green in
isolation), regression suites green in isolation (harness_aligned_tools,
agent_document_tools, agent_unlimited_tools, search_tool, chat_rag).
Gates: uv run pytest green (1501); coverage TOTAL 99% (>90%); ruff +
pyright clean. Carries the still-uncommitted phase-71 todo/ -> complete/
move and both phases' .agent/reports/ (AGENTS.md 8).
The phase-72 iteration loop cleared the database, git-cloned the homelab repo, re-imported 38-51 documents and re-embedded per run — many minutes per iteration against a different KB every time (owner directive 2026-09-04: stop importing the homelab repo on every test run). Replace it with:
- tests/fixtures/agent_kb/: 8 hand-written markdown docs (sources 'deployments'/'homelab') whose specifics (rack7, 10.77.42.0/24, VLAN 130, rbm-8842, 17 2 * * *, obsidian-bor:2026.7.14, 18765, 18443, ...) no model can guess; read targets carry non-topical filenames so their questions do not lexically seed them (the read must actually happen)
- tests/fixtures/test_kb.dump.sql: data-only snapshot (TRUNCATE + INSERTs incl. embeddings, self-contained git_sources rows, static KB overview) — verified by round-trip checksum at build time
- scripts/load_test_kb.py: one-off rebuild (real pipeline + embeddings, ~2s) that also prints the per-question retrieval report (all 10 battery questions must be grounded)
- scripts/restore_test_kb.py: sub-second one-transaction restore (no git clone, no re-embedding)
- scripts/agent_realmodel_check.py: the gate gains --restore / --mode fixture (curated 10-question battery with one unambiguously correct tool behavior per question) / --turns N (12s micro-loop) / --concurrency / per-turn + total wall timing, and a second accuracy metric (contract accuracy: well-formed calls targeting resolvable entities) alongside the phase-72 locked executed ratio — the re-read of a seeded doc is a copy-invariant model behavior (5 variants, 0/15 flipped) that the dedupe refusal counts as a failure
- TOOL_CALLING_TESTING.md: the human-readable methodology (fast loop, design rules, metrics, copy levers + tried-and-reverted table, current standing, open design question)
Measured: restore 0.03s; micro-loop ~12s; full loop ~43-55s; concurrency 2/3 gives no gain (endpoint serializes).
The model treated the combined 'source/path' string (as printed in
search result lines, read-result headers and refusals) as the
document's identity and passed it as 'source' — e.g.
source='homelab/active/container_caddy/caddy.md' instead of
source='homelab', path='active/container_caddy/caddy.md'.
- Rewrite the read_document description with the split rule (source =
before the FIRST '/', path = after it) and a worked example; share
the source/path parameter descriptions between read_document and
search_documents; map search result lines back onto the split.
- New _resolve_document: on a lookup miss with a '/' in source, retry
at the first slash (source names are directory basenames and can
never contain '/'), plus a continuation candidate for a split at a
later slash; a self-corrected combined form for an already-in-context
document is still rejected as ALREADY_IN_CONTEXT.
- A slash-carrying source that matches nothing gets an educational
refusal naming the corrected arguments instead of the generic line
that repeated the combined form.
Verified live against aipi (lite) + the imported homelab KB: A/B on
the exact failure scenario (5 runs each, right after a
combined-source search result) — old descriptions 5/5 combined, new
descriptions 5/5 clean; two live UI turns (Playwright) produced only
clean split arguments, including a multi-hop read of
install_caddy_deskwork.yaml that landed in done.sources. Full suite:
1376 passed, app coverage 99% (agent.py 100%), ruff + pyright clean,
agent/search E2E green in isolation.
meta description -> locked (A3) auto-save string (history.html L6).
page-sub -> locked (A3) string, the <strong>Save</strong> emphasis retired with the button (L106-109).
empty row -> locked (A3) string; colspan=6, hidden, and row id untouched (L163).
h1, the anonymous gate section, and history.js are byte-identical — state language verified accurate.
- task 01: relocate the .chat-actions row (New chat + Share, comments byte-identical with a Phase 65 note) from the top of the column to the bottom of .chat-shell, directly above the composer
- task 02 (owner-locked A1): wrap the row + #composer in ONE sticky .chat-bottom unit (position: sticky; bottom: env(safe-area-inset-bottom, 0), no z-index) — the pills stay at the bottom of the screen at every scroll position and settle into flow above the footer
- task 03 (owner-locked A2): right-align the bottom row to the column's right edge (justify-content: flex-end), mirroring the right-aligned Save-as-doc corner; the five action pills share one 44px / 999px-pill geometry
- task 04: dedicated Playwright suite tests/e2e/test_bottom_chat_actions.py (resting geometry, the A1 pin across the sticky range, A2 alignment + DOM order + mobile stack + 360px overflow bound + 44px touch targets, New chat / Share click-through) — green in isolation
- task 05: regression matrix green in isolation (pinned_composer 4, save_share_ux 5, chat_persistence 4, share_chat 4, chat_history 5, smoke 3); full gate green — unit + integration pass, app/ coverage 99% (>90%), ruff + pyright clean
Phase 65 (TODO.md L3): move the New chat + Share cluster to the pinned
bottom of the chat column (sticky .chat-bottom unit with the composer)
and right-align the row to the Save-as-doc action corner.
Phase 66 (TODO.md L4): History tab copy — every chat saves
automatically; retire the Save-button references.
Fixed: index.html meta description, empty-state sub and composer
placeholder (A1); app/config.py default suggestion chips → the four
neutral A2 defaults (BOR_SUGGESTIONS override unchanged); sources.html
KB page-sub → the current source model (git repos + local dirs +
uploaded archives, Sync pulls/imports); git-sources.html example URL
→ your-repo.git (A3); all 9 footers → neutral default in
span.footer-text (the phase-62 hook); E2E/unit conftests force the
code defaults so a local .env cannot leak corpus copy into tests;
new unit text pins + dedicated E2E suite.
Task 02 verification read-through — no change needed:
- sources.html sync result/error copy (matches the real sync behavior)
- tuning.html page-sub (accurate as written)
- history.html page-sub (accurate as written)
- doc-edit.html page-sub (accurate as written)
- git-sources.html page-sub (accurate as written)
- #sources-gate anonymous copy (accurate as written)
Remove the blanket .agent/ gitignore so the phase roadmap, user
stories, reports, and PLAN.md are versioned with the code. Only
runtime artifacts (.agent/phase-sessions/, .agent/pipeline.log)
remain ignored. Update AGENTS.md git protocol rule to match.
Phase 52's first pass shipped `position: sticky; bottom` on `.composer` and
called the phase done, but the owner's requirement — "the chat message-input
textarea should be at the bottom of the screen" — still failed in the browser:
on an empty/short chat the input rested just under the empty state (~57% of
the viewport) with a dead band down to the footer.
`position: sticky` can only pull a box UP toward the scrollport's bottom edge;
it can never push a box DOWN to meet it, so on a page that does not overflow
it is a no-op. The old story suite only exercised an overflowing conversation
(one test even asserted the buggy resting position as expected), which is why
the half-fix passed.
- `.messages { flex: 1 1 auto }` — absorbs a short page's free space so the
composer's resting in-flow position is the bottom of the full-height column
(body min-height:100dvh -> .app-main flex:1 -> .chat-shell flex:1); basis
stays `auto`, no height cap, no overflow — the document stays the scroller
- `.composer { bottom: env(safe-area-inset-bottom, 0) }` — the explicit 0
fallback replaces the env()-only offset, which degraded to `auto` (no pin)
wherever env() is unsupported
- E2E: `test_empty_chat_composer_sits_in_normal_flow` ->
`..._at_the_screen_bottom` (chrome-only band below the resting composer);
the phone suite now checks the resting position as well as the pinned one
- Unit pins: the flex-grow half and the full-height column are pinned, so the
fix cannot silently regress to sticky-only
Still CSS-only — no DOM change, no JS, no new scroll call site (phase 42
never-auto-scroll contract intact), no z-index.
Verified: 1019 unit/integration tests pass (app/ coverage 99%), ruff and
pyright clean; tests/e2e/test_pinned_composer.py green in isolation (4), plus
the stop/autoscroll/persistence/mobile-nav suites and 14 layout/scroll
neighbours green in isolation.