755 lines
46 KiB
Markdown
755 lines
46 KiB
Markdown
# REMEDIATION_PLAN.md — Vibe Bot "Gold Quality" Roadmap
|
||
|
||
Date of audit: 2026-08-17
|
||
Status: Approved for execution (maintainer confirmed scope decisions in §1)
|
||
|
||
This document is written for a coding agent. Work through it top to bottom:
|
||
§1 (scope), then §2 (the audit, for reference), then §3 (the phases, in order).
|
||
|
||
## 0. Agent Protocol
|
||
|
||
### Commands (run from the repo root)
|
||
|
||
```bash
|
||
uv sync --extra dev # setup (plain `uv sync` already includes ruff+pyright)
|
||
uv run pytest vibe_bot/tests/ -v # tests (must run from repo root)
|
||
uv run ruff check vibe_bot/ # lint
|
||
uv run mypy vibe_bot/ # strict types
|
||
uv run pyright vibe_bot/ # strict types
|
||
uv run black --check vibe_bot/ # format check
|
||
```
|
||
|
||
- `uv run python -m vibe_bot.main` logs into the real Discord server. **Never run it as a smoke test.**
|
||
- A repo-root `.env` with placeholder values is required for imports/tests (see AGENTS.md).
|
||
- If `uv run <tool>` fails with `ModuleNotFoundError`/bad interpreter: `rm -rf .venv && uv sync --extra dev`.
|
||
- Baseline at audit time: 156/156 tests pass; all four static gates pass. Two tests
|
||
(`test_llama_wrapper.py::test_chat_completion_think`, `::test_chat_completion_instruct`)
|
||
make **live network calls** and only pass when `CHAT_ENDPOINT` is reachable.
|
||
|
||
### Per-task rules
|
||
|
||
1. Work **one task at a time**, in the order given within a phase.
|
||
2. After each task, run the full gate set (pytest + ruff + mypy + pyright + black). All must pass before starting the next task.
|
||
3. Add/adjust tests as specified in each task. Tests are the safety net — do not delete existing assertions to make a gate pass.
|
||
4. Check off the task's `[ ]` box in this file when its "Done when" criteria are met.
|
||
5. Do not commit unless the maintainer asks. (Suggested: one commit per task, message `P1-T1: <summary>`.)
|
||
6. Line numbers cited in §2/§3 are from the audit date and will drift as you edit — treat them as orientation, then locate by symbol name.
|
||
|
||
### Severity key
|
||
|
||
Critical = bot harms itself/availability · High = data corruption or dead-end UX · Medium = maintainability/cost · Low = polish.
|
||
|
||
## 1. Scope Decisions (confirmed with maintainer — do not revisit)
|
||
|
||
| Topic | Decision | Consequence for this plan |
|
||
|---|---|---|
|
||
| Multi-user privacy | Private, trusted group. **Shared chat history is a feature.** Anyone may overwrite any bot; anyone may wipe history. | Findings 1.1, 1.2, 5.1 (authorization) are **accepted by design**. Only scoped residuals remain: fix the misleading "It may already exist" message, document shared-memory semantics in README, add cost-abuse cooldowns. |
|
||
| Structural change | **Full restructure approved** (§3 Phase 2). | `main.py` is split into `commands/` + `services/` + `prompts.py`; `llama_wrapper` renamed. |
|
||
| TTS model files in git (325MB `.onnx` + 26MB `.bin`) | **Leave as-is.** | Finding 6.3 dropped; no LFS/build-download work. |
|
||
|
||
## 2. PART I — THE AUDIT REPORT
|
||
|
||
### State of the Codebase
|
||
|
||
~5,900 lines across 8 source files + 6 test files. Hygiene is good (strict mypy+pyright clean,
|
||
ruff/black clean, 156 green tests, env-based secrets, fully parameterized SQL, solid docstrings).
|
||
The core problems are structural: an async runtime doing almost everything synchronously on the
|
||
event loop, a data-retention logic bug, unguarded failure paths, and a 1,276-line god module
|
||
mixing command handlers, LLM orchestration, and 3 KB of prompt constants.
|
||
|
||
Tally: 2 Critical · 6 High · 9 Medium · 8 Low (some Highs downgraded to accepted-by-design in §1).
|
||
|
||
### Findings
|
||
|
||
#### Dimension 1: Code Quality & Correctness
|
||
|
||
- **1.1** Sev: High (→ accepted, scoped). Custom-bot namespace is global; `INSERT OR REPLACE`
|
||
(`database.py` `create_custom_bot`) lets any user silently overwrite another's bot; user is told
|
||
"It may already exist" when it was actually overwritten.
|
||
→ Residual: message must say "updated/replaced".
|
||
- **1.2** Sev: High (→ accepted). `search_similar_messages` is unscoped; RAG context can include
|
||
other users' conversations. Intended shared-memory behavior.
|
||
→ Residual: document explicitly in README; do not "fix".
|
||
- **1.3** Sev: High. `_cleanup_old_messages` (`database.py`): the second DELETE's subquery runs
|
||
**after** the first delete mutated `chat_messages`, so it removes embeddings for the next-oldest
|
||
*live* messages and never the just-deleted ones. Unbounded orphan growth in `message_embeddings`
|
||
+ lost RAG recall.
|
||
- **1.4** Sev: Medium. Magic string `WHERE username != 'vibe-bot'` (`database.py` in
|
||
`search_similar_messages` and `get_user_history`). Silently breaks if the bot's Discord
|
||
username differs; bot rows leak into user search results.
|
||
- **1.5** Sev: Medium. `llama_wrapper.embedding()`: `resp.json()` sits **outside** the try block;
|
||
a non-JSON 2xx body raises `JSONDecodeError` out of `get_conversation_context`, which is called
|
||
outside `handle_chat`'s try → unhandled `CommandError`, user left hanging.
|
||
- **1.6** Sev: High. Doodlebob/retcon failure holes: `image_generation` catches only
|
||
`openai.APIConnectionError` (4xx/5xx propagate, no user feedback); `!retcon` has no
|
||
attachment-count check (empty image list → API error), and empty `image_b64` decodes to `b""`
|
||
and an empty file is uploaded with a success message.
|
||
- **1.7** Sev: Low. `talkforme`: int limit validated **after** the announcement message; the 20-cap
|
||
never announced; first reply sent unchunked (can exceed Discord's 2000-char limit and abort the
|
||
conversation).
|
||
- **1.8** Sev: Low. `!speak` `--voice` parse uses `message.rsplit("--voice ", 1)`; text containing
|
||
the flag mid-string is corrupted.
|
||
- **1.9** Sev: Low. `get_recent_messages` and `get_custom_bot` annotate `datetime` but SQLite
|
||
returns `str` (no row factory).
|
||
- **1.10** Sev: Low. `handle_chat`/`_speak_with_bot` ignore `add_message`'s `False` return —
|
||
silent persistence failure.
|
||
|
||
#### Dimension 2: Architecture & Structure
|
||
|
||
- **2.1** Sev: **Critical.** Synchronous work on the asyncio event loop. `chat_completion_with_tools`
|
||
is `async def` but wraps a **sync** OpenAI client; `doodlebob`/`retcon`/`talkforme`/`speak`/
|
||
`handle_chat` call sync LLM (60s timeout), image (300s timeout), embedding, TTS (CPU), and
|
||
`requests.get` directly inside async handlers. One slow `!doodlebob` can miss Discord heartbeats
|
||
and drop the whole bot offline for every user.
|
||
- **2.2** Sev: High. `main.py` is a 1,276-line god module: command handlers + LLM orchestration +
|
||
prompt constants + chunking + TTS wiring; `bot = commands.Bot(...)` at import time makes the
|
||
module un-importable without full env.
|
||
- **2.3** Sev: Medium. `config.py` raises `RuntimeError` at **import** if env vars are missing;
|
||
`logging.basicConfig` is called in three modules (`config.py`, `database.py`, `main.py`) —
|
||
first wins, rest dead.
|
||
- **2.4** Sev: Medium. `on_message` constructs a `CustomBotManager` (connection + DDL) and runs a
|
||
full `list_custom_bots()` SELECT on **every message**, including non-`!` chatter.
|
||
- **2.5** Sev: Medium. Inconsistent lifetimes: `get_database()` singleton vs per-command
|
||
`CustomBotManager()`; `ChatDatabase.client` (OpenAI) is created but **never used** (dead code —
|
||
embeddings go through raw `requests`); `llama_wrapper` creates a new `openai.OpenAI` per call
|
||
(no connection reuse).
|
||
- **2.6** Sev: Low. Dead config: `COMPLETION_ENDPOINT`/`_KEY`/`_MODEL` are required at import but
|
||
used nowhere; `EMBEDDING_DIMENSION` unused.
|
||
- **2.7** Sev: Medium. Tool-calling scaffolding (tool-def dict + executor + notifier closures)
|
||
copy-pasted between `handle_chat` and `_speak_with_bot`; adding a tool means editing 3+ places.
|
||
|
||
#### Dimension 3: Readability & Maintainability
|
||
|
||
- **3.1** Sev: Medium. Three near-identical chat functions in `llama_wrapper.py`
|
||
(`chat_completion` / `chat_completion_instruct` / `chat_completion_with_history`) differing only
|
||
by `seed=-1` and message shape; `chat_completion` has zero production callers.
|
||
- **3.2** Sev: Low. Misleading names: `llama_wrapper` (no Llama inside); `get_channel_members_impl`
|
||
actually returns **guild-wide** members (Discord has no per-channel membership) and its docstring
|
||
overpromises; `bot_name` loop var in `on_message` shadows the module `bot`.
|
||
- **3.3** Sev: Medium. The 1900-char chunking while-loop is copy-pasted **five** times in
|
||
`main.py` (debug members/whoami/tools, voices, history); two 1000-char mid-word splitters
|
||
(handle_chat, talkforme) can break words and multi-codepoint emoji.
|
||
- **3.4** Sev: Medium. Log noise + PII: INFO-level logs per cursor operation in `database.py`;
|
||
`logger.info("Chat prompts: %s", prompts)` logs full conversation + user profile every message.
|
||
- **3.5** Sev: Low. 11 `type: ignore` suppressions; most hide that `tool_call.function` is
|
||
`Function | None` and should be narrowed with a guard.
|
||
|
||
#### Dimension 4: Documentation
|
||
|
||
- **4.1** Sev: Medium. README stale: file tree missing `tools.py`, `build.sh`, `vibe-bot.container`,
|
||
`.gitea/`, `AGENTS.md`; lists `.env` as a repo file; `!debug` and `!lobotomize` undocumented;
|
||
RAG description omits the shared-memory (cross-user) behavior.
|
||
- **4.2** Sev: Medium. `test_config.py` hardcodes
|
||
`sys.path.insert(0, "/var/home/ducoterra/Projects/vibe_discord_bots")` (two places); passes only
|
||
via cwd fallback.
|
||
- **4.3** Sev: Low. Command reference exists in three places (README, docstrings, in-chat `!debug`)
|
||
with no single source of truth.
|
||
|
||
#### Dimension 5: Robustness & Error Handling
|
||
|
||
- **5.1** Sev: High (→ accepted, scoped). No command authorization or rate limiting anywhere.
|
||
Accepted for a trusted group; residual: per-user cooldowns on paid commands
|
||
(`!doodlebob`, `!talkforme`, `!speak`) as a cost-abuse guard.
|
||
- **5.2** Sev: Medium. No input size bounds: personality (min only), `!speak` text, `!doodlebob`
|
||
prompt, `!talkforme` topic. Unbounded personality is also a prompt-injection/cost surface
|
||
(concatenated verbatim into system prompts).
|
||
- **5.3** Sev: Medium. `!retcon` fetches attachment URLs with no allowlist, no size cap, blocking
|
||
`requests.get`.
|
||
- **5.4** Sev: Medium. No SQLite `WAL`/`busy_timeout` — required once 2.1 is fixed with threads,
|
||
or concurrent writes fail with "database is locked".
|
||
- **5.5** Sev: Low. Swallowed failures with silent degradation: TTS skips failing chunks (partial
|
||
audio, no user notice); `get_channel_members_impl` returns an error *string* the LLM treats as
|
||
real roster data.
|
||
|
||
#### Dimension 6: Performance & Security
|
||
|
||
- **6.1** Sev: Medium. RAG retrieval is O(N): loads **all** embeddings (up to 1000 × 8 KB) into
|
||
Python per query, plus one extra `SELECT` per candidate for the `_response` join (N+1).
|
||
- **6.2** Sev: Medium. Redundant embedding API calls: ~3 per turn, and the **bot response**
|
||
embedding is stored but never retrieved (search excludes bot rows). ~⅓ of embedding spend wasted.
|
||
- **6.3** Sev: Medium (→ **dropped** per §1). 351 MB of model binaries committed to git.
|
||
- **6.4** Sev: Medium. Dependency hygiene: `mypy` (two conflicting versions: main ≥2.1.0, dev
|
||
extra ≥1.17.0), `pytest`, `pytest-env`, `types-requests` are **runtime** deps shipped into the
|
||
container; `[tool.uv] required-environments` pins the lock to linux/x86_64.
|
||
- **6.5** Sev: Low. CI builds/pushes with **no lint/test gate** and only on main/release (no PR
|
||
signal); container runs as root, no healthcheck.
|
||
- **6.6** Security positives (no action): no hardcoded secrets (env only, `.env*` gitignored,
|
||
registry creds via Gitea secrets); all SQL parameterized; no XSS surface. Residuals covered by
|
||
1.1/1.2 (accepted), 3.4 (PII logging), 5.3 (fetch allowlist).
|
||
|
||
## 3. PART II — THE REMEDIATION ROADMAP
|
||
|
||
Phase order is by dependency and risk: Phase 1 removes the ways the bot harms itself
|
||
(behavior-preserving fixes), Phase 2 is the restructure on the stabilized codebase,
|
||
Phase 3 makes the dev loop and CI trustworthy, Phase 4 optimizes the final shapes.
|
||
|
||
---
|
||
|
||
### PHASE 1 — Stability & Foundation
|
||
|
||
**Goal:** The bot can no longer freeze globally, corrupt retention data, or strand users in
|
||
unhandled error states. Zero structural changes; behavior-preserving fixes behind the existing
|
||
test suite.
|
||
|
||
#### P1-T1 — SQLite concurrency preparation (finding 5.4)
|
||
|
||
Prerequisite for P1-T2; land it first.
|
||
|
||
- File: `vibe_bot/database.py` → `ChatDatabase._initialize_database` (and
|
||
`CustomBotManager._initialize_custom_bots_table` if you want both tables covered).
|
||
- After connecting, execute:
|
||
```sql
|
||
PRAGMA journal_mode=WAL;
|
||
PRAGMA busy_timeout=5000;
|
||
```
|
||
(WAL is persistent; busy_timeout is per-connection, so set it wherever connections are opened —
|
||
introduce a small `_connect()` helper returning a configured connection and use it in every
|
||
method instead of raw `sqlite3.connect`.)
|
||
- Tests: none required (pragma-only), but keep the suite green.
|
||
- Done when: [ ] all `sqlite3.connect` calls in `database.py` go through the helper; gates green.
|
||
|
||
#### P1-T2 — Unblock the event loop (finding 2.1 — THE critical fix)
|
||
|
||
- Files: `vibe_bot/main.py`, `vibe_bot/llama_wrapper.py`.
|
||
- Strategy: **`asyncio.to_thread` first** (lowest risk, behavior-preserving). True async clients
|
||
come in Phase 2.
|
||
1. In every async command/handler, wrap blocking calls:
|
||
- `handle_chat`: `db.get_conversation_context(...)`, `llama_wrapper.chat_completion_with_tools(...)`
|
||
(it is `async def` but its inner client calls are sync — either keep `await`ing it after
|
||
converting its inner `.create()` calls to run in a thread, or simplest: change
|
||
`chat_completion_with_tools` to a plain sync function and `await asyncio.to_thread(...)` it),
|
||
`db.add_message(...)` (×2).
|
||
- `doodlebob`: `select_image_layout`, `chat_completion_instruct` (×2: prompt + verify),
|
||
`image_generation`.
|
||
- `retcon`: the `requests.get` download (remove the `# noqa: ASYNC210` once async via
|
||
`to_thread`) and `llama_wrapper.image_edit`.
|
||
- `talkforme`: every `llama_wrapper.chat_completion_with_history` call.
|
||
- `_speak_with_bot` / `_speak_plain`: `db.get_conversation_context`,
|
||
`chat_completion_with_tools`, `db.add_message`, `engine.generate_audio` (CPU-bound TTS).
|
||
2. Keep all function signatures otherwise identical.
|
||
- Tests:
|
||
- New test: while a mocked 0.5s LLM call is in flight inside `handle_chat` (or `doodlebob`),
|
||
a concurrent `asyncio.create_task` sleep completes — i.e. prove the loop is not blocked
|
||
(assert with `asyncio.wait` / timing: the concurrent task's elapsed time < the mocked call
|
||
time by a wide margin, e.g. concurrent task finishes while the call is still running).
|
||
- Done when: [ ] no direct blocking call (openai sync client, `requests`, sqlite, TTS) executes
|
||
on the event loop thread; new concurrency test passes; gates green.
|
||
|
||
#### P1-T3 — Fix `_cleanup_old_messages` (finding 1.3)
|
||
|
||
- File: `vibe_bot/database.py` → `ChatDatabase._cleanup_old_messages`.
|
||
- Rewrite: first `SELECT id, message_id FROM chat_messages ORDER BY timestamp ASC LIMIT ?`
|
||
(where `? = count - MAX_HISTORY_MESSAGES`); derive the affected message_ids **including**
|
||
their `_response` companions in Python (`mid` and `f"{mid}_response"`); then
|
||
`DELETE FROM chat_messages WHERE id IN (...)` and
|
||
`DELETE FROM message_embeddings WHERE message_id IN (...)` — both from the captured set,
|
||
in the same transaction (already the case: caller commits).
|
||
- Tests (in `test_database.py`):
|
||
- Insert `MAX_HISTORY_MESSAGES + 3` user+response rows (with embeddings). Call the internal
|
||
cleanup path via `add_message`. Assert: exactly 3 oldest user rows gone; their embeddings
|
||
gone; their `_response` embeddings gone; the **next** oldest surviving user row still has
|
||
its embedding; `SELECT COUNT(*) FROM message_embeddings` == rows that have a live message.
|
||
- Done when: [ ] regression test passes; no orphaned embeddings possible after cleanup; gates green.
|
||
|
||
#### P1-T4 — Guard `embedding()` JSON parsing (finding 1.5)
|
||
|
||
- File: `vibe_bot/llama_wrapper.py` → `embedding`.
|
||
- Move `data = resp.json()` inside the `try` and catch `ValueError` (JSONDecodeError's base)
|
||
alongside `requests.RequestException` → return `[]`.
|
||
- Test: mock `requests.post` returning a 200 with body `<html>rate limited</html>` → assert
|
||
`embedding(...) == []` and no exception.
|
||
- Done when: [ ] non-JSON 2xx can't escape `embedding()`; test passes; gates green.
|
||
|
||
#### P1-T5 — Close doodlebob/retcon failure holes (finding 1.6, part of 5.3)
|
||
|
||
- Files: `vibe_bot/llama_wrapper.py` (`image_generation`), `vibe_bot/main.py` (`doodlebob`, `retcon`).
|
||
- 1. In `image_generation` and `image_edit`, catch `openai.OpenAIError` (covers
|
||
`APIStatusError`, `APIConnectionError`, `APITimeoutError`) → return `""` (log the error).
|
||
Callers already handle `""`.
|
||
- 2. `retcon`: before calling `image_edit`, if `image_data_list` is empty → send
|
||
"Please attach an image to edit." and return. After `image_edit`, if result is `""` → send
|
||
"Failed to edit the image." and return. Wrap `base64.b64decode` in try/except
|
||
(`binascii.Error`) with a user-facing failure message.
|
||
- 3. `retcon` download: restrict to Discord CDN hosts (`discordcdn.com`, `discordapp.com`,
|
||
`discord.com`, `discord.media`) — reject anything else with a warning log; cap download at
|
||
8 MB (stream with `stream=True`, abort if `Content-Length` or accumulated bytes exceed cap).
|
||
- 4. `doodlebob`: after `select_image_layout`/prompt/verify, the existing `== ""` checks remain;
|
||
no extra try needed once wrappers swallow `OpenAIError`.
|
||
- Tests:
|
||
- retcon with zero attachments → friendly message, `image_edit` not called.
|
||
- retcon with `image_edit` returning `""` → failure message, no file sent.
|
||
- retcon with non-Discord attachment URL → not downloaded.
|
||
- `image_generation` with mocked `openai.APIStatusError` → returns `""`.
|
||
- Done when: [ ] all four tests pass; no `!retcon`/`!doodlebob` path ends without a user message
|
||
on API failure; gates green.
|
||
|
||
#### P1-T6 — Replace the `'vibe-bot'` magic string with a role column (finding 1.4)
|
||
|
||
- File: `vibe_bot/database.py`.
|
||
- 1. In `_initialize_database`, after the existing `bot_name` migration pattern, add: if
|
||
`role` column missing → `ALTER TABLE chat_messages ADD COLUMN role TEXT`.
|
||
Backfill in the same migration:
|
||
`UPDATE chat_messages SET role = 'assistant' WHERE message_id LIKE '%_response' AND role IS NULL;`
|
||
`UPDATE chat_messages SET role = 'user' WHERE role IS NULL;`
|
||
- 2. `add_message`: accept `role: str = "user"` and store it (callers in `main.py` pass
|
||
`role="assistant"` for the bot-response rows).
|
||
- 3. `search_similar_messages`: `WHERE cm.username != 'vibe-bot'` → `WHERE cm.role = 'user'`.
|
||
`get_user_history`: same replacement.
|
||
- 4. `main.py`: update both `add_message` call sites (handle_chat, _speak_with_bot) to pass role.
|
||
- Tests:
|
||
- Fresh DB: add user + response rows, assert `get_user_history` excludes responses and
|
||
`search_similar_messages` only matches user rows.
|
||
- Migration: create a legacy row (no role, `message_id` ending `_response`) by inserting
|
||
directly via sqlite, run `ChatDatabase()`, assert role backfilled to `assistant`.
|
||
- Done when: [ ] no `vibe-bot` string remains in `database.py` (grep); tests pass; gates green.
|
||
|
||
#### P1-T7 — Fix `talkforme` ordering and chunking (finding 1.7)
|
||
|
||
- File: `vibe_bot/main.py` → `talkforme`.
|
||
- 1. Move `int(limit)` parsing/validation **before** the "is going to talk" announcement.
|
||
- 2. Announce the effective cap: `for {min(limit, talk_limit)} replies`.
|
||
- 3. Send the first reply through the same 1000-char loop used for subsequent replies
|
||
(extract a local helper now; the global `split_message` arrives in Phase 2).
|
||
- Tests:
|
||
- non-integer limit → usage/error message, **no** announcement message sent first.
|
||
- first reply > 1000 chars → sent in chunks (assert `ctx.send` call count).
|
||
- Done when: [ ] both tests pass; gates green.
|
||
|
||
#### P1-T8 — Input size bounds (finding 5.2)
|
||
|
||
- File: `vibe_bot/main.py` (constants near `MIN_BOT_NAME_LENGTH`).
|
||
- Add and enforce with friendly rejection messages:
|
||
- `MAX_PERSONALITY_LENGTH = 1000` in `custom_bot`.
|
||
- `MAX_SPEAK_LENGTH = 5000` in `speak` (check the text part after `--voice` parsing).
|
||
- `MAX_IMAGE_PROMPT_LENGTH = 2000` in `doodlebob` and `retcon`.
|
||
- `MAX_TOPIC_LENGTH = 500` in `talkforme`.
|
||
- Tests: one per bound — over-limit input → rejection message, no LLM/DB call.
|
||
- Done when: [ ] four tests pass; gates green.
|
||
|
||
#### P1-T9 — Cost-abuse cooldowns (scoped finding 5.1)
|
||
|
||
- File: `vibe_bot/main.py`.
|
||
- Add `@commands.cooldown(rate, per, type=commands.BucketType.user)` (pick sane values, e.g.
|
||
doodlebob 1/60s, talkforme 1/120s, speak 3/30s) to the three paid commands.
|
||
- Handle `commands.MaxConcurrency`/`CooldownRetry` gracefully: register a
|
||
`on_command_error` handler that sends "You're using that too quickly, try again in Ns."
|
||
for `commands.CooldownRaise`.
|
||
- Test: mock a second immediate invocation → cooldown message, LLM not called.
|
||
- Done when: [ ] test passes; gates green.
|
||
|
||
#### P1-T10 — Honest overwrite message (scoped finding 1.1)
|
||
|
||
- File: `vibe_bot/database.py` (`create_custom_bot`), `vibe_bot/main.py` (`custom_bot`).
|
||
- Make `create_custom_bot` distinguish: check existence first (`SELECT 1 ... WHERE bot_name=?`)
|
||
→ return `"created" | "replaced" | False`. `custom_bot` command: on `"replaced"` send
|
||
"Custom bot **'X'** already existed and has been **replaced**." (Shared namespace is intended.)
|
||
- Test: create same name twice → second call reports replaced; both succeed.
|
||
- Done when: [ ] test passes; no "It may already exist" string remains; gates green.
|
||
|
||
#### P1-T11 — Stop silent TTS partial audio (scoped finding 5.5)
|
||
|
||
- Files: `vibe_bot/tts.py`, `vibe_bot/main.py`.
|
||
- `generate_audio`: track failed-chunk count; if any chunk failed, log a warning and either
|
||
(a) attach a note — simplest: return the audio and let the caller detect via a new
|
||
`partial: bool` attribute on a small result dataclass, or (b) raise `ValueError` listing the
|
||
failed chunk count and let the command send "Some audio chunks failed; audio may be incomplete."
|
||
Choose (b) if the maintainer prefers fail-fast — **default to (a)**: keep the audio, and have
|
||
the command send a short warning line when partial.
|
||
- Test: mock `process_chunk_sequential` to fail on chunk 1 of 2 → audio still produced, warning
|
||
sent (or ValueError raised, matching the chosen behavior).
|
||
- Done when: [ ] test passes; gates green.
|
||
|
||
#### Phase 1 verification & success criteria
|
||
|
||
Run the full gate set, then:
|
||
|
||
- [ ] Concurrency proof: a slow (mocked ≥60s) LLM call no longer blocks a concurrent task
|
||
(P1-T2 test) — the bot survives a slow `!doodlebob` globally.
|
||
- [ ] No orphaned embeddings and correct retention after cleanup (P1-T3 test).
|
||
- [ ] No command path ends without a user-facing message on API failure (mock 4xx/5xx tests:
|
||
doodlebob, retcon, handle_chat).
|
||
- [ ] Full suite green: `uv run pytest vibe_bot/tests/ -v` (156 original + new tests).
|
||
- [ ] `ruff` / `mypy` / `pyright` / `black --check` all pass.
|
||
|
||
---
|
||
|
||
### PHASE 2 — Structural Integrity (full restructure)
|
||
|
||
**Goal:** Kill the god module and duplicated orchestration; commands become thin adapters over
|
||
services; lifetimes explicit; true async clients. Do this **after Phase 1 gates are green**.
|
||
Order matters: utilities and clients first (P2-T1/T2), services next (P2-T3…T5), command split
|
||
after (P2-T6), DB layer last (P2-T7) because services depend on it.
|
||
|
||
#### P2-T1 — `textutil.split_message` (finding 3.3 — do first, Phase 1 code already uses it indirectly)
|
||
|
||
- New file: `vibe_bot/textutil.py`.
|
||
- `def split_message(text: str, limit: int = 1900) -> list[str]`:
|
||
- Split on newlines first (never break a line mid-text when avoidable); if a single line
|
||
exceeds `limit`, split it **codepoint-safely** (Python `str` slicing is already codepoint-safe,
|
||
but never split inside a grapheme cluster — use `regex` package's `\X` only if you add the dep;
|
||
otherwise plain slicing at `limit` is acceptable and matches today's behavior — document the
|
||
choice).
|
||
- Preserve existing behavior for the 5 call sites; replace all five 1900-char loops in
|
||
`main.py` (debug members/whoami/tools, voices, history) and both 1000-char splitters
|
||
(`handle_chat`, `talkforme` → `limit=1000`).
|
||
- Tests: exact-limit text; text with no newlines; emoji at the split boundary (assert no lone
|
||
surrogate/combining char corruption by round-tripping `"".join(chunks) == original`);
|
||
multi-paragraph text respects newline preference.
|
||
- Done when: [ ] `main.py` contains zero hand-rolled chunking loops; `"".join(split_message(t)) == t`
|
||
property test passes for a corpus including emoji; gates green.
|
||
|
||
#### P2-T2 — `llm_client.py`: one core, shared async clients (findings 2.1-completion, 3.1, 2.5, 3.2, 3.5)
|
||
|
||
- New file: `vibe_bot/llm_client.py` (delete `llama_wrapper.py`; update all imports).
|
||
- 1. **Shared clients:** module-level lazy singletons `get_chat_client()`,
|
||
`get_image_gen_client()`, `get_image_edit_client()`, `get_embedding_http_session()` returning
|
||
`openai.AsyncOpenAI(base_url=..., api_key=..., max_retries=0 where currently set)` instances
|
||
built once. Delete the per-call `openai.OpenAI(...)` constructions.
|
||
- Also delete `ChatDatabase.client` (dead code) and its `openai` import in `database.py`;
|
||
drop `db.client.close()` from the `chat_db` test fixture.
|
||
- 2. **One core function:**
|
||
`async def chat_complete(messages: list[ChatCompletionMessageParam], *, model: str, max_tokens: int, seed: int | None = None, tools: ... | None = None, tool_executor: ... | None = None, tool_call_notifier: ... | None = None, max_tool_rounds: int = 5) -> str`
|
||
— contains the (now async) tool loop currently in `chat_completion_with_tools`.
|
||
Replace the `# type: ignore[union-attr]` cluster with a real guard:
|
||
`if tool_call.function is None: continue`.
|
||
- Keep thin compat wrappers *only if* it reduces churn: `chat_completion_instruct(system_prompt, user_prompt, ...)` and `chat_completion_with_history(system_prompt, prompts, ...)` become 3-line adapters over `chat_complete`. Delete `chat_completion` (test-only caller — port its test to `chat_complete`).
|
||
- 3. **`image_generation` / `image_edit`:** async (`await client.images.generate/edit`), catch
|
||
`openai.OpenAIError` → `""` (carries P1-T5 behavior).
|
||
- 4. **`embedding`:** async via `httpx.AsyncClient` (add `httpx` to deps) **or** keep `requests`
|
||
inside `asyncio.to_thread`. Prefer `httpx.AsyncClient` with a shared session; keep the
|
||
OpenAI-style and Ollama-style response handling and P1-T4's JSON guard.
|
||
- 5. `ToolRegistry` (finding 2.7): small class in `llm_client.py` (or `tools.py`):
|
||
- `register(name, description, args_schema, impl)` — seed it with `get_channel_members`
|
||
(schema) → `get_channel_members_impl` (impl).
|
||
- `to_openai_tools() -> list[dict]` and `async execute(name, args, channel) -> str`.
|
||
- The notifier ("is looking at the channel members...") becomes a registry-level callback.
|
||
- Tests: port all `test_llama_wrapper.py` tests to `test_llm_client.py` (rename module + async
|
||
awaits); add: client singleton identity test (same object across calls); tool registry
|
||
execute/dispatch test incl. unknown-tool path; guard test for `function is None` tool call.
|
||
- Done when: [ ] `llama_wrapper.py` deleted; grep shows zero `openai.OpenAI(` per-call
|
||
constructions; zero `type: ignore` in the new client (except the two `import-untyped` in
|
||
`tts.py`); gates green.
|
||
|
||
#### P2-T3 — `prompts.py` (part of finding 2.2)
|
||
|
||
- New file: `vibe_bot/prompts.py` — move all prompt constants from `main.py`:
|
||
`IMAGE_LAYOUT_SYSTEM_PROMPT`, `IMAGE_PROMPT_SYSTEM_PROMPT_TEMPLATE`,
|
||
`IMAGE_PROMPT_VERIFY_SYSTEM_PROMPT`, plus the "Keep your responses under 2-3 sentences."
|
||
suffix and "User Information:" assembly as a helper
|
||
`build_system_prompt(personality, user_info) -> str` (currently duplicated string in
|
||
`handle_chat` and `_speak_with_bot`).
|
||
- Move the layout parsing (`parse_image_layout`, `VALID_IMAGE_LAYOUTS`, `DEFAULT_IMAGE_LAYOUT`,
|
||
`LAYOUT_SIZES`) here too.
|
||
- Tests: existing prompt-related tests in `test_main.py` re-point at `vibe_bot.prompts`;
|
||
add a test that `build_system_prompt` output contains personality + user info + length hint.
|
||
- Done when: [ ] `main.py` (pre-split) contains no prompt text; gates green.
|
||
|
||
#### P2-T4 — `services/` (findings 2.2, 2.4, 2.5, 6.2)
|
||
|
||
New package `vibe_bot/services/` with a dataclass `BotContext` (or pass explicit args):
|
||
|
||
- `services/chat_service.py` — `ChatService.handle(ctx, bot_name, message, system_prompt, response_prefix)`:
|
||
RAG context → prompts → `chat_complete` with tools (via `ToolRegistry`) → persist
|
||
(user row `role="user"`, response row `role="assistant"`; **log a warning when
|
||
`add_message` returns False** — finding 1.10) → chunked reply via `split_message`.
|
||
**Do not embed bot responses** (finding 6.2): only user messages get embeddings — remove the
|
||
response-row embedding by not calling the embedding path for assistant rows
|
||
(`add_message(role="assistant", embed=False)` parameter, or a separate non-embedding insert).
|
||
- `services/image_service.py` — `ImageService.generate(ctx, message)`:
|
||
layout select (`max_tokens=2` — Phase 4 item, do it here since we're touching it anyway) →
|
||
prompt engineering → verify → ETA message → generate → record time → send file.
|
||
`ImageService.edit(ctx, message)`: the P1-T5-hardened retcon flow.
|
||
- `services/speech_service.py` — `SpeechService.speak(ctx, message)`: voice-flag parse
|
||
(replace `rsplit` with a **trailing-anchored regex**: `r"^(?P<text>.*)\s+--voice\s+(?P<voice>\S+)$"`
|
||
— fixes finding 1.8), voice validation, bot-vs-plain dispatch, TTS via `asyncio.to_thread`,
|
||
language lookup via a precomputed `VOICE_LANGUAGES: dict[str, str]` built from `VOICES_LIST`
|
||
at import (replaces the per-call list scan).
|
||
- `services/conversation_service.py` — `ConversationService.run(ctx, bot1, bot2, limit, topic)`:
|
||
the Phase-1-fixed talkforme loop, using `split_message(1000)`.
|
||
- Each service: constructor takes the llm client(s), `ChatDatabase`, `CustomBotManager`,
|
||
`TTSEngine | None`, `ToolRegistry` (dependency injection; no module-level singletons inside).
|
||
- Tests: port `test_main.py` command tests to service-level tests where the logic moved;
|
||
keep a thin command-layer test per command proving wiring (mock service, assert called with
|
||
parsed args). Add: voice-regex tests (flag mid-text is preserved; trailing flag extracted;
|
||
missing flag → None).
|
||
- Done when: [ ] four service modules exist; no service imports `discord` except for
|
||
`ctx`/`channel` type hints under `TYPE_CHECKING`; gates green.
|
||
|
||
#### P2-T5 — `app.py` + composition root (findings 2.2, 2.3, 2.4, 2.5)
|
||
|
||
- New file: `vibe_bot/app.py`:
|
||
- `configure_logging()` — the **only** `logging.basicConfig` call in the codebase (remove the
|
||
other two from `config.py` and `database.py`).
|
||
- `create_app() -> App` dataclass: builds `ChatDatabase`, `CustomBotManager` (singleton —
|
||
no more per-message construction), LLM clients, `ToolRegistry`, `TTSEngine | None`
|
||
(keep the tolerant init-or-None behavior), and the four services.
|
||
- Bot construction moves here: `build_bot(app) -> commands.Bot` (intents, `on_ready`,
|
||
command registration, `on_message`, `on_command_error`).
|
||
- `on_message`: skip bots; **if not `message.content.startswith("!"): return`** before any
|
||
DB/manager touch (finding 2.4); then look up the bot-name cache.
|
||
- Bot-name cache: `app.custom_bots: dict[str, tuple[prompt, creator]]` rebuilt from
|
||
`list_custom_bots()`; invalidated by `custom_bot`/`delete_custom_bot` commands (call
|
||
`app.invalidate_bot_cache()`).
|
||
- `vibe_bot/main.py` shrinks to: `def main() -> None: validate_config(); configure_logging(); app = create_app(); build_bot(app).run(DISCORD_TOKEN)` + `if __name__ == "__main__": main()`.
|
||
- `config.py`: split loading from validation — module only **loads**; add
|
||
`def validate_config() -> None` containing today's `RuntimeError` checks (called from
|
||
`main()` only, so the package becomes importable without full env).
|
||
- Tests: `test_main.py` reorganized: command tests now call `build_bot(app)`-registered
|
||
functions or the service layer; the `mock_discord` fixture is replaced by direct service
|
||
tests (Phase 4 cleanup of leftovers). Import-time behavior test: importing `vibe_bot.config`
|
||
with empty env must NOT raise (only `validate_config()` raises).
|
||
- Done when: [ ] `main.py` < 40 lines; grep shows exactly one `logging.basicConfig` in
|
||
`vibe_bot/`; no `CustomBotManager()` construction outside `app.py`; gates green.
|
||
|
||
#### P2-T6 — `commands/` package (completes finding 2.2, 2.6)
|
||
|
||
- New package `vibe_bot/commands/` — one module per group, each exposing
|
||
`def register(bot: commands.Bot, app: App) -> None`:
|
||
- `custom_bots.py`: `custom_bot`, `list_custom_bots`, `delete_custom_bot` (use
|
||
`app.bot_cache`; keep P1-T10's replaced/created semantics).
|
||
- `chat.py`: nothing registered (custom-bot chat flows through `on_message` →
|
||
`app.services.chat`) — keep the `on_message` dispatch in `app.py` calling
|
||
`app.services.chat.handle(...)`.
|
||
- `speech.py`: `speak`, `voices` (via `SpeechService`).
|
||
- `images.py`: `doodlebob`, `retcon` (via `ImageService`; keep P1-T9 cooldowns).
|
||
- `conversation.py`: `talkforme` (via `ConversationService`).
|
||
- `admin.py`: `lobotomize`, `debug`, `history`.
|
||
- `config.py`: delete dead `COMPLETION_ENDPOINT`/`COMPLETION_ENDPOINT_KEY`/`COMPLETION_MODEL`
|
||
and `EMBEDDING_DIMENSION` (finding 2.6); update README env section in Phase 3.
|
||
- `tools.py`: fix `get_channel_members_impl` docstring → "guild members" (finding 3.2);
|
||
keep the `@tool` stub as the schema source for the registry.
|
||
- Tests: each command module gets a wiring test (mock service + mock ctx); delete now-dead
|
||
fixtures (`mock_env_vars`, `mock_discord` if unused) from `conftest.py`.
|
||
- Done when: [ ] no module in `vibe_bot/` exceeds ~400 lines; every command traceable to a
|
||
single service method; gates green.
|
||
|
||
#### P2-T7 — Database layer cleanup (findings 1.9, 6.1, 2.5)
|
||
|
||
- File: `vibe_bot/database.py`.
|
||
- 1. **Single-JOIN RAG fetch** (6.1): in `search_similar_messages`, replace
|
||
"fetch all embeddings + per-candidate SELECT" with one query:
|
||
```sql
|
||
SELECT cm.message_id, cm.content,
|
||
(SELECT content FROM chat_messages r WHERE r.message_id = cm.message_id || '_response') AS response,
|
||
me.embedding
|
||
FROM chat_messages cm
|
||
JOIN message_embeddings me ON cm.message_id = me.message_id
|
||
WHERE cm.role = 'user'
|
||
```
|
||
(correlated subquery avoids a LEFT JOIN fan-out; if a real LEFT JOIN is cleaner, use
|
||
`LEFT JOIN chat_messages r ON r.message_id = cm.message_id || '_response'`).
|
||
Vectorize the cosine loop: stack blobs into one `np.ndarray` of shape `(n, dim)` and compute
|
||
similarities with one matrix op (normalize query vector once; keep the per-vector norm
|
||
recompute for now — norms get precomputed in Phase 4).
|
||
- 2. **Honest types** (1.9): register a sqlite `datetime` converter
|
||
(`sqlite3.register_adapter`/`detect_types` or a row factory mapping the `timestamp` column)
|
||
so `get_recent_messages`/`get_custom_bot` really return `datetime`; adjust tests that compare
|
||
strings.
|
||
- 3. `add_message`: add `embed: bool = True` (assistant rows skip the embedding call — finding 6.2
|
||
wiring); keep `role` parameter from P1-T6.
|
||
- 4. Delete the dead `OpenAI` import/client (if P2-T2 didn't already).
|
||
- Tests: update RAG tests for the JOIN (assert identical results vs old behavior on a seeded DB:
|
||
same top-k, same ordering); property: one query executed per search (count with
|
||
`sqlite3` tracing or a wrapper).
|
||
- Done when: [ ] `search_similar_messages` issues exactly 1 SELECT per call; type annotations
|
||
match runtime values; gates green.
|
||
|
||
#### Phase 2 verification & success criteria
|
||
|
||
- [ ] No module > ~400 lines (`wc -l vibe_bot/**/*.py`).
|
||
- [ ] `main.py` no longer imports anything except `app`.
|
||
- [ ] Every command testable via mock services without discord module mocking.
|
||
- [ ] Behavior unchanged: original 156 tests (ported) + new tests all green.
|
||
- [ ] Strict gates green with **net-fewer** `type: ignore`s.
|
||
- [ ] grep: zero `logging.basicConfig` outside `app.py`; zero `openai.OpenAI(` per-call
|
||
constructions; zero `'vibe-bot'` in `database.py`; zero `llama_wrapper` references.
|
||
|
||
---
|
||
|
||
### PHASE 3 — Developer Experience & Documentation
|
||
|
||
**Goal:** Hermetic network-free test suite; CI that gates merges; docs that match reality.
|
||
Items 3.1–3.4 can start as soon as Phase 2's P2-T6 lands (they're independent of P2-T7).
|
||
|
||
#### P3-T1 — Hermetic test suite (AGENTS.md baseline retirement)
|
||
|
||
- Files: `vibe_bot/tests/test_llm_client.py` (or `test_llama_wrapper.py` if renamed later).
|
||
- Mark the two live tests `@pytest.mark.live` and add to `pyproject.toml`:
|
||
`[tool.pytest.ini_options] addopts = "-m 'not live'"` (keep `filterwarnings` as-is).
|
||
Document in the Testing section: `uv run pytest -m live` runs them.
|
||
- Done when: [ ] `rm -rf .venv && uv sync --extra dev && uv run pytest vibe_bot/tests/ -v`
|
||
passes on a machine with placeholder `.env` and **no network** to `CHAT_ENDPOINT`
|
||
(verify by pointing `CHAT_ENDPOINT` at an unroutable address for the test run).
|
||
|
||
#### P3-T2 — Test hygiene (finding 4.2)
|
||
|
||
- Files: `vibe_bot/tests/test_config.py`, `conftest.py`, `test_llm_client.py`.
|
||
- Delete both hardcoded `sys.path.insert(0, "/var/home/ducoterra/...")` lines; if `test_config.py`
|
||
then fails on import, add `pythonpath = ["."]` under `[tool.pytest.ini_options]` instead.
|
||
- Delete dead fixtures (`mock_env_vars` if still unused) and `TEMPDIR`.
|
||
- Fix `test_bot_intents_set` to actually assert intents (or rename to match what it asserts).
|
||
- Done when: [ ] grep for `/var/home` in `vibe_bot/tests/` returns nothing; suite green.
|
||
|
||
#### P3-T3 — Coverage for new surface
|
||
|
||
- Add tests (in the matching module):
|
||
- `split_message`: property test `"".join(split_message(t, L)) == t` over a corpus including
|
||
emoji, combining marks, and code spans; exact-multiple-of-limit input.
|
||
- Voice-flag regex: trailing flag, flag mid-text (must be preserved), no flag, flag with
|
||
missing value.
|
||
- Cleanup regression (P1-T3) — already required, verify present.
|
||
- `role` migration backfill (P1-T6) — verify present.
|
||
- Cooldowns (P1-T9): immediate re-invocation → friendly message.
|
||
- Service error paths: `ChatService` with failing `chat_complete` → user sees the error
|
||
message and no rows persisted; `ImageService.edit` with rejected URL.
|
||
- Done when: [ ] all listed tests exist and pass.
|
||
|
||
#### P3-T4 — Dependency hygiene (finding 6.4)
|
||
|
||
- File: `pyproject.toml`.
|
||
- Move to `[dependency-groups] dev` (and delete the now-redundant `[project.optional-dependencies] dev`):
|
||
`mypy` (single version pin — drop the conflicting second one), `pytest`, `pytest-env`,
|
||
`black`, `debugpy`; keep `pyright`, `ruff` where they are.
|
||
- Add `httpx` to runtime deps (P2-T2).
|
||
- Widen or remove `[tool.uv] required-environments` (recommend: remove, accept multi-platform lock).
|
||
- `uv lock` after changes; verify `uv sync --extra dev` and the Containerfile's `uv sync --locked`
|
||
both work (`./build.sh` dry-run optional — do **not** push images).
|
||
- Done when: [ ] `uv run python -c "import mypy"` fails in a prod-only sync
|
||
(`uv sync --no-dev`); container build still passes `uv sync --locked`; gates green.
|
||
|
||
#### P3-T5 — CI gates (finding 6.5)
|
||
|
||
- File: `.gitea/workflows/build-push.yml` (or a new `ci.yml`).
|
||
- New job `test` on `pull_request` (and `push` to main):
|
||
```yaml
|
||
- uses: actions/setup-python@v5 (3.13)
|
||
- name: Install uv && uv sync --extra dev
|
||
- run: uv run ruff check vibe_bot/
|
||
- run: uv run mypy vibe_bot/
|
||
- run: uv run pyright vibe_bot/
|
||
- run: uv run black --check vibe_bot/
|
||
- run: uv run pytest vibe_bot/tests/ -v # hermetic per P3-T1
|
||
```
|
||
(Adapt to the Gitea runner flavor already in use; note: portaudio system dep must be present
|
||
on CI runners — `sudo dnf install portaudio` per AGENTS.md, or install in the job; if the
|
||
runner is apt-based: `apt-get install -y portaudio19-dev`.)
|
||
- Make the build/push job depend on `test` passing.
|
||
- Done when: [ ] a PR with a deliberate type error cannot pass CI (verify locally by simulating
|
||
the job steps); merge is blocked on red `test`.
|
||
|
||
#### P3-T6 — Container hardening (finding 6.5)
|
||
|
||
- File: `Containerfile`.
|
||
- Add a non-root user: `RUN useradd -m bot && chown -R bot /app` … `USER bot`; ensure
|
||
`DB_PATH` default dir is writable (`vibe-bot.container` mounts `/db` — `mkdir -p /db && chown bot /db`).
|
||
- Optional: `HEALTHCHECK` — a Discord bot has no HTTP endpoint; use a lightweight liveness probe
|
||
(e.g. a small asyncio task writing a heartbeat file, checked by `test -f`) **only if** the
|
||
maintainer wants it; otherwise skip and note why.
|
||
- Done when: [ ] `./build.sh` succeeds; `podman run --rm localhost/vibe-bot:latest uv run python -c "print(1)"`
|
||
runs as non-root (verify with `podman run ... id`).
|
||
|
||
#### P3-T7 — Logging discipline (finding 3.4)
|
||
|
||
- Files: `vibe_bot/database.py`, `vibe_bot/services/*`.
|
||
- Per-cursor-operation `logger.info` → `logger.debug` (keep init/migration at INFO).
|
||
- `logger.info("Chat prompts: %s", prompts)` → log metadata only:
|
||
`logger.info("chat: bot=%s user=%s context_msgs=%d", ...)`.
|
||
- Done when: [ ] no log line can contain message content or user profile data
|
||
(spot-check by grepping for `prompts`, `content`, `get_user_info` in `logger.` calls).
|
||
|
||
#### P3-T8 — Documentation (findings 4.1, 4.3)
|
||
|
||
- Files: `README.md`, `AGENTS.md`.
|
||
- README:
|
||
- Regenerate the file tree from the actual post-Phase-2 layout.
|
||
- Document **all** commands including `!debug` and `!lobotomize`.
|
||
- Add a "Design notes" subsection stating the intentional semantics: shared memory (any
|
||
user's history can inform any bot's RAG context), shared bot namespace (creating an existing
|
||
name replaces it), any member may use any bot.
|
||
- Document input limits and cooldowns (P1-T8/P1-T9 values).
|
||
- Remove `COMPLETION_*` from the env table (P2-T6).
|
||
- AGENTS.md: update code map to new layout; retire the fixed gotchas (hardcoded sys.path,
|
||
two live tests → now `-m live`); keep the portaudio/.env notes.
|
||
- Done when: [ ] README tree matches `git ls-files` (ignore dotfiles/binaries); every `@bot.command`
|
||
name appears in the README; AGENTS.md gotcha list only contains still-true items.
|
||
|
||
#### Phase 3 verification & success criteria
|
||
|
||
- [ ] Fresh-clone hermetic run green with no network (P3-T1 command).
|
||
- [ ] A failing lint/type/test on a PR is blocked in CI (P3-T5).
|
||
- [ ] Container runs non-root (P3-T6).
|
||
- [ ] README/AGENTS.md accurate (P3-T8).
|
||
- [ ] Full gate set green.
|
||
|
||
---
|
||
|
||
### PHASE 4 — Optimization & Polish
|
||
|
||
**Goal:** Flatten RAG cost at scale, remove per-call allocations, pick off residual nits.
|
||
Only now — optimizations land in the final shapes.
|
||
|
||
#### P4-T1 — Precomputed vector norms (completes finding 6.1)
|
||
|
||
- File: `vibe_bot/database.py`.
|
||
- Store the L2 norm alongside the embedding (new column `norm REAL` on
|
||
`message_embeddings`, backfilled on migration via numpy over existing blobs) so
|
||
`search_similar_messages` computes `query_norm * (Q @ V.T) / norms` with one matrix multiply
|
||
and no per-vector renormalization.
|
||
- Benchmark script (add `scripts/bench_rag.py`, not committed to the package): seed N messages
|
||
(1k/5k), measure p95 of `get_conversation_context`; record before/after numbers in this file
|
||
next to the task.
|
||
- Done when: [ ] p95 at 5k ≤ p95 at 1k × 1.2 (flat-ish); suite green. (If `sqlite-vec` is
|
||
considered, only then — default is to stay on pure SQLite.)
|
||
|
||
#### P4-T2 — Client reuse verification
|
||
|
||
- Profile/trace a full `!doodlebob` + chat turn: assert exactly one HTTP session per endpoint
|
||
(httpx/AsyncOpenAI client identity stable; no new `AsyncOpenAI(...)` after startup).
|
||
- Done when: [ ] verified (add a unit test asserting `get_chat_client() is get_chat_client()`);
|
||
no per-call construction in traces.
|
||
|
||
#### P4-T3 — Remaining caches & token economy
|
||
|
||
- Confirm `VOICE_LANGUAGES` dict lookup (P2-T4) and `app.bot_cache` (P2-T5) are actually used on
|
||
the hot path (add asserts/measurements in tests if trivial).
|
||
- Verify layout-selection call uses `max_tokens=2` and verify-call a bounded small value
|
||
(already set in P2-T4 — confirm and measure end-to-end `doodlebob` latency before/after;
|
||
record the number next to this task).
|
||
- Done when: [ ] latency numbers recorded; no regressions in suite.
|
||
|
||
#### P4-T4 — Polish sweep
|
||
|
||
- [ ] `!retcon` output size: match source aspect ratio instead of hardcoded `768x768`
|
||
(revisit the original "keep generation time down" rationale with the configurable
|
||
`IMAGE_GEN_SIZE_*` knobs; keep 768 if the endpoint is slow — document the decision in
|
||
`llm_client.py` docstring).
|
||
- [ ] Remove `DEFAULT_VOICE`/`DEFAULT_SPEED` duplication between `tts.py` and `config.py`
|
||
(single source in `config.py`; `tts.py` defaults reference it).
|
||
- [ ] Delete any leftover `# noqa: ASYNC210` (should be gone after P1-T2/P2-T4).
|
||
- [ ] Remove the two `type: ignore[import-untyped]` in `tts.py` if `kokoro-tts`/`soundfile`
|
||
stubs now exist; otherwise keep and add a comment that stubs are upstream-missing.
|
||
- [ ] Sweep for dead code: `get_recent_messages` (has callers? if not, delete or keep with a
|
||
docstring noting it's a utility), `deactivate_custom_bot` (unused? wire into delete or
|
||
delete), unused test fixtures.
|
||
- Done when: [ ] all sub-items checked; `ruff` clean (no unused imports warnings suppressed).
|
||
|
||
#### Phase 4 verification & success criteria
|
||
|
||
- [ ] RAG p95 flat 1k→5k (P4-T1 numbers recorded).
|
||
- [ ] Embedding API calls per chat turn == 2 (user + query) — assert via mock call count in a
|
||
service test.
|
||
- [ ] `!doodlebob` latency improved vs Phase 1 baseline (numbers recorded).
|
||
- [ ] Full gate set green; no module > ~400 lines; zero dead fixtures/imports.
|
||
|
||
---
|
||
|
||
## 4. Final Acceptance (end of Phase 4)
|
||
|
||
- [ ] `uv run pytest vibe_bot/tests/ -v` green, hermetic (no network).
|
||
- [ ] `uv run ruff check vibe_bot/` / `mypy` / `pyright` / `black --check` all green.
|
||
- [ ] CI green on a PR (lint+test job gates merge).
|
||
- [ ] Container builds and runs as non-root.
|
||
- [ ] README + AGENTS.md match the code.
|
||
- [ ] No Critical or High finding from §2 is open (accepted-by-design items 1.1/1.2/5.1 have
|
||
their scoped residuals complete and are documented in the README).
|