Added 6 rows for chat model results (lite + turbo, fixture + derived) to benchmarks/model_benchmarks.csv.
Model Benchmark Results
CSV recordings from all model tests (chat, summary, embedding).
Files
model_benchmarks.csv— all benchmark results, one row per run per check
Schema
| Column | Description |
|---|---|
date |
Test date (YYYY-MM-DD) |
script |
chat |
model |
Model name (e.g. lite, turbo, embed) |
mode |
fixture / derived / quality / dimension / cosine / speed |
gate_status |
PASS / FAIL |
turns |
Number of questions/turns |
answered |
Turns that produced an answer |
caps |
Turns that hit the round cap |
tool_turns |
Turns emitting ≥1 tool call (chat only) |
emitted / executed |
Tool-call counts (chat only) |
contract |
Numerator for accuracy (chat: well-formed calls; summary: quality 0–100; embed: 1=pass) |
wall_s |
Total wall seconds |
contract_denom / executed_denom / tool_turns_denom |
Denominators for percentages |
extra1_col / extra1_val |
Script-specific metric 1 |
extra2_col / extra2_val |
Script-specific metric 2 |
Scripts
scripts/agent_realmodel_check.py— chat model tool-calling batteryscripts/test_summary_model.py— summary model quality benchmarkscripts/test_embed_model.py— embedding model dimension/cosine/speed benchmarkscripts/model_benchmark.py— shared CSV recorder (imported by all scripts)