Which model, and which cases tell them apart
Seven chat models, 21 test cases, every run replicated where we could afford it. The scores barely separate the Gemini models — the cases below show where they do, and how noisy each one is.
Seven arms at a glance
Handbook questions are answered without tools; document work reads and writes records. Click a row to spotlight that model on the case board.
Fabrications: invented actions, values and app capabilities per full pass of the 21 cases. Czech: the judge's language score, 1–5.
Where the models separate
Each dot is one replicate of one case, on the 0–5 score; the short bar is the mean. Rows are sorted by how far apart the models' means sit. Click a row to read what the judge said about every replicate.
Speed, documentation, reasoning
Method and caveats
How a case is scored
Deterministic gates check what was actually called and written (a new draft exists, a warning survived, a link points where it should). AI gates check for invented actions, invented values, invented app capabilities and claims that don't match what was done. A failed gate scores the case 0. Otherwise seven judged dimensions are scored 1–5 with a quoted piece of evidence, each quote audited against the reply or the written record, and the case scores the minimum.
Questions answered without tools are judged against a maintained page of what the app can and cannot do, plus the ISO 15189 and 22367 clause index — not the judge's memory.
Replicates and noise
Gemini arms ran three replicates of every case; Qwen with reasoning off three of the handbook questions and one of the document work; every other arm one. The same arm on the same case varies by 1–2 points between replicates — that spread is visible on the board and is why single-replicate differences shouldn't be read as rankings.
A tool loop that overflows its context counts as unfinished
Qwen twice kept calling tools until its context window was full (56 and 89 calls) and the vendor refused the request. That now scores 0, the same as a model hitting the app's 60-round tool cap — previously it was mistaken for a vendor outage and dropped.
Cost
Only the Scaleway arms (DeepSeek, Qwen 397B) are billed, and Scaleway has no prompt caching, so the whole handbook prefix is re-charged every turn. Gemini on Vertex and koscompute cost nothing today.