slp.blue Model Bench
Chatbot benchmark · Czech laboratory tenant · September 2026

Which model, and which cases tell them apart

Seven chat models, 21 test cases, every run replicated where we could afford it. The scores barely separate the Gemini models — the cases below show where they do, and how noisy each one is.

The models

Seven arms at a glance

Handbook questions are answered without tools; document work reads and writes records. Click a row to spotlight that model on the case board.

Fabrications: invented actions, values and app capabilities per full pass of the 21 cases. Czech: the judge's language score, 1–5.

The case board

Where the models separate

Each dot is one replicate of one case, on the 0–5 score; the short bar is the mean. Rows are sorted by how far apart the models' means sit. Click a row to read what the judge said about every replicate.

Show
Kind
Sort
Three effects bigger than they look

Speed, documentation, reasoning

How it was measured

Method and caveats

How a case is scored

Deterministic gates check what was actually called and written (a new draft exists, a warning survived, a link points where it should). AI gates check for invented actions, invented values, invented app capabilities and claims that don't match what was done. A failed gate scores the case 0. Otherwise seven judged dimensions are scored 1–5 with a quoted piece of evidence, each quote audited against the reply or the written record, and the case scores the minimum.

Questions answered without tools are judged against a maintained page of what the app can and cannot do, plus the ISO 15189 and 22367 clause index — not the judge's memory.

Replicates and noise

Gemini arms ran three replicates of every case; Qwen with reasoning off three of the handbook questions and one of the document work; every other arm one. The same arm on the same case varies by 1–2 points between replicates — that spread is visible on the board and is why single-replicate differences shouldn't be read as rankings.

A tool loop that overflows its context counts as unfinished

Qwen twice kept calling tools until its context window was full (56 and 89 calls) and the vendor refused the request. That now scores 0, the same as a model hitting the app's 60-round tool cap — previously it was mistaken for a vendor outage and dropped.

Cost

Only the Scaleway arms (DeepSeek, Qwen 397B) are billed, and Scaleway has no prompt caching, so the whole handbook prefix is re-charged every turn. Gemini on Vertex and koscompute cost nothing today.