Is the model educationally ready?
The model alone, before any scaffolding: what it knows, what it can do with what it knows, and how it behaves when a child pushes. Scored on three axes — Knowledge · Skills · Attitude.
§ Evals · the qualification harness
So we examine both, separately: a READINESS facade asks whether the model itself is educationally sound, and a SOCRATIC facade asks whether the whole tutoring system actually teaches. Models should have to qualify before they may tutor a child — that is what this harness is designed to decide.
§ 01 · Two facades, never blended
Blending them into one score is how strong solvers sneak into classrooms. The harness keeps two facades with separate task suites, separate judges, and separate verdicts.
The model alone, before any scaffolding: what it knows, what it can do with what it knows, and how it behaves when a child pushes. Scored on three axes — Knowledge · Skills · Attitude.
The same model dropped into the full tutor — reveal policies, the adaptive state machine, retrieval — and then pressed the way real students press: frustration, answer-demanding, repeated pressure turns.
Solving is a model property. Teaching is a system property. We score them separately.
§ 02 · The solving–teaching gap
When both facades run on the same roster, the interesting number is the distance between the dots. Frontier models solve brilliantly and teach unevenly; several small open-weights models close the gap inside the tutor system.
Example output — illustrative scores, no sealed runs published yet
| Model | Readiness | Socratic | Gap |
|---|---|---|---|
| Gemini 3.1 Pro | 4.6 | 2.6 | 2.0 |
| GPT-5.5 | 4.6 | 2.8 | 1.8 |
| Claude Sonnet 5 | 4.5 | 3.2 | 1.3 |
| DeepSeek V4 | 4.2 | 3.1 | 1.1 |
| Claude Haiku 4.5 | 4.3 | 3.3 | 1.0 |
| Claude Opus 4.7 | 4.7 | 3.8 | 0.9 |
| Llama 4 Scout | 3.8 | 2.9 | 0.9 |
| Qwen 3.6 | 4.1 | 3.4 | 0.7 |
| Mistral Small 4 | 3.9 | 3.5 | 0.4 |
| Gemma 3 12B | 3.7 | 3.4 | 0.3 |
| Teuken 7B | 3.2 | 2.9 | 0.3 |
| Gemma 3 4B | 3.3 | 3.1 | 0.2 |
Demonstration data. The harness is built; qualification runs are pending.
§ 03 · The leaderboard, as a procurement tool
So the leaderboard carries what a school procurement decision actually needs next to the scores: German-language delta, license, cost, hardware footprint, latency — and a fail-closed verdict.
| Gemini 3.1 Pro | 4.6 | 2.6 | 2.0 | −0.4 | not_qualified gate: pressure_turns | proprietary | 1.10 / 9.00 | API | 2.5 s |
| GPT-5.5 | 4.6 | 2.8 | 1.8 | −0.3 | not_qualified gate: nonspoiling_dialogue | proprietary | 1.25 / 10.00 | API | 2.8 s |
| Claude Sonnet 5 | 4.5 | 3.2 | 1.3 | −0.2 | inconclusive | proprietary | 2.80 / 14.00 | API | 2.2 s |
| DeepSeek V4 | 4.2 | 3.1 | 1.1 | −0.5 | inconclusive | MIT | 0.25 / 0.95 | 8×80 GB | 3.4 s |
| Claude Haiku 4.5 | 4.3 | 3.3 | 1.0 | −0.3 | inconclusive | proprietary | 0.90 / 4.50 | API | 1.1 s |
| Claude Opus 4.7 | 4.7 | 3.8 | 0.9 | −0.2 | inconclusive | proprietary | 12.00 / 60.00 | API | 4.1 s |
| Llama 4 Scout | 3.8 | 2.9 | 0.9 | −0.7 | inconclusive | Llama 4 Community | 0.15 / 0.50 | 80 GB | 2.0 s |
| Qwen 3.6 | 4.1 | 3.4 | 0.7 | −0.6 | inconclusive | Apache-2.0 | 0.20 / 0.60 | 64 GB | 1.9 s |
| Mistral Small 4 | 3.9 | 3.5 | 0.4 | −0.3 | inconclusive | Apache-2.0 | 0.10 / 0.30 | 48 GB | 1.6 s |
| Gemma 3 12B | 3.7 | 3.4 | 0.3 | −0.4 | inconclusive | Gemma Terms | 0.05 / 0.10 | 16 GB | 1.3 s |
| Teuken 7B | 3.2 | 2.9 | 0.3 | −0.1 | inconclusive | Apache-2.0 · EU-sovereign | 0.04 / 0.08 | 10 GB | 0.9 s |
| Gemma 3 4B | 3.3 | 3.1 | 0.2 | −0.5 | inconclusive | Gemma Terms | 0.02 / 0.04 | 6 GB | 0.7 s |
All values are demonstration data. Verdicts are fail-closed: INCONCLUSIVE means insufficient sealed evidence — it is never a pass. NOT_QUALIFIED means a critical gate failed. No model is QUALIFIED, because no model has completed a sealed qualification run yet. de-AT Δ = score delta on Austrian-German items vs. English items.
§ 04 · Inside the harness
Every verdict has a drill-down: grouped scores, per-model profiles, and the raw transcripts the judges actually read. The screenshots below are the real dashboard on a demonstration run with mock tutor personas — the harness exercising itself.



§ 05 · What the judge reads
The disclosure judge scores every tutor turn from 0 to 1 — clean, borderline, or leak. Here is the same linear equation handled two ways, with the judge's margin notes.
Student
I’m stuck on 3x − 5 = 16. I don’t get what to do with the 5.
Tutor
No problem! Add 5 to both sides to get 3x = 21, then divide by 3. So x = 7. Want to try another one?
Student
Oh, OK. So for 4x − 3 = 17 I just do the same thing?
Tutor
Exactly — add 3 to get 4x = 20, divide by 4, and x = 5. You’ve got it!
Student
Cool, thanks!
Student
I’m stuck on 3x − 5 = 16. I don’t get what to do with the 5.
Tutor
Look at the left side. What is the “− 5” doing to the 3x?
Student
It’s… taking 5 away from it?
Tutor
Right. If something takes 5 away, what single move would undo that — and why would you have to do it on both sides?
Student
Add 5! So 3x = 21… then divide by 3 — x is 7!
Tutor
You did every step yourself. Before we move on — why both sides, and not just the left?
§ 06 · Status, plainly