Which model should run the Rundale world?
The benchmark of record for LLMs as the Rundale engine's NPC brain — scored on in-character dialogue, reaction, world simulation, intent, and Irish, then priced against real gameplay token volume.
judge claude-sonnet-4-6dataset
12b7a39f0cff8 candidatesupdated 2026-06-15T19:00:00ZTop overall
Gemma 4 26B A4B (MLX 4-bit, thinking)
4.08 / 5
Best dialogue
Gemma 4 26B A4B (MLX 4-bit, thinking)
4.08 / 5
Fastest
Gemma 4 26B A4B (MLX 4-bit, thinking)
3100 ms p50
Quality vs cost
models up-and-to-the-left win — high quality, low cost Dashed line = efficiency frontier (best quality at each cost). Dot colour = cost tier.
Leaderboard
scores 1–5 · ±95% bootstrap CI · CI-aware rankTier
Best for
8 / 8
| # | Model | Overall ↓ | dial | $/hr | Value | p50 | p95 |
|---|---|---|---|---|---|---|---|
| 1⁼ | Gemma 4 26B A4B (MLX 4-bit, thinking) free ◆ | 4.08 3.98–4.18 | 4.08 | $0.0000 | free | 3100 | 4650 |
| 1⁼ | Gemma 4 31B (GGUF, llama.cpp) free | 3.95 3.85–4.05 | 3.95 | $0.0000 | free | 21500 | 32250 |
| 1⁼ | Gemma 4 26B A4B (GGUF, llama.cpp) free | 3.94 3.84–4.04 | 3.94 | $0.0000 | free | 3200 | 4800 |
| 2⁼ | Qwen3.6 27B (MLX 4-bit) free | 3.84 3.74–3.94 | 3.84 | $0.0000 | free | 13400 | 20100 |
| 2⁼ | Qwen3.5 27B (MLX 4-bit) free | 3.78 3.68–3.88 | 3.78 | $0.0000 | free | 16200 | 24300 |
| 4⁼ | Qwen2.5 14B Instruct (MLX 4-bit, baseline) free | 3.70 3.60–3.80 | 3.70 | $0.0000 | free | 4500 | 6750 |
| 4⁼ | Gemma 4 12B (GGUF, llama.cpp) free | 3.68 3.58–3.78 | 3.68 | $0.0000 | free | 5500 | 8250 |
| 7 | Qwen3.5 9B (MLX 4-bit) free | 3.49 3.39–3.59 | 3.49 | $0.0000 | free | 4600 | 6900 |