Which model should run the Rundale world?

The benchmark of record for LLMs as the Rundale engine's NPC brain — scored on in-character dialogue, reaction, world simulation, intent, and Irish, then priced against real gameplay token volume.

judge claude-sonnet-4-6dataset 12b7a39f0cff8 candidatesupdated 2026-06-15T19:00:00Z
Top overall
Gemma 4 26B A4B (MLX 4-bit, thinking)
4.08 / 5
Best dialogue
Gemma 4 26B A4B (MLX 4-bit, thinking)
4.08 / 5
Fastest
Gemma 4 26B A4B (MLX 4-bit, thinking)
3100 ms p50

Quality vs cost

models up-and-to-the-left win — high quality, low cost
Dashed line = efficiency frontier (best quality at each cost). Dot colour = cost tier.
3.03.54.04.5$ / game-hour (log)overall quality (1–5)Gemma 4 26B A4B (…

Leaderboard

scores 1–5 · ±95% bootstrap CI · CI-aware rank
Tier
Best for
8 / 8
#ModelOverall ↓dial$/hrValuep50p95
1⁼Google Gemma 4 26B A4B (MLX 4-bit, thinking) free 4.08 3.98–4.184.08$0.0000free31004650
1⁼Google Gemma 4 31B (GGUF, llama.cpp) free 3.95 3.85–4.053.95$0.0000free2150032250
1⁼Google Gemma 4 26B A4B (GGUF, llama.cpp) free 3.94 3.84–4.043.94$0.0000free32004800
2⁼QWen Qwen3.6 27B (MLX 4-bit) free 3.84 3.74–3.943.84$0.0000free1340020100
2⁼QWen Qwen3.5 27B (MLX 4-bit) free 3.78 3.68–3.883.78$0.0000free1620024300
4⁼QWen Qwen2.5 14B Instruct (MLX 4-bit, baseline) free 3.70 3.60–3.803.70$0.0000free45006750
4⁼Google Gemma 4 12B (GGUF, llama.cpp) free 3.68 3.58–3.783.68$0.0000free55008250
7QWen Qwen3.5 9B (MLX 4-bit) free 3.49 3.39–3.593.49$0.0000free46006900