Methodology

The v2 (promptfoo) benchmark of record. Every row is read directly from promptfoo/leaderboard/leaderboard.jsonl — this site does no scoring of its own.

Overall score

Scores are on a 1–5 scale. Overall is the gameplay-token-weighted mean over the categories a candidate was evaluated on, with weights renormalised to those present. Weights track the real per-minute token volume the engine spends per inference category:

CategoryWeightMeasures
intent18%tool/label selection — deterministic match
dialogue28%in-character reply quality (5 axes); folds in multiturn
reaction10%in-character reaction to events
simulation40%plausible world simulation (tier2 + tier3)
gaeilge4%Irish-language fluency within dialogue

Confidence intervals & rank

Each score carries a 95% bootstrap CI (2000 resamples, deterministic seed) over the per-item scores. Ranking isCI-aware, following the LMArena scheme: a model's rank rises by one for every model whose lower CI exceeds this model'supper CI. Models no one dominates by CI share a rank(a tie, marked ) — overlapping CIs mean the gap is not statistically real.

Cost, value & the frontier

$/game-hour prices the candidate's per-Mtok rate against the normal-play token profile (a five-minute vLLM-MLX demo run, gameplay categories only). Value is overall ÷ $/game-hour; free / flat-rate providers carry no per-token cost and rank by throughput. The efficiency frontier (dashed line on the scatter, ◆ in the table) is the Pareto set: models no other model beats on both quality and cost. Cost tier (free / budget / mid / premium) comes from the enumerated candidate catalog.

Provenance

Judge model (pinned): claude-sonnet-4-6. Dataset merkle:12b7a39f0cffe796ae5dcd78129ee597c558f9feaa0dfdb11d297e16b9601cf8. Every row stamps the judge model + dataset merkle, so scores stay comparable across re-runs on the same frozen dataset.