Dialogue qualification funnel
Every cloud profile tested through Rundale's production dialogue path. Structural reliability, guard use, and request reliability are hard gates; player-facing latency ranks the profiles that survive.
Screening policy
Reliability is gated. Speed is ranked among the survivors.Blinded dialogue quality ranking
Blinded multi-family consensus · self-family scores excluded · 18 retained production outputs per judge| Rank | Model | Verdict | Overall | Minimum critical axis | Hard failures | Unusable outputs | Speed rank | Judge cost |
|---|---|---|---|---|---|---|---|---|
| #1 | openai/gpt-5.6-luna | qualified | 4.56 | 4.44 | 0 | 0 | #11 | $0.909 |
| #2 | gemini-3.7-flash | qualified | 4.54 | 4.47 | 0 | 0 | #16 | $1.407 |
| #3 | google/gemini-3.6-flash | qualified | 4.45 | 4.08 | 0 | 0 | #9 | $1.407 |
| #4 | google/gemini-3.7-flash | qualified | 4.44 | 4.06 | 0 | 0 | #15 | $1.463 |
| #5 | openai/gpt-5.6-luna | rejected | 4.43 | 4.00 | 1 | 1 | #10 | $1.798 |
| #6 | google/gemini-3.7-flash | qualified | 4.42 | 4.11 | 0 | 0 | #2 | $1.465 |
| #7 | google/gemini-3.5-flash-lite | qualified | 4.41 | 4.17 | 0 | 0 | #6 | $1.452 |
| #8 | gemini-3.7-flash | qualified | 4.34 | 4.00 | 0 | 0 | #4 | $1.395 |
| #9 | moonshotai/kimi-k2-0905 | qualified | 4.34 | 3.89 | 0 | 0 | #5 | $1.791 |
| #10 | gemini-3.7-flash | qualified | 4.33 | 4.08 | 0 | 0 | #14 | $1.396 |
| #11 | google/gemini-3.7-flash | qualified | 4.33 | 4.03 | 0 | 0 | #13 | $1.416 |
| #12 | openai/gpt-4.1-mini | rejected | 3.77 | 2.89 | 6 | 0 | #3 | $1.780 |
| #13 | qwen/qwen2.5-vl-72b-instruct | rejected | 3.71 | 3.00 | 9 | 0 | #17 | $1.715 |
Qualified cloud speed ranking
Lower speed index is better. Quality judgment still determines whether dialogue is usable.| Rank | Model | Reasoning | Max tokens | Warm TTFT p95 | Warm completion p95 | Speed index | Error rate |
|---|---|---|---|---|---|---|---|
| #1 | moonshotai/kimi-k2-0905:nitro | provider default | 768 | 1336 ms | 1344 ms | 1338 ms | 0.0% |
| #2 | google/gemini-3.7-flash | low | 4096 | 1843 ms | 2712 ms | 2060 ms | 0.0% |
| #3 | openai/gpt-4.1-mini | provider default | 768 | 2416 ms | 5716 ms | 3241 ms | 0.0% |
| #4 | gemini-3.7-flash | low | 4096 | 3278 ms | 3655 ms | 3372 ms | 0.0% |
| #5 | moonshotai/kimi-k2-0905 | provider default | 768 | 2950 ms | 4790 ms | 3410 ms | 0.0% |
| #6 | google/gemini-3.5-flash-lite | low | 2048 | 3597 ms | 4007 ms | 3700 ms | 0.0% |
| #7 | x-ai/grok-4.3 | provider default | 768 | 942 ms | 12485 ms | 3828 ms | 0.0% |
| #8 | x-ai/grok-4.3:nitro | provider default | 768 | 2022 ms | 9654 ms | 3930 ms | 0.0% |
| #9 | google/gemini-3.6-flash | low | 768 | 4548 ms | 4967 ms | 4653 ms | 0.0% |
| #10 | openai/gpt-5.6-luna | none | 768 | 4870 ms | 6175 ms | 5196 ms | 0.0% |
| #11 | openai/gpt-5.6-luna | high | 4096 | 5496 ms | 6365 ms | 5713 ms | 0.0% |
| #12 | z-ai/glm-5.2 | provider default | 768 | 5896 ms | 8454 ms | 6536 ms | 0.0% |
| #13 | google/gemini-3.7-flash | medium | 4096 | 6722 ms | 7402 ms | 6892 ms | 0.0% |
| #14 | gemini-3.7-flash | medium | 4096 | 7145 ms | 7336 ms | 7193 ms | 0.0% |
| #15 | google/gemini-3.7-flash | high | 4096 | 7992 ms | 8514 ms | 8123 ms | 0.0% |
| #16 | gemini-3.7-flash | high | 4096 | 8506 ms | 8663 ms | 8545 ms | 0.0% |
| #17 | qwen/qwen2.5-vl-72b-instruct | provider default | 768 | 6361 ms | 20904 ms | 9997 ms | 0.0% |
| #18 | openai/gpt-5.6-sol | none | 768 | 11374 ms | 13470 ms | 11898 ms | 0.0% |
| #19 | deepseek/deepseek-v4-flash-0731 | medium | 4096 | 11906 ms | 14207 ms | 12481 ms | 0.0% |
| #20 | z-ai/glm-4.7 | off | 768 | 2593 ms | 48244 ms | 14006 ms | 0.0% |
| #21 | deepseek/deepseek-v4-flash-0731 | max | 4096 | 99967 ms | 104620 ms | 101130 ms | 0.0% |
Production screening
Select a run to inspect the complete evidence trail| Date | Model | Status | Preflight | Guards | Quality | Speed rank | Warm TTFT p95 | Warm completion p95 | Decision | |
|---|---|---|---|---|---|---|---|---|---|---|
| 2026-08-14 | gemini-3.7-flash | qualified | 12/12100.0% | 8.3% | 4.33 | #14 | 7145 ms | 7336 ms | quality passed; quality rank #10 of 13; speed rank #14 of 21 | InspectRun qualified2026-08-14/gemini-3.7-flash-google-native-medium-reasoningDecisionquality passed; quality rank #10 of 13; speed rank #14 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-medium-reasoning-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-medium-reasoning-perf.json Quality judgment
openai-sol-high 4.25 · passhash retainedanthropic-sonnet-low 4.40 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-14__gemini-3.7-flash-google-native-medium-reasoning.json |
| 2026-08-14 | gemini-3.7-flash | qualified | 12/12100.0% | 0.0% | 4.34 | #4 | 3278 ms | 3655 ms | quality passed; quality rank #8 of 13; speed rank #4 of 21 | InspectRun qualified2026-08-14/gemini-3.7-flash-google-native-low-reasoning-attempt-2Decisionquality passed; quality rank #8 of 13; speed rank #4 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-low-reasoning-attempt-2-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-low-reasoning-attempt-2-perf.json Quality judgment
openai-sol-high 4.31 · passhash retainedanthropic-sonnet-low 4.37 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-14__gemini-3.7-flash-google-native-low-reasoning-attempt-2.json |
| 2026-08-14 | gemini-3.7-flash | qualified | 12/12100.0% | 0.0% | 4.54 | #16 | 8506 ms | 8663 ms | quality passed; quality rank #2 of 13; speed rank #16 of 21 | InspectRun qualified2026-08-14/gemini-3.7-flash-google-native-high-reasoningDecisionquality passed; quality rank #2 of 13; speed rank #16 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-high-reasoning-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-high-reasoning-perf.json Quality judgment
openai-sol-high 4.66 · passhash retainedanthropic-sonnet-low 4.42 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-14__gemini-3.7-flash-google-native-high-reasoning.json |
| 2026-08-13 | google/gemini-3.7-flash | qualified | 12/12100.0% | 0.0% | 4.33 | #13 | 6722 ms | 7402 ms | quality passed; quality rank #11 of 13; speed rank #13 of 21 | InspectRun qualified2026-08-13/gemini-3.7-flash-medium-reasoningDecisionquality passed; quality rank #11 of 13; speed rank #13 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-medium-reasoning-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-medium-reasoning-perf.json Quality judgment
openai-sol-high 4.30 · passhash retainedanthropic-sonnet-low 4.35 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-13__gemini-3.7-flash-medium-reasoning.json |
| 2026-08-13 | google/gemini-3.7-flash | qualified | 12/12100.0% | 8.3% | 4.42 | #2 | 1843 ms | 2712 ms | quality passed; quality rank #6 of 13; speed rank #2 of 21 | InspectRun qualified2026-08-13/gemini-3.7-flash-low-reasoningDecisionquality passed; quality rank #6 of 13; speed rank #2 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-low-reasoning-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-low-reasoning-perf.json Quality judgment
openai-sol-high 4.41 · passhash retainedanthropic-sonnet-low 4.42 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-13__gemini-3.7-flash-low-reasoning.json |
| 2026-08-13 | google/gemini-3.7-flash | qualified | 12/12100.0% | 0.0% | 4.44 | #15 | 7992 ms | 8514 ms | quality passed; quality rank #4 of 13; speed rank #15 of 21 | InspectRun qualified2026-08-13/gemini-3.7-flash-high-reasoningDecisionquality passed; quality rank #4 of 13; speed rank #15 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-high-reasoning-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-high-reasoning-perf.json Quality judgment
openai-sol-high 4.49 · passhash retainedanthropic-sonnet-low 4.39 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-13__gemini-3.7-flash-high-reasoning.json |
| 2026-08-08 | moonshotai/kimi-k2.5:nitro | stopped | 2/2100.0% | 0.0% | — | — | — | — | stopped after 2/12 calls | InspectRun stopped2026-08-08/kimi-k2.5-nitroDecisionstopped after 2/12 calls
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/kimi-k2.5-nitro-preflight.partial.jsonl Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 2 individual API calls2 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__kimi-k2.5-nitro.json |
| 2026-08-08 | deepseek/deepseek-v4-pro | rejected | 11/11100.0% | 18.2% | — | — | — | — | guard intervention rate (early stop) | InspectRun rejected2026-08-08/deepseek-v4-pro-no-reasoningDecisionguard intervention rate (early stop)
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-pro-no-reasoning-preflight.partial.jsonl Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 11 individual API calls11 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__deepseek-v4-pro-no-reasoning.json |
| 2026-08-08 | deepseek/deepseek-v4-flash-0731 | rejected | 1/250.0% | 0.0% | — | — | — | — | structural reliability (early stop) | InspectRun rejected2026-08-08/deepseek-v4-flash-0731-max-reasoningDecisionstructural reliability (early stop)
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-max-reasoning-preflight.partial.jsonl Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 2 individual API calls2 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__deepseek-v4-flash-0731-max-reasoning.json |
| 2026-08-08 | moonshotai/kimi-k2.5:nitro | rejected | 12/12100.0% | 16.7% | — | — | — | — | guard intervention rate | InspectRun rejected2026-08-08/kimi-k2.5-nitro-no-reasoningDecisionguard intervention rate
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/kimi-k2.5-nitro-no-reasoning-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__kimi-k2.5-nitro-no-reasoning.json |
| 2026-08-08 | moonshotai/kimi-k2-0905:nitro | rejected | 12/12100.0% | 16.7% | — | — | — | — | guard intervention rate | InspectRun rejected2026-08-08/kimi-k2-0905-nitro-fundedDecisionguard intervention rate
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/kimi-k2-0905-nitro-funded-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__kimi-k2-0905-nitro-funded.json |
| 2026-08-08 | openai/gpt-5.6-sol | needs judgment | 12/12100.0% | 0.0% | — | #18 | 11374 ms | 13470 ms | awaiting independent judges (0/2); cloud speed rank #18 of 21 | InspectRun needs judgment2026-08-08/gpt-5.6-sol-no-reasoningDecisionawaiting independent judges (0/2); cloud speed rank #18 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-sol-no-reasoning-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-sol-no-reasoning-perf.json Quality judgment
openai-sol-high · excluded 4.17 · passhash retainedReview 49 individual API calls12 preflight · 19 performance · 18 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__gpt-5.6-sol-no-reasoning.json |
| 2026-08-08 | openai/gpt-5.6-luna | quality rejected | 12/12100.0% | 8.3% | 4.43 | #10 | 4870 ms | 6175 ms | quality screen failed; quality rank #5 of 13; speed rank #10 of 21 | InspectRun quality rejected2026-08-08/gpt-5.6-luna-no-reasoningDecisionquality screen failed; quality rank #5 of 13; speed rank #10 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-no-reasoning-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-no-reasoning-perf.json Quality judgment
openai-sol-high · excluded 4.17 · passhash retainedanthropic-sonnet-low 4.14 · failhash retainedgoogle-pro-high 4.71 · failhash retainedReview 85 individual API calls12 preflight · 19 performance · 54 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__gpt-5.6-luna-no-reasoning.json |
| 2026-08-08 | openai/gpt-5.6-luna | rejected | 12/12100.0% | 16.7% | — | — | — | — | guard intervention rate | InspectRun rejected2026-08-08/gpt-5.6-luna-max-reasoning-4096Decisionguard intervention rate
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-max-reasoning-4096-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__gpt-5.6-luna-max-reasoning-4096.json |
| 2026-08-08 | openai/gpt-5.6-luna | qualified | 12/12100.0% | 0.0% | 4.56 | #11 | 5496 ms | 6365 ms | quality passed; quality rank #1 of 13; speed rank #11 of 21 | InspectRun qualified2026-08-08/gpt-5.6-luna-high-reasoning-4096Decisionquality passed; quality rank #1 of 13; speed rank #11 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-high-reasoning-4096-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-high-reasoning-4096-perf.json Quality judgment
anthropic-sonnet-low 4.27 · passhash retainedgoogle-pro-high 4.86 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__gpt-5.6-luna-high-reasoning-4096.json |
| 2026-08-08 | z-ai/glm-4.7 | needs judgment | 12/12100.0% | 0.0% | 3.69 | #20 | 2593 ms | 48244 ms | awaiting independent judges (1/2); cloud speed rank #20 of 21 | InspectRun needs judgment2026-08-08/glm-4.7-no-reasoningDecisionawaiting independent judges (1/2); cloud speed rank #20 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/glm-4.7-no-reasoning-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/glm-4.7-no-reasoning-perf.json Quality judgment
openai-sol-high 3.69 · failhash retainedReview 49 individual API calls12 preflight · 19 performance · 18 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__glm-4.7-no-reasoning.json |
| 2026-08-08 | google/gemini-3.6-flash | qualified | 12/12100.0% | 8.3% | 4.45 | #9 | 4548 ms | 4967 ms | quality passed; quality rank #3 of 13; speed rank #9 of 21 | InspectRun qualified2026-08-08/gemini-3.6-flash-low-reasoningDecisionquality passed; quality rank #3 of 13; speed rank #9 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.6-flash-low-reasoning-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.6-flash-low-reasoning-perf.json Quality judgment
openai-sol-high 4.44 · passhash retainedanthropic-sonnet-low 4.45 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__gemini-3.6-flash-low-reasoning.json |
| 2026-08-08 | google/gemini-3.5-flash-lite | invalid profile | 10/1283.3% | 0.0% | — | — | — | — | superseded: insufficient completion budget | InspectRun invalid profile2026-08-08/gemini-3.5-flash-lite-low-reasoningDecisionsuperseded: insufficient completion budget
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.5-flash-lite-low-reasoning-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 14 individual API calls12 preflight · 0 performance · 0 judgment · 2 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__gemini-3.5-flash-lite-low-reasoning.json |
| 2026-08-08 | google/gemini-3.5-flash-lite | qualified | 12/12100.0% | 0.0% | 4.41 | #6 | 3597 ms | 4007 ms | quality passed; quality rank #7 of 13; speed rank #6 of 21 | InspectRun qualified2026-08-08/gemini-3.5-flash-lite-low-reasoning-2048Decisionquality passed; quality rank #7 of 13; speed rank #6 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.5-flash-lite-low-reasoning-2048-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.5-flash-lite-low-reasoning-2048-perf.json Quality judgment
openai-sol-high 4.44 · passhash retainedanthropic-sonnet-low 4.37 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__gemini-3.5-flash-lite-low-reasoning-2048.json |
| 2026-08-08 | deepseek-v4-flash | rejected | 12/12100.0% | 16.7% | — | — | — | — | guard intervention rate | InspectRun rejected2026-08-08/deepseek-v4-flash-direct-high-reasoning-4096Decisionguard intervention rate
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-direct-high-reasoning-4096-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__deepseek-v4-flash-direct-high-reasoning-4096.json |
| 2026-08-08 | deepseek/deepseek-v4-flash-0731 | rejected | 12/12100.0% | 25.0% | — | — | — | — | guard intervention rate | InspectRun rejected2026-08-08/deepseek-v4-flash-0731-no-reasoningDecisionguard intervention rate
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-no-reasoning-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__deepseek-v4-flash-0731-no-reasoning.json |
| 2026-08-08 | deepseek/deepseek-v4-flash-0731 | needs judgment | 12/12100.0% | 0.0% | 4.37 | #19 | 11906 ms | 14207 ms | awaiting independent judges (1/2); cloud speed rank #19 of 21 | InspectRun needs judgment2026-08-08/deepseek-v4-flash-0731-medium-reasoning-4096Decisionawaiting independent judges (1/2); cloud speed rank #19 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-medium-reasoning-4096-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-medium-reasoning-4096-perf.json Quality judgment
anthropic-sonnet-low 4.37 · failhash retainedReview 49 individual API calls12 preflight · 19 performance · 18 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__deepseek-v4-flash-0731-medium-reasoning-4096.json |
| 2026-08-08 | deepseek/deepseek-v4-flash-0731 | needs adjudication | 12/12100.0% | 0.0% | 4.50 | #21 | 99967 ms | 104620 ms | judge disagreement requires adjudication; speed rank #21 of 21 | InspectRun needs adjudication2026-08-08/deepseek-v4-flash-0731-max-reasoning-4096Decisionjudge disagreement requires adjudication; speed rank #21 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-max-reasoning-4096-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-max-reasoning-4096-perf.json Quality judgment
anthropic-sonnet-low 4.28 · failhash retainedgoogle-pro-high 4.72 · passhash retainedReview 67 individual API calls12 preflight · 19 performance · 36 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__deepseek-v4-flash-0731-max-reasoning-4096.json |
| 2026-08-08 | deepseek/deepseek-v4-flash-0731 | rejected | 10/1283.3% | 25.0% | — | — | — | — | structural reliability | InspectRun rejected2026-08-08/deepseek-v4-flash-0731-low-reasoning-4096Decisionstructural reliability
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-low-reasoning-4096-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__deepseek-v4-flash-0731-low-reasoning-4096.json |
| 2026-08-08 | deepseek/deepseek-v4-flash-0731 | rejected | 12/12100.0% | 16.7% | — | — | — | — | guard intervention rate | InspectRun rejected2026-08-08/deepseek-v4-flash-0731-high-reasoning-4096Decisionguard intervention rate
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-high-reasoning-4096-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-08-08__deepseek-v4-flash-0731-high-reasoning-4096.json |
| 2026-07-27 | qwen/qwen2.5-vl-72b-instruct | quality rejected | 12/12100.0% | 0.0% | 3.71 | #17 | 6361 ms | 20904 ms | quality screen failed; quality rank #13 of 13; speed rank #17 of 21 | InspectRun quality rejected2026-07-27/qwen25-vl-72bDecisionquality screen failed; quality rank #13 of 13; speed rank #17 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/qwen25-vl-72b-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/qwen25-vl-72b-perf.json Quality judgment
openai-sol-high 3.71 · failhash retainedanthropic-sonnet-low 3.64 · failhash retainedgoogle-pro-high 3.89 · failhash retainedReview 85 individual API calls12 preflight · 19 performance · 54 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__qwen25-vl-72b.json |
| 2026-07-27 | nvidia/nemotron-3-ultra-550b-a55b | rejected | 4/1233.3% | 25.0% | — | — | — | — | structural reliability | InspectRun rejected2026-07-27/nemotron-3-ultraDecisionstructural reliability
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/nemotron-3-ultra-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__nemotron-3-ultra.json |
| 2026-07-27 | mistralai/mistral-medium-3.1 | rejected | 12/12100.0% | 41.7% | — | — | — | — | guard intervention rate | InspectRun rejected2026-07-27/mistral-medium-3.1Decisionguard intervention rate
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/mistral-medium-3.1-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__mistral-medium-3.1.json |
| 2026-07-27 | moonshotai/kimi-k2-0905 | qualified | 12/12100.0% | 0.0% | 4.34 | #5 | 2950 ms | 4790 ms | quality passed; quality rank #9 of 13; speed rank #5 of 21 | InspectRun qualified2026-07-27/kimi-k2-0905Decisionquality passed; quality rank #9 of 13; speed rank #5 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/kimi-k2-0905-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/kimi-k2-0905-perf.json Quality judgment
openai-sol-high 4.17 · passhash retainedanthropic-sonnet-low 4.34 · passhash retainedgoogle-pro-high 4.37 · passhash retainedReview 85 individual API calls12 preflight · 19 performance · 54 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__kimi-k2-0905.json |
| 2026-07-27 | moonshotai/kimi-k2-0905:nitro | needs adjudication | 12/12100.0% | 0.0% | 4.14 | #1 | 1336 ms | 1344 ms | judge disagreement requires adjudication; speed rank #1 of 21 | InspectRun needs adjudication2026-07-27/kimi-k2-0905-nitroDecisionjudge disagreement requires adjudication; speed rank #1 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/kimi-k2-0905-nitro-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/kimi-k2-0905-nitro-perf-expanded.json Quality judgment
openai-sol-high 3.92 · failhash retainedanthropic-sonnet-low 4.14 · failhash retainedgoogle-pro-high 4.76 · passhash retainedReview 115 individual API calls12 preflight · 49 performance · 54 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__kimi-k2-0905-nitro.json |
| 2026-07-27 | x-ai/grok-4.3 | needs adjudication | 12/12100.0% | 0.0% | 4.12 | #7 | 942 ms | 12485 ms | judge disagreement requires adjudication; speed rank #7 of 21 | InspectRun needs adjudication2026-07-27/grok-4.3Decisionjudge disagreement requires adjudication; speed rank #7 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/grok-4.3-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/grok-4.3-perf.json Quality judgment
openai-sol-high 3.13 · failhash retainedanthropic-sonnet-low 4.12 · passhash retainedgoogle-pro-high 4.21 · passhash retainedReview 85 individual API calls12 preflight · 19 performance · 54 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__grok-4.3.json |
| 2026-07-27 | x-ai/grok-4.3:nitro | needs adjudication | 12/12100.0% | 0.0% | 3.94 | #8 | 2022 ms | 9654 ms | judge disagreement requires adjudication; speed rank #8 of 21 | InspectRun needs adjudication2026-07-27/grok-4.3-nitroDecisionjudge disagreement requires adjudication; speed rank #8 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/grok-4.3-nitro-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/grok-4.3-nitro-perf.json Quality judgment
openai-sol-high 3.33 · failhash retainedanthropic-sonnet-low 3.94 · failhash retainedgoogle-pro-high 4.28 · passhash retainedReview 85 individual API calls12 preflight · 19 performance · 54 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__grok-4.3-nitro.json |
| 2026-07-27 | openai/gpt-oss-120b:nitro | rejected | 9/1275.0% | 0.0% | — | — | — | — | structural reliability | InspectRun rejected2026-07-27/gpt-oss-120b-nitroDecisionstructural reliability
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gpt-oss-120b-nitro-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__gpt-oss-120b-nitro.json |
| 2026-07-27 | openai/gpt-4o | rejected | 12/12100.0% | 16.7% | — | — | — | — | guard intervention rate | InspectRun rejected2026-07-27/gpt-4oDecisionguard intervention rate
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gpt-4o-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__gpt-4o.json |
| 2026-07-27 | openai/gpt-4.1-mini | quality rejected | 12/12100.0% | 0.0% | 3.77 | #3 | 2416 ms | 5716 ms | quality screen failed; quality rank #12 of 13; speed rank #3 of 21 | InspectRun quality rejected2026-07-27/gpt-4.1-miniDecisionquality screen failed; quality rank #12 of 13; speed rank #3 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gpt-4.1-mini-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gpt-4.1-mini-perf-expanded.json Quality judgment
openai-sol-high · excluded 3.66 · failhash retainedanthropic-sonnet-low 3.41 · failhash retainedgoogle-pro-high 4.13 · failhash retainedReview 115 individual API calls12 preflight · 49 performance · 54 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__gpt-4.1-mini.json |
| 2026-07-27 | z-ai/glm-5.2 | needs adjudication | 12/12100.0% | 0.0% | 4.18 | #12 | 5896 ms | 8454 ms | judge disagreement requires adjudication; speed rank #12 of 21 | InspectRun needs adjudication2026-07-27/glm-5.2Decisionjudge disagreement requires adjudication; speed rank #12 of 21
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/glm-5.2-preflight.json Performance evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/glm-5.2-perf.json Quality judgment
openai-sol-high 3.04 · failhash retainedanthropic-sonnet-low 4.18 · failhash retainedgoogle-pro-high 4.72 · passhash retainedReview 85 individual API calls12 preflight · 19 performance · 54 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__glm-5.2.json |
| 2026-07-27 | z-ai/glm-4.5v | rejected | 10/1283.3% | 33.3% | — | — | — | — | structural reliability | InspectRun rejected2026-07-27/glm-4.5vDecisionstructural reliability
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/glm-4.5v-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__glm-4.5v.json |
| 2026-07-27 | google/gemini-2.5-flash | rejected | 11/1291.7% | 8.3% | — | — | — | — | structural reliability | InspectRun rejected2026-07-27/gemini-2.5-flashDecisionstructural reliability
Request profile
Preflight evidence
docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gemini-2.5-flash-preflight.json Performance evidenceNot run: this profile did not clear preflight. Quality judgmentNot judged: this profile did not clear deterministic screening. Review 12 individual API calls12 preflight · 0 performance · 0 judgment · 0 diagnostic replay Open this section to load the immutable call evidence. qualification-calls/2026-07-27__gemini-2.5-flash.json |