Dialogue qualification funnel

Every cloud profile tested through Rundale's production dialogue path. Structural reliability, guard use, and request reliability are hard gates; player-facing latency ranks the profiles that survive.

38 runs1 invalid profiles15 rejected1 stopped10 quality qualified3 quality rejected3 awaiting judgment5 awaiting adjudication

Screening policy

Reliability is gated. Speed is ranked among the survivors.
immutable cloud-dialogue qualification receipts
Preflight12 calls
Valid responses100.0%
Guard interventions≤ 10.0%
Error rate≤ 0.5%
Performance sample15+ total
Cold / warm sample5 / 10+
Latencyranking only
Speed index75% TTFT + 25% completion
Judge panel3 families · median
Independent votes2+ · self-family excluded
Quality overall≥ 3.5
Critical axes≥ 3.0
Hard failures≤ 0

Blinded dialogue quality ranking

Blinded multi-family consensus · self-family scores excluded · 18 retained production outputs per judge
RankModelVerdictOverallMinimum critical axisHard failuresUnusable outputsSpeed rankJudge cost
#1openai/gpt-5.6-lunaqualified4.564.4400#11$0.909
#2gemini-3.7-flashqualified4.544.4700#16$1.407
#3google/gemini-3.6-flashqualified4.454.0800#9$1.407
#4google/gemini-3.7-flashqualified4.444.0600#15$1.463
#5openai/gpt-5.6-lunarejected4.434.0011#10$1.798
#6google/gemini-3.7-flashqualified4.424.1100#2$1.465
#7google/gemini-3.5-flash-litequalified4.414.1700#6$1.452
#8gemini-3.7-flashqualified4.344.0000#4$1.395
#9moonshotai/kimi-k2-0905qualified4.343.8900#5$1.791
#10gemini-3.7-flashqualified4.334.0800#14$1.396
#11google/gemini-3.7-flashqualified4.334.0300#13$1.416
#12openai/gpt-4.1-minirejected3.772.8960#3$1.780
#13qwen/qwen2.5-vl-72b-instructrejected3.713.0090#17$1.715

Qualified cloud speed ranking

Lower speed index is better. Quality judgment still determines whether dialogue is usable.
RankModelReasoningMax tokensWarm TTFT p95Warm completion p95Speed indexError rate
#1moonshotai/kimi-k2-0905:nitroprovider default7681336 ms1344 ms1338 ms0.0%
#2google/gemini-3.7-flashlow40961843 ms2712 ms2060 ms0.0%
#3openai/gpt-4.1-miniprovider default7682416 ms5716 ms3241 ms0.0%
#4gemini-3.7-flashlow40963278 ms3655 ms3372 ms0.0%
#5moonshotai/kimi-k2-0905provider default7682950 ms4790 ms3410 ms0.0%
#6google/gemini-3.5-flash-litelow20483597 ms4007 ms3700 ms0.0%
#7x-ai/grok-4.3provider default768942 ms12485 ms3828 ms0.0%
#8x-ai/grok-4.3:nitroprovider default7682022 ms9654 ms3930 ms0.0%
#9google/gemini-3.6-flashlow7684548 ms4967 ms4653 ms0.0%
#10openai/gpt-5.6-lunanone7684870 ms6175 ms5196 ms0.0%
#11openai/gpt-5.6-lunahigh40965496 ms6365 ms5713 ms0.0%
#12z-ai/glm-5.2provider default7685896 ms8454 ms6536 ms0.0%
#13google/gemini-3.7-flashmedium40966722 ms7402 ms6892 ms0.0%
#14gemini-3.7-flashmedium40967145 ms7336 ms7193 ms0.0%
#15google/gemini-3.7-flashhigh40967992 ms8514 ms8123 ms0.0%
#16gemini-3.7-flashhigh40968506 ms8663 ms8545 ms0.0%
#17qwen/qwen2.5-vl-72b-instructprovider default7686361 ms20904 ms9997 ms0.0%
#18openai/gpt-5.6-solnone76811374 ms13470 ms11898 ms0.0%
#19deepseek/deepseek-v4-flash-0731medium409611906 ms14207 ms12481 ms0.0%
#20z-ai/glm-4.7off7682593 ms48244 ms14006 ms0.0%
#21deepseek/deepseek-v4-flash-0731max409699967 ms104620 ms101130 ms0.0%

Production screening

Select a run to inspect the complete evidence trail
DateModelStatusPreflightGuardsQualitySpeed rankWarm TTFT p95Warm completion p95Decision
2026-08-14gemini-3.7-flashqualified12/12100.0%8.3%4.33#147145 ms7336 msquality passed; quality rank #10 of 13; speed rank #14 of 21
Inspect
Run2026-08-14/gemini-3.7-flash-google-native-medium-reasoning
qualified

Decision

quality passed; quality rank #10 of 13; speed rank #14 of 21

Stage
judgment
Candidate
gemini-3.7-flash@https://generativelanguage.googleapis.com/v1beta/openai#env:GOOGLE_API_KEY

Request profile

Reasoning effort
medium
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
1 (8.3%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-medium-reasoning-preflight.jsond89b5ac1ebc6f9c544f190089a8833ed3497df37a90f7b68406b637e0448f69b

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#14 of 21
Speed index
7193 ms
Cold TTFT p95
5488 ms
Warm TTFT p95
7145 ms
Cold completion p95
5602 ms
Warm completion p95
7336 ms
Throughput p50
0.0 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-medium-reasoning-perf.jsonc7d2c126f1627ef04135389d9dd035e6d8f16df566b7f009c115a67edabce8e8

Quality judgment

Consensus / rank
4.33 · #10 of 13
Verdict
qualified
Independent votes
2 pass · 0 fail · 2/2 present
Self-family excluded
0
Judge spread
0.15
Usable outputs
18/18
Panel cost
$1.3965
openai-sol-high4.25 · passhash retained
anthropic-sonnet-low4.40 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-14__gemini-3.7-flash-google-native-medium-reasoning.json0d30eb6775d3a43c6f7242626b8821332d6179eca5ffcfbb6b6e7f880bb8881a

2026-08-14gemini-3.7-flashqualified12/12100.0%0.0%4.34#43278 ms3655 msquality passed; quality rank #8 of 13; speed rank #4 of 21
Inspect
Run2026-08-14/gemini-3.7-flash-google-native-low-reasoning-attempt-2
qualified

Decision

quality passed; quality rank #8 of 13; speed rank #4 of 21

Stage
judgment
Candidate
gemini-3.7-flash@https://generativelanguage.googleapis.com/v1beta/openai#env:GOOGLE_API_KEY

Request profile

Reasoning effort
low
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-low-reasoning-attempt-2-preflight.jsonc22f8e444f08fe6a2ecbcb589544cd5cb023085e3755bb74e97ab2c1f6937744

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#4 of 21
Speed index
3372 ms
Cold TTFT p95
9777 ms
Warm TTFT p95
3278 ms
Cold completion p95
10064 ms
Warm completion p95
3655 ms
Throughput p50
0.0 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-low-reasoning-attempt-2-perf.json575e2b4750fa704cb10e76e33393d700a2863e877ad99eaaaffe97c15e182692

Quality judgment

Consensus / rank
4.34 · #8 of 13
Verdict
qualified
Independent votes
2 pass · 0 fail · 2/2 present
Self-family excluded
0
Judge spread
0.06
Usable outputs
18/18
Panel cost
$1.3949
openai-sol-high4.31 · passhash retained
anthropic-sonnet-low4.37 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-14__gemini-3.7-flash-google-native-low-reasoning-attempt-2.jsonf784db4a38a4c41bdd3f9eb9caeb753c5b16169a984b14a18a28ce4b2c307d58

2026-08-14gemini-3.7-flashqualified12/12100.0%0.0%4.54#168506 ms8663 msquality passed; quality rank #2 of 13; speed rank #16 of 21
Inspect
Run2026-08-14/gemini-3.7-flash-google-native-high-reasoning
qualified

Decision

quality passed; quality rank #2 of 13; speed rank #16 of 21

Stage
judgment
Candidate
gemini-3.7-flash@https://generativelanguage.googleapis.com/v1beta/openai#env:GOOGLE_API_KEY

Request profile

Reasoning effort
high
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-high-reasoning-preflight.jsonf69be041d493dad86257b3f2a43f587ec9da3486c28f1edf5e0ea50bd80cf704

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#16 of 21
Speed index
8545 ms
Cold TTFT p95
6527 ms
Warm TTFT p95
8506 ms
Cold completion p95
6688 ms
Warm completion p95
8663 ms
Throughput p50
0.0 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-14/gemini-3.7-flash-google-native-high-reasoning-perf.json9e193d0388efad4c0f6cb51650a23d3ab505574b08175628cacb5d5c43ccf02a

Quality judgment

Consensus / rank
4.54 · #2 of 13
Verdict
qualified
Independent votes
2 pass · 0 fail · 2/2 present
Self-family excluded
0
Judge spread
0.24
Usable outputs
18/18
Panel cost
$1.4067
openai-sol-high4.66 · passhash retained
anthropic-sonnet-low4.42 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-14__gemini-3.7-flash-google-native-high-reasoning.json956a1cd29328cedb1a2c56a5fc8128eb4fdd3dcee3e5ff7488b7deb0a27aafee

2026-08-13google/gemini-3.7-flashqualified12/12100.0%0.0%4.33#136722 ms7402 msquality passed; quality rank #11 of 13; speed rank #13 of 21
Inspect
Run2026-08-13/gemini-3.7-flash-medium-reasoning
qualified

Decision

quality passed; quality rank #11 of 13; speed rank #13 of 21

Stage
judgment
Candidate
google/gemini-3.7-flash@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
medium
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-medium-reasoning-preflight.json591efc5ce06871bd20e94369df45081c1fc79f206ae0cfedbc580b71697b8db5

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#13 of 21
Speed index
6892 ms
Cold TTFT p95
5242 ms
Warm TTFT p95
6722 ms
Cold completion p95
5598 ms
Warm completion p95
7402 ms
Throughput p50
1218.6 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-medium-reasoning-perf.jsonf58a6de6172fb16d9239efc554d2d0de74502ab1441f1b678a43d98b24a9a5d8

Quality judgment

Consensus / rank
4.33 · #11 of 13
Verdict
qualified
Independent votes
2 pass · 0 fail · 2/2 present
Self-family excluded
0
Judge spread
0.05
Usable outputs
18/18
Panel cost
$1.4160
openai-sol-high4.30 · passhash retained
anthropic-sonnet-low4.35 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-13__gemini-3.7-flash-medium-reasoning.json4a336547b388582ac59b17e46ac147ad0818d26523875d563407f1bfc1458fdf

2026-08-13google/gemini-3.7-flashqualified12/12100.0%8.3%4.42#21843 ms2712 msquality passed; quality rank #6 of 13; speed rank #2 of 21
Inspect
Run2026-08-13/gemini-3.7-flash-low-reasoning
qualified

Decision

quality passed; quality rank #6 of 13; speed rank #2 of 21

Stage
judgment
Candidate
google/gemini-3.7-flash@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
low
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
1 (8.3%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-low-reasoning-preflight.json93e52754ea2908a5467de7a77482b2adbd98d692f3a6466db95845452931ebc0

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#2 of 21
Speed index
2060 ms
Cold TTFT p95
1869 ms
Warm TTFT p95
1843 ms
Cold completion p95
2444 ms
Warm completion p95
2712 ms
Throughput p50
169.1 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-low-reasoning-perf.json924488549d522673f9bd4e3343dfd7724fe6902f9ec6b2edecadf887fdb6fce5

Quality judgment

Consensus / rank
4.42 · #6 of 13
Verdict
qualified
Independent votes
2 pass · 0 fail · 2/2 present
Self-family excluded
0
Judge spread
0.00
Usable outputs
18/18
Panel cost
$1.4653
openai-sol-high4.41 · passhash retained
anthropic-sonnet-low4.42 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-13__gemini-3.7-flash-low-reasoning.json8973a0c6abcff803c0a20e8c4bde7a183b75c4d5ad05e66e8537a87088b674de

2026-08-13google/gemini-3.7-flashqualified12/12100.0%0.0%4.44#157992 ms8514 msquality passed; quality rank #4 of 13; speed rank #15 of 21
Inspect
Run2026-08-13/gemini-3.7-flash-high-reasoning
qualified

Decision

quality passed; quality rank #4 of 13; speed rank #15 of 21

Stage
judgment
Candidate
google/gemini-3.7-flash@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
high
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-high-reasoning-preflight.json1ec45d752a40cb71326ecb8b41e7e751e3103eb0ad3fedd01e1c1af72a062074

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#15 of 21
Speed index
8123 ms
Cold TTFT p95
7103 ms
Warm TTFT p95
7992 ms
Cold completion p95
7710 ms
Warm completion p95
8514 ms
Throughput p50
1826.3 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-13/gemini-3.7-flash-high-reasoning-perf.json09f256c23fa24e21d0347c3c692e4ef16140fcdf3b68f20dbd2870e2c6abc8ae

Quality judgment

Consensus / rank
4.44 · #4 of 13
Verdict
qualified
Independent votes
2 pass · 0 fail · 2/2 present
Self-family excluded
0
Judge spread
0.10
Usable outputs
18/18
Panel cost
$1.4633
openai-sol-high4.49 · passhash retained
anthropic-sonnet-low4.39 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-13__gemini-3.7-flash-high-reasoning.jsonecfa2158e37451d8d7f86dfe49d95a881dbc86de982ea04fbc4e86a80d05d826

2026-08-08moonshotai/kimi-k2.5:nitrostopped2/2100.0%0.0%stopped after 2/12 calls
Inspect
Run2026-08-08/kimi-k2.5-nitro
stopped

Decision

stopped after 2/12 calls

Stage
preflight
Candidate
moonshotai/kimi-k2.5:nitro@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
not captured
JSON mode
not captured
Temperature
Max tokens
Frequency penalty

Preflight evidence

Valid responses
2/2 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
26282 ms–36346 ms
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/kimi-k2.5-nitro-preflight.partial.jsonl70f21dff7ba52c7c3fc0b61b55c52b74c8fa44996ad6bbc89416699594abb2df

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 2 individual API calls

2 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__kimi-k2.5-nitro.jsonc84de2cbdd02066228c5874293ca9dcdc818c9c6f96a46eeb3d8fd5d2431a518

2026-08-08deepseek/deepseek-v4-prorejected11/11100.0%18.2%guard intervention rate (early stop)
Inspect
Run2026-08-08/deepseek-v4-pro-no-reasoning
rejected

Decision

guard intervention rate (early stop)

Stage
preflight
Candidate
deepseek/deepseek-v4-pro@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
not captured
JSON mode
not captured
Temperature
Max tokens
Frequency penalty

Preflight evidence

Valid responses
11/11 (100.0%)
Guard interventions
2 (18.2%)
Observed elapsed
2569 ms–4773 ms
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-pro-no-reasoning-preflight.partial.jsonl8bcd7ceac1cf4397ee188a30ac5f3b8aa6b09c1d56d1b6cb6fdcca34bffda668

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 11 individual API calls

11 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__deepseek-v4-pro-no-reasoning.jsone3a81702bfd4e874c40a6f64e8d57a38dfc59774b8420e6c21d644b4332c0fde

2026-08-08deepseek/deepseek-v4-flash-0731rejected1/250.0%0.0%structural reliability (early stop)
Inspect
Run2026-08-08/deepseek-v4-flash-0731-max-reasoning
rejected

Decision

structural reliability (early stop)

Stage
preflight
Candidate
deepseek/deepseek-v4-flash-0731@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
not captured
JSON mode
not captured
Temperature
Max tokens
Frequency penalty

Preflight evidence

Valid responses
1/2 (50.0%)
Guard interventions
0 (0.0%)
Observed elapsed
10407 ms–15348 ms
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-max-reasoning-preflight.partial.jsonl4aa848ed336de905669b57fb7cdcf73fc5ff57b4c77f53af933121278b9641e7

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 2 individual API calls

2 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__deepseek-v4-flash-0731-max-reasoning.jsond9dd9047723163a9c88b589144ee4d40b4d97caf4552c4275244518087f73aa7

2026-08-08moonshotai/kimi-k2.5:nitrorejected12/12100.0%16.7%guard intervention rate
Inspect
Run2026-08-08/kimi-k2.5-nitro-no-reasoning
rejected

Decision

guard intervention rate

Stage
preflight
Candidate
moonshotai/kimi-k2.5:nitro@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
off
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
2 (16.7%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/kimi-k2.5-nitro-no-reasoning-preflight.json5c6715172e73285a83e5df6fafc4b9484c5b56d599b9f31b46296ecceda9d1d4

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__kimi-k2.5-nitro-no-reasoning.json9ccbba330c58163eb1389af7956f8d064bb6a73fe5bc9d91e0f425c5371908e0

2026-08-08moonshotai/kimi-k2-0905:nitrorejected12/12100.0%16.7%guard intervention rate
Inspect
Run2026-08-08/kimi-k2-0905-nitro-funded
rejected

Decision

guard intervention rate

Stage
preflight
Candidate
moonshotai/kimi-k2-0905:nitro@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
2 (16.7%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/kimi-k2-0905-nitro-funded-preflight.jsonafc70bb4906009023505b50eea72f5f924be7dd8343b51888fa3b2af5800da55

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__kimi-k2-0905-nitro-funded.jsonbee63fd3ffdb1f1a35616bcfa041453d8e5f7ecec3b3e2fef47140f4cb889abf

2026-08-08openai/gpt-5.6-solneeds judgment12/12100.0%0.0%#1811374 ms13470 msawaiting independent judges (0/2); cloud speed rank #18 of 21
Inspect
Run2026-08-08/gpt-5.6-sol-no-reasoning
needs judgment

Decision

awaiting independent judges (0/2); cloud speed rank #18 of 21

Stage
judgment
Candidate
openai/gpt-5.6-sol@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
none
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-sol-no-reasoning-preflight.json9c10c160528b5c49831e088346300cc5602db563dadcf5b5e9a9aed333905a5a

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#18 of 21
Speed index
11898 ms
Cold TTFT p95
8907 ms
Warm TTFT p95
11374 ms
Cold completion p95
10159 ms
Warm completion p95
13470 ms
Throughput p50
137.1 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-sol-no-reasoning-perf.json043e6eae64d2caab3d4cf4f20f9d98e69dcaffb249e1d585867e557fca302b87

Quality judgment

Consensus / rank
pending · unranked
Verdict
panel incomplete
Independent votes
0 pass · 0 fail · 0/2 present
Self-family excluded
1
Judge spread
Usable outputs
0/0
Panel cost
$0.8796
openai-sol-high · excluded4.17 · passhash retained
Review 49 individual API calls

12 preflight · 19 performance · 18 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__gpt-5.6-sol-no-reasoning.json14a335bfc3d3277052b4c1ff775e0520919b190e1680bf1b33b1cc66cbc6da1c

2026-08-08openai/gpt-5.6-lunaquality rejected12/12100.0%8.3%4.43#104870 ms6175 msquality screen failed; quality rank #5 of 13; speed rank #10 of 21
Inspect
Run2026-08-08/gpt-5.6-luna-no-reasoning
quality rejected

Decision

quality screen failed; quality rank #5 of 13; speed rank #10 of 21

Stage
judgment
Candidate
openai/gpt-5.6-luna@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
none
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
1 (8.3%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-no-reasoning-preflight.json1cc2f609c8ff59b5b8b62af75bf8b85ff51efda5e5dfdabb41450a66a9cf973d

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#10 of 21
Speed index
5196 ms
Cold TTFT p95
4105 ms
Warm TTFT p95
4870 ms
Cold completion p95
4618 ms
Warm completion p95
6175 ms
Throughput p50
526.6 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-no-reasoning-perf.json4a8aa87147b548ceb04c3578d5892b841944ecc7f060033eb31d60f7bf06a0b1

Quality judgment

Consensus / rank
4.43 · #5 of 13
Verdict
rejected
Independent votes
0 pass · 2 fail · 2/2 present
Self-family excluded
1
Judge spread
0.56
Usable outputs
17/18
Panel cost
$1.7976
openai-sol-high · excluded4.17 · passhash retained
anthropic-sonnet-low4.14 · failhash retained
google-pro-high4.71 · failhash retained
Review 85 individual API calls

12 preflight · 19 performance · 54 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__gpt-5.6-luna-no-reasoning.json67c98ca2023eb396c01d35e92a9ed2f854e3ab0d7c46d43c6bb67f148ab51d47

2026-08-08openai/gpt-5.6-lunarejected12/12100.0%16.7%guard intervention rate
Inspect
Run2026-08-08/gpt-5.6-luna-max-reasoning-4096
rejected

Decision

guard intervention rate

Stage
preflight
Candidate
openai/gpt-5.6-luna@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
max
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
2 (16.7%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-max-reasoning-4096-preflight.jsona41ce3d9172d9192be9528da688b9d001ae40d0e477fbf988e7fced55aa356af

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__gpt-5.6-luna-max-reasoning-4096.jsone2186aa21dd5ca4ecf17ecb03a11af73131f91e93d5177812f1b3102a02b9fa7

2026-08-08openai/gpt-5.6-lunaqualified12/12100.0%0.0%4.56#115496 ms6365 msquality passed; quality rank #1 of 13; speed rank #11 of 21
Inspect
Run2026-08-08/gpt-5.6-luna-high-reasoning-4096
qualified

Decision

quality passed; quality rank #1 of 13; speed rank #11 of 21

Stage
judgment
Candidate
openai/gpt-5.6-luna@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
high
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-high-reasoning-4096-preflight.jsona1f9cd53066eb20e5bf999333cfcad766356b87c5834e6432bd74c401487d842

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#11 of 21
Speed index
5713 ms
Cold TTFT p95
6884 ms
Warm TTFT p95
5496 ms
Cold completion p95
7844 ms
Warm completion p95
6365 ms
Throughput p50
1057.1 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gpt-5.6-luna-high-reasoning-4096-perf.json4f81be963b60621d93fa771665f81a82f4a5465d8e2b7fb5dd0341a7c113c3cf

Quality judgment

Consensus / rank
4.56 · #1 of 13
Verdict
qualified
Independent votes
2 pass · 0 fail · 2/2 present
Self-family excluded
0
Judge spread
0.59
Usable outputs
18/18
Panel cost
$0.9085
anthropic-sonnet-low4.27 · passhash retained
google-pro-high4.86 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__gpt-5.6-luna-high-reasoning-4096.jsoncc542987234778fa757304a230b3903dc8eb129e09381e611adbec3239998bd7

2026-08-08z-ai/glm-4.7needs judgment12/12100.0%0.0%3.69#202593 ms48244 msawaiting independent judges (1/2); cloud speed rank #20 of 21
Inspect
Run2026-08-08/glm-4.7-no-reasoning
needs judgment

Decision

awaiting independent judges (1/2); cloud speed rank #20 of 21

Stage
judgment
Candidate
z-ai/glm-4.7@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
off
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/glm-4.7-no-reasoning-preflight.json21b0e0092a456e190a2319b243b0c83f6c273edbc3176ff1ea1ec782c01f3eb6

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#20 of 21
Speed index
14006 ms
Cold TTFT p95
30590 ms
Warm TTFT p95
2593 ms
Cold completion p95
48121 ms
Warm completion p95
48244 ms
Throughput p50
52.6 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/glm-4.7-no-reasoning-perf.json6379e39c9c8228a161d15a39b5e38f762810a2b31ffdecf3595a794a3c2c8415

Quality judgment

Consensus / rank
3.69 · unranked
Verdict
panel incomplete
Independent votes
0 pass · 1 fail · 1/2 present
Self-family excluded
0
Judge spread
Usable outputs
10/18
Panel cost
$0.9085
openai-sol-high3.69 · failhash retained
Review 49 individual API calls

12 preflight · 19 performance · 18 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__glm-4.7-no-reasoning.json5fe07183b35014638fa26b47a45b1e71015940ed5e3f0fa4789f298710b21a42

2026-08-08google/gemini-3.6-flashqualified12/12100.0%8.3%4.45#94548 ms4967 msquality passed; quality rank #3 of 13; speed rank #9 of 21
Inspect
Run2026-08-08/gemini-3.6-flash-low-reasoning
qualified

Decision

quality passed; quality rank #3 of 13; speed rank #9 of 21

Stage
judgment
Candidate
google/gemini-3.6-flash@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
low
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
1 (8.3%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.6-flash-low-reasoning-preflight.json13cf0a3b8e158456ccc29ee421a4182db64097822e4819a5cb35008ddb898851

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#9 of 21
Speed index
4653 ms
Cold TTFT p95
2834 ms
Warm TTFT p95
4548 ms
Cold completion p95
3368 ms
Warm completion p95
4967 ms
Throughput p50
724.4 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.6-flash-low-reasoning-perf.json91d1857058339851ecc5791595b83da57e8f515b36ded953b689a06b06ff23a8

Quality judgment

Consensus / rank
4.45 · #3 of 13
Verdict
qualified
Independent votes
2 pass · 0 fail · 2/2 present
Self-family excluded
0
Judge spread
0.01
Usable outputs
18/18
Panel cost
$1.4074
openai-sol-high4.44 · passhash retained
anthropic-sonnet-low4.45 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__gemini-3.6-flash-low-reasoning.json2b27236ee3f8b482565893cdd643f1c67de1f6dead87da904ab7463be7185fd4

2026-08-08google/gemini-3.5-flash-liteinvalid profile10/1283.3%0.0%superseded: insufficient completion budget
Inspect
Run2026-08-08/gemini-3.5-flash-lite-low-reasoning
invalid profile

Decision

superseded: insufficient completion budget

Stage
configuration
Candidate
google/gemini-3.5-flash-lite@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
low
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
10/12 (83.3%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.5-flash-lite-low-reasoning-preflight.json26f3bfeb1d5577aedd0c3e814e0cea64b4fec9e9940ffd55627e3a4fff81d539

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 14 individual API calls

12 preflight · 0 performance · 0 judgment · 2 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__gemini-3.5-flash-lite-low-reasoning.jsonb2a7f94345e687c4d3d76cb7ebfac35e9f6c1620d224723230208e97c37193ef

2026-08-08google/gemini-3.5-flash-litequalified12/12100.0%0.0%4.41#63597 ms4007 msquality passed; quality rank #7 of 13; speed rank #6 of 21
Inspect
Run2026-08-08/gemini-3.5-flash-lite-low-reasoning-2048
qualified

Decision

quality passed; quality rank #7 of 13; speed rank #6 of 21

Stage
judgment
Candidate
google/gemini-3.5-flash-lite@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
low
JSON mode
on
Temperature
0.7
Max tokens
2048
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.5-flash-lite-low-reasoning-2048-preflight.jsonb6478393bfc69ff9c97a1605c9cb5d7733d96170bc3807c69b486120a7551161

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#6 of 21
Speed index
3700 ms
Cold TTFT p95
3641 ms
Warm TTFT p95
3597 ms
Cold completion p95
3751 ms
Warm completion p95
4007 ms
Throughput p50
1574.5 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/gemini-3.5-flash-lite-low-reasoning-2048-perf.json7a028b4da4f76e403882cf17d71329a6f4cfd23ce6fad642bb34f2993ed117ac

Quality judgment

Consensus / rank
4.41 · #7 of 13
Verdict
qualified
Independent votes
2 pass · 0 fail · 2/2 present
Self-family excluded
0
Judge spread
0.07
Usable outputs
18/18
Panel cost
$1.4522
openai-sol-high4.44 · passhash retained
anthropic-sonnet-low4.37 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__gemini-3.5-flash-lite-low-reasoning-2048.json5d6963f2c4d34cc66393f9cfcce4dccab7da4bff3185ecb2bc8ba99a8337c4e4

2026-08-08deepseek-v4-flashrejected12/12100.0%16.7%guard intervention rate
Inspect
Run2026-08-08/deepseek-v4-flash-direct-high-reasoning-4096
rejected

Decision

guard intervention rate

Stage
preflight
Candidate
deepseek-v4-flash@https://api.deepseek.com/v1#env:DEEPSEEK_API_KEY

Request profile

Reasoning effort
high
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
2 (16.7%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-direct-high-reasoning-4096-preflight.json6163d9e14fbc564dc2ee59ac6054da0f7971b13e23fefc0ad7960b46c6737f4e

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__deepseek-v4-flash-direct-high-reasoning-4096.jsondc7adbc0596f81e4e18e99c1a0849e2cdfa88a688a051263c329653ae95a2631

2026-08-08deepseek/deepseek-v4-flash-0731rejected12/12100.0%25.0%guard intervention rate
Inspect
Run2026-08-08/deepseek-v4-flash-0731-no-reasoning
rejected

Decision

guard intervention rate

Stage
preflight
Candidate
deepseek/deepseek-v4-flash-0731@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
off
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
3 (25.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-no-reasoning-preflight.jsone3574e4f1f343ed74b2f09198bb2ba46f88840d3b3fe7e00ded4f4c6e60306ac

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__deepseek-v4-flash-0731-no-reasoning.jsond00dc324c47749fd7570a8a43423125d2d7775c6ed0f33dbae322c4bac8a501c

2026-08-08deepseek/deepseek-v4-flash-0731needs judgment12/12100.0%0.0%4.37#1911906 ms14207 msawaiting independent judges (1/2); cloud speed rank #19 of 21
Inspect
Run2026-08-08/deepseek-v4-flash-0731-medium-reasoning-4096
needs judgment

Decision

awaiting independent judges (1/2); cloud speed rank #19 of 21

Stage
judgment
Candidate
deepseek/deepseek-v4-flash-0731@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
medium
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-medium-reasoning-4096-preflight.json6330fe5aa8ed64be084628c86679f0124b0eecc7fcf4cbf846a82e2462a748cf

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#19 of 21
Speed index
12481 ms
Cold TTFT p95
29019 ms
Warm TTFT p95
11906 ms
Cold completion p95
31051 ms
Warm completion p95
14207 ms
Throughput p50
202.5 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-medium-reasoning-4096-perf.jsoneb911f619608a3bea35a1e84169901def20cba8e37c1388f568d33daa969f2cf

Quality judgment

Consensus / rank
4.37 · unranked
Verdict
panel incomplete
Independent votes
0 pass · 1 fail · 1/2 present
Self-family excluded
0
Judge spread
Usable outputs
17/18
Panel cost
$0.5363
anthropic-sonnet-low4.37 · failhash retained
Review 49 individual API calls

12 preflight · 19 performance · 18 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__deepseek-v4-flash-0731-medium-reasoning-4096.json1c4db6c4ce5ac77b85dab78878e7a1b2644735879f43974d0899bac7a7410fd5

2026-08-08deepseek/deepseek-v4-flash-0731needs adjudication12/12100.0%0.0%4.50#2199967 ms104620 msjudge disagreement requires adjudication; speed rank #21 of 21
Inspect
Run2026-08-08/deepseek-v4-flash-0731-max-reasoning-4096
needs adjudication

Decision

judge disagreement requires adjudication; speed rank #21 of 21

Stage
judgment
Candidate
deepseek/deepseek-v4-flash-0731@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
max
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-max-reasoning-4096-preflight.json1d822e6cd8d55981961aee05612cbbc37ae4f44728a37cc4aa3b2c851c9cbbc2

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#21 of 21
Speed index
101130 ms
Cold TTFT p95
123127 ms
Warm TTFT p95
99967 ms
Cold completion p95
124790 ms
Warm completion p95
104620 ms
Throughput p50
409.1 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-max-reasoning-4096-perf.jsoncbb65974797d08374a2dbaa02a78954a96e57be961d53c0250964b740cda6f7a

Quality judgment

Consensus / rank
4.50 · unranked
Verdict
adjudication required
Independent votes
1 pass · 1 fail · 2/2 present
Self-family excluded
0
Judge spread
0.43
Usable outputs
18/18
Panel cost
$0.8898
anthropic-sonnet-low4.28 · failhash retained
google-pro-high4.72 · passhash retained
Review 67 individual API calls

12 preflight · 19 performance · 36 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__deepseek-v4-flash-0731-max-reasoning-4096.json2d596d86d95e89ff7532cdac29a0e04bbcf3eb54729e1ab999b4938715c88e43

2026-08-08deepseek/deepseek-v4-flash-0731rejected10/1283.3%25.0%structural reliability
Inspect
Run2026-08-08/deepseek-v4-flash-0731-low-reasoning-4096
rejected

Decision

structural reliability

Stage
preflight
Candidate
deepseek/deepseek-v4-flash-0731@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
low
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
10/12 (83.3%)
Guard interventions
3 (25.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-low-reasoning-4096-preflight.jsond50acf3319f44964a578aed95bacd04ac401388c6b935b2cf4005f6e0f08dbde

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__deepseek-v4-flash-0731-low-reasoning-4096.json999da849729e787c352683c92fefdae9926d9102d3c2d834b2f2f89c1bf3a05e

2026-08-08deepseek/deepseek-v4-flash-0731rejected12/12100.0%16.7%guard intervention rate
Inspect
Run2026-08-08/deepseek-v4-flash-0731-high-reasoning-4096
rejected

Decision

guard intervention rate

Stage
preflight
Candidate
deepseek/deepseek-v4-flash-0731@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
high
JSON mode
on
Temperature
0.7
Max tokens
4096
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
2 (16.7%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-08-08/deepseek-v4-flash-0731-high-reasoning-4096-preflight.jsonf1bc20c3089aa71859ad840a47ee7344e359a2187ab61994412eea0c3c252764

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-08-08__deepseek-v4-flash-0731-high-reasoning-4096.json60aba7a21f2e42728d0e959d4340fa9fd5ab096919995643f6dd05f613a0d0c2

2026-07-27qwen/qwen2.5-vl-72b-instructquality rejected12/12100.0%0.0%3.71#176361 ms20904 msquality screen failed; quality rank #13 of 13; speed rank #17 of 21
Inspect
Run2026-07-27/qwen25-vl-72b
quality rejected

Decision

quality screen failed; quality rank #13 of 13; speed rank #17 of 21

Stage
judgment
Candidate
qwen/qwen2.5-vl-72b-instruct@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/qwen25-vl-72b-preflight.json97f31fad3663428c2c5200352e9ce61941a5a06d0202c35da9c8766f1790ccf5

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#17 of 21
Speed index
9997 ms
Cold TTFT p95
4494 ms
Warm TTFT p95
6361 ms
Cold completion p95
19988 ms
Warm completion p95
20904 ms
Throughput p50
8.1 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/qwen25-vl-72b-perf.json1256dc2692f28e7c7a02913111544d287d63293eb5ccaa77c88b63c24ac2659a

Quality judgment

Consensus / rank
3.71 · #13 of 13
Verdict
rejected
Independent votes
0 pass · 3 fail · 3/2 present
Self-family excluded
0
Judge spread
0.24
Usable outputs
18/18
Panel cost
$1.7151
openai-sol-high3.71 · failhash retained
anthropic-sonnet-low3.64 · failhash retained
google-pro-high3.89 · failhash retained
Review 85 individual API calls

12 preflight · 19 performance · 54 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__qwen25-vl-72b.json19a996901b382e357fdcd9191d7d431220fdd0b8d49c76e6cfdba2f4e1cae62e

2026-07-27nvidia/nemotron-3-ultra-550b-a55brejected4/1233.3%25.0%structural reliability
Inspect
Run2026-07-27/nemotron-3-ultra
rejected

Decision

structural reliability

Stage
preflight
Candidate
nvidia/nemotron-3-ultra-550b-a55b@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
4/12 (33.3%)
Guard interventions
3 (25.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/nemotron-3-ultra-preflight.json841804932eda683370b609b05bbecda02911be129ff41bb17e808491868ec051

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__nemotron-3-ultra.json67f2085c44cb271a33f10dddf2532fd769b9cad3e6f313b74476bf3093bcbfc1

2026-07-27mistralai/mistral-medium-3.1rejected12/12100.0%41.7%guard intervention rate
Inspect
Run2026-07-27/mistral-medium-3.1
rejected

Decision

guard intervention rate

Stage
preflight
Candidate
mistralai/mistral-medium-3.1@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
5 (41.7%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/mistral-medium-3.1-preflight.json2982f5f6e25603ceb85500a9c0a893f5f3a8fd20637f5a0653085a3913c7dd4c

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__mistral-medium-3.1.json8fb2efa0e01f8a4c6f1d5bb52fc7b2550c12b66c7efa6a00c7639694296ef74d

2026-07-27moonshotai/kimi-k2-0905qualified12/12100.0%0.0%4.34#52950 ms4790 msquality passed; quality rank #9 of 13; speed rank #5 of 21
Inspect
Run2026-07-27/kimi-k2-0905
qualified

Decision

quality passed; quality rank #9 of 13; speed rank #5 of 21

Stage
judgment
Candidate
moonshotai/kimi-k2-0905@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/kimi-k2-0905-preflight.json4c10ed7a7639e6d552e069ea3eca09f8a94349721a96c19af14f65de8c5bd61e

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#5 of 21
Speed index
3410 ms
Cold TTFT p95
2912 ms
Warm TTFT p95
2950 ms
Cold completion p95
5510 ms
Warm completion p95
4790 ms
Throughput p50
40.4 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/kimi-k2-0905-perf.json0d2cddcbf6a059beb3206c12242e0aa50e8623f6ce7770165040563717d8d55f

Quality judgment

Consensus / rank
4.34 · #9 of 13
Verdict
qualified
Independent votes
3 pass · 0 fail · 3/2 present
Self-family excluded
0
Judge spread
0.21
Usable outputs
18/18
Panel cost
$1.7905
openai-sol-high4.17 · passhash retained
anthropic-sonnet-low4.34 · passhash retained
google-pro-high4.37 · passhash retained
Review 85 individual API calls

12 preflight · 19 performance · 54 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__kimi-k2-0905.jsone4f54d1aa51bd966e9aee380cd01d953bb1c98f33cbee544c18954b7932a362e

2026-07-27moonshotai/kimi-k2-0905:nitroneeds adjudication12/12100.0%0.0%4.14#11336 ms1344 msjudge disagreement requires adjudication; speed rank #1 of 21
Inspect
Run2026-07-27/kimi-k2-0905-nitro
needs adjudication

Decision

judge disagreement requires adjudication; speed rank #1 of 21

Stage
judgment
Candidate
moonshotai/kimi-k2-0905:nitro@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/kimi-k2-0905-nitro-preflight.jsona693e679b53abff4ff555220671625d39e44fd7478428318e7d465d713755104

Performance evidence

Measurements
48
Cold / warm
6 / 42
Speed rank
#1 of 21
Speed index
1338 ms
Cold TTFT p95
1664 ms
Warm TTFT p95
1336 ms
Cold completion p95
1681 ms
Warm completion p95
1344 ms
Throughput p50
6000.0 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/kimi-k2-0905-nitro-perf-expanded.json1984ba432a4cbb82e60be1a1fc52a2fc36828d2a2654a41229fd4e052d258133

Quality judgment

Consensus / rank
4.14 · unranked
Verdict
adjudication required
Independent votes
1 pass · 2 fail · 3/2 present
Self-family excluded
0
Judge spread
0.84
Usable outputs
18/18
Panel cost
$1.8702
openai-sol-high3.92 · failhash retained
anthropic-sonnet-low4.14 · failhash retained
google-pro-high4.76 · passhash retained
Review 115 individual API calls

12 preflight · 49 performance · 54 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__kimi-k2-0905-nitro.json92eb221e48ab8bb8785cc86ad7a5dfa261eb3928b330456345fa3522096be7a3

2026-07-27x-ai/grok-4.3needs adjudication12/12100.0%0.0%4.12#7942 ms12485 msjudge disagreement requires adjudication; speed rank #7 of 21
Inspect
Run2026-07-27/grok-4.3
needs adjudication

Decision

judge disagreement requires adjudication; speed rank #7 of 21

Stage
judgment
Candidate
x-ai/grok-4.3@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/grok-4.3-preflight.json2c86351404f17832eb303ed2516c9b91877ec04f6bb7af6807e1a7e8501f1308

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#7 of 21
Speed index
3828 ms
Cold TTFT p95
1402 ms
Warm TTFT p95
942 ms
Cold completion p95
13085 ms
Warm completion p95
12485 ms
Throughput p50
87.4 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/grok-4.3-perf.json21c39b8de9ac7094b9b363307bf4814b95296b53abce0d6afc3af6d862567d4b

Quality judgment

Consensus / rank
4.12 · unranked
Verdict
adjudication required
Independent votes
2 pass · 1 fail · 3/2 present
Self-family excluded
0
Judge spread
1.08
Usable outputs
18/18
Panel cost
$1.8462
openai-sol-high3.13 · failhash retained
anthropic-sonnet-low4.12 · passhash retained
google-pro-high4.21 · passhash retained
Review 85 individual API calls

12 preflight · 19 performance · 54 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__grok-4.3.jsonbbb27a87fd7e7707ee9b67fd13574a985feab15d7009def148263d5a341e0635

2026-07-27x-ai/grok-4.3:nitroneeds adjudication12/12100.0%0.0%3.94#82022 ms9654 msjudge disagreement requires adjudication; speed rank #8 of 21
Inspect
Run2026-07-27/grok-4.3-nitro
needs adjudication

Decision

judge disagreement requires adjudication; speed rank #8 of 21

Stage
judgment
Candidate
x-ai/grok-4.3:nitro@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/grok-4.3-nitro-preflight.json8ad22e9ce388b6be2ac2153583bf9578080341f1789f70b4d15dc67be320a30c

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#8 of 21
Speed index
3930 ms
Cold TTFT p95
1086 ms
Warm TTFT p95
2022 ms
Cold completion p95
11562 ms
Warm completion p95
9654 ms
Throughput p50
103.9 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/grok-4.3-nitro-perf.json28458bc124df89f5cc37175ae044b4792eb0ffec126725c6da23e8795681948a

Quality judgment

Consensus / rank
3.94 · unranked
Verdict
adjudication required
Independent votes
1 pass · 2 fail · 3/2 present
Self-family excluded
0
Judge spread
0.95
Usable outputs
18/18
Panel cost
$1.8816
openai-sol-high3.33 · failhash retained
anthropic-sonnet-low3.94 · failhash retained
google-pro-high4.28 · passhash retained
Review 85 individual API calls

12 preflight · 19 performance · 54 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__grok-4.3-nitro.jsona0311521ecdff6c4a4f12fb7ba426b31a2f260bcba7e851675684b0a372092a3

2026-07-27openai/gpt-oss-120b:nitrorejected9/1275.0%0.0%structural reliability
Inspect
Run2026-07-27/gpt-oss-120b-nitro
rejected

Decision

structural reliability

Stage
preflight
Candidate
openai/gpt-oss-120b:nitro@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
9/12 (75.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gpt-oss-120b-nitro-preflight.json267c0bb8c91dfb5d892fb8966144c4a3c9f2ebfa72ca077ea5a51f08fe679325

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__gpt-oss-120b-nitro.json1d4ad179de0984e3fd68c14b9923dc54ada4ff33a94058d2b4439debc334f884

2026-07-27openai/gpt-4orejected12/12100.0%16.7%guard intervention rate
Inspect
Run2026-07-27/gpt-4o
rejected

Decision

guard intervention rate

Stage
preflight
Candidate
openai/gpt-4o@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
2 (16.7%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gpt-4o-preflight.jsonbcd5308204ea1c01d3dfa6ee675f0b1908d720b28fc262090e50cb21e481acc3

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__gpt-4o.json9f80067136a99b7baaf440d5454d2175730624399c3be5df31ab600db390bef4

2026-07-27openai/gpt-4.1-miniquality rejected12/12100.0%0.0%3.77#32416 ms5716 msquality screen failed; quality rank #12 of 13; speed rank #3 of 21
Inspect
Run2026-07-27/gpt-4.1-mini
quality rejected

Decision

quality screen failed; quality rank #12 of 13; speed rank #3 of 21

Stage
judgment
Candidate
openai/gpt-4.1-mini@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gpt-4.1-mini-preflight.json9e3801a1c127ebf228c59acde1edc85d88f83a207be481b04b4aa9685d05957e

Performance evidence

Measurements
48
Cold / warm
6 / 42
Speed rank
#3 of 21
Speed index
3241 ms
Cold TTFT p95
1116 ms
Warm TTFT p95
2416 ms
Cold completion p95
3605 ms
Warm completion p95
5716 ms
Throughput p50
73.2 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gpt-4.1-mini-perf-expanded.json9b7a22a2670f17d09cd502984ee682e81e0997c1ceff19ff6830d16035df080b

Quality judgment

Consensus / rank
3.77 · #12 of 13
Verdict
rejected
Independent votes
0 pass · 2 fail · 2/2 present
Self-family excluded
1
Judge spread
0.72
Usable outputs
18/18
Panel cost
$1.7797
openai-sol-high · excluded3.66 · failhash retained
anthropic-sonnet-low3.41 · failhash retained
google-pro-high4.13 · failhash retained
Review 115 individual API calls

12 preflight · 49 performance · 54 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__gpt-4.1-mini.json2b77583cb395cb3cd30215632ac56d2e3c4ee3a341589e67e9106321fab4a5f9

2026-07-27z-ai/glm-5.2needs adjudication12/12100.0%0.0%4.18#125896 ms8454 msjudge disagreement requires adjudication; speed rank #12 of 21
Inspect
Run2026-07-27/glm-5.2
needs adjudication

Decision

judge disagreement requires adjudication; speed rank #12 of 21

Stage
judgment
Candidate
z-ai/glm-5.2@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
12/12 (100.0%)
Guard interventions
0 (0.0%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/glm-5.2-preflight.json6c7095353754934c61fd6f8471d616445006c6b187f5c8ff1a64b7586ecd88c1

Performance evidence

Measurements
18
Cold / warm
6 / 12
Speed rank
#12 of 21
Speed index
6536 ms
Cold TTFT p95
2166 ms
Warm TTFT p95
5896 ms
Cold completion p95
10476 ms
Warm completion p95
8454 ms
Throughput p50
61.3 tok/s
Error rate
0.0%
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/glm-5.2-perf.jsonbe49dbe1898676b4c19fe866be95e8641a0f0a510053dbea16fe800e6d6fcb88

Quality judgment

Consensus / rank
4.18 · unranked
Verdict
adjudication required
Independent votes
1 pass · 2 fail · 3/2 present
Self-family excluded
0
Judge spread
1.69
Usable outputs
18/18
Panel cost
$1.8963
openai-sol-high3.04 · failhash retained
anthropic-sonnet-low4.18 · failhash retained
google-pro-high4.72 · passhash retained
Review 85 individual API calls

12 preflight · 19 performance · 54 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__glm-5.2.json2d6225758ed4f93c2f473e96c577610bdb72794ff8eab1b329bfd91720d2ff80

2026-07-27z-ai/glm-4.5vrejected10/1283.3%33.3%structural reliability
Inspect
Run2026-07-27/glm-4.5v
rejected

Decision

structural reliability

Stage
preflight
Candidate
z-ai/glm-4.5v@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
10/12 (83.3%)
Guard interventions
4 (33.3%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/glm-4.5v-preflight.json525e1ec5f33b05372d050ca7cddde3e7cd7d239c278f37d61cacd9487e03744f

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__glm-4.5v.jsonee29de33fedb9cd913fd48a72a30e1ff2271aaf5fb5e3ab26a62bd2a435f4f9e

2026-07-27google/gemini-2.5-flashrejected11/1291.7%8.3%structural reliability
Inspect
Run2026-07-27/gemini-2.5-flash
rejected

Decision

structural reliability

Stage
preflight
Candidate
google/gemini-2.5-flash@https://openrouter.ai/api/v1#env:OPENROUTER_API_KEY

Request profile

Reasoning effort
provider default
JSON mode
on
Temperature
0.7
Max tokens
768
Frequency penalty
0.5

Preflight evidence

Valid responses
11/12 (91.7%)
Guard interventions
1 (8.3%)
Observed elapsed
not captured
Raw receipt archived outside Git

docs/proofs/cloud-dialogue-qualification/runs/2026-07-27/gemini-2.5-flash-preflight.jsond01665f8eafd611e0f8be34dafbae53b036e5d798f85c82541f82f0469a86198

Performance evidence

Not run: this profile did not clear preflight.

Quality judgment

Not judged: this profile did not clear deterministic screening.

Review 12 individual API calls

12 preflight · 0 performance · 0 judgment · 0 diagnostic replay

Open this section to load the immutable call evidence.

qualification-calls/2026-07-27__gemini-2.5-flash.jsonae392e2e4ded77403a24773c171ba5528bc027d643295302dc91b6ac2dc9ed84