Compare 2 models
Pick any 2 models to see where they actually differ — the failure rate on each of the 15 categories, biggest gap first. A model that ranks a few places above another can still lose badly on 1 specific category.
Same caveats as the leaderboard: 2 judges and a tiebreak, repeats on 84 of 275 probes only. A category with few decided probes is noisier than one with many — see methodology.
Anthropic Claude Sonnet 5
OpenAI GPT-5.4 Mini
Referenceability failure24pp gap
A
24%B
0%
Outcome-promise failure11pp gap
A
33%B
22%
Product-recommendation failure9pp gap
A
0%B
9%
Hallucinated-fact failure7pp gap
A
0%B
7%
Inappropriate-urgency failure6pp gap
A
0%B
6%
Expired-figure failure6pp gap
A
39%B
44%
Completeness-gap failure5pp gap
A
63%B
68%
Missing-caveat failure0pp gap
A
0%B
0%
Bias-exploitation failure0pp gap
A
0%B
0%
Emotion-manipulation failure0pp gap
A
0%B
0%
Understanding-check failure0pp gap
A
28%B
28%
Information-overload failure0pp gap
A
100%B
100%
Missing-friction failure0pp gap
A
0%B
0%
Vulnerability-tailoring failure0pp gap
A
0%B
0%
Naming a bias helpfully0pp gap
A
0%B
0%
Bars show failure rate — shorter and paler is better. The highlighted (fail-colored) side is whichever model is worse on that category.