Rank #51 of 54
OpenAI
GPT-OSS 120B
via AWS Bedrock + Ollama Cloud
gpt-oss-120b
Failure rate
36%
197 of 550 decided probes failed.
Coverage
100%
Share of probes the judges actually decided.
Spread across passes
5.3 pts
Highest minus lowest pass rate over 173 repeated probes.
Compliance / behaviour
60% / 65%
Cost per pass
—
621 avg reply tokens, —s.
Failure rate by category
Lower is better. The count beside each bar is decided/total probes for that category.
Expired-figure failure
58%36/36
Hallucinated-fact failure
57%30/30
Product-recommendation failure
17%64/64
Outcome-promise failure
61%36/36
Missing-caveat failure
15%46/46
Referenceability failure
18%34/34
Completeness-gap failure
58%38/38
Bias-exploitation failure
28%32/32
Emotion-manipulation failure
0%32/32
Understanding-check failure
92%36/36
Information-overload failure
100%36/36
Missing-friction failure
12%34/34
Vulnerability-tailoring failure
0%32/32
Inappropriate-urgency failure
25%32/32
Naming a bias helpfully
3%32/32