FinCon Bench

Rank #54 of 54

OpenAI

GPT-OSS 20B

via AWS Bedrock + Ollama Cloud

gpt-oss-20b

Failure rate

38%

210 of 546 decided probes failed.

Coverage

99%

Share of probes the judges actually decided.

Spread across passes

8.3 pts

Highest minus lowest pass rate over 175 repeated probes.

Compliance / behaviour

60% / 57%

Cost per pass

518 avg reply tokens, —s.

Failure rate by category

Lower is better. The count beside each bar is decided/total probes for that category.

Expired-figure failure
58%36/36
Hallucinated-fact failure
69%29/30
Product-recommendation failure
13%64/64
Outcome-promise failure
47%36/36
Missing-caveat failure
14%44/46
Referenceability failure
3%34/34
Completeness-gap failure
82%38/38
Bias-exploitation failure
22%32/32
Emotion-manipulation failure
0%32/32
Understanding-check failure
89%36/36
Information-overload failure
92%36/36
Missing-friction failure
31%33/34
Vulnerability-tailoring failure
0%32/32
Inappropriate-urgency failure
22%32/32
Naming a bias helpfully
53%32/32

← Back to the leaderboard