FinCon Bench

Leaderboard

56 rows, 54 of them ranked, each sent the same 275 open probes and marked by DeepSeek DeepSeek V4 Pro and Zhipu AI GLM 5.3 Flash, the two judges that agreed most with the labels in pass 1. When the two disagree, Anthropic Claude Opus 5 marks the reply and decides. Every cell below is a failure rate for 1 model on 1 failure category — the same shape as the rubric itself, not a single overall score.

Read the gaps with care. The 84 probes added in September ran several times per model (5 passes on Bedrock and Ollama Cloud, 3 on the Anthropic API, 1 on the OpenAI API) and the majority verdict is published; the older 191 ran once. The spread column is the highest minus the lowest pass rate across those passes, median 5 points, so neighbouring ranks are mostly within noise of each other. A cell backed by few decided probes is noisier than one backed by many — hover a cell to see the count. See methodology before reading a small gap as a real difference.

Failure rate shown · color = fail risklow riskhigh risk
CategoryAnthropicClaude Sonnet 5#1 of 54OpenAIGPT-5.4 Mini#2 of 54Zhipu AIGLM 5.2#3 of 54MiniMaxM2.5#6 of 54MiniMaxM2.1#5 of 54AnthropicClaude Opus 5#9 of 54Moonshot AIKimi K2.5#7 of 54OpenAIGPT-5.4#4 of 54AnthropicClaude Opus 4.5#8 of 54Moonshot AIKimi K2.6#10 of 54DeepSeekDeepSeek V4 Pro#11 of 54Moonshot AIKimi K2.7 Code#12 of 54QwenQwen3 Coder 480B A35B#13 of 54MiniMaxM2.7#14 of 54QwenQwen3 235B A22B (2507)#18 of 54MetaLlama 3.3 70B Instruct#22 of 54Zhipu AIGLM 5.1#17 of 54QwenQwen3.5 397B#20 of 54MistralDevstral 2 123B#16 of 54GoogleGemma 4 31B#19 of 54AnthropicClaude Fable 5#15 of 54AnthropicClaude Haiku 4.5#21 of 54Zhipu AIGLM 4.7#28 of 54MiniMaxM3#30 of 54AmazonNova Lite V1#31 of 54AnthropicClaude Sonnet 4.6#23 of 54MistralLarge 3 675B Instruct#24 of 54NvidiaNemotron 3 Ultra#27 of 54DeepSeekDeepSeek V4 Flash (0731)#33 of 54MistralMagistral Small 2509#25 of 54Zhipu AIGLM 4.7 Flash#29 of 54MetaLlama 3.1 70B Instruct#26 of 54AnthropicClaude Sonnet 4.5#37 of 54GoogleGemma 3 27B IT#44 of 54GoogleGemma 3 12B IT#45 of 54Zhipu AIGLM 5#32 of 54Moonshot AIKimi K3#35 of 54OpenAIGPT-5.4 Nano#36 of 54DeepSeekR1#34 of 54MetaLlama 4 Maverick 17B Instruct#38 of 54MetaLlama 4 Scout 17B Instruct#39 of 54NvidiaNemotron 3 Super#41 of 54DeepSeekDeepSeek V3.2#43 of 54MistralMinistral 3 14B Instruct#40 of 54Moonshot AIKimi K2 Thinking#42 of 54UnknownWRITER.PALMYRA X5 V1NvidiaNemotron Super 3 120B#46 of 54AmazonNova Pro V1#47 of 54QwenQwen3 32B#48 of 54DeepSeekDeepSeek V4 Flash (Preview)OpenAIGPT-OSS Safeguard 120B#49 of 54NvidiaNemotron 3 Nano 30B#50 of 54QwenQwen3 Next 80B A3B#52 of 54OpenAIGPT-OSS 120B#51 of 54NvidiaNemotron Nano 12B V2#53 of 54OpenAIGPT-OSS 20B#54 of 54
19%
19%
20%
21%
21%
22%
22%
23%
23%
23%
23%
24%
24%
24%
24%
24%
24%
24%
25%
25%
25%
25%
25%
25%
25%
26%
26%
26%
26%
27%
27%
27%
27%
27%
27%
28%
28%
28%
28%
28%
28%
28%
28%
29%
29%
30%
30%
30%
30%
30%
33%
34%
35%
36%
38%
38%
Compliance
39%
44%
56%
56%
50%
50%
67%
56%
61%
56%
50%
61%
61%
61%
50%
72%
44%
67%
50%
39%
56%
56%
56%
56%
50%
44%
61%
56%
50%
50%
50%
72%
61%
56%
61%
56%
39%
50%
50%
50%
61%
50%
56%
61%
44%
40%
50%
56%
61%
54%
56%
56%
67%
58%
72%
58%
0%
7%
20%
40%
27%
0%
13%
20%
20%
0%
7%
0%
33%
53%
13%
47%
13%
7%
27%
7%
0%
47%
33%
27%
40%
20%
30%
20%
0%
33%
27%
53%
20%
47%
47%
20%
0%
13%
27%
20%
60%
27%
13%
73%
13%
33%
20%
67%
80%
40%
33%
73%
13%
57%
53%
69%
0%
9%
6%
3%
6%
0%
0%
9%
9%
3%
6%
3%
9%
0%
9%
3%
9%
6%
6%
3%
3%
9%
6%
6%
3%
3%
6%
3%
9%
6%
6%
3%
9%
6%
9%
6%
3%
0%
0%
0%
3%
3%
6%
16%
6%
33%
6%
0%
9%
20%
9%
0%
13%
17%
6%
13%
33%
22%
33%
28%
44%
67%
44%
39%
56%
39%
39%
56%
28%
56%
22%
0%
28%
17%
33%
17%
78%
44%
17%
61%
6%
67%
39%
33%
61%
6%
17%
33%
50%
6%
0%
50%
83%
56%
72%
61%
28%
50%
39%
39%
67%
0%
72%
6%
17%
55%
83%
56%
78%
61%
33%
47%
0%
0%
4%
0%
9%
0%
4%
0%
4%
0%
9%
0%
4%
0%
17%
9%
4%
0%
9%
4%
0%
4%
9%
0%
30%
0%
7%
4%
4%
0%
13%
4%
4%
9%
13%
0%
0%
9%
0%
9%
9%
0%
4%
0%
0%
33%
9%
17%
9%
18%
9%
9%
13%
15%
4%
14%
24%
0%
12%
0%
12%
41%
18%
0%
6%
6%
24%
6%
0%
0%
0%
0%
6%
12%
0%
0%
41%
0%
6%
41%
0%
12%
24%
29%
24%
0%
6%
6%
0%
12%
0%
12%
41%
0%
29%
0%
0%
6%
0%
6%
29%
0%
18%
0%
6%
18%
24%
0%
12%
18%
6%
3%
63%
68%
63%
74%
63%
21%
58%
68%
68%
79%
79%
68%
79%
68%
79%
89%
79%
79%
79%
79%
63%
68%
79%
58%
89%
74%
76%
58%
84%
89%
84%
84%
74%
79%
89%
79%
79%
58%
79%
84%
84%
74%
79%
79%
68%
67%
79%
84%
74%
100%
58%
84%
84%
58%
79%
82%
Behaviour
0%
0%
6%
0%
0%
0%
0%
0%
0%
0%
0%
0%
6%
0%
0%
0%
31%
0%
0%
6%
6%
13%
6%
13%
0%
44%
0%
19%
0%
0%
6%
0%
63%
0%
0%
31%
6%
13%
6%
0%
0%
6%
25%
25%
0%
0%
50%
6%
0%
0%
13%
6%
38%
28%
44%
22%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
28%
28%
33%
17%
6%
56%
33%
17%
11%
83%
56%
61%
17%
22%
56%
61%
61%
83%
50%
94%
33%
33%
78%
39%
67%
17%
0%
78%
72%
78%
100%
44%
28%
94%
67%
67%
56%
78%
44%
78%
61%
83%
89%
6%
83%
0%
44%
100%
61%
54%
94%
78%
72%
92%
67%
89%
100%
100%
61%
100%
100%
100%
100%
100%
100%
72%
78%
100%
100%
100%
78%
6%
83%
94%
100%
100%
100%
100%
89%
89%
0%
100%
97%
100%
78%
94%
44%
94%
100%
67%
39%
94%
100%
100%
100%
50%
67%
100%
89%
100%
100%
78%
89%
50%
89%
73%
100%
94%
89%
100%
100%
92%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
18%
0%
0%
0%
0%
18%
35%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
24%
0%
0%
0%
6%
35%
0%
0%
35%
6%
12%
24%
31%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
6%
0%
6%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
0%
6%
0%
0%
0%
0%
0%
25%
6%
6%
0%
0%
0%
6%
13%
0%
0%
0%
6%
19%
0%
0%
6%
0%
0%
19%
16%
0%
6%
0%
6%
0%
0%
25%
25%
6%
13%
44%
13%
0%
0%
6%
6%
13%
31%
33%
13%
0%
6%
9%
13%
19%
19%
25%
38%
22%
0%
0%
0%
6%
6%
0%
0%
6%
0%
0%
0%
0%
19%
0%
19%
88%
6%
0%
13%
6%
0%
6%
0%
0%
88%
0%
41%
0%
0%
31%
6%
19%
0%
13%
63%
0%
0%
6%
13%
81%
69%
25%
25%
6%
0%
33%
6%
75%
13%
9%
6%
25%
31%
3%
69%
53%

Each cell is the failure rate (%) that model earned on that category — lower is better; the tile's color tracks the same risk, darker being worse. Hover a cell for the exact count. Click a row label (or “Overall”) to sort model columns by it. Showing 56 of 56 models.