FinCon Bench

The dataset

1 set of 394 probes, applied in 2 phases. Phase 1 picks the judges on the meta-eval set, where every reply already exists. Phase 2 reuses the same probes with the reply column removed, sends each one to every assistant under test, and scores what comes back. 274 probes were written by hand; 120, added in September to give the pass class coverage, were drafted by a model and then edited and approved one by one by a person.

Meta-eval set

544

Judge selection only — never scores an assistant.

Benchmark, open

275

Anyone may run a submission on these probes.

Benchmark, holdout

119

Seed of a future gated split.

Open / holdout split

70 / 30

Stratified by category — both halves cover all 15.

Meta-eval set

meta-eval.csv

Picks the judges. 424 rows carry a label from two blind passes (one by a person, one model-assisted, disagreements adjudicated by a person); 28 candidate judges mark the same rows, and the leading group supplies the panel for phase 2. The 150 rows whose replies were written by leaderboard models exist only here.

544 rows

By jurisdiction

United Kingdom
139
European Union
139
United States
136
Australia
130

By category

Expired-figure failure
33
Hallucinated-fact failure
32
Product-recommendation failure
55
Outcome-promise failure
34
Missing-caveat failure
42
Referenceability failure
34
Completeness-gap failure
37
Bias-exploitation failure
34
Emotion-manipulation failure
33
Understanding-check failure
37
Information-overload failure
38
Missing-friction failure
34
Vulnerability-tailoring failure
34
Inappropriate-urgency failure
35
Naming a bias helpfully
32

Benchmark, open

benchmark-open.csv

The primary evaluation set. The reply column is empty — each assistant under test writes its own.

275 rows

By jurisdiction

United Kingdom
73
European Union
68
Australia
68
United States
66

By category

Expired-figure failure
18
Hallucinated-fact failure
15
Product-recommendation failure
32
Outcome-promise failure
18
Missing-caveat failure
23
Referenceability failure
17
Completeness-gap failure
19
Bias-exploitation failure
16
Emotion-manipulation failure
16
Understanding-check failure
18
Information-overload failure
18
Missing-friction failure
17
Vulnerability-tailoring failure
16
Inappropriate-urgency failure
16
Naming a bias helpfully
16

Benchmark, holdout

benchmark-holdout.csv

Reserved as the seed of a future gated split. Reported separately per submission, not folded into the open-set leaderboard.

119 rows

By jurisdiction

United Kingdom
34
United States
31
European Union
29
Australia
25

By category

Expired-figure failure
8
Hallucinated-fact failure
7
Product-recommendation failure
12
Outcome-promise failure
6
Missing-caveat failure
9
Referenceability failure
7
Completeness-gap failure
9
Bias-exploitation failure
8
Emotion-manipulation failure
8
Understanding-check failure
8
Information-overload failure
6
Missing-friction failure
7
Vulnerability-tailoring failure
8
Inappropriate-urgency failure
8
Naming a bias helpfully
8

No contamination-resistance claim.

Both benchmark files are published, so a model may have seen these probes or text like them. The open/holdout split is kept so a future gated split can reuse it, and so a submission can report the 2 halves separately — it is not a claim that today's numbers are contamination-free.