Preliminary results. One judge, one run, no repeat for variance, and the frontier hosted models are not in it yet. What this does and does not show →
FinCom Bench

Dataset

One set of 274 probes, applied in two phases. The counts and breakdowns below are computed from the actual CSV files in the repository at build time, not typed by hand.

Phase 1 — meta-eval (choose the judge)

datasets/meta-eval.csv

274 probes, each with a pre-written reply. Two human labellers mark each reply pass or fail; candidate judge models mark the same rows with no sight of the human labels. The model that agrees most with the humans becomes the judge. The human labels are never published.

274 rows · 10 columns

By jurisdiction

United Kingdom77
European Union67
United States67
Australia63

By category

Product recommendation36
Missing caveat24
Completeness gap20
Failing to check understanding18
Expired figure18
Exploiting bias16
Manipulating emotion16
Missing friction16
Not tailoring to vulnerability16
Inappropriate urgency16
Naming a bias helpfully16
Outcome promise16
Referenceability failure16
Information overload16
Hallucinated fact14

Phase 2 — benchmark, open split

datasets/benchmark-open.csv

The primary evaluation set. Anyone may run a submission on these probes. Reply and label columns are empty — the runner sends each probe to an assistant and the chosen judge grades what comes back.

191 rows · 8 columns

By jurisdiction

United Kingdom52
European Union47
Australia47
United States45

By category

Product recommendation25
Missing caveat17
Completeness gap14
Failing to check understanding13
Expired figure13
Exploiting bias11
Not tailoring to vulnerability11
Naming a bias helpfully11
Referenceability failure11
Information overload11
Manipulating emotion11
Missing friction11
Inappropriate urgency11
Outcome promise11
Hallucinated fact10

Phase 2 — benchmark, holdout split

datasets/benchmark-holdout.csv

Reserved as the seed of a future gated split; reported separately per submission. Published today alongside the open split — the benchmark makes no contamination-resistance claim either way.

83 rows · 8 columns

By jurisdiction

United Kingdom25
United States22
European Union20
Australia16

By category

Product recommendation11
Missing caveat7
Completeness gap6
Manipulating emotion5
Missing friction5
Inappropriate urgency5
Outcome promise5
Failing to check understanding5
Information overload5
Expired figure5
Exploiting bias5
Naming a bias helpfully5
Referenceability failure5
Not tailoring to vulnerability5
Hallucinated fact4

See methodology for how a phase-1 run picks the judge, and how phase 2 scores an assistant against it.