The dataset
1 set of 394 probes, applied in 2 phases. Phase 1 picks the judges on the meta-eval set, where every reply already exists. Phase 2 reuses the same probes with the reply column removed, sends each one to every assistant under test, and scores what comes back. 274 probes were written by hand; 120, added in September to give the pass class coverage, were drafted by a model and then edited and approved one by one by a person.
Meta-eval set
544
Judge selection only — never scores an assistant.
Benchmark, open
275
Anyone may run a submission on these probes.
Benchmark, holdout
119
Seed of a future gated split.
Open / holdout split
70 / 30
Stratified by category — both halves cover all 15.
Meta-eval set
meta-eval.csvPicks the judges. 424 rows carry a label from two blind passes (one by a person, one model-assisted, disagreements adjudicated by a person); 28 candidate judges mark the same rows, and the leading group supplies the panel for phase 2. The 150 rows whose replies were written by leaderboard models exist only here.
544 rows
By jurisdiction
By category
Benchmark, open
benchmark-open.csvThe primary evaluation set. The reply column is empty — each assistant under test writes its own.
275 rows
By jurisdiction
By category
Benchmark, holdout
benchmark-holdout.csvReserved as the seed of a future gated split. Reported separately per submission, not folded into the open-set leaderboard.
119 rows
By jurisdiction
By category
No contamination-resistance claim.
Both benchmark files are published, so a model may have seen these probes or text like them. The open/holdout split is kept so a future gated split can reuse it, and so a submission can report the 2 halves separately — it is not a claim that today's numbers are contamination-free.