Governance register

What each benchmark actually measures.

Every documented case of a leaderboard result not meaning what it appeared to mean was found by reading a policy or comparing an artefact, not by running a detector. So this is a register of checkable facts: what a test measures, how much room is left in it, what stops contamination, who may republish it, and what has gone wrong with it.

Every incident carries a primary source and a date. An entry with neither is not published here. Where a figure has not been measured, the register says so rather than printing a zero.

LMArena Text Arena

May republish

Crowd preference over blind pairwise responses to open-ended text prompts, reported as an Elo-style rating with an interval.

CC BY 4.0Continuous, snapshotted0 incidents recorded

No fixed ceiling

Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it. No bar is drawn, because any bar would have to invent a ceiling in order to have a length.

Human preference between anonymously generated front-end websites.

CC BY 4.0Continuous, snapshotted0 incidents recorded

No fixed ceiling

Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it. No bar is drawn, because any bar would have to invent a ceiling in order to have a length.

Human preference over responses to image-bearing prompts.

CC BY 4.0Continuous, snapshotted0 incidents recorded

No fixed ceiling

Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it. No bar is drawn, because any bar would have to invent a ceiling in order to have a length.

Aider Polyglot

May republish

Instruction-following code editing across 225 Exercism exercises in six languages, scored on pass rate after one retry and on whether the edit was well formed.

Apache-2.0Maintainer-run; new models do not appear automatically0 incidents recorded
not evaluatedSaturation not assessed76.9 pts · headroom unknown
Text equivalentAider Polyglot: best reproduced result 76.9 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

tau²-bench airline

May republish

Multi-turn airline customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.

MITManual verified runs; models do not appear automatically0 incidents recorded
not evaluatedSaturation not assessed80.3 pts · headroom unknown
Text equivalenttau²-bench airline: best reproduced result 80.3 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

tau²-bench retail

May republish

Multi-turn retail customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.

MITManual verified runs; models do not appear automatically0 incidents recorded
not evaluatedSaturation not assessed79.7 pts · headroom unknown
Text equivalenttau²-bench retail: best reproduced result 79.7 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

tau²-bench telecom

May republish

Multi-turn telecom customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.

MITManual verified runs; models do not appear automatically0 incidents recorded
not evaluatedSaturation not assessed89.5 pts · headroom unknown
Text equivalenttau²-bench telecom: best reproduced result 89.5 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

GPQA Diamond

May republish

Expert-level multiple-choice reasoning in physics, chemistry and biology.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessed93.9 pts · headroom unknown
Text equivalentGPQA Diamond: best reproduced result 93.9 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

FrontierMath

May republish

Research-level mathematical problem solving on original, unpublished expert-authored problems.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them2 incidents recorded
not evaluatedSaturation not assessed47.2 pts · headroom unknown
Text equivalentFrontierMath: best reproduced result 47.2 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

FrontierMath Tier 4

May republish

The hardest FrontierMath tier, using restricted materials rather than a public answer key.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessed31.3 pts · headroom unknown
Text equivalentFrontierMath Tier 4: best reproduced result 31.3 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

MATH Level 5

May republish

The hardest tier of the MATH competition-mathematics dataset.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentMATH Level 5: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SWE-bench Verified

May republish

Resolving real Python repository issues by generating patches that must pass repository tests.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them2 incidents recorded
not evaluatedSaturation not assessed83.5 pts · headroom unknown
Text equivalentSWE-bench Verified: best reproduced result 83.5 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SimpleQA Verified

May republish

Short-answer factual accuracy and hallucination resistance on fact-seeking questions.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessed71.6 pts · headroom unknown
Text equivalentSimpleQA Verified: best reproduced result 71.6 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

Fiction.LiveBench

May republish

Recall and comprehension over long serialized fiction.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentFiction.LiveBench: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SimpleBench

May republish

Everyday commonsense, spatial and social reasoning with adversarial linguistic traps.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentSimpleBench: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

Instruction-following code editing, as collected by Epoch from external reports.

CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentAider Polyglot (externally reported): no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SWE-rebench

May republish

Coding-agent issue resolution on continuously collected recent GitHub issue/PR pairs under a fixed scaffold, repeated five times.

CC BY 4.0Monthly rolling splits1 incident recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentSWE-rebench: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

NIAH v2

May republish

Long-context retrieval across context length and needle depth, including multi-fact and UUID-chain tasks.

MIT0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentNIAH v2: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

GraphWalks

May republish

Multi-hop reasoning over graphs represented as edge lists in long context.

MIT0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentGraphWalks: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

RULER

May republish

The effective context size of long-context models via multi-hop tracing and aggregation beyond simple retrieval.

Apache-2.00 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentRULER: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

IFEval-FC

May republish

Strict instruction and format adherence, scored by regex and JSON-schema validation rather than a judge.

Apache-2.00 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentIFEval-FC: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

SafetyBench

May republish

Safety understanding across seven categories, in Chinese and English.

MIT0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentSafetyBench: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

HaluEval

May republish

Whether a model recognises and avoids hallucinated content, against human-annotated samples.

MIT0 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentHaluEval: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

Terminal-Bench 2.0

May republish

Agents completing realistic command-line tasks in containerised environments, checked by deterministic tests.

Apache-2.0Submission-driven; entries are audited rather than published automatically2 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentTerminal-Bench 2.0: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

BrowseComp

May republish

Persistent agentic web research on questions whose answers are easy to verify and hard to locate.

MIT (evaluation tooling)No unified cadence; scores appear when vendors or independent evaluators publish runs1 incident recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentBrowseComp: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.

DeepSWE

May republish

Whether a coding agent can complete original, long-horizon engineering tasks in real repositories, verified by hand-written tests that check behaviour rather than implementation details.

Apache-2.00 incidents recorded
not evaluatedSaturation not assessedno reproduced result · headroom unknown
Text equivalentDeepSWE: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Ingested, not redistributed

What we hold but may not republish.

These benchmarks inform our own reading and never appear as a figure anywhere on this site. The licence is the reason in every case, and each one links to where that licence was read. Naming them is the point: a register listing only what it may publish would look complete while hiding its own edges.

This list is generated from every restricted record we hold, not maintained by hand. A hand-maintained list of the things you are not showing is the one list nobody notices is wrong.

EQ-Bench Slop Score

Not republishedNot stated

No licence is stated on the results repository, so we have no permission to republish its figures and do not assume one. Nothing about the benchmark is implied: an unstated licence is a fact about paperwork, not about the test.

Licence read at github.com/EQ-bench/eqbench-leaderboard-results

EQ-Bench Judgemark

Not republishedNot stated

No licence is stated on the results repository, so we have no permission to republish its figures and do not assume one. Nothing about the benchmark is implied: an unstated licence is a fact about paperwork, not about the test.

Licence read at github.com/EQ-bench/eqbench-leaderboard-results

L-Eval

Not republishedGPL-3.0

Licensed GPL-3.0. Its reciprocal terms are not ones we are in a position to satisfy for published output, so we hold it for reference and publish no figure from it.

Licence read at github.com/OpenLMLab/LEval

AIME 2025

Not republishedCC BY-NC-SA 4.0

Licensed CC BY-NC-SA 4.0, which does not grant commercial redistribution. We may read it and we may say the set exists; we may not put its numbers on a page. Note this is MathArena, not Epoch AI OTIS Mock AIME, which carries a different licence and is published under its own entry.

Licence read at huggingface.co/datasets/MathArena/aime_2025

MathArena Apex

Not republishedCC BY-NC-SA 4.0

Licensed CC BY-NC-SA 4.0, which does not grant commercial redistribution. We may read it and we may say the set exists; we may not put its numbers on a page. Note this is MathArena, not Epoch AI OTIS Mock AIME, which carries a different licence and is published under its own entry.

Licence read at huggingface.co/datasets/MathArena/aime_2025

Design Arena

Not republishedNo licence granted; automated access prohibited

No licence is granted to users over the platform content, and §17(a) assigns rights to Arcada Labs. Republishing leaderboard data would need express written consent.

Licence read at designarena.ai/terms-and-conditions

The practice record

Facts about benchmarking, not about one benchmark.

These are the governance facts that make the case for a register like this one. Each says what was found, by whom, and when. None of them attributes a motive, because intent is not observable from outside and this product publishes measurements.

2025-04-29

Private pre-release testing on a public leaderboard

A peer-reviewed audit found that undisclosed private testing lets providers try several variants before a public release and withdraw scores they do not like, and estimated a gain of up to 112 per cent relative on ArenaHard from the additional arena data. The finding is about the distributional consequence of a policy, which the audit quantified and the policy itself did not.

Singh et al., NeurIPS 2025 (peer-reviewed)

2025-05-01

The operator’s reply, published beside the audit

LMArena responded that the policy had been publicly available since March 2024, that any provider may submit as many public and private variants as capacity allows, and that larger labs submit more models because they build more. It also announced changes: clearer marking of retired models, and provisional scoring until fresh post-release votes accumulate where many models underwent parallel pre-release testing.

LMArena (first-party)

2025-04-06

The evaluated system was not the released system

Meta submitted a variant described as experimental and tuned for human preference, which ranked second, rather than the weights it released for download. Meta has said it described the entry as an experimental chat version. The checkable claim is the artefact difference between what was evaluated and what shipped, which nobody has disputed.

TechCrunch (reporting)

2026-05-12

A benchmark that could not be built

MathArena reported that it could not ship a 2026 set at all, and said the earlier set could by then be contaminated. The maintainers withdrew a planned benchmark on their own assessment of its problem supply. That is task-source exhaustion reported by the people building the tasks, and no detector was involved in establishing it.

MathArena (first-party)