Every documented case of a leaderboard result not meaning what it appeared to mean was found by reading a policy or comparing an artefact, not by running a detector. So this is a register of checkable facts: what a test measures, how much room is left in it, what stops contamination, who may republish it, and what has gone wrong with it.
Every incident carries a primary source and a date. An entry with neither is not published here. Where a figure has not been measured, the register says so rather than printing a zero.
Crowd preference over blind pairwise responses to open-ended text prompts, reported as an Elo-style rating with an interval.
CC BY 4.0Continuous, snapshotted0 incidents recorded
No fixed ceiling
Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it. No bar is drawn, because any bar would have to invent a ceiling in order to have a length.
Human preference between anonymously generated front-end websites.
CC BY 4.0Continuous, snapshotted0 incidents recorded
No fixed ceiling
Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it. No bar is drawn, because any bar would have to invent a ceiling in order to have a length.
Human preference over responses to image-bearing prompts.
CC BY 4.0Continuous, snapshotted0 incidents recorded
No fixed ceiling
Scored as a rating rather than a percentage, so there is no maximum to approach and saturation is not well defined for it. No bar is drawn, because any bar would have to invent a ceiling in order to have a length.
Instruction-following code editing across 225 Exercism exercises in six languages, scored on pass rate after one retry and on whether the edit was well formed.
Apache-2.0Maintainer-run; new models do not appear automatically0 incidents recorded
Text equivalentAider Polyglot: best reproduced result 76.9 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Multi-turn airline customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.
MITManual verified runs; models do not appear automatically0 incidents recorded
Text equivalenttau²-bench airline: best reproduced result 80.3 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Multi-turn retail customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.
MITManual verified runs; models do not appear automatically0 incidents recorded
Text equivalenttau²-bench retail: best reproduced result 79.7 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Multi-turn telecom customer-service agent tasks against a mutable simulated world state, scored by deterministic outcome checks.
MITManual verified runs; models do not appear automatically0 incidents recorded
Text equivalenttau²-bench telecom: best reproduced result 89.5 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Expert-level multiple-choice reasoning in physics, chemistry and biology.
CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
Text equivalentGPQA Diamond: best reproduced result 95.8 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Research-level mathematical problem solving on original, unpublished expert-authored problems.
CC BY 4.0Maintainer-run; models are added when Epoch evaluates them2 incidents recorded
Text equivalentFrontierMath: best reproduced result 47.2 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
The hardest FrontierMath tier, using restricted materials rather than a public answer key.
CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
Text equivalentFrontierMath Tier 4: best reproduced result 31.3 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Resolving real Python repository issues by generating patches that must pass repository tests.
CC BY 4.0Maintainer-run; models are added when Epoch evaluates them2 incidents recorded
Text equivalentSWE-bench Verified: best reproduced result 83.5 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Short-answer factual accuracy and hallucination resistance on fact-seeking questions.
CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
Text equivalentSimpleQA Verified: best reproduced result 71.6 points. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Instruction-following code editing, as collected by Epoch from external reports.
CC BY 4.0Maintainer-run; models are added when Epoch evaluates them0 incidents recorded
Text equivalentAider Polyglot (externally reported): no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
Whether a coding agent can complete original, long-horizon engineering tasks in real repositories, verified by hand-written tests that check behaviour rather than implementation details.
Apache-2.00 incidents recorded
Text equivalentDeepSWE: no reproduced result. Remaining headroom has not been measured, so how much room is left is unknown — not zero.
The practice record
Facts about benchmarking, not about one benchmark.
These are the governance facts that make the case for a register like this one. Each says what was found, by whom, and when. None of them attributes a motive, because intent is not observable from outside and this product publishes measurements.
2025-04-29
Private pre-release testing on a public leaderboard
A peer-reviewed audit found that undisclosed private testing lets providers try several variants before a public release and withdraw scores they do not like, and estimated a gain of up to 112 per cent relative on ArenaHard from the additional arena data. The finding is about the distributional consequence of a policy, which the audit quantified and the policy itself did not.
LMArena responded that the policy had been publicly available since March 2024, that any provider may submit as many public and private variants as capacity allows, and that larger labs submit more models because they build more. It also announced changes: clearer marking of retired models, and provisional scoring until fresh post-release votes accumulate where many models underwent parallel pre-release testing.
Meta submitted a variant described as experimental and tuned for human preference, which ranked second, rather than the weights it released for download. Meta has said it described the entry as an experimental chat version. The checkable claim is the artefact difference between what was evaluated and what shipped, which nobody has disputed.
MathArena reported that it could not ship a 2026 set at all, and said the earlier set could by then be contaminated. The maintainers withdrew a planned benchmark on their own assessment of its problem supply. That is task-source exhaustion reported by the people building the tasks, and no detector was involved in establishing it.