Featured
Benchmarks23 Feb 20265 min readOpenAI quietly stopped using SWE-bench Verified. So did we.
Their own audit of 138 hard tasks found 59.4% carried material test or prompt defects, and every frontier model they tested showed signs of having seen some of them. Here is what we replaced it with, and why SWE-rebench is the only coding leaderboard still doing real work.
Read itEverything else, newest first
The score you’re quoted isn’t the score you get
Vendors benchmark at max effort. OpenAI ships you medium by default, Grok 4.3 ships you low. We measured the gap across nine model families and it is wider than we expected.
Read it Method2 Jun 20266 min readEvery benchmark we threw away, and why
Saturated, defective, or a licence we cannot honour. The list halved, and the refusals turned out to be more informative than the rankings. Here is what went and on what grounds.
Read it