The working

What we found out

We publish the working, including the bits we could not stand up. Every figure below is linked to the source it came from, and where the evidence ran out we say so rather than rounding it up.

3 piecesLatest 11 Jul 2026No sponsored placement

Featured

Benchmarks23 Feb 20265 min read

OpenAI quietly stopped using SWE-bench Verified. So did we.

Their own audit of 138 hard tasks found 59.4% carried material test or prompt defects, and every frontier model they tested showed signs of having seen some of them. Here is what we replaced it with, and why SWE-rebench is the only coding leaderboard still doing real work.

Read it

Everything else, newest first