Read asAgentic codingclear
Nobody has measured any of these on this kind of task, so we are not going to name a best one. What is below is what exists, drawn as the gaps they are.
| Attribute | Kimi K2.6Moonshot AI | Grok 4.5xAI | Gemini 3.5 FlashGoogle |
|---|---|---|---|
| In your plans | Not held · Kimi Moderato, $19/mo | Not held · SuperGrok, $30/mo | Not held · Google AI Pro, $19.99/mo |
| Effort you get | maxthe vendor benchmarks here too | highthe vendor benchmarks here too | highthe vendor benchmarks here too |
| Evidence | Nothing measured | Nothing measured | Nothing measured |
| This task | $0.058k in · 1.5k out · agentic ×4 | $0.108k in · 1.5k out · agentic ×4 | $0.108k in · 1.5k out · agentic ×4 |
| Context | 262K | 500K | 1049K |
Scores are shown at the effort a normal caller gets, not the effort the vendor benchmarked at. A hatched cell means nobody has run it — a gap in what is known, never a zero.
Don’t take our word for it
Run your prompt on the shortlist and read what comes back. A ranking is an argument from other people’s measurements; this is the thing itself.
What you are about to read is model output, not measurement. These are single generations, unreplicated and ungraded. They are here so you can judge the writing yourself — they are not evidence, they carry no score, and nothing below feeds the ranking above.