The evaluation queue never empties now
Eleven-plus models in twenty days, adoption within hours, and cost per intelligence unit halving. The bottleneck has moved from building models to knowing whether they are any good — and it is not moving back.
More than eleven releases in twenty days across five-plus providers, with anonymous frontier models reaching production within hours. Careful independent evaluation of one model takes days to weeks. The arithmetic does not close, and it will not close by trying harder.
What a permanent backlog changes
A temporary backlog is a resourcing problem. A permanent one is a structural change in who gets to make claims. If independent evaluation cannot reach a model before the market does, then at the moment of decision the only available numbers are the vendor's — which is precisely the condition independent benchmarking was invented to end.
We have been here before, in other industries. It resolves one of two ways: accredited third-party testing that buyers pay for, or buyers doing their own testing and keeping the results. Nothing suggests the first is close.
The pricing half makes it worse
Cost per unit of intelligence fell roughly 50% this month across several tiers at once. A simultaneous fall across tiers is a cost structure changing, not one vendor discounting.
Cheap models get adopted casually. Casual adoption is exactly the case where nobody evaluates, because the perceived stake is low — and the aggregate of many low-stakes integrations is an organisation whose behaviour depends on models nobody has assessed.
The evaluation that survives
Not the rigorous public benchmark; it cannot run often enough. What survives is small, private and repeatable: a hundred examples from your own distribution, a pass criterion you actually care about, and a runtime measured in an afternoon.
It is scientifically weaker than a proper benchmark and operationally far stronger, because it can be run every time something new appears — and something new now appears every other day.
The claim to stop making
"State of the art on X" has a shelf life measured in days. Worse, as the ICLR safety track keeps finding, single-turn scores do not transfer to long multi-step work — which is what agents are. A leaderboard position is now a weaker signal than at any point since leaderboards existed.
Axis Intelligence — AI Model Release Tracker 2026 → · LLM Stats — AI Updates Today (August 2026) — Latest AI Model Releases → · AI Release Tracker — Latest AI Model Releases — August 2026 → · Local AI Zone — Latest AI Developments: August 2026 Update →