ARC-AGI-3 moves the target from answers to agency
The third iteration is framed as a challenge for frontier agentic intelligence rather than for question answering. Benchmarks that saturate get replaced by benchmarks that measure something harder to fake.
ARC-AGI-3 has been published as a challenge aimed at frontier agentic intelligence — a deliberate shift from the earlier iterations' framing around abstract reasoning over static puzzles.
The move follows a familiar arc. A benchmark is proposed as a measure of general capability, models improve on it, the improvement turns out to be partly benchmark-specific, and a successor is written to close the gap between what is measured and what was meant.
Agentic framing is harder to game in one specific way: an agent has to sequence actions in an environment that responds, so there is no single output to optimise toward. It is also harder to score, harder to reproduce, and far more sensitive to scaffolding — which means the number depends on engineering choices that have nothing to do with the model.
That is the tension the whole evaluation field is sitting in. The measurements closest to what we care about are the least reliable, and the reliable ones measure things that stopped being informative. ARC-AGI-3 chooses relevance over reliability, which is defensible, and the community should expect the reproducibility arguments that come with it.
Related, from earlier this week: the field has been benchmarking its own benchmarks.
arXiv — ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence → · Skycrumbs — AI Research Highlights: The Breakthroughs of August 2026 →