// blog · analysis · research-papers2026-08-16source: arXiv

Automating the search

A system that runs the research loop without a human at each step. The question is not whether it finds anything — it is what happens to a literature that already cannot verify what it publishes.

AutoSOTA describes an end-to-end automated research system for discovering state-of-the-art models — hypothesis, experiment, evaluation, without a person in the loop at each step.

The second-order question is the interesting one

Whether it finds something a human would not is a narrow technical matter. What it does to the literature if it works is not.

Automated research systems produce results at machine throughput. The literature has no mechanism for absorbing results faster than humans can read them, and machine learning already has a replication problem driven by volume.

Increasing generation without increasing verification makes the existing failure worse in exactly the dimension it is already failing.

The optimistic reading

Automated search is genuinely best suited to the work humans are worst at — exhaustive architecture and hyperparameter exploration, where the bottleneck is patience rather than insight. Automating that frees people for the part that requires taste.

That is a real benefit and it is not the whole picture. Which effect dominates depends on whether verification automates as fast as generation, and right now it is not close.

The measurement side is having the same argument

ARC-AGI-3 has been framed as a challenge for agentic intelligence rather than for question answering — the familiar arc where a benchmark is proposed as a general measure, models improve on it, the improvement proves partly benchmark-specific, and a successor gets written.

Agentic framing is harder to game, because an agent must sequence actions in a responsive environment and there is no single output to optimise toward. It is also harder to score, harder to reproduce, and far more sensitive to scaffolding — which means the number depends on engineering choices that have nothing to do with the model.

The tension the whole field sits in

The measurements closest to what we care about are the least reliable. The reliable ones measure things that stopped being informative years ago.

ARC-AGI-3 chose relevance over reliability, which is defensible, and the reproducibility arguments will follow. AutoSOTA chose throughput, and the verification arguments will follow that. Both are the same problem viewed from opposite ends: we can generate faster than we can check, and nothing on the horizon changes the ratio.

arXiv — AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery → · arXiv — ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence → · Skycrumbs — AI Research Highlights: The Breakthroughs of August 2026 →