// news · research-papers2026-08-02source: arxiv / arxiv

PaperBench asks whether AI can replicate AI research — and a companion paper maps where reasoning collapses

PaperBench evaluates an AI system's ability to replicate published AI research — reproducing results from a paper end to end — while 'Logical Phase Transitions' studies the point at which LLM logical reasoning abruptly collapses. Together they measure two edges of the field: how far automated research can reach, and where reasoning reliably breaks.

PaperBench targets a capability with unusual leverage. A system that can replicate AI research is a system that can accelerate AI research, so measuring it is measuring how close the field is to a feedback loop where models help build models. Framing replication as a benchmark makes that progress trackable rather than anecdotal, and honest about how far it still has to go.

The logical-phase-transitions work is the sobering counterweight. Studying the point at which a model's logical reasoning collapses — a sharp transition from competent to failing as problem complexity rises — maps the boundary of reliable reasoning rather than celebrating its peak. Knowing where a capability breaks is as valuable as knowing how far it extends.

Read together, the two papers describe a field measuring its own edges. One asks how much of research the models can take over; the other asks where their reasoning stops being trustworthy. Both are the apparatus of a science that has moved past demonstrating capability and into characterising it — the same reproducibility turn visible across the interpretability and reasoning work this year.

See our analysis →

arXiv — PaperBench: evaluating AI's ability to replicate AI research → · arXiv — Logical phase transitions: understanding collapse in LLM logical reasoning →