A paper argues the RL rig is disproportionate to the problem
REASONMAXXER matches full reinforcement learning on clean comparisons, with costs reported against on-demand GPU pricing. If heavy post-training infrastructure is oversized for the task, a lot of spending is being misallocated.
A paper rethinking reinforcement learning for LLM reasoning reports that its method, REASONMAXXER, matches full RL on clean comparisons, and concludes that heavy RL infrastructure for post-training may be disproportionate to the complexity of the problem.
The methodological choice worth noting is that the authors report estimated monetary cost, measured on RTX Pro 6000 GPUs at RunPod on-demand pricing. Most post-training papers report compute in GPU-hours or omit it, which makes results incomparable across labs with different hardware access. Dollars at a public rate is a number anyone can check.
"Clean comparisons" is carrying weight in that claim, and deserves scrutiny. RL comparisons are notoriously sensitive to tuning budget, and a method that matches full RL under matched-and-limited tuning has demonstrated something narrower than a method that matches it at each approach's best. The paper's framing is careful; secondhand summaries of it will not be.
If the result holds, the implication is uncomfortable for the current structure of the field. Elaborate post-training pipelines are one of the moats large labs are assumed to have. A cheaper method reaching the same place moves capability toward whoever has good data and good judgement, rather than whoever has the largest cluster.
That is the same direction as work getting small models off the floor with classical search traces. Both point at post-training being under-optimised rather than under-resourced.
arXiv — Rethinking RL for LLM Reasoning → · Sebastian Raschka — LLM Research Papers: The 2026 List (January to May) →