The “Age of LLM” benchmark pits models against each other under fog of war, testing diplomacy and reliability rather than recall
A turn-based 1v1 benchmark places two models in direct competition under incomplete information, with diplomacy and reliability as measured dimensions. It is part of a broader move away from static question sets toward evaluations where the adversary is another model.
Static benchmarks have a structural problem that has become impossible to ignore: they are finite, they leak into training data, and a model can be optimised against them without becoming better at anything. An opponent-based evaluation has no fixed answer key to memorise, because the correct move depends on what the other model does.
Fog of war is the design choice that makes it a reliability test rather than a reasoning test. Under incomplete information a model must decide how much to trust its own inference and when to act anyway — which is precisely the failure mode that shows up when agents run unsupervised over long horizons.
The limitation is that relative results do not straightforwardly convert into absolute capability claims. Model A beating Model B tells you about the pairing. Building a rating system on top of that has worked for chess and for Arena-style comparisons, and it will need the same volume of matches here before the numbers mean much.
arXiv — Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability under Fog of War → · arXiv — Machine Learning listings, July 2026 →