// news · research-papers · frontier-models2026-08-09source: arXiv preprints

Deployment-aware evaluation, and rubrics instead of holistic scores

One paper evaluates open reasoning models under the conditions they actually run in. Another argues for structured rubrics over holistic judgement. Both are attacks on the same weakness: benchmark numbers that do not survive contact with production.

Deployment-aware evaluation is the more directly useful. A model benchmarked at unlimited context, unlimited latency and full precision is not the model anyone runs, and the ranking can invert once quantisation and latency budgets are applied.

The rubric argument attacks the grader rather than the task. Holistic scoring by a judge model is fast, cheap and correlates with almost anything you did not intend to measure. Structured criteria are slower and make disagreement legible, which is the point.

Both sit in the same turn the measurement literature has been taking all year — auditing the instruments rather than adding to the leaderboard.

See our analysis →

arXiv — Unified deployment-aware evaluation of open reasoning language models → · arXiv — From holistic evaluation to structured criteria: rubrics across the evolving LLM landscape →