BusinessCaseBench measures judgement under uncertainty, not question answering
A case-grounded benchmark of knowledge work tests synthesis of complex information, judgement under uncertainty and strategic thinking across multi-stakeholder settings. Submitted in July and revised on 3 August, it targets the analytical work white-collar professionals do rather than the tasks benchmarks usually measure.
The gap it addresses is the one every enterprise buyer runs into. A model can top coding and mathematics leaderboards and still be unreliable at the thing an analyst is paid for — weighing incomplete evidence, reconciling stakeholders with different interests, committing to a recommendation.
Case-grounded is the design decision that makes it hard to game. Cases have no clean answer key, which forces graded judgement over string matching, and it is exactly that property that makes the benchmark expensive to run and difficult to saturate.
Whether it survives contact with the field is a separate question. The measurement literature is increasingly turning on itself, and a benchmark's half-life is now counted in months.