// news · alignment2026-08-18source: arXiv

A framework for measuring uplift, which is the only question that matters

Not whether a model can describe something dangerous, but whether it meaningfully helps someone who wants to do it. Uplift is the right question and the hardest one to measure.

A framework for evaluating human-AI safety proposes measuring harmful capability uplift — the marginal advantage a model provides to a person attempting harm, relative to what they could achieve without it.

Uplift is the correct framing and the field has been slow to adopt it. Most capability evaluations ask whether a model can produce dangerous content. That question has been answerable since search engines existed, and answering it yes tells you almost nothing about risk, because the relevant comparison was never zero.

The design problem is that measuring uplift honestly requires a control group of people attempting the harmful task without the model, which is exactly the study nobody can run. The substitutes — expert judgement, literature comparison, time-to-completion proxies — are all weaker, and the framework's value depends heavily on which one it defends.

There is also a moving baseline problem that gets less attention than it should. Uplift is measured against publicly available alternatives, and those alternatives now include other models. As open-weight capability rises, the marginal uplift of any single frontier model falls, which means the same model can become "safer" by this metric without changing at all.

That is not a reason to reject the framing — it is the right question and a moving baseline is a real feature of the risk, not an artifact. But it does mean uplift figures need a date and a stated comparison set, or they measure the ecosystem rather than the model.

See our analysis →

arXiv — Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift → · arXiv — Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods →