// blog · analysis · alignment2026-08-18source: Research analysis

A floor, not a certificate

An evaluation that finds nothing tells you what your team could not find. Formalising that is awkward, because every regulatory regime being drafted assumes otherwise.

A passing red-team evaluation is a lower bound on dangerous capability, not a safety certificate. Formal analysis now says so citably.

The logic is not subtle

Red-teaming searches a space of inputs for one that produces bad behaviour. Finding one proves the capability exists. Failing to find one proves the search did not find it — which is a statement about the search, its budget, and the people running it.

Practitioners knew. What changes is that governance can no longer cite the report as evidence of the thing it was being used for.

Why that is a problem right now

Because evaluations are what every emerging regime leans on. Obligations arriving in 2027 and 2028 presuppose that pre-deployment assessment establishes something. If the strongest instrument returns a lower bound, the regime inherits the gap.

Ask the better question

Uplift — the marginal advantage a model gives someone attempting harm — is the right frame, and it is much harder to measure than capability.

Measuring it honestly needs a control group attempting the task without the model, which nobody can run. And the baseline moves: as open-weight capability rises, any single frontier model's marginal uplift falls, so a model can become safer by this metric without changing at all.

That is not a flaw in the framing. It is a real feature of the risk. But it means uplift figures need a date and a stated comparison set, or they measure the ecosystem instead of the model.

What this asks of you

Report evaluations the way measurements are reported: search budget, methods, coverage. We found nothing means very little. We ran these attacks with this budget and found nothing means something.

And take the internal-signal work seriously — a signature in the representation is not a search over inputs, which is precisely why it might answer the complaint.

TechTimes — AI Safety Evaluations Are Not Safety Certificates: Formal Analysis → · arXiv — Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift →