// news · alignment2026-08-18source: Research analysis

A passing red-team evaluation is a lower bound, not a safety certificate

Formal analysis makes explicit what practitioners already suspected: an evaluation that finds nothing tells you what your team could not find, which is a statement about your team.

Recent formal analysis of AI safety evaluations states the position directly: a passing red-team evaluation is a lower bound on dangerous capability, not a safety certificate. A separate position paper formalised the same concern from a governance angle.

The logic is not subtle and that is what makes it awkward. Red-teaming searches a space of inputs for one that produces bad behaviour. Finding one proves the capability exists. Failing to find one proves that the search did not find it — a statement about the search, its budget, and the people running it.

Practitioners have known this. What changes when it is formalised is that governance frameworks can no longer treat an evaluation report as evidence of safety, because there is now a citable argument that it is not the kind of evidence it was being used as.

That is a problem, because evaluations are what every emerging regulatory regime leans on. Obligations to assess risk before deployment presuppose that assessment can establish something. If the strongest available instrument yields a lower bound, then "we evaluated and found nothing" is compatible with a serious hazard, and the regime built on it inherits that gap.

The constructive reading is that evaluations should be reported the way measurements are — with the search budget, the methods and the coverage stated, so a reader knows what was looked for. "We found nothing" means very little; "we ran these attacks with this budget and found nothing" means something. That points the same direction as work looking for internal signatures rather than behavioural ones.

See our analysis →

TechTimes — AI Safety Evaluations Are Not Safety Certificates: Formal Analysis → · arXiv — Reasons to Doubt the Impact of AI Risk Evaluations →