It refused nothing
An independent evaluator ran every offensive cyber and dual-use biology prompt it had at an open-weight frontier model. The model completed all of them. That is the finding, and the capability number is the smaller half of it.
SaferAI evaluated GLM-5.2 against the four systemic risk areas in the EU GPAI Code of Practice — loss of control, cyber offence, CBRN, harmful manipulation. Cyber capability landed within two to four months of the closed frontier. Refusals: zero.
The comparison that isolates the variable
The same report notes Claude Opus 4.7 refused so consistently that CyberGym could not be completed against it at all.
Same evaluator, same tasks, same window. One model has frontier-adjacent capability. The other has frontier-adjacent capability plus a safety layer. That is a cleaner controlled comparison than most published safety work manages, and it makes the safety layer visible as a distinct component rather than an assumed property.
Why months is the number that matters
Every voluntary framework in operation rests on one assumption: that dangerous capability appears first in a model somebody controls. A lab can hold a release, restrict access, grant a government evaluation window.
That assumption survives only while the frontier is closed. At two to four months, it is a scheduling detail. Whatever triggers a hold at a closed lab shows up shortly afterwards in weights anyone can download, at which point holding is not a lever anybody possesses.
And the governance gap is documented, not inferred
No published safety framework. No pre-deployment testing commitments. No risk assessment. Those are absences of record, not accusations.
Which lands directly on a live regime. Europe is enforcing GPAI obligations now, and this evaluation was run against Europe's own code of practice — on a developer outside its jurisdiction. Two thousand proposals are in flight worldwide and none of them reaches the machine this was tested on.
The honest caveat: capability evaluations in these domains are sensitive to elicitation, and a gap stated in months is a claim about a moving target. But a refusal rate of zero requires no calibration to interpret.
SaferAI — GLM-5.2 risk evaluation report → · SaferAI — GLM-5.2 risk evaluation report, 2 August 2026 (PDF) → · The Next Web — Open-weight AI caught the frontier on capability. On safety, it didn't →