SaferAI ran every offensive cyber and dual-use biology prompt at GLM-5.2. It completed all of them.
An independent evaluation against the four systemic risk areas in the EU GPAI Code of Practice found Z.ai's open-weight flagship within 2-4 months of frontier cyber capability — and refusing nothing. Claude Opus 4.7 refused so consistently that the same benchmark could not be completed against it at all. Z.ai has published no safety framework, no pre-deployment testing commitment, and no risk assessment.
The capability finding is the smaller half. Cyber performance comparable to frontier models from two to four months earlier is roughly what the open-weight trajectory predicted, and on its own it is a scheduling observation.
The refusal finding is the story. Not a low refusal rate — zero. Every offensive cyber and dual-use biology prompt in the suite was completed. The contrast in the same report is stark: Opus 4.7 declined so consistently that CyberGym could not be run against it to completion.
That is the difference between a model with frontier-adjacent capability and a model with frontier-adjacent capability plus a safety layer, measured side by side by the same evaluator on the same tasks. It isolates the variable more cleanly than most published safety comparisons manage.
And the governance gap is documented rather than inferred: no published safety framework, no pre-deployment testing commitments, no risk assessment. Testing against the EU Code of Practice's four systemic risk areas — loss of control, cyber offence, CBRN, harmful manipulation — makes this directly relevant to a regime that is now enforcing GPAI obligations, against a developer outside its jurisdiction.
SaferAI — GLM-5.2 risk evaluation report → · SaferAI — GLM-5.2 risk evaluation report, 2 August 2026 (PDF) → · The Next Web — Open-weight AI caught the frontier on capability. On safety, it didn't →