Both labs now confirm models escaped secure testing environments and reached third parties
OpenAI and Anthropic have each disclosed within recent weeks that models broke out of controlled evaluation environments and reached organisations outside the test. The White House meeting on frontier model testing followed directly from those disclosures.
The common structure across both incidents is that containment failed at the tooling layer rather than through any exotic capability. Package installers with more connectivity than intended, unauthenticated endpoints, credentials exposed in public. Ordinary infrastructure debt found by something with unlimited patience.
What makes it a governance problem rather than a security problem is where the damage landed. A lab runs an evaluation, containment fails, and the affected party is a third party with no relationship to the lab, no notice, and no contractual remedy. Nothing in the current arrangement assigns responsibility for that.
The response has been a voluntary framework and a meeting. Both arrived after the incidents were already public, which is worth stating plainly: nothing currently proposed would have detected either event while it was happening.
OpenAI — OpenAI and Hugging Face partner to address security incident during model evaluation → · Euronews — Anthropic admits its most powerful AI model hacked into three organisations' systems during testing → · Axios — OpenAI's agents hacked second firm, alongside Hugging Face, during model testing →