// news · alignment2026-08-04source: openai

OpenAI and Hugging Face jointly disclose a security incident that occurred during model evaluation

OpenAI and Hugging Face published a joint disclosure of a security incident that arose during model evaluation on the platform. The details matter less than the precedent: two of the ecosystem's central institutions treating an evaluation-time failure as something the public gets told about, together, on the record.

Evaluation infrastructure is infrastructure, and it just had its first jointly-disclosed incident. Model evaluation now involves running powerful, partially-understood systems against shared platforms at scale — an attack surface nobody priced in when leaderboards were toy benchmarks. An incident arising in that seam, between a frontier lab and the platform hosting the ecosystem's models, was close to inevitable.

The joint disclosure is the notable behavior. The reflex in both corporate security and AI development has historically been silence; a co-authored public account instead treats evaluation-time failures the way mature industries treat aviation incidents — as information the whole field needs. It follows a summer in which labs disclosed models breaching real organizations during sanctioned security tests, and it extends that norm from capability findings to operational failures.

The uncomfortable read is that the disclosures are accelerating because the incidents are. As models gain the agentic capability to act on real systems and evaluations increasingly exercise that capability, the line between 'testing a model' and 'operating a powerful system in production' has dissolved. Evaluation pipelines now need the same security engineering as deployments — that is the lesson both companies are implicitly conceding.

See our analysis →

OpenAI — OpenAI and Hugging Face partner to address security incident during model evaluation →