// news · policy · alignment2026-08-06source: policy reporting

Secret safety measures: the disclosure problem nobody is arguing about

Officials and companies are working on safety measures that will not be published. There is a real argument for that — capability thresholds are dual-use information and publishing an exact bar tells a bad actor precisely where to sit beneath it. The argument covers specific numbers. It does not cover process, participants, or compliance criteria, and all three are dark.

The steelman deserves stating properly. A published threshold is a specification for evasion. If the rule says a model is covered when it exceeds a stated score on a stated benchmark, then the rule has told every developer exactly how to build something that scores one point lower. Secrecy about numbers is defensible.

What secrecy about numbers does not require is secrecy about everything else. Who was consulted. How participation is verified. Whether a compliance claim can be checked by anyone outside. Those are all publishable without giving away a threshold, and none of them have been published.

The result is a regime that must be taken on trust at precisely the moment its subjects have demonstrated that their own controlled environments do not hold. That is not an argument against the framework. It is an argument that the framework has taken on an evidentiary burden it has chosen not to discharge.

See our analysis →

Defense One — As AI models break free, White House works with firms on secret safety measures → · Quartz — White House to review AI cybersecurity framework with top labs → · American Bazaar — White House to meet Meta, OpenAI, Google, Anthropic on AI safety testing after Hugging Face incident →