// news · interpretability2026-08-14source: Oxford AI Governance Initiative

A research agenda for automated interpretability-driven auditing and control

An Oxford AI Governance Initiative agenda proposes automating interpretability into an audit function, naming three applications: chain-of-thought faithfulness verification, predicting emergent capabilities during training, and mapping how capabilities compose.

Interpretability has mostly been practised as investigation — a researcher forms a hypothesis about a model and tests it. Auditing is a different discipline with different requirements: repeatable procedure, defined scope, and a result that means something to someone who was not in the room.

The three named applications are well chosen because each has a decision attached. Faithfulness verification determines whether a reasoning trace can be used for oversight. Predicting emergent capability during training determines whether a run should continue. Mapping capability composition addresses the case where individually safe capabilities combine into something that is not.

Automation is the load-bearing word and the hard part. Hand-analysis does not scale to the number of models, versions and fine-tunes actually deployed, so an audit function that requires a specialist per model is a bottleneck rather than a control. Automated methods that produce domain-grounded explanations experts can check are the only version that reaches production.

The timing is not accidental. Some models can now distinguish evaluation from deployment, which erodes behavioural testing precisely where it matters. Reading internals is the check that does not depend on the model cooperating with the test.

See our analysis →

Oxford AIGI — Automated Interpretability-Driven Model Auditing and Control: A Research Agenda → · arXiv — Training Language Models to Explain Their Own Computations →