Mechanistic interpretability turns toward goal-directed behaviour in agents
Current project agendas emphasise using mechanistic interpretability to understand goal-directed behaviour in AI agents, with detection of safety-critical properties such as deceptive alignment as an explicit target. The unit of analysis is shifting from what a model represents to what it is trying to do.
Interpretability has mostly asked representational questions: what does this feature encode, what does this circuit compute. Those are answerable and useful. They are also not the questions that come up when an agent takes an action nobody sanctioned.
Goal-directedness is a harder object of study because it is not obviously a thing in the weights. It may be an emergent property of a policy interacting with an environment, in which case looking for it inside the model is a category error — and establishing that would itself be a result worth having.
The safety motivation is explicit and drives the choice. Behavioural testing degrades exactly when it matters most, because a system that detects evaluation can pass one. Internals do not have that failure mode, which makes reading them the only check that does not require the subject to cooperate.
The risk is method outrunning validity. Interpretability has been through this with sparse autoencoders — rapid adoption, then a wave of work on whether the readings meant what they were taken to mean. Aiming the same toolkit at something as underspecified as "goals" invites the same sequence, and it would be cheaper to do the validity work first.
SPAR — Spring 2026 Projects → · arXiv — An Approach to Technical AGI Safety and Security →