// news · interpretability2026-08-07source: arXiv

Steering robustness into world-action models using interpretability and optimal control

A line of work combines mechanistic interpretability with optimal control to steer robustness into world-action models — using knowledge of internal structure to shape behaviour rather than only to describe it.

Combining interpretability with control theory is a more promising pairing than it might sound. Control theory is the mature discipline for steering a system toward a desired trajectory given a model of its dynamics, and interpretability is what supplies the model of the dynamics.

The practical appeal is that it addresses robustness rather than accuracy. A policy that performs well on the test distribution and degrades unpredictably outside it is the standard failure, and average-case training does nothing about it.

It also inherits the field's central uncertainty. Steering demonstrates that a direction in activation space is causally relevant; it does not establish that it is the only one, or that suppressing it removes a capability rather than one expression of it.

See our analysis →

arXiv — Understanding emergent misalignment via feature superposition geometry → · arXiv — Probing and steering evaluation awareness of language models → · arXiv — Evaluation awareness is not one capability: evidence from open language models →