The shape of a jailbreak
Under attack, the latent space gets structurally simpler — across six models and two attack types. If that generalises, it is a different kind of defence.
Why generalisation would matter
Nearly every jailbreak defence so far has been specific: trained on known attacks, defeated by the next phrasing. A geometric property that changes under attack regardless of the attack's form detects the state rather than the input. That is a different category of defence.
The intuition is plausible. An adversarial prompt pushes the model toward a narrow region, and a representation collapsing toward one region is a simpler object.
Two things to push on
Six models and two attack types is a small grid to found a universality claim on, and two adversarial conditions is where a sceptic starts. Any monitor published as generalisable also becomes a target — an attacker who knows the detector measures topological simplicity has a stated objective to optimise against.
The direction is right regardless
This is interpretability producing a runtime signal instead of a post-hoc explanation, which is the same move as adaptive steering audits that test the mechanism rather than the surface.
Both answer the complaint that behavioural testing only returns a lower bound. Neither is a certificate either — but measuring the representation is not the same activity as searching the input space, and the failure modes are different.
The uncomfortable adjacent claim
Separately, there is an argument that a different architecture would simply be more legible — that opacity is a property of the design, not a difficulty to be tooled around.
If that holds, much of the current effort is recovering structure a different design would never have destroyed. The objection is that nobody has shown a legible architecture reaching the frontier. The trade has been asserted from both sides and measured by neither.
What this asks of you
Take your best explanation of a behaviour and try to change the behaviour with it. If you cannot, you have a lens. The field keeps selling levers.
arXiv — The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology → · arXiv — Enhancing AI Interpretability and Safety through Localised Architectures →