// blog · analysis · interpretability2026-08-18source: arXiv

The shape of a jailbreak

Under attack, the latent space gets structurally simpler — across six models and two attack types. If that generalises, it is a different kind of defence.

Persistent homology finds a shared topological compression signature across six models under two adversarial conditions.

Why generalisation would matter

Nearly every jailbreak defence so far has been specific: trained on known attacks, defeated by the next phrasing. A geometric property that changes under attack regardless of the attack's form detects the state rather than the input. That is a different category of defence.

The intuition is plausible. An adversarial prompt pushes the model toward a narrow region, and a representation collapsing toward one region is a simpler object.

Two things to push on

Six models and two attack types is a small grid to found a universality claim on, and two adversarial conditions is where a sceptic starts. Any monitor published as generalisable also becomes a target — an attacker who knows the detector measures topological simplicity has a stated objective to optimise against.

The direction is right regardless

This is interpretability producing a runtime signal instead of a post-hoc explanation, which is the same move as adaptive steering audits that test the mechanism rather than the surface.

Both answer the complaint that behavioural testing only returns a lower bound. Neither is a certificate either — but measuring the representation is not the same activity as searching the input space, and the failure modes are different.

The uncomfortable adjacent claim

Separately, there is an argument that a different architecture would simply be more legible — that opacity is a property of the design, not a difficulty to be tooled around.

If that holds, much of the current effort is recovering structure a different design would never have destroyed. The objection is that nobody has shown a legible architecture reaching the frontier. The trade has been asserted from both sides and measured by neither.

What this asks of you

Take your best explanation of a behaviour and try to change the behaviour with it. If you cannot, you have a lens. The field keeps selling levers.

arXiv — The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology → · arXiv — Enhancing AI Interpretability and Safety through Localised Architectures →