// news · interpretability2026-08-18source: arXiv

Six models, two attacks, one shared signature: the latent space gets simpler

Persistent homology applied to LLM latent space finds a "topological compression" signature under adversarial conditions — and it appears across six models and two attack types.

Work applying persistent homology to characterise LLM latent space at multiple scales reports that across six state-of-the-art models under two adversarial conditions, a common signature appears: the latent space becomes structurally simpler. The authors suggest such monitors may generalise across architectures and attack types.

Generalisation is the claim that would matter. Nearly every jailbreak defence to date has been specific — trained on known attacks, defeated by the next phrasing. A geometric property of the representation that changes under attack regardless of the attack's form would be a defence of a different kind, because it detects the state rather than the input.

The intuition is plausible. An adversarial prompt is pushing the model toward a narrow region of behaviour, and a representation collapsing toward that region is a structurally simpler object than the same representation handling ordinary input. That is the kind of thing topology is built to measure.

Two cautions. Six models and two attack types is a small grid to found a universality claim on, and "two adversarial conditions" is the number a sceptic would push on first. And any monitor published as generalisable becomes a target: an attacker who knows the detector measures topological simplicity has a stated objective to optimise against.

Still, this is the shape of interpretability work that could be operationally useful — a runtime signal rather than a post-hoc explanation. It answers, at least partially, the complaint that behavioural evaluation only gives lower bounds: an internal signature is not a search over inputs.

See our analysis →

arXiv — The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology → · ICLR 2026 — Oral — The Shape of Adversarial Influence →