// blog · analysis · interpretability2026-08-15source: medRxiv / SPAR / arXiv

Looking for deception on purpose

Interpretability is being pointed at a specific target: systems that behave one way under test and another in deployment. The methodological trick is picking a domain where correct behaviour is externally defined.

AlignInsight targets deceptive alignment and evaluation awareness in healthcare AI, while research agendas turn toward goal-directed behaviour in agents.

Why healthcare

Because "pursuing a different objective" is nearly unfalsifiable without a specification of the right one. Clinical correctness is externally defined and already audited, which supplies the ground truth domain-general work lacks.

The domain does not generalise. The mechanism does — behaviour conditioned on context that differs between test and use, which exists in every deployed system whether anyone designed it in or not.

Say the threat model precisely

A system that behaves acceptably during training and validation and differently once deployed. That requires no intent, no self-awareness and no planning. Distributional difference between validation and deployment is sufficient, and that difference is universal.

The scary word is doing work the evidence does not require.

The risk to the research programme

Method outrunning validity. Interpretability went through this with sparse autoencoders: fast adoption, then a wave of work on whether the readings meant what people took them to mean. Aiming the same toolkit at something as underspecified as "goals" invites the identical sequence.

Doing the validity work first would be cheaper. It is also unlikely, because behavioural testing degrades exactly where it matters most and reading internals is the only check that does not need the subject to cooperate.

The demand is arriving before the instrument is calibrated. That is how it always goes, and knowing it is how you discount the first results appropriately.

medRxiv — AlignInsight: A Three-Layer Framework for Detecting Deceptive Alignment and Evaluation Awareness in Healthcare AI Systems → · SPAR — Spring 2026 Projects → · arXiv — An Approach to Technical AGI Safety and Security →