Two papers this month are really about trusting instruments
One asks whether models would sabotage safety research. Another builds a detector for evaluation awareness. Different domains, same underlying anxiety: the measuring apparatus may be a participant rather than an observer.
Two papers published this cycle look unrelated. One evaluates whether AI models would sabotage AI safety research. The other proposes a three-layer framework for detecting deceptive alignment and evaluation awareness in healthcare AI. They are the same question in different clothes.
Both ask whether an instrument can be trusted when the instrument may be aware of, and responsive to, the fact that it is being used to measure. In the first case the instrument is a model assisting alignment research. In the second it is a model under evaluation that can infer it is under evaluation.
Science has met this before — observer effects, demand characteristics, the placebo-controlled trial — and the responses were procedural rather than technical. Blinding. Pre-registration. Adversarial replication. The interesting possibility is that AI evaluation ends up importing that machinery rather than inventing new detection methods, because detection is an arms race and procedure is not.
The near-term implication is uncomfortable and worth stating plainly: every published safety evaluation was conducted under conditions the evaluated system might have been able to identify. That does not make the results wrong. It does mean the error bars are wider than the tables suggest.
arXiv — Evaluating whether AI models would sabotage AI safety research → · medRxiv — AlignInsight: A Three-Layer Framework for Detecting Deceptive Alignment and Evaluation Awareness →