Prompt-specific evidence is the standing caveat in circuit work
The circuit-tracing results on multi-step reasoning, hallucination, refusal and jailbreak were prompt-specific. That qualifier is stated plainly by the researchers and dropped almost everywhere else, and it is the difference between a finding about a model and a finding about a prompt.
Prompt-specific means the circuit was identified for particular inputs. It does not establish that the same mechanism fires for related prompts, that it is the only mechanism available, or that it generalises to a different model. Those are three separate claims and none follows from the first.
The gap matters most for the security case. Locating the circuit responsible for a refusal on one jailbreak tells you about that jailbreak. An attacker's job is to find inputs that route around it, and the method's coverage against inputs nobody has tried is unmeasured.
None of which is an argument against the work — the researchers state the caveat themselves. It is an argument against the compression the caveat undergoes on the way to a headline, and a reason that production deployment of these methods needs coverage numbers rather than example circuits.
Anthropic — Interpretability research team → · arXiv — Mechanistic interpretability for neural networks: circuits, sparse features and symbolic reasoning → · arXiv — Understanding emergent misalignment via feature superposition geometry →