Six years of a workshop, read as a history of the field's ambitions
A retrospective covering six years of TrustNLP traces the arc from interpretability to control. The title is the argument: understanding was never the destination.
A retrospective paper covering six years of the TrustNLP workshop traces the field's movement from interpretability toward control.
Workshop retrospectives are underrated as evidence. Conference proceedings show what was accepted; a six-year workshop history shows what a community kept choosing to talk about, including the directions that were abandoned. That record is harder to reconstruct from citation counts.
The stated arc — interpretability to control — names something the field has been slow to say directly. Understanding a model's internals was rarely the terminal goal. It was instrumental: the reason to know which circuit produces a behaviour is to be able to change the behaviour without retraining the model.
Saying it out loud has consequences for how the work is evaluated. An interpretability result judged as understanding is assessed on whether the explanation is faithful. The same result judged as control is assessed on whether intervening on the identified structure actually changes the output in the predicted direction — a considerably harder bar, and one many published explanations have not been tested against.
It also raises a question the retrospective format is well suited to answer and that summaries of it tend to skip: which early directions were dropped, and were they dropped because they failed or because attention moved. The design-principle argument is the same claim from the other end.
arXiv — From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop → · arXiv — Open Problems in Mechanistic Interpretability →