// news · research-papers2026-08-14source: arXiv

Faithfulness scores move depending on which classifier you score them with

Work on classifier sensitivity in chain-of-thought evaluation finds that measured faithfulness depends materially on the classifier used to measure it — which means published faithfulness numbers are only comparable when the measurement apparatus is identical.

Faithfulness evaluations typically use a classifier to decide whether a reasoning trace acknowledges an influence. If that classifier's choice changes the score, then the score is a property of the pair — model and classifier — and not of the model alone.

The consequence for the literature is immediate and unglamorous: two papers reporting different faithfulness numbers for the same model may not disagree about anything except instrumentation. Without reported apparatus, the comparison cannot be made.

This is the same lesson interpretability learned about sparse autoencoders — when the instrument has free parameters, the reading inherits them, and validity work has to happen before the results are used as evidence. Fields tend to learn it once per instrument.

It is a good-faith result, not a debunking. Faithfulness remains worth measuring, and the gap between what a reasoning trace acknowledges and what the answer states is large enough to survive a lot of measurement noise. The instruction is to report the apparatus, not to stop measuring.

See our analysis →

arXiv — Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation → · arXiv — Counterfactual Simulation Training for Chain-of-Thought Faithfulness →