“LLM Reasoning Is Latent, Not the Chain of Thought” has reshaped how the field talks about reasoning
The paper argues that the reasoning driving a model's answer happens in latent space, and that the emitted chain of thought is not that process. It has become one of the more influential results in the reasoning community, and the second half of 2026 has been organised around its implications.
The claim cuts against a comfortable assumption. If chain-of-thought text were the reasoning, then improving reasoning would mean improving that text, and monitoring reasoning would mean reading it. If the text is a rendering of a process happening elsewhere, both activities are aimed at the wrong object.
The follow-on work is visible in the literature: token-wise latent-explicit routing, adaptive information control for search-augmented reasoning, and a general shift toward treating the latent trajectory as the thing to optimise. That is a research agenda reorganising around a single load-bearing result.
The caution worth registering is that “latent, not the chain of thought” is a strong claim and a convenient one — it licenses ignoring an inconvenient monitoring surface. The result deserves the attention it is getting and also deserves the replication scrutiny that influential results in this field have not always received.
arXiv — LLM Reasoning Is Latent, Not the Chain of Thought → · Sebastian Raschka — LLM Research Papers: The 2026 List →