Grounded world models and the gap between predicting and planning
Work on grounded world models for semantically generalizable planning targets the difference between a model that can predict what happens next and one that can plan toward a goal it has never been given before.
Prediction and planning look adjacent and are not. A model that accurately forecasts the next frame has learned dynamics. A model that can plan has learned dynamics plus a way to search over them toward an objective, and the second capability does not fall out of the first.
Semantic generalisation is the harder half. Planning within a known task distribution is well-studied; planning toward a goal described in language that the training set never contained requires the world model and the language grounding to share a representation.
The reason this matters beyond robotics is that it is the same architecture question facing any agent operating over long horizons in an environment it cannot fully observe.
arXiv — Grounded world model for semantically generalizable planning → · arXiv — Agents' Last Exam → · arXiv — LLM reasoning is latent, not the chain of thought →