// news · research-papers2026-08-07source: arXiv

Grounded world models and the gap between predicting and planning

Work on grounded world models for semantically generalizable planning targets the difference between a model that can predict what happens next and one that can plan toward a goal it has never been given before.

Prediction and planning look adjacent and are not. A model that accurately forecasts the next frame has learned dynamics. A model that can plan has learned dynamics plus a way to search over them toward an objective, and the second capability does not fall out of the first.

Semantic generalisation is the harder half. Planning within a known task distribution is well-studied; planning toward a goal described in language that the training set never contained requires the world model and the language grounding to share a representation.

The reason this matters beyond robotics is that it is the same architecture question facing any agent operating over long horizons in an environment it cannot fully observe.

See our analysis →

arXiv — Grounded world model for semantically generalizable planning → · arXiv — Agents' Last Exam → · arXiv — LLM reasoning is latent, not the chain of thought →