// blog · analysis · multimodal2026-08-01source: marktechpost / arxiv

Action models and the collapse of the divide between generating a video and driving a body

Image models, video models, and control policies were separate research programs with separate architectures. The 2026 turn is the recognition that they are doing structurally the same thing — predicting how a scene evolves — and can share a model. That collapse dissolves the line between a generative model and an agent.

The multimodal frontier is converging on a single object: a model that takes video, audio, or text and outputs not just pixels but actions. A video model predicting the next frame and a control policy predicting the next action are doing structurally similar things, and 2026 is the year they merged.

The equivalence made explicit

Research framing exocentric video generation as humanoid control treats 'generate a video of the motion' and 'execute the motion' as two readouts of one model. A model that can generate a plausible video of an action is most of the way to a model that can perform it — which is why the image-video-action convergence is the multimodal through-line of the half.

The substrate underneath is the world model. A comprehensive 2026 survey marks world models moving to the center of robot learning — a multimodal prediction engine turned inward, generating not content for a human to watch but dynamics for an agent to plan against. Progress in video generation and progress in robot learning are increasingly the same curve.

Generator and agent, dissolving

The consequence is that the boundary between a generative model and an agent is disappearing. When predicting a scene's evolution and acting within it become one capability, the multimodal model stops being a thing that makes media and becomes a thing that can do — which is a far larger claim than 'better image quality.'

The pixels were never the point. Predicting what happens next was, and prediction is most of the way to action.

MarkTechPost — Google DeepMind ships three physical AI models for whole-body control → · arXiv — World model for robot learning: a comprehensive survey →