Action models collapse the divide between generating an image, generating a video, and controlling a body
The multimodal frontier is converging on a single object: a model that takes video, audio, or text and outputs not just pixels but actions. Google DeepMind's Gemini Robotics 2 accepts multimodal input directly and produces whole-body control, and research on exocentric video generation as humanoid control treats generating a video and driving a robot as the same problem.
The conceptual merge is the story. For years image models, video models, and control policies were separate research programs with separate architectures. The 2026 turn is the recognition that a video model that predicts the next frame and a control policy that predicts the next action are doing structurally similar things — predicting how a scene evolves — and can share a model.
Gemini Robotics 2 is the clearest commercial instance: multimodal in, whole-body control out, no separate perception-then-action pipeline. Academic work on exocentric video generation as generalisable humanoid control makes the equivalence explicit, framing 'generate a video of the motion' and 'execute the motion' as two readouts of one model.
The consequence for the multimodal stack is that the boundary between a generative model and an agent is dissolving. A model that can generate a plausible video of an action is most of the way to a model that can perform it, and that collapse is what makes the image-video-action convergence the through-line of multimodal in H2 2026.
MarkTechPost — Google DeepMind ships three physical AI models for whole-body control, dexterity, multi-robot collaboration → · arXiv — ExoActor: exocentric video generation as generalizable interactive humanoid control →