Multimodal AI turns from generating media to driving bodies
The center of gravity in multimodal is shifting from producing pixels to producing actions. The same architectures that generate video are being repurposed as world models and control policies for robots, folding perception, prediction, and action into one system — the through-line connecting this year's video and robotics advances.
The repurposing is the key move. A video model that predicts the next frame and a control policy that predicts the next action are doing structurally similar things — modeling how a scene evolves — so the machinery built for generation transfers to control. Multimodal generation and embodied action are turning out to be two readouts of the same capability.
The pull toward embodiment is where the durable value sits. Generating a clip is useful; driving a robot that does physical work is transformational, and the labs that built strong video models have exactly the substrate — learned dynamics of how the world moves — that embodied control needs. The generative frontier and the robotics frontier are becoming one curve.
The consequence is that the line between a generative model and an agent is dissolving. A model that can generate a plausible video of an action is most of the way to a model that can perform it, and 2026's advances — unified audio-video, tactile policies, world models — are all steps along that single path from depicting the world to acting in it.
RoboZaps — Future of humanoid robots: what 2026 is proving → · IEEE Spectrum — Video Friday: humanoid robot production, Mars rovers, and more →