// blog · analysis · multimodal2026-08-03source: ieee / robozaps

Multimodal grows a sense of touch — and leaves the screen

Text, image, video were always about depicting the world. Touch is about acting in it. When a multimodal policy learns to feel contact, multimodal AI stops being a media generator and starts being a body.

The Humanoid Transformer with Touch Dreaming fuses vision, distributed tactile sensing, and reinforcement-learned control into one policy. Adding touch is what multimodal has to do to leave the screen: vision and language describe a scene, but manipulating the world requires feeling contact, force, and slip. Touch is a first-class modality here, not an afterthought.

The same machinery, pointed at action

This is one instance of a larger turn. The architectures that generate video are being repurposed as world models and control policies — because predicting the next frame and predicting the next action are structurally the same problem. Multimodal generation and embodied control are two readouts of one capability.

Where the value migrates

Generating a clip is useful; driving a robot that does physical work is transformational. The labs that built strong video models hold exactly the substrate embodied control needs — learned dynamics of how the world moves — which is why the generative frontier and the robotics frontier are becoming one curve. Touch is what grounds it in the physical world.

The line between a model that depicts the world and one that acts in it is dissolving. A model that can imagine an action realistically is most of the way to performing it — and touch is the sense that closes the gap.

IEEE Spectrum — Video Friday: humanoid robot production, Mars rovers, and more → · RoboZaps — Future of humanoid robots: what 2026 is proving →