'Touch Dreaming' brings tactile-visual multimodal policies to humanoid control
Robotics research is fusing senses into one policy: Humanoid Transformer with Touch Dreaming (HTD) combines vision, distributed tactile sensing, and reinforcement-learned control, alongside VR teleoperation and dexterous-hand retargeting. Multimodal is expanding past text-image-video into touch — the modality embodied systems can't do without.
Adding touch is what multimodal has to do to leave the screen. Vision and language are enough to describe a scene; manipulating the physical world requires feeling contact, force, and slip, which is why distributed tactile sensing folded into a single control policy is the frontier for embodied AI. HTD treats touch as a first-class modality, not an afterthought.
The 'dreaming' part is the world-model connection. Learning a predictive model that includes tactile outcomes lets a policy rehearse contact-rich manipulation internally, reducing the expensive real-world trials that make robot learning slow — the same data-efficiency argument driving world models across embodied AI, now extended to the sense of touch.
The convergence with the rest of multimodal is the through-line. As generative video models learn the dynamics of how scenes evolve and control policies learn to act in them, the boundary between a model that perceives, one that predicts, and one that acts keeps dissolving. Touch is the modality that grounds that convergence in the physical world.
IEEE Spectrum — Video Friday: humanoid robot production, Mars rovers, and more → · Humanoid Press — Humanoid robots news: AI breakthroughs, robotics trends →