// news · robotics2026-08-01source: marktechpost / ieee

Gemini Robotics 2 extends foundation-model control to the whole body, dexterity, and multi-robot collaboration

Google DeepMind shipped three physical-AI models, led by Gemini Robotics 2, which extends control to whole-body motion for the first time — beyond prior models that only drove the upper body — and adds dexterity and multi-robot collaboration. It takes multimodal video, audio, or text directly and outputs control.

Whole-body control is the capability line being crossed. Upper-body manipulation — a fixed base with moving arms — is a constrained problem; whole-body control means coordinating locomotion and manipulation together, which is the difference between a robot that can reach and a robot that can go somewhere and then do something. Extending a foundation model to that scope is the substantive step.

The multimodal-in, control-out design is the same convergence visible across the multimodal frontier: no separate perception module feeding a separate controller, but one model mapping video, audio, or language to action. That architecture is what lets a single foundation model generalise across tasks rather than being trained per behaviour.

Multi-robot collaboration is the quiet expansion. A model that can coordinate several robots is aiming past the single-humanoid demo toward fleets that divide work — which is where the commercial value of physical AI actually sits, and which pairs naturally with the cross-agent coordination protocols emerging on the software side.

See our analysis →

MarkTechPost — Google DeepMind Gemini Robotics 2: whole-body control, dexterity, multi-robot collaboration → · IEEE Spectrum — Video Friday: physical AI robotics, robot hands, and more →