Whole-body control and the foundation-model robotics merger — the leading platforms agree on the shape of the answer
When three leading platforms converge on the same architecture, a field has left the exploratory phase. GR00T, Figure 02, and Gemini Robotics 2 are three routes to one design: a transformer that maps multimodal input to joint control. Agreement on the shape of the solution is how you know the exploration is over.
Gemini Robotics 2 extends foundation-model control to whole-body motion for the first time — beyond the upper-body manipulation prior models were limited to — and adds dexterity and multi-robot collaboration, taking multimodal video, audio, or text directly to control.
Whole-body is the capability line
Upper-body manipulation is a constrained problem: a fixed base with moving arms. Whole-body control means coordinating locomotion and manipulation together — the difference between a robot that can reach and one that can go somewhere and then do something. Extending a foundation model to that scope is the substantive step, not a demo polish.
The convergence is on architecture, not vendor. GR00T learning from video and simulation, Figure 02's built-in multimodal model, and full-body sensing platforms are three routes to the same transformer-based design. When the leading platforms agree on the shape of the solution, the field has consolidated.
Pilots, not deployment
The honest framing is 'commercial piloting' — 2026 is the year of paid pilots, not deployment at scale, with a market Goldman pegs at $38 billion by 2035. The models and the bodies are ready to be tried in the field. The field will decide which survive contact with actual work, and that verdict is the next chapter, not this one.
The exploration phase ends when everyone builds the same thing. Robotics just reached that point.
MarkTechPost — Google DeepMind Gemini Robotics 2: whole-body control, dexterity, multi-robot → · Meta Intelligence — Humanoid robots 2026: Figure 02 & NVIDIA Isaac status →