// news · robotics · multimodal2026-08-08source: company announcement

Gemini Robotics ER 2 is positioned as the high-level brain, not the controller

Google's embodied reasoning model handles real-time spatial reasoning, multi-step task planning and coordination between different robots. The division of labour it assumes — foundation model plans, dedicated controllers execute — is becoming the default architecture for the field.

The split is doing real work. Low-level control needs deterministic latency and hard safety guarantees that a large model cannot provide. Planning needs world knowledge and language grounding that a controller has no way to acquire. Putting a foundation model above the control stack rather than inside it resolves the conflict instead of trading one failure for another.

Collaboration between different robots is the more forward-looking claim. A planning layer that coordinates heterogeneous hardware makes the model the integration point for a mixed fleet, which is what most real facilities actually operate.

It also inherits the multi-agent problem wholesale. Several robots coordinated by a shared planner is exactly the interaction surface DeepMind's own safety call describes as under-researched — with the added property that failures here have mass and momentum.

See our analysis →

Google — Introducing Gemini Robotics ER 2 → · The Robot Report — Boston Dynamics, Google reunite on next-gen Atlas humanoid →