// news · research-papers · robotics2026-08-06source: arXiv and code releases

WholeBodyVLA learns whole-body control from action-free video, beating GR00T by 21.3%

An ICLR 2026 paper presents a unified latent vision-language-action framework for whole-body loco-manipulation that learns latent actions from action-free egocentric video, and reports outperforming GR00T by 21.3% on AgiBot X2. Learning from video without action labels attacks the field's central bottleneck directly.

Action-free is the load-bearing word. Robot learning is constrained by paired data — video plus the exact joint commands that produced it — which requires teleoperation or instrumented demonstration and is therefore expensive and scarce. Egocentric video without action labels is comparatively abundant.

Learning a latent action space from unlabelled video and then grounding it in a real robot is not a new idea, and the reason it keeps returning is that the payoff is enormous if it works. A 21.3% margin on a specific platform is a real result and a narrow one; whether the latent space transfers across embodiments is the question that decides whether this is a technique or a milestone.

Whole-body loco-manipulation is also exactly where the gap sits. Manipulation and locomotion have each been solved reasonably well in isolation, and the behaviours that need both at once — bracing, leveraging body weight, catching balance while carrying — fall between the two controllers.

See our analysis →

OpenDriveLab — WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control → · arXiv — GR00T N1: an open foundation model for generalist humanoid robots → · arXiv — Vision-language-action models: concepts, progress, applications and challenges →