LingBot-VLA: an industry-scale model trained on 20,000 hours of real dual-arm data
Ant Group's LingBot-VLA is trained on twenty thousand hours of real dual-arm robot data across nine hardware configurations. Where academic work attacks the data bottleneck with clever architecture, this attacks it with money and time, and both approaches are now visible in the same quarter.
Twenty thousand hours is roughly two and a quarter robot-years of continuous operation, and it is real data rather than simulated. That is a capital expenditure disguised as a dataset, and it is only available to an organisation willing to run robot farms for the purpose.
Nine hardware configurations is the more technically interesting number. Training across multiple embodiments is what separates a policy that works on one arm from a foundation model, and it is also where most cross-embodiment efforts break down, because the mapping between different kinematics is not learned for free.
Set against learning latent actions from unlabelled video, the field now has two live answers to the same bottleneck: buy the data, or learn without it. They are not exclusive and the winner is probably whichever gets combined first.
arXiv — Vision-language-action models: concepts, progress, applications and challenges → · arXiv — SmolVLA: a vision-language-action model for affordable and efficient robotics → · arXiv — Nautilus: from one prompt to plug-and-play robot learning →