The driving model is a reasoning model now
Detection and trajectory prediction handle the common case. What breaks autonomy is the situation that requires working out why the car ahead has stopped.
NVIDIA released Alpamayo 2 Super, a 34B vision-language-action model, under a licence permitting commercial inspection, fine-tuning and deployment.
The premise is that perception is no longer the bottleneck
Autonomy stacks have historically been pipelines: detect, predict, plan. That decomposition handles the overwhelming majority of miles and fails on the rare case that needs causal reasoning before action — why that vehicle is stopped, what the person at the kerb is about to do, whether the obstruction is temporary.
Folding scene understanding, reasoning and planning into one model is a bet that the long tail is a reasoning problem wearing a perception costume.
Giving it away is the strategy
NVIDIA does not need model revenue. It needs the autonomy industry standardising on a stack that runs on its silicon, and open weights are the cheapest possible way to buy that. Half a million downloads for the family says it is working.
Inspect, fine-tune, modify, deploy — for free, on our hardware.
Benchmark hygiene
A 23.2-point lead over GPT-4o on LingoQA is a statement about a general-purpose model answering a specialist question, not about driving. LingoQA measures reasoning over driving scenes. The comparison a fleet operator needs is against the systems currently in vehicles, and that number is not in the release.
The other end of robotics stayed resolutely physical this week: Avatar Robotics raised $6.5 million having shipped 900,000 products, with reducing remote-operator dependence named as the use of proceeds. Reasoning models and teleoperation budgets are the same problem measured differently.
NVIDIA Newsroom — NVIDIA Launches Alpamayo 2 Super Open Reasoning Model for Robotaxis → · The AI Insider — Avatar Robotics Raises $6.5M in Seed Funding →