Molmo2 releases open weights and open data for vision-language models
A family of VLMs with video understanding and point-driven grounding across single-image, multi-image and video tasks — released with the data as well as the weights, which is a materially stronger form of openness than the phrase usually indicates.
Open weights let you run a model and fine-tune it. Open data lets you understand why it behaves as it does, reproduce the training, and audit what went in. Only the second supports the scientific claims that open release is usually justified by.
The technical focus is grounding: point-driven reference across single images, multiple images and video. Grounding is the capability that connects language to specific locations in a visual scene, and it is what separates a model that describes an image from one that can be instructed about it — which is the difference between a captioner and something a robot can use.
Releasing data is rarer than releasing weights for reasons that are mostly not technical. Training corpora carry licensing exposure, provenance questions and competitive information, and most labs have concluded the risk outweighs the benefit. A release that includes it is choosing a different trade.
It is a useful contrast with the week's other releases: a model announced as open-weight whose weights have not shipped, and models shipping under bespoke or revenue-share terms. Weights, data, and licence are three separate axes, and the word "open" is being asked to carry all three.
arXiv — Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding → · BentoML — Multimodal AI: The Best Open-Source Vision Language Models in 2026 →