Video learns to make its own sound — and it is open
A frontier-class video model that generates synchronized stereo audio, takes any modality in, and ships with open weights. The bar for AI video just moved from picture to picture-plus-sound-plus-control.
MiniMax open-sourced H3, an omni-modal video model generating 2K clips with native 32kHz stereo audio, ranking second on Artificial Analysis video board. Generating synchronized sound with the picture, rather than dubbing it on, means the model learned the joint distribution of how a scene looks and sounds — a genuinely richer object.
The bar moved to control and modality
Because the leaders have converged on quality. Native audio and in-model editing are becoming table stakes across MiniMax H3, Gemini Omni Flash, Seedance, Veo, and Sora, so the contest is now controllability — motion transfer, reference-driven generation, editing — rather than raw fidelity. Silent clips are no longer competitive.
Open weights arrive in video
The most consequential part is that a number-two-ranked video model shipped open. The same open-weight pressure that reshaped language models is now arriving in video, where compute and data costs are even higher — which means frontier video generation is no longer a closed-lab preserve, and studios can run and control the engine themselves.
Omni-modal in, omni-modal out, open to run. That combination is what turns AI video from a demo you watch into a tool you build production pipelines on.
MarkTechPost — MiniMax releases MiniMax H3, an omni-modal video model with native stereo audio → · Hedra — The AI image and video models that defined the first half of 2026 →