// blog · analysis · multimodal2026-08-11source: MiniMax / Google

Video generation goes open

The modality where the closed-weights advantage looked most durable just produced a second-place model that anyone can download.

MiniMax open-sourced H3 — 2K clips of four to fifteen seconds with native stereo audio, ranked second on Artificial Analysis's video board, 3.3 Elo behind Gemini Omni Flash.

Native audio is the technically hard part

Most video models produce silent clips and add sound afterwards, which falls apart on anything where the sound has to match the motion — footsteps, impacts, a mouth forming words. Generating both from one model makes alignment a property of the output rather than a post-production problem someone else has to solve.

Why this modality was supposed to stay closed

The argument was economic: video training is expensive enough that nobody would give the result away. Three point three Elo behind a closed frontier model, released openly, retires that argument.

What replaces it is a question about who needs weights rather than an endpoint. At $0.13 per second for 2K hosted, most users will call the API. Weights matter to the people who cannot send footage anywhere — studios under NDA, anyone managing likeness rights, and researchers who need to modify the model rather than prompt it.

Meanwhile, the useful boring capability

Google's image models hit GA with video as an input: pass a clip, get a thumbnail or a poster frame. That is comprehension dressed as generation, and every video platform on earth currently solves it with a human scrubbing a timeline.

One of these stories will get the attention. The other will get deployed.

MarkTechPost — MiniMax Releases MiniMax H3: An Omni-Modal Video Model → · Google AI for Developers — Gemini API release notes →