// news · multimodal · open-source2026-08-05source: video model coverage

MiniMax open-sources H3 — 15-second 2K video with native stereo audio

MiniMax released its next-generation video model H3 as open source on 3 August. It takes text, images, audio and existing video as input, outputs 4 to 15 seconds at 24 FPS with 32 kHz stereo sound, and reaches 2K resolution through a dedicated regeneration path. It went live in the platform API and the Hailuo consumer app on 31 July.

Native audio is the part that changes the workflow. Generated video has largely been silent, with sound added afterwards and synchronisation left as the user's problem. Producing 32 kHz stereo jointly with the frames means lip movement, footsteps and impacts are generated in the same pass as the images that motivate them — which is the difference between a clip and a shot.

Open-sourcing it is the more consequential decision. The commercial video models this competes with are closed, and this is the strongest omni-modal system to ship with weights available. That single fact resets what the self-hosted floor looks like, and everything built on top of it now has an alternative to an API bill.

The fifteen-second ceiling remains the honest limit. That is a shot, not a scene, and coherent minutes remain unsolved. But arriving in the same week that EU deepfake labelling became enforceable makes the provenance question immediate rather than theoretical — open weights and machine-readable marking obligations are not naturally compatible.

See our analysis →

MarkTechPost — MiniMax releases MiniMax H3, an omni-modal video model with native stereo audio → · AIBase — Visual large models receive a major open-source announcement → · Precedence Research — MiniMax launches H3 multimodal AI for video creation →