A video model where the audio was not bolted on afterwards
H3 generates 4 to 15 seconds at 24fps with native 32kHz stereo, up to 2K through a dedicated regeneration path, across 21:9, 16:9, 4:3 and 1:1. Coverage notes the sound is generated with the picture rather than in a second stage.
The specification reads like a product sheet: 4 to 15 seconds of output, 24 frames per second, 32kHz stereo, aspect ratios from 21:9 down to 1:1, a 768-pixel short side by default and up to 2K through a dedicated H3-Regenerate-2K path.
One line in the coverage carries more weight than the rest: unlike most open video releases before it, the audio is not a separate stage bolted on afterwards. That is an architectural claim, not a feature.
Generating sound in a second pass means the model matches audio to picture it has already committed to. Generating them together means the two can be consistent by construction — footsteps landing on the frame where the foot lands, a voice whose room matches the room on screen. Anyone who has watched a dubbed clip knows how quickly the eye catches the mismatch.
The 4-to-15-second window tells you what this is for. It is the length of an advertisement, a product loop, a title sequence — and the stated targets are advertising, branding, e-commerce, product design, UI, and games. This is tooling for commercial short-form, not a step toward long-form film.
The catch is who may use it. The licence is reported to exclude four of the largest advertising markets on earth, which is an odd place to land for a model built for advertising.
MiniMax — MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities → · AIbase — Visual Large Models Receive a Major Open-Source Announcement → · ComfyUI Wiki — MiniMax H3: Open Omni-Modal Video Generation Model →