Sound was the missing half
Generated video has been silent for three years and everyone treated that as normal. H3 generates 32 kHz stereo jointly with the frames — and open-sources the whole thing in the same week Europe started requiring synthetic media to be labelled.
MiniMax open-sourced H3 on 3 August: text, image, audio and video in, 4 to 15 seconds at 24 FPS with 32 kHz stereo out, 2K through a dedicated regeneration path. Two things about that are more consequential than the resolution.
Joint generation is not post-production
Adding sound afterwards means solving synchronisation as a separate problem, and synchronisation is where the uncanny lives — a footstep landing two frames late reads as wrong before a viewer can say why. Generating audio in the same pass as the frames that motivate it means the impact and the sound share a cause instead of being aligned after the fact.
That is why silent output is starting to read as a deficiency rather than a norm. Feature baselines shift in a recognisable pattern — differentiator for one release cycle, then the thing reviewers mention when it is missing. Native audio is entering phase two.
Open-sourcing it is the bigger decision
The commercial video models this competes with are closed. This is the strongest omni-modal system to ship with weights available, which resets what the self-hosted floor looks like and gives everything built on top an alternative to an API bill. The aspect-ratio range — 21:9 through 1:1 — tells you the intended user is a production pipeline, not a research demo.
The timing is not comfortable
Article 50 became enforceable on 2 August, requiring synthetic output to be machine-readable and detectable, and deepfakes labelled. Open weights and mandatory provenance marking are not naturally compatible: anyone who can run the model locally can strip whatever the reference implementation embeds.
That tension is not MiniMax's to solve, and pretending otherwise would be dishonest — the same tension applies to every open image model already in circulation. But it is the first time a genuinely frontier-class open video model has landed into a live labelling regime, and the answer so far is that the obligation attaches to providers while the capability attaches to anyone with a GPU.
What is still missing
Fifteen seconds is a shot, not a scene. Coherent minutes remain unsolved, and the harder problem underneath is identity and continuity — whether the same character, lighting and space survive from one clip to the next. Joint audio-video fixes synchronisation within a clip and says nothing about what happens between them. That gap is what still separates generated footage from a finished sequence.
MarkTechPost — MiniMax releases MiniMax H3, an omni-modal video model with native stereo audio → · AIBase — Visual large models receive a major open-source announcement → · Digital Strategy EU — Code of Practice on Transparency of AI-generated Content →