// blog · analysis · multimodal2026-08-05source: video model coverage

Sound was the missing half

Generated video has been silent for three years and everyone treated that as normal. H3 generates 32 kHz stereo jointly with the frames — and open-sources the whole thing in the same week Europe started requiring synthetic media to be labelled.

MiniMax open-sourced H3 on 3 August: text, image, audio and video in, 4 to 15 seconds at 24 FPS with 32 kHz stereo out, 2K through a dedicated regeneration path. Two things about that are more consequential than the resolution.

Joint generation is not post-production

Adding sound afterwards means solving synchronisation as a separate problem, and synchronisation is where the uncanny lives — a footstep landing two frames late reads as wrong before a viewer can say why. Generating audio in the same pass as the frames that motivate it means the impact and the sound share a cause instead of being aligned after the fact.

That is why silent output is starting to read as a deficiency rather than a norm. Feature baselines shift in a recognisable pattern — differentiator for one release cycle, then the thing reviewers mention when it is missing. Native audio is entering phase two.

Open-sourcing it is the bigger decision

The commercial video models this competes with are closed. This is the strongest omni-modal system to ship with weights available, which resets what the self-hosted floor looks like and gives everything built on top an alternative to an API bill. The aspect-ratio range — 21:9 through 1:1 — tells you the intended user is a production pipeline, not a research demo.

The timing is not comfortable

Article 50 became enforceable on 2 August, requiring synthetic output to be machine-readable and detectable, and deepfakes labelled. Open weights and mandatory provenance marking are not naturally compatible: anyone who can run the model locally can strip whatever the reference implementation embeds.

That tension is not MiniMax's to solve, and pretending otherwise would be dishonest — the same tension applies to every open image model already in circulation. But it is the first time a genuinely frontier-class open video model has landed into a live labelling regime, and the answer so far is that the obligation attaches to providers while the capability attaches to anyone with a GPU.

What is still missing

Fifteen seconds is a shot, not a scene. Coherent minutes remain unsolved, and the harder problem underneath is identity and continuity — whether the same character, lighting and space survive from one clip to the next. Joint audio-video fixes synchronisation within a clip and says nothing about what happens between them. That gap is what still separates generated footage from a finished sequence.

MarkTechPost — MiniMax releases MiniMax H3, an omni-modal video model with native stereo audio → · AIBase — Visual large models receive a major open-source announcement → · Digital Strategy EU — Code of Practice on Transparency of AI-generated Content →