Native audio becomes table stakes for video generation
H3's joint audio-video generation, with support for 21:9, 16:9, 4:3 and 1:1 across 4 to 15 seconds, marks the point at which silent output starts reading as a deficiency rather than a norm. The competitive baseline for the category has moved.
Feature baselines shift in a recognisable pattern: a capability is a differentiator for roughly one release cycle, then its absence becomes the thing reviewers mention. Native audio is entering that second phase. A model launching without it now has to explain the omission rather than treat it as out of scope.
The aspect-ratio range is a quieter signal about who the intended user is. Supporting 21:9 through 1:1 is not a research consideration, it is a distribution consideration — cinematic, broadcast, and vertical social all in one model. That is a production tool being specified, not a demo.
The harder problem remains identity and continuity across shots. Joint audio-video solves synchronisation within a clip; nothing about it addresses whether the same character, lighting and space survive from one clip to the next, which is what stands between generated footage and a finished sequence.
Precedence Research — MiniMax launches H3 multimodal AI for video creation → · AIBase — Multimodal generation upgraded: 2K HD audio and video → · Digital Strategy EU — Code of Practice on Transparency of AI-generated Content →