Video generation goes omni-modal: native audio and in-model editing become table stakes
The multimodal video race has a new bar. With MiniMax H3 joining Gemini Omni Flash, Seedance, Veo, and Sora, the leaders now take mixed inputs — text, image, video, audio — and output synchronized sound with the picture, plus motion transfer and generative editing. Generating silent clips is no longer competitive; omni-modal is the standard.
The bar moved from picture to picture-plus-sound-plus-control. A year ago a strong text-to-video model was frontier; now the leaders accept up to a dozen mixed inputs and emit synchronized audio, and the differentiators are controllability — motion transfer, reference-driven generation, in-model editing — rather than raw fidelity. The models have converged on quality, so the contest is control and modality breadth.
Native audio changes what these models are for. A clip that arrives with matched sound is usable in a production pipeline without a separate audio pass, which is what turns a generation model from a novelty into a tool creative teams actually build on. Omni-modal output is the feature that moves video AI from demo to workflow.
The competitive frame is a crowded frontier where open and closed sit side by side. Gemini Omni Flash leads a board on which an open model, H3, sits second — so the video contest is being fought on modality and control across both open and closed offerings simultaneously. For creators, that means a frontier-class omni-modal engine is available on both sides of the open line.
Hedra — The AI image and video models that defined the first half of 2026 → · Let's Data Science — Multimodal AI news: image, video and vision models →