MiniMax open-sources H3, an omni-modal video model with native stereo audio
MiniMax released H3 (Hailuo 3.0) as an open-source, general-purpose multimodal video model that takes text, image, video, and audio in and generates 2K clips of 4–15 seconds with native 32kHz stereo sound. It ranks #2 on Artificial Analysis's video board, just behind Gemini Omni Flash — a frontier-class video model with open weights.
Native stereo audio is the technical leap. Generating synchronized 32kHz stereo sound with the video, rather than dubbing audio onto silent footage, means the model learned the joint distribution of how a scene looks and sounds — a harder and more useful object. Omni-modal input (text, image, video, and audio together) plus omni-modal output is the direction the whole field is converging on.
Open-sourcing a #2-ranked video model resets expectations. Sitting just behind Gemini Omni Flash on Artificial Analysis's board while shipping open weights means frontier-class video generation is no longer a closed-lab preserve — the same open-weight pressure that reshaped language models is now arriving in video, where the compute and data costs are even higher.
The pricing and features signal a production tool, not a demo. Per-second billing at $0.13 for 2K and $0.09 for 768p, with motion transfer, reference-driven generation, and generative editing, is a model built to sit inside real creative pipelines. Combined with open weights, it gives studios and developers a frontier video engine they can run and control themselves.
MarkTechPost — MiniMax releases MiniMax H3, an omni-modal video model with native stereo audio → · AIbase — Visual large models: multimodal generation upgraded with 2K HD audio and video →