// news · multimodal · open-source2026-08-11source: MiniMax / reporting

MiniMax open-sources H3, its 2K video model with native stereo audio

H3 takes text, image, video and audio as input and outputs 4-to-15-second 2K clips with native stereo sound, plus motion transfer and generative video editing. Artificial Analysis ranks it second on its video board, 3.3 Elo behind Gemini Omni Flash.

Native audio is the part that is hard. Most video generators produce silent clips and bolt sound on afterwards, which fails on anything where the sound has to match the motion — footsteps, impacts, speech. Generating both from one model means the alignment is a property of the output rather than a post-production problem.

Second place by 3.3 Elo against a closed frontier model, then released openly, is the more consequential fact. Video generation has been the modality where the closed-weights gap looked most durable, on the theory that the compute and data requirements were prohibitive for anyone giving the result away.

Priced at $0.13 per second for 2K and $0.09 for 768p in hosted form, the open release mostly matters to people who cannot send their footage to an API — studios under NDA, anyone with likeness rights to manage, and the long tail of research that needs to modify the model rather than call it.

See our analysis →

MarkTechPost — MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio → · AIbase — Visual Large Models Receive a Major Open-Source Announcement →