// news · multimodal2026-08-06source: release coverage

H3 generates 32 kHz stereo in the same pass as the frames

MiniMax H3 takes text, image, audio and video as input and produces 4 to 15 seconds at 24 FPS with 32 kHz stereo, reaching 2K through a dedicated regeneration path. Generating sound jointly with the images rather than adding it afterwards is the part that changes what the output is usable for.

Adding audio after the fact makes synchronisation a separate problem, and synchronisation is where the uncanny lives — a footstep landing two frames late reads as wrong before a viewer can say why. Producing both from the same pass means the impact and the sound share a cause instead of being aligned afterwards.

The input side is the underrated half. Accepting existing video and audio as conditioning, not just text and images, is what turns a generator into something that can extend or revise footage rather than only originate it. That is the difference between a demo and a step in a pipeline.

Fifteen seconds remains the honest ceiling, and the harder unsolved problem sits between clips rather than inside them: whether the same character, lighting and space survive from one shot to the next. Joint audio-video does nothing for that.

See our analysis →

MarkTechPost — MiniMax releases H3: 15-second 2K clips with native stereo audio → · Hugging Face — What is MiniMax H3 (Hailuo 3.0)? The omni-modal video model, explained → · Swisher Post — MiniMax H3 AI video model arrives with native audio →