The audio was always the hard part
Generated video has been silent for three years and the industry treated that as normal. Producing sound in the same pass as the frames changes what the output is for — and hands the provenance people a four-channel problem.
MiniMax H3 produces 4 to 15 seconds at 24 FPS with 32 kHz stereo, generated jointly with the frames, taking text, image, audio and video as input.
Why joint generation is not post-production
Adding sound afterwards makes synchronisation a separate problem, and synchronisation is where the uncanny lives. A footstep landing two frames late reads as wrong before a viewer can articulate why. Generating both in one pass means the impact and the sound share a cause rather than being aligned after the fact.
The input side is the underrated half. Accepting existing video and audio as conditioning — not only text and images — turns a generator into something that can extend and revise footage rather than only originate it. That is the difference between a demo and a step in a pipeline.
The provenance problem just got harder
Article 50 requires synthetic audio, image, video and text to be machine-readable and detectable as generated. Watermarking discussion has overwhelmingly concerned images, where techniques are mature and attacks catalogued. Audio provenance is less developed and has different physics — resampling, compression and re-recording degrade signal differently from how cropping degrades pixels.
Joint generation raises a question with no settled answer: one mark or two? Mark the channels separately and an attacker strips one and keeps the other. Mark them jointly and the mark must survive both pipelines, which are routinely processed by different tools at different stages.
The December deadline for products already on the market is the regulator conceding that retrofitting provenance is a different order of work from adding a banner. Realistic — and it also means most synthetic media circulating this year carries no machine-readable mark at all.
What is still unsolved
Fifteen seconds is a shot, not a scene. The harder problem sits between clips rather than inside them: whether the same character, lighting and space survive from one shot to the next. Joint audio-video fixes synchronisation within a clip and does nothing for continuity across them.
That gap is the remaining distance between generated footage and a finished sequence, and no amount of resolution or audio fidelity closes it.
MarkTechPost — MiniMax releases H3: 15-second 2K clips with native stereo audio → · European Commission — Code of Practice on Transparency of AI-generated Content → · Precedence Research — MiniMax launches H3 multimodal AI for video creation →