// blog · analysis · multimodal2026-08-08source: product announcements and Commission guidance

One model, all the modalities, and a marking problem

Joint training across modalities removes the seam that pipelines fall apart at. It also produces exactly the output that the new transparency rules are hardest to apply to.

FLUX 3 generates images or combined audio-video clips up to 20 seconds from one prompt, jointly trained across modalities rather than assembled from separate systems. Gemini Omni reasons across image, audio, video and text together and has shipped into the Gemini app, YouTube Shorts and Flow.

The seam is the thing being removed

Most audio-video systems generate picture and sound separately and align them afterwards, which is why lip sync and impact timing drift. A single model trained across modalities has no seam to come apart at.

The same argument applies on the input side. A pipeline that transcribes audio and captions images before reasoning has discarded timing, prosody and spatial relationships — most of what was actually in the signal. Reasoning over raw modalities keeps it.

Twenty seconds with audio is a working length rather than a demo length. It clears a full shot, which is also long enough that temporal consistency failures stop being hideable.

Now the regulatory collision

The EU transparency rules that began enforcement on 2 August require machine-readable marking of generated or altered content.

Marking a still image is a solved engineering problem. Marking a jointly generated audio-video stream so that the mark survives re-encoding, cropping, platform transcoding and the audio being separated from the picture is not. And a model that generates all modalities in one pass has no natural boundary at which to insert one.

Consumer distribution sharpens it. Omni Flash in YouTube Shorts puts omni-modal generation in front of an audience that never opted into an AI product — the exact population the disclosure rules were written for, receiving the content through the pipeline where marking is hardest to preserve.

VentureBeat — Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio → · Google — Introducing Gemini Omni → · European Commission — Commission starts enforcing AI Act rules and new transparency requirements on 2 August →