// news · multimodal · tools2026-08-08source: product launch reporting

FLUX 3 generates 20-second video with audio from one jointly trained model

Black Forest Labs' first public video model understands and generates images or combined audio-video clips up to 20 seconds from a single prompt. The architectural claim is that it was trained jointly across modalities rather than assembled from separate image, video and audio systems.

Joint training is the whole claim, and it is a real distinction. Most audio-video systems generate the picture and the sound in separate passes and align them afterwards, which is why lip sync and impact timing tend to drift. A single model trained across modalities has no seam to fall apart at.

Twenty seconds with audio is a meaningful working length rather than a demo length. It clears a full shot, and it is enough that temporal consistency failures become obvious instead of hideable.

The release is limited to start, which is now standard for video-capable systems and reflects an unresolved provenance question. The EU transparency rules that began enforcement last week require machine-readable marking of generated content — and marking a jointly generated audio-video stream is harder than marking a still.

See our analysis →

VentureBeat — Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio →