// news · multimodal2026-08-17source: Model comparisons

One video model accepts twelve mixed inputs in a single generation

Nine images, three video clips and three audio files — twelve conditioning inputs at once. The frontier in video generation has moved from output fidelity to how much you can specify.

Among the four video models that currently matter — Seedance 2.0, Kling 3.0, Veo 3.1 and the now-deprecated Sora 2 — the differentiator is no longer resolution. All are diffusion transformers, all reach 1080p natively, all bill by the second. What separates them is what they accept as input.

Sora 2 and Kling 3.0 take one to two image references. Veo 3.1 takes one to two images plus one to two video clips. Seedance 2.0 takes up to nine images, up to three video clips and up to three audio files, for twelve mixed conditioning inputs in a single generation.

That is a different product. One reference image is a style hint; twelve mixed inputs is a specification — this character, this location, this camera move, this audio timing. The unit of work shifts from prompting toward assembling, which is much closer to how anyone actually producing video already thinks.

The catch is that conditioning inputs interact. Twelve constraints can be mutually inconsistent in ways a text prompt cannot easily be, and the failure modes are correspondingly harder to debug — you get a result that honours some inputs and quietly drops others, with no indication of which.

Input richness is nonetheless the right axis to compete on, because it is the one that determines whether these models enter production pipelines or stay in demos. Convergence on output quality was always going to happen; convergence on control has not.

See our analysis →

Teamday — Best AI Video Models 2026: Veo, Runway, Kling, Sora Ranked → · WaveSpeed — AI Video Generation Models: 2026 Complete Guide →