// blog · analysis · multimodal2026-08-17source: Model comparisons

More inputs, fewer endpoints

One video model takes twelve conditioning inputs at once. Another is being switched off next month. Both facts describe the same transition.

Nine images, three video clips, three audio files — twelve mixed conditioning inputs in a single generation. The competitors take one or two image references. All four models are diffusion transformers, all reach 1080p, all bill by the second.

Output quality stopped being the axis

When every model in a category clears the same resolution bar, the differentiator moves to control. One reference image is a style hint. Twelve mixed inputs is a specification — this character, this location, this camera move, this audio timing — and that is much closer to how anyone producing video actually works.

The unit of work shifts from prompting to assembling.

The failure mode shifts with it. Twelve constraints can contradict each other in ways a text prompt cannot, and the result honours some inputs while quietly dropping others, with no indication which. Richer control means harder debugging, and nobody has built the tooling for that yet.

And the category's founder is being retired

Sora 2's consumer product ended in April; the API stops on 24 September. The model that made video generation a mainstream topic is being switched off while its competitors ship releases.

Defining a category and holding it are separate capabilities, and the second one is less about research than about staying interested in a product after the announcement.

The pattern this week keeps producing

Google retired three Imagen 4 model IDs this morning. OpenAI retires Sora's API next month. Specialised endpoints are collapsing into general models, and every collapse is a forced migration for whoever built on the specialised one.

What this asks of you

Budget migration as a recurring line item, not an incident. A hosted endpoint is a dependency with a termination date you do not control, and the last two years have established the cadence clearly enough that surprise is no longer an excuse.

Then pick on inputs, not output samples. Sample reels converge; what a model will accept as a constraint is what decides whether it survives contact with a real pipeline.

Teamday — Best AI Video Models 2026: Veo, Runway, Kling, Sora Ranked → · WaveSpeed — AI Video Generation Models: 2026 Complete Guide →