// news · multimodal2026-08-14source: Google DeepMind / reporting

Gemini Omni unifies text, image and video generation in one model

Omni Flash creates and edits video conversationally from image, audio, video and text inputs — the first top-tier system to unify all three generation modes in a single model rather than routing between specialists. Flash-tier clips cap at 10 seconds, which Google describes as a deployment decision.

Every multimodal product until now has been a router with a good interface: a text model, an image model and a video model behind one prompt box. Unification is an architectural claim, and the observable consequence is conversational editing — changing a video by describing the change, because the same model holds the whole thing.

The input side matters as much as the output. Accepting image, audio, video and text in one context is what makes "make the second shot match the first, but slower" a tractable instruction. Routed systems lose that because each specialist sees only its own slice.

The 10-second cap on the Flash tier is described as a deployment decision rather than a model limit, which is a specific and checkable claim. If true, the constraint is serving cost and the cap moves as inference gets cheaper. Distribution is immediate: the Gemini app, Google Flow, and YouTube Shorts and Create at no cost.

The competitive frame is unavoidable — open-weight video is now within a few Elo of the top of the board, so unification and distribution are what a closed lab has left to differentiate on. SynthID being on by default is the other half of that answer.

See our analysis →

The Next Web — Google launches Gemini Omni Flash, a conversational video-generation model → · Digital Applied — Gemini Omni: Multimodal Video Creation Guide for 2026 →