// news · multimodal2026-07-31source: dentro / buildez

Google DeepMind's Gemini Omni creates and edits video from any mix of image, audio, video and text

Gemini Omni is a multimodal family that takes any combination of image, audio, video and text as input and produces or edits video. Omni Flash rolls out first across the Gemini app, Google Flow and YouTube Shorts for paid tiers, with API access following. The interesting claim is the absence of a fixed input signature.

Multimodal has mostly meant a fixed set of accepted input types with a fixed output type. A model defined by arbitrary input combinations is a different abstraction: the prompt becomes a bundle of whatever material is at hand rather than a structured call. That is closer to how people actually assemble a brief.

Shipping into YouTube Shorts before the API is a deliberate ordering. It puts the model in front of the largest possible volume of real usage, under a surface where output quality is judged by an audience rather than an eval, and it defers the harder questions that a public API creates about provenance and abuse at scale.

The distribution point is the competitive one. A video model reaching Shorts, Flow and the Gemini app simultaneously starts with an install base no independent video lab can approach. For H2 2026 the constraint on video generation is shifting from model quality to shelf space, and shelf space is not a research problem.

See our analysis →

Dentro AI — AI News — July 2026: Key Events & Releases → · BuildEZ — AI New Model: July 2026 Developments & What's Next →