// news · multimodal · frontier-models2026-08-08source: company announcement and reporting

Gemini Omni reasons across image, audio, video and text to produce one consistent output

Google's omni-modal system lets users combine inputs of any type and reasons across all of them together rather than routing each to a specialist. Omni Flash has rolled out to the Gemini app, YouTube Shorts and the Flow creative studio.

Reasoning across modalities is a stronger claim than accepting them. A pipeline that transcribes audio, captions images and then reasons over text has thrown away most of what was in the original signal — timing, prosody, spatial relationships. A model reasoning over the raw modalities together has not.

The distribution is what makes it consequential rather than impressive. Omni Flash in YouTube Shorts puts omni-modal generation in front of a consumer audience that never opted into an AI product, which is a different adoption curve from a developer API.

Google's model line has fragmented accordingly — Gemini 3.6 Flash, 3.5 Flash-Lite and a 3.5 Flash Cyber variant shipped together in July, with Robotics ER 2 covering the embodied case. One frontier model has become a catalogue.

See our analysis →

Google — Introducing Gemini Omni → · TechCrunch — Google's Gemini Omni turns images, audio, and text into video → · TechCrunch — Google releases three new Gemini models — but no 3.5 Pro →