// news · multimodal · tools2026-08-11source: Google

Google's image models hit GA, and video becomes an input to image generation

Gemini 3.1 Flash Image and Gemini 3 Pro Image reached general availability. The Flash model can now take a video file as multimodal context alongside a text prompt, for thumbnails, posters and summary infographics.

Video-in, image-out is a narrower capability than it sounds and a more useful one. The task is not generation from scratch — it is comprehension: watch the clip, understand what it is about, and produce a single frame that represents it. That is a summarisation problem where the output happens to be pixels.

The demand is unambiguous. Every video platform needs thumbnails, every editor needs a poster frame, and the current process is a human scrubbing a timeline. Automating it well is worth more to more people than another point of photorealism.

GA status is the part developers care about. Preview endpoints carry no stability commitment, and the gap between a capability existing and a capability being safe to build a product on is measured in exactly this announcement.

See our analysis →

Google AI for Developers — Gemini API release notes →