Black Forest Labs' FLUX 3 claims to outperform Seedance 2.0, Gemini Omni and Grok Imagine — the multimodal flow-model race sharpens
FLUX 3 from Black Forest Labs claims wins over Seedance 2.0, Gemini Omni and Grok Imagine across multimodal flow modelling — models that move from images to video and toward robotics-style action generation in the same architecture.
The architectural claim is more interesting than the benchmark claim. Flow models that span images, video and action are converging on a single representation for things that used to need separate systems. If that holds, the boundary between a media model and a robotics policy becomes an output-head detail rather than a different discipline.
That convergence is visible elsewhere this month. Mistral's Robostral Navigate doing embodied navigation from a single RGB camera is the same collapse from the other end — vision models that used to describe a scene are now acting in it.
The Neuron — Everything That Happened in AI Today → · AI Release Tracker — Latest AI Model Releases — July 2026 →