Black Forest Labs' FLUX 3 and Meta's Muse push image labs into the multimodal frontier
The image-generation labs are becoming multimodal frontier labs. Black Forest Labs announced FLUX 3 on 23 July — its first multimodal frontier model, with the video variant in early access — while Meta launched Muse Image on 7 July and previewed Muse Video. The boundary between an image model and a full multimodal system is dissolving.
FLUX 3 being described as a 'frontier' model is the notable escalation. Black Forest Labs built its name on open image generation; declaring a multimodal frontier model, video variant included, is a claim to compete at the top of the capability range rather than in a specialist niche. The image labs are no longer content to own one modality.
Meta's Muse pairing follows the same logic from the other direction. Launching Muse Image and previewing Muse Video as a matched set treats image and video as one product surface, not two, which is how a lab signals it sees generation as a single multimodal capability to be scaled rather than a set of separate models to be maintained.
The pattern across both is convergence on one object: a model that takes prompts across modalities and outputs pixels, frames, or both. As the labs that started with images add video and the labs that started with language add generation, the field is collapsing toward a shared architecture — the same convergence visible in the action-model and world-model work on the research side.
Digital Applied — Seven days, seven model releases: the new AI normal → · ThursdAI — July 2026 AI releases: OpenAI, Anthropic, Google DeepMind, Cognition →