Skip to content
News TechTimes Jul 2026

Black Forest Labs: FLUX 3 unifies image, video, and audio generation

Black Forest Labs released FLUX 3 on July 23, 2026, marking a significant architectural step from their earlier work. Where previous FLUX releases focused on image generation, FLUX 3 is a single multimodal foundation model trained jointly on images, video, and audio from the start. The unified architecture means each modality strengthens the others during training rather than being stitched together after the fact.

The initial early access release covers the video and action prediction components. The image generation release follows in the coming weeks, with an open-weight developer version planned for later in 2026. Video generation produces clips up to 20 seconds long with native audio — the sound is generated alongside the visuals rather than added separately. Early evaluations highlight strong performance in human facial expression capture, sound-to-physical-event association, and multilingual output.

For designers, the most immediate relevance is in the tooling ecosystem. Adobe Photoshop and Picsart have already integrated earlier FLUX models, and FLUX 3’s expanded capabilities in video and audio make it a more complete substrate for AI-assisted creative production. The action prediction capability — which allows the model to predict robot movements — is separate from design work, but it signals the depth of generalization the architecture can achieve.

The robotics application is operational rather than experimental. Through a partnership with Mimic Robotics, FLUX 3 is driving robots at Audi production lines, with a real-world reaction time of 101 milliseconds. The company has positioned this as part of a longer strategy toward physical AI, where the same model that generates images and video also guides machines acting in the physical world.

CEO Robin Rombach described the unified architecture as essential: “Joint training within one unified architecture is what will get us there, because each training modality strengthens the others.” That design choice distinguishes FLUX 3 from multimodal systems that combine separately trained components.