Black Forest Labs unveiled FLUX 3 on Thursday, marking the first time its flagship model has moved beyond still-image generation into video.

The German AI lab, known for its FLUX image models, said the new system was trained on images, video, and audio together inside one shared model.
One model for images, video, and audio
That setup is what the industry calls multimodality: a single model learning from several kinds of information at once instead of relying on separate systems joined together afterward.
Video is the centerpiece of the release. FLUX 3 can generate clips up to 20 seconds long, with audio produced alongside the visuals and synchronized to on-screen events, including dialogue, sound effects, and ambient noise.
In early evaluations, human reviewers preferred FLUX 3 over Runway Gen-4.5 in 77% of head-to-head comparisons and over Luma Ray 3.2 in 93%. Against Gemini Omni and Seedance, the model appeared to hold a smaller advantage, winning 52% of the evaluations.
BFL described those results as preference testing rather than a fixed scoring system. Evaluators simply watched two clips and chose the one that looked and sounded more convincing, and the company counted how often FLUX 3 was selected.
Still-image generation remains part of the package
The company also presented FLUX 3 as a capable still-image model. Based on the images BFL shared, the system appears able to work across a broad range of styles and not just photorealistic output.

From media generation to robotics
BFL is positioning the release as more than a content tool. “A model that only learns images can only generate images,” co-founder and CEO Robin Rombach said.
The company’s thesis is that learning to predict video also teaches a model the physics underneath motion, including weight, contact, and timing. BFL argues that those are the same ingredients machines need to move through the physical world.
That idea is now packaged into FLUX-mimic, built with Zurich-based mimic robotics. The system takes FLUX 3’s video-prediction engine and adds a lightweight decoder, a small component that converts the model’s internal understanding of motion into actual robot actions.
Audi is already testing the system on jobs such as fitting flexible door seals, work that conventional automation has struggled with.
“Audi represents the kind of manufacturing partner we built FLUX-mimic for,” mimic co-founder Stephan-Daniel Gravert said. Audi’s Christoph Schneider said the robots now “solve complex soft-body manipulation work” that older machines could not handle.
BFL said the full system responds in about 101 milliseconds, putting it near the range of human visual reflexes.

How FLUX got here
FLUX’s rise came quickly. Black Forest Labs was founded in August 2024 by veteran researchers who had helped build the original Stable Diffusion models at Stability AI. Its early Flux releases went on to outperform MidJourney and also surpassed Stability’s weaker-than-expected Stable Diffusion 3.
The open-source Flux Dev and Schnell models then took over the “best open source image generator” title, a position many AI artists had expected Stable Diffusion 3.5, Stability’s second attempt, to reclaim.
That never happened. FLUX 1.1 Pro later topped the Artificial Analysis image arena in October, although that model was not open source.
BFL released FLUX.2 in November 2025, but it did not gain the same traction. The original Flux line kept its open-source lead until late 2025, when Alibaba’s Z-Image Turbo overtook it by matching its quality on lower-end consumer GPUs. At the time, one CivitAI user wrote, “This is what SD3 was supposed to be.”
Access and release timeline
BFL is presenting FLUX 3 as a return to form, but the model is not fully open yet. The Video and Action products are available in early access through APIs and selected partners, including mimic robotics.
Image generation, according to BFL, will follow “in the coming weeks.” The open-weight Dev release, the only version the company plans to make available for local use, is not expected until later in 2026.

