Quick Facts
- FLUX 3 generates images, video up to 20 seconds, and audio jointly in a single model pass, not through separate pipelines.
- Early access partners include Canva, Picsart, and Krea; pricing has not been published.
- Black Forest Labs raised $300 million in December 2025 at a $3.25 billion valuation, backed by Salesforce Ventures, a16z, and NVIDIA.
Black Forest Labs launched FLUX 3 on July 23, positioning the model as a single unified system for image generation, video creation, audio synthesis, and robotic control. The Freiburg, Germany-based lab says the model was jointly trained across all those modalities from the start.
That architecture choice separates FLUX 3 from competitors like Runway and Luma, which attach audio as a separate step after video generation. Black Forest Labs trained audio and video together, which the company says produces better causal alignment between sound and image frames.
How It Works
The model runs on an architecture Black Forest Labs calls Self-Flow, a self-supervised flow matching framework developed in-house. The company trained it on tens of millions of hours of general video to learn broad world dynamics, and on hundreds of thousands of hours of footage focused on human and robot manipulation tasks.
FLUX 3 supports text-to-video, image-to-video, video-to-video, and keyframe-controlled transitions. Users can chain generated shots through agent-controlled transitions to build longer sequences from a single prompt.
Benchmark Claims and Caveats
Black Forest Labs published internal benchmarks showing FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen 4.5 in 77%. Results were closer against Seedance 2.0 and Gemini Omni Flash.
The company itself labels the benchmark chart a “preliminary evaluation of an early FLUX 3 candidate,” meaning the results come from a pre-release checkpoint. The numbers have not been independently verified, and the shipping model is distinct from what was tested.
Robotics Already in Production
The launch included a robotics application. Black Forest Labs partnered with Mimic Robotics to build FLUX-mimic, a video-action model combining the FLUX 3 backbone with Mimic’s robot learning capabilities. Audi is testing it for production tasks including kitting parts, inserting electronic control units, and handling cables and soft materials.
Some Audi tasks were fine-tuned using as little as 30 minutes of robot data. Black Forest Labs reports FLUX-mimic achieves up to 10 times the sample efficiency of previous vision-language-action models, reaching a target success rate in half the training steps.
Rollout Schedule
The video and audio generation module entered early access on launch day. Image synthesis and editing through APIs will follow in the coming weeks. FLUX 3 Dev, an open-weight multimodal backbone, is planned for later in 2026. Pricing has not been announced for any tier.
Early access partners include Canva, Burda, Magnific, Krea, and Picsart. API access and private weight access will expand through a phased rollout over the next several months.
What This Means for Enterprise Buyers
Black Forest Labs CEO Robin Rombach made the company’s ambition clear at launch: “A model that only learns images can only generate images.” The company is framing FLUX 3 not as a media tool but as foundational infrastructure spanning generative content, simulation, and physical robotics.
For software and technology companies evaluating generative media vendors, FLUX 3 introduces a different buying question. The model’s joint training approach means enterprises could potentially access creative generation, computer use, and robotics control from a single API. Whether benchmark performance holds against independently verified tests will determine how quickly enterprise buyers move beyond early access.
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
