German AI company Black Forest Labs (BFL) today released Flux 3, a multimodal foundation model that generates videos up to 20 seconds long with native audio, and introduced Flux-mimic, a video-action model for robotics being tested at Audi. The release marks a significant step in the company’s push toward what it calls “real-world visual intelligence.”
Flux 3 learns from images, video, and audio simultaneously, a departure from models trained on a single modality. BFL argues that no single modality captures reality in full: images show spatial structure, video captures change over time, and audio reveals links between mechanical events and sounds. Training on all three together lets them fill gaps for one another, the company says.
The model is based on Self-Flow, BFL’s approach for teaching one model to generate and understand content at the same time. BFL says this unified learning process delivers better results than the previously standard flow-matching method, both in generation quality and in the model’s grasp of the physical world.
Video Generation With Native Audio
Flux 3 generates videos with native audio for the first time, supporting text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips for longer multi-shot sequences. BFL says the model is especially good at human facial expressions and matching sounds to physical events.
In early evaluations using 10-second clips at 720p, BFL reports that Flux 3 was preferred over eight rival models. The preference rate was 93% versus Luma Ray 3.2, 77% versus Runway Gen-4.5, 69% versus Grok Imagine Video, 60% versus Kling v3 Pro, 59% versus Happy Horse v1, 57% versus Happy Horse 1.1, 52% versus Seedance 2.0, and 52% versus Gemini Omni Flash. BFL says results are preliminary, and no independent tests are available yet.
The article notes that matching leading systems such as Seedance and Gemini Omni Flash would put Flux 3 among the top video models. Seedance has already reached Hollywood.
Flux-mimic and Robotics Applications
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
BFL also introduced Flux-mimic, a video-action model developed with Mimic Robotics. The model is being tested on production tasks at Audi. BFL says Flux 3 can also predict actions based on its understanding of the world. A component for actions provides the foundation for robotics applications.
Flux 3 uses a multimodal transformer with dedicated components to convert images, video, and audio into a shared internal representation and then turn it back into outputs. BFL defines real-world visual intelligence as models that can “perceive, predict, and act across physical and digital environments.”
Rollout and Availability
Flux 3 Video is already available. BFL plans a phased rollout with early-access phases for feedback and safety testing. Flux 3 Image is set to follow in the coming weeks, with early access within the next few weeks. Action prediction will initially be offered through select partners.
BFL plans to release open-weight access to the multimodal backbone under the name “Flux 3 Dev.” Longer term, BFL is working on next-generation models that aim to combine perception, action, and language prediction in a single model. BFL expects Flux 3 to improve image generation, especially for complex prompts and accurate text rendering in multiple languages.
The release comes as BFL is part of a broader push to build so-called world models. The company’s approach, Self-Flow, is designed to teach one model to generate and understand content simultaneously, a method BFL says outperforms the previously standard flow-matching technique.

