German AI company Black Forest Labs has released an early version of FLUX 3, marking the first time its flagship model has expanded from still image generation to video generation. The model can generate videos up to 20 seconds long and simultaneously output audio that matches the visuals. The company is also applying the same underlying capabilities to robot control, and Audi has begun testing the system on its production line.
From image modeling to multimodal generation
Black Forest Labs has garnered attention for its FLUX image generation model. The newly released FLUX 3 no longer handles only single image tasks, but instead trains images, videos, and audio within the same system. According to the company, this means the model not only generates images but also learns object motion, contact, and temporal relationships.
Video is the core feature of this update. FLUX 3 can now generate video clips up to 20 seconds long, with audio output synchronized with the video, including dialogue, ambient sounds, and sound effects.
- The preference rate for Runway Gen-4.5 is 77%.
- The preference rate for LumaRay 3.2 is 93%.
- The win rate against Gemini Omni and Seedance is 52%.
However, these results are subjective preference tests, not fixed scores under a standardized scale.
The same underlying model enters the robot scene
Black Forest Labs has further applied this video prediction capability to robotic systems, launching FLUX-mimic in collaboration with Zurich-based mimic robotics. This system adds a decoding component to FLUX 3 to translate the model's internal representations of actions and motions into actual robotic operations.
The company believes that being able to predict motion in videos means that the model is closer to understanding weight, contact, and time differences in the real world, which is exactly the capability that robots need to perform physical tasks.
Audi is currently testing this system on its production line, with applications including the installation of flexible door seals. Tasks involving the handling of soft materials are typically more difficult to accomplish with traditional automation solutions. Black Forest Labs states that the system's response time is approximately 101 milliseconds, close to the level of human visual reflexes.
Open strategies remain limited
Black Forest Labs was founded in August 2024 by researchers who participated in the development of the first-generation Stable Diffusion model for Stability AI. The company had previously gained rapid attention in the open-source image generation field with models such as Flux Dev and Schnell.
However, FLUX 3 is not fully open this time. Currently, the Video and Action features are only available through the API and to select partners in early access, while image generation functionality will be available in the coming weeks. The company plans to release an open-weighted Dev version later in 2026, which is currently the only version planned to support on-premises deployment.











