Black Forest Labs Unveils FLUX 3: AI Model Generates Video with Synced Audio
Black Forest Labs has launched FLUX 3 in early access, its first model that generates video. This marks a significant departure from previous models, which only produced still images.
The new system trains on images, video, and audio simultaneously, leveraging multimodality to create more convincing and lifelike content. In head-to-head comparisons, human reviewers preferred FLUX 3's output over Runway Gen-4.5 in 77% of cases and Luma Ray 3.2 in 93%.
FLUX 3 produces clips up to 20 seconds long with synced audio, generating a broad range of styles beyond photorealism. The model's versatility is evident in its ability to produce high-quality images as well.
Black Forest Labs sees FLUX 3 as more than just a content tool, it's a stepping stone towards machine learning and robotics. The company has developed FLUX-mimic, a robotics model that leverages the video-prediction engine of FLUX 3 to perform complex tasks like fitting flexible door seals.
Audi is already testing FLUX-mimic on its production line, with promising results. The system reacts in about 101 milliseconds, comparable to human visual reflexes. While FLUX 3's video capabilities are currently restricted to APIs and partner access, the open-weight Dev version will be released later in 2026.