Black Forest Labs launches Flux 3, its new multimodal model for video and audio generation

Black Forest Labs launches Flux 3, its new multimodal model for video and audio generation

German AI startup Black Forest Labs has launched Flux 3, its first multimodal model capable of generating videos up to 20 seconds long with native synchronized audio. The foundation model is trained jointly on images, video, and audio.

Flux 3 supports text to video, image to video, video to video, and keyframe controlled transitions. It can generate multilingual dialogue, match audio to physical events, and produce more natural facial expressions. Users can also connect generated shots through agent controlled transitions to create longer video sequences.

In Black Forest Labs' internal benchmarks, Flux 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen 4.5 in 77%, and several other models including Grok Imagine Video, Kling v3 Pro, and Happy Horse. Results were closer against Seedance 2.0 and Gemini Omni Flash, and have not yet been independently verified. Flux 3 Video is available now, while Flux 3 Image will enter early access soon. The company also plans to release the model's open weight backbone as Flux 3 Dev.

by Mauricio B. Holguin

Add as a preferred source on Google
No comments so far, maybe you want to be first?
Gu