VentureBeat
Follow
Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start
Black Forest Labs has launched FLUX 3, a multimodal AI model capable of generating images, audio, and video clips up to 20 seconds from a single prompt. This new model extends its architecture to robotic vision and actions, aiming to unify creative generation, simulation, and robotics under "visual intelligence." FLUX 3 will be offered through four product lines: Video, Image, Action, and the open-source Dev version. Early Access for FLUX 3 Video and Action is now available, with FLUX 3 Image rolling out soon.The company highlights FLUX 3's joint training across modalities, differentiating it from models assembled from separate components. While BFL claims FLUX 3 outperforms competitors in preliminary video generation tests, specific pricing, service commitments, and comprehensive benchmarks are not yet public. Downloadable weights and an open-source license will be available later this year with the FLUX 3 Dev release.FLUX 3 Video supports text-to-video, image-to-video, and video-to-video generation with native audio. A key claimed capability is agentic chaining of clips to produce sequences lasting several minutes, addressing video continuity challenges. The model also reportedly excels at human facial expressions and multilingual output. BFL is also developing FLUX-mimic, a video-action model based on FLUX 3, for robotic action prediction. The unified architecture aims to improve data efficiency for robotics by leveraging pre-trained motion and behavior understanding.