Model
Black Forest Labs Releases FLUX 3: A Multimodal Foundation Model Unifying Image, Video, Audio, and Robot Action Prediction
Black Forest Labs releases FLUX 3, the first multimodal foundation model that jointly learns images, videos, and audio in a single architecture and outputs video, audio, and action predictions from the same set of weights. FLUX 3 Video can generate up to 20-second videos with native audio in a single pass, and in human preference tests for 10-second 720p text-to-video generation, it beats Luma Ray 3.2 with a 93% preference rate.
Read the original (opens in a new tab)
News stream data aggregated by AI HOT