ModelsMedia Generation 🇺🇸 26.07.2026 21:02

Black Forest Labs Releases FLUX 3: Multimodal Model for Images, Video, Audio, and Robot Action Prediction

Luma AILuma AI RunwayRunway xAIxAI Google/DeepMindGoogle/DeepMind
Black Forest Labs (BFL) has introduced FLUX 3, a multimodal foundation model trained on images, video, and audio within a single architecture. The model generates up to 20-second videos with synchronized audio and achieves high user preference results, surpassing Luma Ray 3.2 in 93% of cases and Runway Gen-4.5 in 77%.
The Black Forest Labs (BFL) team has released FLUX 3 — a multimodal foundation model trained on images, video, and audio within a single architecture. This is the first FLUX model capable of generating video, audio, and predicting actions from a single set of weights. BFL researchers argue that no single modality provides a complete description of the world: images capture spatial structure at a single moment, video restores time and reveals physical dynamics, and audio uncovers causal relationships between mechanical events and sound. Training on all modalities simultaneously forces them to constrain each other — sound must match collisions, motion must obey mass. The model is built on the Self-Flow method, which combines a flow-matching objective with a self-supervised feature reconstruction goal. The reference implementation on GitHub (ImageNet 256×256) uses SiT-XL/2 with per-frame temporal conditioning, a 25% mask, and self-distillation from an EMA teacher at layer 20 to a student at layer 8. For FLUX 3, BFL significantly scaled up compute and data resources, training the model on video, images, and audio simultaneously. FLUX 3 Video generates clips of up to 20 seconds in a single pass with native audio, supporting text-to-video, image-to-video, video-to-video, keyframe-to-video, and generative video-audio extension modes. BFL also highlights capabilities for multilingual dialogue, agentic assembly of clips into multi-scene sequences, and generation of high-quality typography with animation. The model is particularly strong in generating human facial expressions and associating sounds with physical events. In preliminary user preference tests, FLUX 3 beat Luma Ray 3.2 in 93% of cases, Runway Gen-4.5 in 77%, Grok Imagine Video up to 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, and Happy Horse 1.1 in 57%. Against Seedance 2.0 and Gemini Omni Flash, the result was 52%. Video prediction consumes over 95% of training compute resources, while audio uses less than 0.5% of tokens. The same foundation is used for FLUX-mimic — a robot policy running under 80 ms on a single RTX 5090. Access is restricted: video and actions are in early access, images come later, and open weights last.
Source: MarkTechPost — original
Our earlier posts on this topic ↓
Fresh news