Black Forest Labs previews multimodal model Flux 3
Black Forest Labs has previewed Flux 3, a unified multimodal foundation model designed to natively integrate image creation, audio synthesis, 720p video generation with up to 20 seconds of synchronized sound, and robotics action prediction. Early access features text-to-video, image-to-video, and keyframe transitions, with an open-weight community release planned.
Unifying video generation, synchronized audio, and robotics action prediction into a single foundation model marks a major shift towards true cross-modal physical and generative AI intelligence.
* Combining audio and video generation natively eliminates secondary lip-syncing and sound design pipelines.
* Incorporating robotics action prediction signals a broader move from digital content creation into embodiment and real-world interaction.
* Extending video duration to 20 seconds at 720p resolution sets a strong benchmark for production-ready generative clips.
DISCOVERED
1h ago
2026-07-26
PUBLISHED
1h ago
2026-07-26
RELEVANCE
AUTHOR
AI Search