Black Forest Labs Launches FLUX 3: First End-to-End Jointly Trained Model for Image, 20-Second Video, and Audio
What happened
Black Forest Labs (BFL), creators of the FLUX model series, launched FLUX 3 on July 23, 2026 β a jointly trained multimodal model capable of generating images or combined video/audio clips up to 20 seconds from a single prompt. The model launched in limited access.
Context and impact
FLUX 3 is architecturally distinct from prior generative pipelines that assemble separate image, video, and audio models. End-to-end joint training across modalities is considered a meaningful step toward truly unified multimodal generation. BFL directly competes with Sora (OpenAI), Veo (Google), Runway, and Kling β but with a single unified model rather than multiple specialized ones.
Details
- Video: up to 20 seconds from a text prompt
- Training: end-to-end joint training (not an assembled pipeline of separate models)
- Modalities: image, video, audio β all in one model
- Availability: limited access at launch, broader access planned
- Black Forest Labs: German startup, creators of FLUX 1.x and FLUX 2.x series
- Competitors: Sora, Veo 3, Runway Gen-4, Kling
Open original source
VentureBeat