FLUX 3 Is Out and It's Beating Runway, Kling, and Seedance in Early Benchmarks
Black Forest Labs just launched FLUX 3 — a unified multimodal model for image, video, and audio trained on one architecture. Early benchmarks show it preferred over Runway Gen-4.5 in 77% of comparisons. Here's what's real and what still needs independent verification.
Black Forest Labs — the German AI company behind the FLUX image generation models that have become a staple of the creative AI community — just launched FLUX 3, and it's a fundamentally different kind of model than anything they've shipped before. FLUX 3 is in early access now, and the preliminary benchmark numbers are worth paying attention to.
What makes FLUX 3 different from everything else
Every major AI video model right now is trained on one modality and extended to others: Runway trains a video model, then adds image and audio tools around it. Kling trains a video model with audio support. Veo 3.1 generates video and audio together. The underlying architecture in each case is a video model.
FLUX 3 is built differently. It's a multimodal foundation model trained simultaneously on images, video, and audio — not sequentially, not with separate models stitched together, but as one unified architecture learning from all three at once. Black Forest Labs' thesis: what the model needs to learn is not any individual modality but a representation of the world itself. "Images capture spatial structures at a specific point in time. Videos restore the dimension of time. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect."
When you train a model on all three at once, their mutual constraints inform each other. The sound has to match the impact. The motion has to obey the mass. The future has to follow from the past. The result, according to BFL, is a model with a more grounded understanding of physical reality than single-modality models trained separately.
The benchmark numbers

The numbers BFL is publishing are preliminary — the company is explicit that these are early evaluations and further improvements are expected during the early access phase. With that caveat clearly stated: FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons, over Luma Ray 3.2 in 93%, over Kling 3.0 Pro in 60%, over HappyHorse 1.1 in 57%, over Seedance 2.0 and Gemini Omni Flash in 52%.
Those are human-preference comparisons, not objective metrics — they reflect which output evaluators preferred in blind testing. The Runway comparison in particular is striking: 77% preference over the platform we've rated as the best professional workflow tool for filmmakers is a significant claim. Independent testing will be the real check on these numbers.
What FLUX 3 Video can actually do

FLUX 3 Video is the piece available in early access now. It generates video and audio jointly — native synchronized audio in every output, not as a separate step. Key capabilities:
- Text-to-video and image-to-video generation
- Video-to-video generation, carrying characters or elements from a source clip into a new scene or context
- Generative video-audio continuation from input video and audio
- Keyframe-to-video for controlled transitions between defined moments
- Multilingual dialogue
- Agentic chaining of individual clips into longer multi-shot sequences
- Up to 20 seconds per generation in a single pass
- Style range from "candid camcorder footage to animation and cinematics"
- Strong typography generation and animated designs
Preliminary evaluations used 10-second 720p clips. Resolution and duration limits will likely expand during and after the early access phase.
The action prediction angle — why it matters beyond filmmaking

The part of FLUX 3 that goes furthest beyond what any other video model has attempted: action prediction. BFL has used FLUX 3's video understanding as a foundation for robotics and physical AI — in partnership with mimic robotics, they've built FLUX-mimic, a video-action model being tested on real production tasks at Audi. The thesis is that video generation and physical AI run on the same foundation: a model that understands how things move and how physics works can both generate footage of that movement and predict what real robotic systems should do next.
For the filmmaking audience: the robotics angle is a separate story, but it's evidence that the multimodal architecture is producing real-world physical understanding rather than just better video output.
What's coming next
FLUX 3 Image — the image generation component — is in early access for a selected group and will open up in the coming weeks. FLUX 3 Dev (open weights for the multimodal backbone) is also coming, along with API access for all capabilities. BFL's stated goal is to eventually unify perceptual, action, and language prediction in the same model.
The honest caveats
Two things worth being clear about before this lands in your production workflow: these are BFL's own preliminary benchmarks, not independent evaluations. The 77% Runway preference number in particular will need confirmation from neutral third-party testing before it can be treated as settled. And FLUX 3 Video is in early access — not all features are at full capability yet, and the pricing structure hasn't been announced.