Black Forest Labs Enters the AI Video Race With FLUX 3, But Can it Actually Challenge the Top Models?

Black Forest Labs, the team behind FLUX images, launched its first AI video model today. Here's how FLUX 3 Video actually performs against Runway, Kling, and Seedance.

Share
Black Forest Labs Enters the AI Video Race With FLUX 3, But Can it Actually Challenge the Top Models?

Black Forest Labs made its name in AI image generation, FLUX models already power features inside Adobe Photoshop and Picsart, and the company counts filmmaker Martin Scorsese among its cited professional users. Today, the company entered AI video for the first time, launching FLUX 3 Video generally available through its API and select partners.

Who Black Forest Labs Actually Is

Founded in 2024 by former Stability AI researchers Robin Rombach, Andreas Blattmann, and Patrick Esser, Black Forest Labs built FLUX into one of the most downloaded text-to-image models on Hugging Face, ranking near the top of image generation benchmarks against much larger competitors. The company now runs roughly 100 people across Freiburg, Germany, and San Francisco. FLUX 3 Video marks its first public video generation model, extending an architecture the company also plans to apply to robotics.

What FLUX 3 Video Actually Does

The initial release generates clips up to 20 seconds long at HD (720p), with Full HD (1080p) available via upscaling, and native audio generated alongside the video in the same pass. Capabilities available today include text-to-video, image-to-video with start frames, end frames, or multiple ordered keyframes, video continuation (extending up to 4 seconds of existing footage while maintaining motion and audio continuity), multi-shot generation within a single continuous video, and dialogue generation across more than a dozen languages, including English, Chinese, Spanish, French, German, Japanese, and Hindi, with lip-sync accuracy.

A genuinely useful, lower-friction feature: Draft Mode. A rough preview generates at $0.06 per second, and if the direction is right, the same prompt renders at full quality, retaining the same subjects, composition, and motion, starting at $0.17 per second. That's a real cost-control mechanism for iteration, letting creators test ideas cheaply before committing to full-quality renders, addressing exactly the kind of expensive trial-and-error that drives up cost on other AI video platforms.

The Actual Benchmark Numbers

Black Forest Labs published human preference testing results comparing 10-second, 720p text-to-video clips with audio against existing models. FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, over Runway Gen-4.5 in 77%, over Grok Imagine Video in 69%, over Kling v3 Pro in 60%, and over Seedance 2.0 and Gemini Omni Flash in 52%, close to a coin flip, indicating genuine parity rather than a clear win at the top of the field.

Worth being direct about this: these are the company's own published results, not independently verified third-party benchmarks. They're a reasonable signal of real competitiveness, especially the wide margins against Runway and Luma specifically, but should be treated as a strong claim worth testing hands-on rather than settled fact.

What's Genuinely Different About the Approach

Most competing models handle video, audio, and other modalities as separate systems stitched together behind one interface. Black Forest Labs' pitch is that FLUX 3 is trained jointly across image, video, and audio in a single architecture, meaning the model learns real physical relationships between modalities during training, sound matching physical impact, motion respecting mass and momentum, rather than approximating those relationships after the fact.

Whether that architectural distinction produces a meaningfully different creative result versus a well-engineered multi-model pipeline is difficult to judge from a company blog post alone, but it's a genuinely different technical bet than most competitors are making, and one worth watching as more hands-on testing comes in.

Competitive Context

FLUX 3 Video enters a field that's consolidated fast around table-stakes features, native audio, character consistency, multi-shot support, this year, as covered in BRC's breakdown of the broader AI video app landscape.

What differentiates Black Forest Labs specifically is an existing, credible reputation in image generation and real enterprise distribution through Adobe and Picsart, giving it a plausible path to adoption beyond a typical first-generation video model launch.

The Signal in the Noise

The honest read on whether FLUX 3 can "challenge the top models" is genuinely mixed, and that's not a knock, it's actually a strong debut. Wide preference margins against Runway and Luma suggest real competitiveness in the mid-tier of the field. A near-coin-flip result against Seedance 2.0 suggests it's not yet a clear leader over the field's current strongest performers.

Draft Mode is arguably the most immediately useful feature for working creators, regardless of how the raw generation quality ultimately compares, since cheap iteration before a full-quality render solves a real, practical cost problem every AI video platform shares.

Specs & Pricing

  • Max clip length: 20 seconds
  • Resolution: HD (720p) native, Full HD (1080p) via upscaling
  • Draft Mode: $0.06/second
  • Full quality: from $0.17/second
  • Languages supported: 12+, including English, Chinese, Spanish, French, German, Japanese, Hindi
  • Availability: BFL API and select partners now; 2K, 4K, and Open Weights variants planned

Resources & Reads