Black Forest Labs Releases FLUX 3 Video: Up to 20 Seconds of Native 1080p, Outperforms Seedance 2.0 in Benchmarks

8.6 On August 5, Black Forest Labs, the German AI lab renowned for its open-source FLUX series of image models, officially released its video generation model FLUX 3 Video to the general public. The model can generate clips up to 20 seconds long in native 1080p resolution with synchronized audio. In internal human evaluations, it outperformed competitors including ByteDance's Seedance 2.0 and MiniMax's H3.

On August 5, Black Forest Labs, the German AI lab renowned for its open-source FLUX series of image models, officially released its video generation model FLUX 3 Video to the general public. The model can generate clips up to 20 seconds long in native 1080p resolution with synchronized audio. In internal human evaluations, it outperformed competitors including ByteDance's Seedance 2.0 and MiniMax's H3.

690195_638197_4222

I. Unified Architecture: From "Generating Images" to "Understanding the World"

FLUX 3 is not an isolated video model but a multimodal foundation model trained on images, video, and audio within a single unified architecture. The core philosophy is that a model must learn from multiple modalities simultaneously to build a deep understanding of the physical world, rather than merely generating stylized pixels.

As Black Forest Labs co-founder and CEO Robin Rombach put it: "a model that only learns images can only generate images"—arguing that learning from video and audio forces the model to build a working representation of how objects hold together, how things move, and how events sound. Based on this philosophy, FLUX 3 Video demonstrates several key capabilities:

  • Multilingual Dialogue with Lip-Sync: Supports over a dozen languages, including English, Chinese, Spanish, Japanese, Hindi, and more, generating dialogue with natural accents and precise lip-syncing.

  • Multi-Shot Storytelling: Can generate multiple scenes and camera angles within a single request, maintaining narrative coherence across transitions.

  • Accurate Text Rendering: Capable of rendering specified typography naturally within the scene.

II. Performance Leadership: Strong Showing in Human Preference Tests

In internal human evaluations, FLUX 3 Video demonstrated advantages over existing competitors across several dimensions:

Comparison ModelText-to-Video Win RateImage-to-Video Win Rate
Luma Ray 3.293%—
Runway Gen-4.577%—
Grok Imagine Video69%—
Kling v3 Pro60%—
Seedance 2.052%52%
Gemini Omni Flash52%—
MiniMax H353%52%

In specific scoring, FLUX 3 Video achieved 1,135 points for text-to-video, surpassing Gemini Omni Flash and MiniMax H3; for image-to-video, its 1,051 points also outperformed Seedance 2.0 and MiniMax H3. It should be noted that these results are from the developer's internal evaluations; independent third-party benchmark tests have yet to be published.

III. Draft Mode and Pricing: Lowering the Cost of Creative Iteration

A distinctive feature of FLUX 3 Video is its "Draft Mode" . Users can first generate a fast, low-cost preview to iterate on ideas. Once satisfied with a preview, the model renders the final video at full quality, preserving the same subjects, composition, and motion as the approved version. This mode aims to solve the "high trial-and-error cost" pain point in video generation, allowing creators to explore different directions confidently.

The model charges per second of generated video, with different price tiers:

ModeResolutionText/Image-to-Video (per sec)Video Continuation (per sec)
Draft Mode720p$0.06$0.12
Standard720p$0.17$0.43
Standard1080p$0.29$0.54

IV. Extending to Physical AI: From Content Creation to Robotics

Black Forest Labs' strategic vision for FLUX 3 extends beyond content generation. The lab is collaborating with robotics company mimic to develop FLUX-mimic, a video-action model based on FLUX 3's video backbone, for predicting robot actions.

The system is reportedly being tested on Audi factory production lines for tasks such as assembling flexible components like door seals—applications previously difficult to automate. The system responds in approximately 101 milliseconds, claimed to be comparable to human reflexes, and can be fine-tuned for some tasks with as little as 30 minutes of robot data. This move positions FLUX 3 not just as a creative tool but potentially as a general-purpose multimodal engine connecting the digital and physical worlds.

V. Availability and Roadmap

FLUX 3 Video is now available via the BFL API and partner platforms like OpenRouter. Black Forest Labs has outlined its roadmap: FLUX 3 Image for image generation and editing is expected in the coming weeks, with an open-weight version, FLUX 3 Dev, planned for later in 2026. Developers have confirmed that 2K and 4K resolution support for FLUX 3 Video is also on the roadmap.

The release of FLUX 3 Video signals that the video generation race has moved from "can it generate" to "how well, how long, and how controllably." Its Draft Mode precisely targets the core pain point of creators—the high cost of trial and error. Black Forest Labs' broader strategy of unifying video, audio, and even robotic action prediction within a single architecture offers a compelling model for the commercialization of "world models." For video creators, FLUX 3 Video provides another tool capable of functioning in real production workflows.