✳ the wire · analysis
FLUX 3 does 20-second video with native audio — and we missed it for eleven days because we track no media models

Black Forest Labs announced FLUX 3 on July 23, 2026. FLUX 3 Video is in early access now: it generates up to 20 seconds in a single generation, and every output carries native audio rather than a soundtrack added afterwards. FLUX 3 Image follows in "the following weeks"; FLUX-mimic and FLUX 3 Action go to selected research and commercial partners; and FLUX 3 Dev, an open-weight multimodal backbone, is promised but undated.
The comparison figures are Black Forest Labs' own human evaluations, run on 720p ten-second text-to-video clips, and they are lopsided. Evaluators preferred FLUX 3 over Luma Ray 3.2 in 93% of comparisons, over Runway Gen-4.5 in 77%, over Grok Imagine Video in up to 69%, over Kling v3 Pro in 60%, and over Seedance 2.0 and Gemini Omni Flash in 52% — a coin flip against those last two. Treat all of it as vendor-reported. The company says full benchmarks and methodology will follow alongside broader availability, and until they do there is nothing here that anyone outside the lab can reproduce.
One detail is harder to wave away, because it is not a score. FLUX-mimic, built on the same backbone with the robotics company mimic, is described as running on production lines at Audi. A single set of weights trained jointly on images, video and audio, then extended to predict robot actions, is a materially different claim from "our video model rates well".
Why you are reading this eleven days late, stated plainly: our watchlist had no entry for Black Forest Labs, and no category for video, image, audio or voice at all — twenty-two vendors, zero media models. Nothing failed. The radar was pointed somewhere else. FLUX 3, FLUX 3 Image and FLUX 3 Dev are now tracked entries, so the next move gets detected rather than researched.
Source: Black Forest Labs (official announcement) ↗ · FLUX 3 tracker · the bench index


