What Black Forest Labs actually shipped

On July 23 Black Forest Labs, the German lab founded by the researchers behind Stable Diffusion, unveiled FLUX 3. It is not another image generator. FLUX 3 is a single model trained jointly on image, video, audio and action, so one system holds a shared representation of how the world looks, sounds and moves.

The headline piece is FLUX 3 Video. It generates clips up to 20 seconds long with native audio, the first time the lab has produced sound and picture together rather than bolting audio on afterwards. It handles text-to-video, image-to-video and video-to-video, keyframe transitions, on-screen typography and multilingual dialogue, and it can chain clips into longer sequences under an agent. The lab says its strengths are human facial expressions and matching a sound to the physical event that should cause it.

FLUX 3 launched in a gated Early Access programme for its Video and Action tiers. Anyone can apply, but Black Forest Labs decides who gets in.

The part that already runs on a factory floor

The same backbone powers FLUX-mimic, a video-action model built for robotics, and it is being tested on production tasks at Audi right now. That is the detail worth pausing on. A video-action model learns to predict the next movement the way a video model predicts the next frame, so the system that drafts a 20-second advertising clip and the system that guides a robot arm are the same model wearing two hats.

Audi running FLUX-mimic on a real line, not a research bench, is the signal that this is production-grade rather than a demo reel. For an operator it collapses two separate procurement decisions, creative generation and industrial automation, onto one European foundation model. That is unusual: the frontier labs have mostly kept their creative and their robotics ambitions in separate products.

Why a European frontier model changes your options

For two years the frontier of generated video has been American and Chinese ground: OpenAI's Sora, Google's Veo, and China's Kling and Seedance. FLUX 3 puts a European lab into that conversation. On Black Forest Labs' own preference tests of 10-second, 720p clips, viewers picked FLUX 3 over Runway Gen-4.5 in 77 percent of comparisons, over Kling v3 Pro 60 percent of the time, and over Seedance 2.0 and Google's Gemini Omni Flash 52 percent of the time. The lab is candid that no independent tests exist yet, so treat those figures as a claim from an interested party.

The score matters less than the licence. FLUX 3 Dev, an open-weight version of the multimodal backbone, is planned for later in 2026. An openly licensed European model you can run on your own hardware sidesteps the data-residency, export-control and vendor-continuity exposure that comes with routing your visual and robotics pipelines through a US or Chinese API. In a week when Washington publicly accused a Chinese lab of dodging chip export controls, a home-grown open option looks more valuable, not less.

What to actually do now

Do not rebuild a production pipeline on this yet. The Video and Action tiers are gated, and a service you have to apply for and can be de-prioritised from is not infrastructure. The date that decides everything is the FLUX 3 Dev open-weight release and, more precisely, its licence terms: whether it permits commercial use, self-hosting and fine-tuning is what separates a genuine sovereignty option from a marketing line.

Until then, treat FLUX 3 as a watch item with one concrete proof point. Audi putting FLUX-mimic on a production task tells you the robotics claims are real enough for a carmaker to trust on a line. If your roadmap has any dependence on generated video or on learned-action robotics, put the Dev licence date in your calendar and read the terms the day they land, before a US or Chinese contract locks you in for another cycle.