Black Forest Labs released FLUX 3 and FLUX-mimic on Thursday, pairing a new multimodal foundation model with the German lab's first model built to control physical machines. It is a sharp turn for a company best known for the image models behind a large share of Europe's generative tooling.

FLUX 3

The headline capability is video with native audio up to 20 seconds in a single generation — not a video pass followed by a separate soundtrack pass. It also handles multilingual dialogue and keyframe-to-video. Preliminary preference evaluations published by the lab put it at 77% against Runway Gen-4.5 and 93% against Luma Ray 3.2. Open weights for a FLUX 3 Dev variant are planned once early access closes.

FLUX-mimic

The robotics model is a video-action system — it predicts what should happen next in a scene and emits the actions to make it happen. It was built with Mimic Robotics, a Zurich-based manipulation company, and is in production deployment testing at Audi. That named industrial deployment is the part that separates it from most physical-AI announcements.

Why a video lab does robots

The bet is that a model trained to predict the next frame of the physical world has already learned most of what a robot policy needs — object permanence, contact, consequence. Video generation and robot control end up being the same prediction problem with different output heads. It is the same wager Google and Nvidia are making with world models, arrived at from the opposite direction.

The European angle

Black Forest Labs is the clearest non-US entrant in the physical-AI race, and its open-weight track record matters: the FLUX image models became default infrastructure partly because the weights shipped. Whether FLUX 3 Dev arrives on the same terms will determine if that repeats for video and control.