FLUX 3
FLUX 3: one model for image, 20-second video with audio, and robot action prediction
Black Forest Labs launched FLUX 3 in early access — its first model generating video, audio, and robot action-prediction from one set of weights, alongside image generation (a FLUX.3 Mimic variant was announced with it). FLUX 3 Video produces clips with native audio up to 20 seconds from text, images or footage, with continuation, keyframe transitions, multilingual dialogue and clip chaining; the same architecture is already teaching robots tasks on an Audi assembly line. BFL's own preference tests: 77% wins vs Runway Gen-4.5, 93% vs Luma Ray 3.2. Image generation and the open-weights FLUX 3 Dev come in later rollout phases.