New Models
FLUX 3
FLUX 3: one model for image, 20-second video with audio, and robot action prediction
Black Forest Labs launched FLUX 3 in early access — its first model generating video, audio, and robot action-prediction from one set of weights, alongside image generation (a FLUX.3 Mimic variant was announced with it). FLUX 3 Video produces clips with native audio up to 20 seconds from text, images or footage, with continuation, keyframe transitions, multilingual dialogue and clip chaining; the same architecture is already teaching robots tasks on an Audi assembly line. BFL's own preference tests: 77% wins vs Runway Gen-4.5, 93% vs Luma Ray 3.2. Image generation and the open-weights FLUX 3 Dev come in later rollout phases.
20 sec max video-with-audio clip length77% / 93% preference wins vs Runway Gen-4.5 / Luma Ray 3.2
New Models
MAI-Image-2.5-Pro & MAI-Voice-2-Flash
Microsoft ships MAI-Image-2.5-Pro and MAI-Voice-2-Flash, launched live during the ThursdAI broadcast
Microsoft AI released two in-house models the same morning as the Jul 23 live show: MAI-Image-2.5-Pro, the flagship high-fidelity tier of its image line ($5/$8 per 1M text/image-input tokens, $106 per 1M image-output tokens), and MAI-Voice-2-Flash, a voice tier Microsoft pitches as 2x faster than MAI-Voice-2 and 32% cheaper at $15 per 1M characters — continuing the build-out of first-party MAI models alongside the OpenAI partnership.
$106 MAI-Image-2.5-Pro per 1M image-output tokens2x / -32% MAI-Voice-2-Flash speed / cost vs MAI-Voice-2
New Models
Reve 2.1
Reve 2.1 takes #2 on the Text-to-Image Arena with layer-based generation
Released a month after Reve 2.0 (and mid-way through the ThursdAI live show), Reve 2.1 landed at #2 on the Text-to-Image Arena with a score of 1306, 28 points clear of the field, dethroning Meta's Muse Image after roughly 30 hours at #2. Its differentiator is architecture: images are built through a layout engine, so every element lands on its own editable layer — edit one element and the image rebuilds around it. Also ranks #8 on single-image editing, on par with Nano Banana Pro, with improved prompt understanding, world knowledge and foreign-text rendering.
1306 #2 Text-to-Image Arena score+28 Points clear of next-best~30h How long Muse Image held #2
New Models
Seedream 5.0 Pro
ByteDance releases Seedream 5.0 Pro with precision editing and layer separation
The flagship tier of the Seedream 5 line pitches a shift from image generator to design tool: interactive precision editing (point, lasso, sketch), intelligent layer separation that decomposes an image into editable layers, dense infographic rendering, and native text in 10+ languages. Rollout is enterprise-first via the BytePlus API, Dreamina and Magnific, with Seedance 2.5 video pre-announced for roughly ten days later.
4K Max native resolution10+ Languages for native text
New Models
Muse Image & Muse Video
Meta Superintelligence Labs ships Muse Image and previews Muse Video
MSL's first media-generation models: Muse Image is live in the Meta AI app, Instagram Stories (US) and WhatsApp, with agentic generation that calls web search and code execution, multi-reference composition, and Instagram social-context conditioning. Muse Video shares the same pretraining base and adds native audio, debuting at #3 on Arena text-to-video while Muse Image lands #2 on image. There is no public API, and public Instagram accounts are opted in to @-mention remixing by default.
#2 Arena text-to-image debut#3 Arena text-to-video debut1280 Arena image score
New Models
NanoBanana 2 Lite
NanoBanana 2 Lite: sub-4-second images at ~3¢ per 1,000
Google's NanoBanana 2 Lite generates images in under four seconds starting at $0.034 per 1,000 images, with quality above the original NanoBanana. The Interactions API hit GA the same week.
3¢ per 1,000 images<4s generation time