FLUX 3 Image
FLUX 3 Image: 4K, bounding-box layouts and 10 reference images
Black Forest Labs released FLUX 3 Image with 4K output, bounding-box layouts and up to 10 reference images.
Image generation and editing models and creative visual tools. — 71 releases covered on the show.
FLUX 3 Image: 4K, bounding-box layouts and 10 reference images
Black Forest Labs released FLUX 3 Image with 4K output, bounding-box layouts and up to 10 reference images.
Google Nano Banana 2.1 at $0.0336 per 1K image
Google released Nano Banana 2.1 at $0.0336 per 1K image. Every infographic on the show was made with it, and Alex found it really good on high reasoning.
Qwen Image 2.1 (7B) adds native transparency
Qwen Image 2.1 is a 7B image generation model with native transparency support.
OpenAI launches ChatGPT Images 2.5 with Flare and Sunburst API models
OpenAI released ChatGPT Images 2.5 with two new API models: GPT-Image-2.5 Flare for speed and volume and GPT-Image-2.5 Sunburst for precise editing with native transparent backgrounds, both at $30 per million image output tokens and up to 50% lower latency than Images 2.0. ChatGPT gains Sketch (type @Sketch and draw), comment-on-image editing and templates. Alex made this week's thumbnails with Sunburst via Fal; after fixing prompts written for GPT-image-2 and dropping the AI-generated reference photo, the second round beat Nano Banana Pro on likeness and tile text.
Meta Muse Image lands on the Meta Model API at $0.01 per image
Meta launched Muse Image on the Meta Model API at $0.01 per image, opening up what was previously only available through Meta AI surfaces. The standout is its agentic reasoning pipeline — it plans, runs web searches, generates code, and self-checks before producing the image. Also available on fal, Runway, and OpenRouter.
xAI Imagine Image 2.0 lands #2 on Arena for T2I and editing
xAI released Imagine Image 2.0, its updated image generation and editing model. It ranks #2 on Arena for both text-to-image and editing.
FLUX 3: one model for image, 20-second video with audio, and robot action prediction
Black Forest Labs launched FLUX 3 in early access — its first model generating video, audio, and robot action-prediction from one set of weights, alongside image generation (a FLUX.3 Mimic variant was announced with it). FLUX 3 Video produces clips with native audio up to 20 seconds from text, images or footage, with continuation, keyframe transitions, multilingual dialogue and clip chaining; the same architecture is already teaching robots tasks on an Audi assembly line. BFL's own preference tests: 77% wins vs Runway Gen-4.5, 93% vs Luma Ray 3.2. Image generation and the open-weights FLUX 3 Dev come in later rollout phases.
Microsoft ships MAI-Image-2.5-Pro and MAI-Voice-2-Flash, launched live during the ThursdAI broadcast
Microsoft AI released two in-house models the same morning as the Jul 23 live show: MAI-Image-2.5-Pro, the flagship high-fidelity tier of its image line ($5/$8 per 1M text/image-input tokens, $106 per 1M image-output tokens), and MAI-Voice-2-Flash, a voice tier Microsoft pitches as 2x faster than MAI-Voice-2 and 32% cheaper at $15 per 1M characters — continuing the build-out of first-party MAI models alongside the OpenAI partnership.
Reve 2.1 takes #2 on the Text-to-Image Arena with layer-based generation
Released a month after Reve 2.0 (and mid-way through the ThursdAI live show), Reve 2.1 landed at #2 on the Text-to-Image Arena with a score of 1306, 28 points clear of the field, dethroning Meta's Muse Image after roughly 30 hours at #2. Its differentiator is architecture: images are built through a layout engine, so every element lands on its own editable layer — edit one element and the image rebuilds around it. Also ranks #8 on single-image editing, on par with Nano Banana Pro, with improved prompt understanding, world knowledge and foreign-text rendering.
ByteDance releases Seedream 5.0 Pro with precision editing and layer separation
The flagship tier of the Seedream 5 line pitches a shift from image generator to design tool: interactive precision editing (point, lasso, sketch), intelligent layer separation that decomposes an image into editable layers, dense infographic rendering, and native text in 10+ languages. Rollout is enterprise-first via the BytePlus API, Dreamina and Magnific, with Seedance 2.5 video pre-announced for roughly ten days later.
Meta Superintelligence Labs ships Muse Image and previews Muse Video
MSL's first media-generation models: Muse Image is live in the Meta AI app, Instagram Stories (US) and WhatsApp, with agentic generation that calls web search and code execution, multi-reference composition, and Instagram social-context conditioning. Muse Video shares the same pretraining base and adds native audio, debuting at #3 on Arena text-to-video while Muse Image lands #2 on image. There is no public API, and public Instagram accounts are opted in to @-mention remixing by default.
NanoBanana 2 Lite: sub-4-second images at ~3¢ per 1,000
Google's NanoBanana 2 Lite generates images in under four seconds starting at $0.034 per 1,000 images, with quality above the original NanoBanana. The Interactions API hit GA the same week.
Ideogram 4.0 becomes the top open-weight text-to-image model
Ideogram released Ideogram 4.0, a 9.3B-parameter text-to-image model with open weights under a non-commercial license. It leads open-weight image models on typography and layout, with bounding-box/layout-style prompting that trades casual generation ease for precise structured control.
Reve 2.0 hits #2 on Text-to-Image Arena with layout-first editing
Reve 2.0 jumped to second place on Text-to-Image Arena (around 1200 ELO) with native 4K output, code-like layout control, and precise editing. Alex's live tests found inconsistent portrait identity, but the layout-first editor is the real differentiator for graphic and image iteration workflows.
Microsoft MAI-Image-2.5 jumps to #3 on Arena text-to-image
MAI-Image-2.5 jumped to number two on Arena's image-to-image leaderboard shortly after launch, with notable strength in image cleanup, backgrounds, documents, and diagrams. Hands-on tests on the show were mixed, and it is publicly accessible through playground.microsoft.ai.
PrismML's 1-bit Bonsai Image 4B runs local image gen under 1GB
PrismML released 1-bit and ternary versions of Bonsai Image 4B, a sub-1GB diffusion transformer for local image generation. The quantized model even runs in-browser via WebGPU and ships with an iOS app and a Hugging Face demo.
Pruna AI's P-Image-Upscale hits 128 megapixel outputs
Pruna AI released P-Image-Upscale, an image upscaling model that reaches 128 megapixel outputs with fast generation and predictable pricing. It is available through Pruna's API and on Replicate.
Runway launches Project Luxo for solo-creator short films
Runway launched Project Luxo, claiming AI-generated video has crossed the uncanny valley for solo-creator short films. The pitch is that a single creator can now produce watchable short-form films end to end with Runway's stack.
Gemini Omni: 'create anything from anything' conversational video editor
Google DeepMind launched Gemini Omni, a multimodal 'create anything from anything' model debuting as Google's first conversational video editor. Unlike pure text-to-video systems, Omni is an iterative multi-turn editing model that combines Gemini intelligence, world knowledge, multimodal inputs and generative media, in the same way Nano Banana brought Gemini to interactive image editing. It is available in the Gemini app, Google Flow and YouTube, with API support coming soon.
Krea 2: Krea's first from-scratch foundation image model
Krea released Krea 2, its first foundation image model trained from scratch, built over six to seven months by nearly half the company. It focuses on aesthetic diversity, style control with up to 4 reference images, and moodboard-driven workflows, generating images in roughly 15 seconds. Co-founder and CEO Victor Perez joined the show to walk through it.
Anthropic ships Claude Design research preview, Figma stock drops 7%
Anthropic released Claude Design as a research preview running on Opus 4.7 at claude.ai/design, and Figma stock dropped 7% on the news. Alex generated a full ThursdAI brand kit including logo, design tokens, and the episode opener videos end-to-end inside Claude Design, then had Codex pick up the kit and produce a GPT-5.5 launch video in 9 minutes. Anthropic also added a new usage meter to Claude Max settings.
Baidu ERNIE-Image: 8B DiT ranks #1 on GenEval among open models
Baidu released ERNIE-Image, an 8B diffusion transformer that ranks #1 on GenEval among open models and features precise multilingual text rendering. It is part of this week's wave of Chinese open releases in image and 3D generation.
OpenAI's GPT-Image-2 leaks on LM Arena under three codenames
OpenAI's GPT-Image-2 posted the biggest single jump ever recorded on Arena, sitting 200+ ELO points above the previous top image model even on medium reasoning. The thinking/reasoning image model generates functioning QR codes, pixel-perfect infographics, 4K output, multi-image character consistency, and equirectangular 360-degree images that Peter Gostev stitched into a walkable street-view reconstruction of ancient Babylon. It even produces screenshots of IDEs containing SVG code that actually renders, enabling a new design-then-implement meta with Codex.
Alibaba Wan2.7-Image unifies generation, editing, and text rendering
Alibaba's Wan team released Wan2.7-Image, a unified image model covering generation, editing, text rendering, and multi-image consistency. The panel covered it in the open ecosystem round-up alongside the Qwen updates.
Microsoft MAI releases MAI-Image-2 image generation model
MAI-Image-2 is Microsoft's new in-house image generation model, debuting at #3 in image-gen rankings as part of the MAI three-model release. The panel compared its positioning against specialist image products and foundation-model APIs.
Luma Labs Uni-1 thinks and generates pixels simultaneously, #1 preference Elo
Luma Labs released Uni-1, an LLM-based image model that thinks and generates pixels simultaneously and claims the number-one human preference Elo. Unlike traditional diffusion workflows you converse with it and iterate together toward results, and it can also generate infographics; a surprising pivot from Luma's video focus.
Modular 26.2 runs FLUX.2 in under a second, 99% cheaper than Nano Banana
Modular shipped its 26.2 release with state-of-the-art image generation, running FLUX.2 in under one second (sub-300ms claims) at 99% lower cost than Nano Banana, plus upgraded AI coding with Mojo. Alex noted the surprise of an inference platform releasing model-level optimization and hoped the approach spreads to all image generation.
Phota Labs launches Phota Studio + API with identity-preserving personalization
Phota Labs launched Phota Studio and an API around a photography-focused image model with identity-preserving personalization: upload a batch of your photos, it trains a personal model, and the generated images actually resemble you. Alex flagged the personalization as a real capability jump over the crowd of photo startups, for professional shots, photo fixes, and adding people to photos.
NVIDIA DLSS 5 adds a generative AI filter for photo-realistic lighting
Announced at GTC, NVIDIA's DLSS 5 introduces a new generative AI filter bringing photo-realistic lighting to RTX 50-series GPUs. It applies generative models to real-time game rendering, extending DLSS beyond upscaling and frame generation.
Black Forest Labs introduces Self-Flow
Black Forest Labs published Self-Flow, new research from the FLUX makers in the AI art and diffusion space. It was included in the week's AI Art & Diffusion roundup.
Google DeepMind launches Nano Banana 2 image model mid-show
Google DeepMind announced Nano Banana 2 during the show, a Flash-quality tier of its image model line. Alex broke in mid-TLDR to describe near-Pro image quality at roughly half the price, plus a new image search capability.
Quiver tackles SVG generation with Arrow 1.0
Quiver released Arrow 1.0, pitched as solving SVG generation. It was included in the week's AI art and diffusion roundup as a notable niche release for vector graphics.
Alibaba launches Qwen-Image-2.0 with native 2K resolution
Alibaba's Qwen team launched Qwen-Image-2.0, a 7B-parameter image generation model with native 2K resolution output and superior text rendering. Available to try on chat.qwen.ai.
Tongyi Lab releases Z-Image generation model
Alibaba's Tongyi Lab released Z-Image, a new image generation model, with support landing in the open-source DiffSynth-Studio toolkit on GitHub. Covered in the AI Art segment alongside HunyuanImage 3.0.
Tencent launches HunyuanImage 3.0-Instruct image model
Tencent's Hunyuan team launched HunyuanImage 3.0-Instruct, an instruction-tuned version of its image generation model. Covered briefly in the AI Art segment alongside other new image models this week.
xAI launches Grok Imagine API with video generation
xAI released the Grok Imagine API, exposing its image and video generation capabilities to developers through the xAI console. The show subtitle notes Grok Imagine ranking #1 among generation models this week.
Black Forest Labs drops Flux 2 Klein, fast open-weights image model
Wolfram broke the news mid-show: Black Forest Labs released Flux 2 Klein, a fast 4B/9B image generation model with open weights under Apache 2.0. It is designed for near-real-time editing and style iteration, and Alex used it minutes later in his live Claude Cowork demo.
Qwen Edit 2512 optimized by PrunaAI: high-res images in under 7s
PrunaAI released an optimized version of Qwen Edit 2512 that generates high-resolution realistic images in under 7 seconds. The optimized model is available to run on Replicate.
Flux 3 becomes the new gold standard for image generation
Flux 3 dropped in August and immediately became the gold standard for image generation, landing three years almost to the day after Stable Diffusion first went public. Wolfram used it as the yardstick for how far image AI traveled in those three years.
GPT-4o native image generation sparks Ghibli-mania
OpenAI shipped native image generation in GPT-4o, producing the viral Ghibli-style image wave and bringing AI image creation to the ChatGPT mainstream. Wolfram cited the 2025 paradigm shift in image generation as his release of the year.
Reve ships a 4-in-1 image creation and editing platform
Reve (rendered as 'RevA' in the episode) emerged in September as a four-in-one image creation and editing platform. Alex said he still uses it daily, making it one of the year's sleeper product hits.
OpenAI GPT Image 1.5: 4x faster, 20% cheaper, #1 on LMSYS Image Arena
OpenAI released GPT Image 1.5, an upgraded image generation model that is 4x faster and 20% cheaper than its predecessor. It debuted at #1 on the LMSYS Image Arena leaderboard, part of OpenAI's rapid-fire release week.
SeeDream 4.5 adds multi-reference fusion and stronger text rendering
ByteDance's SeeDream 4.5 image model shipped with emphasis on multi-reference fusion and improved text rendering, an area the panel noted remains a key differentiator among image generators.
Kling O1 Image expands Kling into image generation
Alongside its video update, Kling shipped O1 Image, expanding the company's generation stack into still images. The release rounds out Kling's multimodal offering beyond its core video models.
Pruna P-Image promises sub-second image generation at $0.005
Pruna AI promoted P-Image, an image generation offering with sub-second generation times at roughly $0.005 per image. The release fit the week's diffusion theme of competing on speed and cost efficiency rather than just quality.
Tongyi's Z-Image Turbo brings sub-second open image generation
Alibaba's Tongyi lab released Z-Image Turbo, a 6B-parameter open image generation model that produces images in under a second. It pushes open-source image generation toward real-time speeds at a fraction of the size of competing models.
Black Forest Labs releases FLUX.2, a 32B multi-reference image model
Black Forest Labs released FLUX.2, a 32B-parameter image model with open weights (FLUX.2-dev) that supports multi-reference image editing. It lets users combine multiple reference images and prompt edits with variables, a step up in controllable image editing.
Nano Banana Pro generates 4K images with perfect text
Google's upgraded image model dropped as breaking news mid-show, adding visible thinking traces, 4K resolution output, and SynthID watermarking with C2PA metadata. Alex demoed it live by one-shotting an 8MB AI-news infographic with flawless text and pixel-accurate logos across the entire image. It also powers generative UIs in Gemini, building interactive dashboards with real data on the fly.
Qwen Image Edit gains Multi-Angle LoRA for camera control
A Multi-Angle LoRA for Qwen Image Edit landed, enabling camera-control style edits that re-render a scene from new angles. Available as a Hugging Face space and on fal, it shows the fast-moving open ecosystem building on Qwen's image editing models.
NVIDIA releases ChronoEdit-14B Upscaler LoRA
NVIDIA released an Upscaler LoRA for its ChronoEdit-14B image editing model, available on Hugging Face with Diffusers pipeline support. It adds high-quality upscaling to the ChronoEdit physics-aware editing stack.
DiT360: SOTA panoramic image generation with hybrid training
DiT360 is a diffusion-transformer approach to panoramic image generation that uses hybrid training across perspective and panoramic data to reach state-of-the-art quality. The project page and GitHub release make the work reproducible.
Riverflow 1 tops the image-editing leaderboard
Sourceful's Riverflow 1 image-editing model took the top spot on the image-editing leaderboard. It is a notable result from a smaller lab in a category dominated by big-name image models.
Reve launches 4-in-1 AI visual platform taking on Nano Banana and Seedream
Reve launched a 4-in-1 AI visual creation platform combining image generation, editing, and related visual workflows in one app. The panel spends real time on it as a serious challenger to Nano Banana and Seedream in the AI image tooling race.
Hunyuan SRPO: preference optimization that supercharges diffusion models
Tencent Hunyuan published SRPO (Semantic Relative Preference Optimization), a post-training technique that significantly improves the output quality of diffusion image models. The team released weights on Hugging Face along with a project page and striking before/after comparisons.
Black Forest Labs drops FLUX.1 Kontext, SOTA image editing
Black Forest Labs, creators of Flux, released Kontext: three models (Pro, Max, and a 12B open-weights Dev in private preview) for consistent, context-aware text and image editing. Unlike GPT-image or VEO-style regeneration, Kontext keeps identity consistent across edits, adding what you ask for without changing your face every generation. Broke as news during the show.
HiDream E1: open-weights image model with standout Ghibli style
HiDream released E1, an open-weights image editing/generation model (Apache 2.0-style licensing) noted for beautiful Ghibli-style outputs. It ranks #4 on the Artificial Analysis image arena leaderboard, sitting among top contenders like Google Imagen and ReCraft.
Runway References brings character and scene consistency to Gen-4
Runway launched References for Gen-4 on all paid plans, letting creators supply reference images (characters, outfits, locations, even selfies) and use tags in prompts to keep those elements consistent across generations. It tackles AI video's biggest pain point, frame-to-frame identity drift, at no extra credit cost per run.
OpenAI's GPT Image generation lands in the API as gpt-image-1
OpenAI's powerful image generation capabilities, previously locked inside ChatGPT, are now available to developers via API under the official name gpt-image-1. This was the big one many developers were waiting for, opening up the viral image generation and editing capabilities for building AI art and image editing applications.
Tencent's Hunyuan 3D 2.5 jumps to 10B params with PBR textures and rigging
Tencent updated its 3D generation model to Hunyuan 3D 2.5, now boasting 10 billion parameters, up from 1B. They highlight massive leaps in precision with 1024-resolution geometry, high-quality textures with PBR support, and improved skeletal rigging for animation.
ByteDance Seedream 3.0: bilingual 2K text-to-image model
ByteDance's Seed team announced Seedream 3.0, a powerful bilingual (Chinese/English) text-to-image model that generates native 2048x2048 images with fast inference of around 3 seconds for a 1K image on an A100. It challenges the top closed image generation models.
HiDream-I1-Dev: 17B MIT-licensed image model surpasses Flux 1.1 [pro]
HiDream released HiDream-I1-Dev, a 17B parameter open-weights image generation model under an MIT license. It became the new leading open-weights image generator, surpassing Flux 1.1 [pro] on quality benchmarks.
Runway Gen-4 announced with major gains in video consistency
Runway announced Gen-4, its next-generation video model focused on character and world consistency across shots. Example videos showed notably coherent characters and scenes, pushing AI video further toward usable filmmaking.
Ideogram 3.0 launches with strong text, logos, and style references
Ideogram launched version 3.0 of its image generation model with another SOTA claim. It is particularly strong on text and logo rendering, photorealism, and style references, continuing Ideogram's edge in typography-heavy image generation.
OpenAI enables native image generation in GPT-4o, internet goes Ghibli
OpenAI finally enabled GPT-4o's native auto-regressive image generation in ChatGPT, sparking the biggest mainstream AI buzz of the week as the internet ghiblified itself. Launched right after Gemini 2.5, it excels at instruction following, text rendering, and multi-turn editing, with viral demos ranging from ad mockups to a full Lord of the Rings trailer.
Reve emerges with SOTA diffusion image generation claims
Reve launched a new diffusion image generation model claiming state-of-the-art quality, reportedly beating heavyweights like Midjourney and Flux at roughly a penny per image. The previously low-profile lab made a splash with strong prompt adherence and image quality.
Gemini Co-Drawing demo uses native image output to help you draw
A Hugging Face space demo, Gemini Co-Drawing, uses Gemini's native image generation output to collaboratively complete and enhance your sketches as you draw. It showcases the new native image-output capability of Gemini 2.0 Flash in an interactive tool.
ByteDance unveils Seedream 2.0 bilingual image generation foundation model
ByteDance released Seedream 2.0, a native Chinese-English bilingual image generation foundation model, alongside a technical paper. It emphasizes excellent text rendering (especially Chinese), cultural nuance, and human preference alignment, generating high-quality, culturally relevant images from prompts in either language.
Gemini Flash gains native image generation and conversational editing
Google enabled native image generation in Gemini Flash Experimental, letting users generate and iteratively edit images conversationally inside the same multimodal model. The crew demoed it live on stream, editing photos of themselves with natural-language instructions, and saw it as a preview of how creative tools like Photoshop will work.
MiniMax launches Image-01 text-to-image model at 1/10 the cost
MiniMax released Image-01, a versatile text-to-image model the company positions at roughly one tenth the cost of competing image generation offerings. It is available through MiniMax's hosted platform.
Zhipu AI open-sources CogView 4, a 6B text-to-image model
Zhipu AI released CogView 4, a 6B-parameter open text-to-image model in the CogView family, with code available on GitHub. It is notable as an open-weights image generation option with strong Chinese and English prompt support.
DeepSeek Janus Pro: open multimodal models in 1.5B and 7B
Amid the R1 frenzy, DeepSeek also released Janus Pro, unified multimodal models at 1.5B and 7B parameters that handle both image understanding and image generation. The open release added to DeepSeek's week of dominating AI news headlines.
Follow Image Generation and everything else in AI — live every Thursday.