EmbeddingGemma 2
EmbeddingGemma 2, an open multimodal embedding model
Google released EmbeddingGemma 2, an open multimodal embedding model, the same week Perplexity opened pplx-embed-v2.
Models that natively combine text, image, audio, or video as inputs or outputs. — 66 releases covered on the show.
EmbeddingGemma 2, an open multimodal embedding model
Google released EmbeddingGemma 2, an open multimodal embedding model, the same week Perplexity opened pplx-embed-v2.
Liquid AI opens d1: d1-3B and d1-omni-600M decision models
Liquid AI opened its d1 decision models: d1 behind an API, the open d1-3B with text and vision, and d1-omni-600M, which also takes audio. d1-3B answers in 8ms on a GPU and about 50ms on a Jetson Orin Nano, tops Liquid's Decision Index under 10B parameters, and the API is a drop-in replacement for Jev.
Perplexity opens pplx-embed-v2 multimodal embedders
Perplexity released pplx-embed-v2, open multimodal embedding models on Hugging Face, with a write-up on multimodal embeddings beyond a single vector.
Reka Rho-1, a 19B research-preview omni model
Reka released Rho-1, a 19B research-preview omni model that collapses the multimodal stack into one model.
Qwen 3.8 Omni Flash with 1M context and Qwen 3.8 Live Translate
Alibaba's Qwen team released Qwen 3.8 Omni Flash with a 1M-token context window, alongside Qwen 3.8 Live Translate.
Gemini 3.8 Live claims #1 on the speech-to-speech index with 97 languages and async tool calls
Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, real-time speech-to-speech models scoring 82.6 to take #1 on the speech-to-speech quality index across 97 languages with asynchronous tool calls. Extended Thinking brings longer reasoning into a live model without breaking real-time interaction, and the release reaches everyone on Android and Google Search rather than just API users.
Gemini Omni 1.1 Flash tops Arena text-to-video with voice-consistent scene extension
Google's new video model dropped during the show: #1 on Arena's text-to-video leaderboard and #2 on image-to-video. It analyzes up to 10 seconds of previous footage to extend scenes while keeping character identity, voice, and lighting locked; adds first/last-frame control and infinite loops; and offers 360p draft generations with built-in upscaling. Rolling out in Google AI Studio, Flow, and Gemini Enterprise with API access.
dots3-note preview: 280B MoE omni model with TEMPO RL for long-horizon agents
Xiaohongshu's dots studio previewed dots3-note: a 280B MoE with 16B active parameters handling text, vision and audio with 512K context, released under Apache 2.0. It uses TEMPO RL training aimed at long-horizon agent tasks.
Gemini 3.7 Flash: 300+ tok/s mid-tier multimodal with a 50% price cut
Google dropped Gemini 3.7 Flash mid-show, right as Artificial Analysis co-founder George Cameron was on air. The mid-tier model clocks over 300 tokens/sec, comes with a 50% price cut through end of year, beats Muse Spark 1.2 on DeepSWE, and lands near the Pareto frontier for cost per task — instantly #1 on Artificial Analysis' model recommender for the cost/speed/intelligence trade-off. It is also strong at multimodal, one of the few models that can watch videos.
Wan 3.0 drops into public beta live during the show, with native 30-second generation
Alibaba's Tongyi Lab pushed Wan 3.0 into public beta minutes before ThursdAI went live: native 30-second single-shot generation and Omni-Reference conditioning on text, images, audio, and video together. Given the Wan line's open-weight track record, this is the drop the open source community most wants weights for.
ByteDance's SeedRealtime: a native audio-visual full-duplex LLM, free on Doubao
A single end-to-end model natively fusing audio, video, and text, replacing the cascaded ASR-VLM-TTS pipelines behind most voice agents: it listens, watches, and speaks simultaneously with turn-taking inside the model and no external VAD, cutting conversational pacing failures by 50% in human evals. It ties voices to faces in noisy rooms and speaks up proactively on scene changes, live for free in the Doubao app and its 300M+ users, ByteDance's first large-scale audio-visual full-duplex deployment.
MiniMax opens H3's weights, and the community ships LoRAs and Apple Silicon in 48 hours
H3 (Hailuo 3.0), a 33B omni-modal transformer generating up to 15 seconds at 2K with native stereo audio from unified text/image/video/audio context, landed on Hugging Face days after its announcement, per Victor Su Ortiz the first open-weight state-of-the-art omni video model. Within 48 hours the community shipped LoRA support and Apple Silicon inference, neither of which MiniMax optimized for, and X filled with recreated episodes of The Office. The panel also dug into the community license's litigation-linked restrictions on US use, the gap between downloadable and cleared.
Thinking Machines releases Inkling-Small: 276B/12B open MoE that beats its 975B sibling on agentic coding
The efficient sibling previewed alongside Inkling ships as open weights: 276B total with 12B active, natively multimodal with an encoder-free architecture (images via hierarchical patch encoding, audio via dMel spectrograms straight into the decoder), and variable thinking effort. On-policy distillation from Inkling plus two extra weeks of agentic-coding RL let it beat the 975B teacher on SWE-Bench Verified (80.2% vs 77.6%) and ARC-AGI-2 (40.1% vs 36.5%), though factual recall regressed hard (SimpleQA 20.6% vs 43.9%). Priced at $0.30/$1.20 per million tokens, roughly 3-4x cheaper than Inkling, with day-zero SGLang, Unsloth GGUF, and Baseten support. Dropped just after the July 30 show aired.
Lyria 3.5 generates full 3-minute songs with BPM and key control inside Flow Music
Google's flagship music model now produces cohesive three-minute songs with tempo and key-signature control in the prompt, more expressive multilingual vocals, style-transfer covers that preserve a track's structure, and lip-synced music videos via Gemini Omni Flash — plus a new Flow Music iOS app. All output is SynthID-watermarked. Google published no benchmarks against Suno or Udio, and early testers say paid Suno 5.5 still edges it, but as a free end-to-end create-to-publish stack it's a real move.
Alibaba previews Qwen3.8-Max at a claimed 2.4T parameters, 'second only to Fable 5'
Alibaba's Qwen team previewed Qwen3.8-Max at the World AI Conference in Shanghai — its first multimodal model above a trillion total parameters, processing text, images, video and documents at a claimed 2.4T. Alibaba shares rose as much as 5.4% in Hong Kong on the news. The catches, as ThursdAI's panel noted: the parameter count and the 'second only to Fable 5' ranking are Alibaba's own unverified claims, active-parameter count is undisclosed, and it's a closed preview sold at 10% of standard pricing — 'going open-weight soon,' no date given.
OpenMOSS open-sources MOSS-VL-Realtime, an 11B VLM that decides when to speak — and when to stay silent
OpenMOSS released MOSS-VL-Realtime, an open-source 11B-parameter vision-language model (~22.7GB) built for real-time streaming video that proactively speaks up or deliberately stays silent instead of only answering prompts. ThursdAI reported it as state-of-the-art on all three open proactivity benchmarks, with the base model included in the release.
Thinking Machines releases Inkling, a 975B open-weights MoE trained on 45T multimodal tokens
Mira Murati's Thinking Machines shipped Inkling, a 975B-total/41B-active Mixture-of-Experts transformer pretrained from scratch on 45 trillion tokens of text, images, audio and video, released under Apache 2.0. The ThursdAI panel called it the top US open-weights model right now — 41 on the Artificial Analysis Index — with encoder-free native reasoning over text, image and audio and a 1M-token context window. A leaner Inkling-Small (276B/12B active) was previewed alongside, and both run on the Tinker platform at a limited-time 50% discount.
PrismML compresses a full 27B model to 3.9GB so it runs on a phone
PrismML released Bonsai 27B, extreme quantizations of Qwen 3.6 27B under Apache 2.0: a 1-bit build at 3.9GB keeping ~90% of full-precision quality — small enough for an iPhone 17 Pro's memory budget — and a ternary build at 5.9GB keeping ~95%. Both stay multimodal with the full 262K-token context window. Nisten demoed it live on the show running on a phone and on a 6GB GTX 1660 Ti.
Google DeepMind debuts OmniFlash, first of the any-to-any Omni family
OmniFlash — first of Google's any-to-any Omni family — generates videos up to 10 seconds with precise conversational multi-turn editing via the Interactions API: say 'make it daytime' and it redoes light, sky and shadows. Editing Elo 1087 at $0.10 per second of output.
xAI launches Grok Imagine Video 1.5 with faster generation and native audio
xAI launched Grok Imagine Video 1.5 with nearly 2x faster generation, native audio, and a claimed #1 leaderboard position. The episode grouped it with Gemini Omni as part of the week’s video-generation frontier.
Google drops Gemma 4 12B, an encoder-free multimodal local model
Google released Gemma 4 12B, an encoder-free multimodal model under Apache 2.0 that targets 16GB VRAM local setups. Instead of bolting separate vision or audio encoders onto a language model, it uses one unified network, which LDJ and Yam argued makes smaller multimodal models cheaper, cleaner, and easier to run locally.
Gemini Omni: 'create anything from anything' conversational video editor
Google DeepMind launched Gemini Omni, a multimodal 'create anything from anything' model debuting as Google's first conversational video editor. Unlike pure text-to-video systems, Omni is an iterative multi-turn editing model that combines Gemini intelligence, world knowledge, multimodal inputs and generative media, in the same way Nano Banana brought Gemini to interactive image editing. It is available in the Gemini app, Google Flow and YouTube, with API support coming soon.
Meta launches Muse Spark voice conversations across its apps and glasses
Meta rolled out Muse Spark-powered voice conversations across the Meta AI app, WhatsApp, Instagram, Facebook, and Ray-Ban Meta glasses. The feature includes real-time image generation, live camera AI, and instant Reels/maps integration. Alex tested it live and called it surprisingly good, the first big consumer ship from Meta Superintelligence Labs.
Thinking Machines Lab drops Interaction Models: real-time multimodal 276B MoE
Mira Murati's Thinking Machines Lab released Interaction Models, a 276B-parameter MoE (12B active) trained from scratch for native real-time multimodal collaboration. It supports full-duplex audio/video/text with 0.40s turn-taking latency and scores 77.8 on FD-bench v1.5. The demo can react live to events like another person entering the camera frame.
NVIDIA Nemotron 3 Nano Omni: hybrid Transformer-Mamba MoE
NVIDIA released Nemotron 3 Nano Omni, a 30B-total/3B-active hybrid Transformer-Mamba MoE with 256K context. It delivers 9x throughput on consumer hardware.
SenseTime open-sources SenseNova U1 unified multimodal MoE
SenseTime open-sourced SenseNova U1, a unified multimodal MoE model with 8B total and 3B active parameters that handles understanding and generation with no separate encoder or VAE. The architecture builds on a paper the team presented at ICLR last year.
Qwen 3.6-35B-A3B: Apache 2.0 MoE with 3B active hits 73.4% SWE-Verified
Alibaba Qwen open-sourced Qwen 3.6-35B-A3B under Apache 2.0 the same morning Opus 4.7 dropped: a 35B MoE with only 3B active parameters that scores 73.4% on SWE-bench Verified, rivaling models 10x its size. It is natively multimodal with 262K context extensible to 1M, and the crew called it the strongest mid-size LLM on nearly all benchmarks, putting to rest doubts about Qwen's open-source commitment after Junyang Ling's departure.
Meta launches Muse Spark, first model from Meta Superintelligence Labs
Meta dropped Muse Spark mid-show, the debut model from Meta Superintelligence Labs. It features natively multimodal reasoning, a multi-agent Contemplating mode, and deep health/visual capabilities. Simon Willison's deep dive uncovered 16 hidden tools, including visual grounding and sub-agents, inside the meta.ai chat UI.
Alibaba open-sources Qwen3.5-Omni, a 397B native omni-modal model
Qwen3.5-Omni is Alibaba's natively omni-modal open model handling text, image, audio, and video, with 397B total parameters and 17B active. It extends the Qwen family's open-source momentum into unified multimodal workloads.
Google releases Gemma 4 open-weights family under Apache 2.0
Google DeepMind's Gemma 4 launch crossed 10M+ downloads with over 1,000 Gemma-4-based fine-tunes on Hugging Face; the Gemma family totals 500M+ downloads. Omar Sanseviero says Gemma is the foundation for the next generation of Gemini Nano shipping on Pixel and Samsung, with the AI Edge gallery letting people run it locally on Android and iOS. It punched above its size on Arena's Pareto curve and is now live on W&B Inference.
Luma Labs Uni-1 thinks and generates pixels simultaneously, #1 preference Elo
Luma Labs released Uni-1, an LLM-based image model that thinks and generates pixels simultaneously and claims the number-one human preference Elo. Unlike traditional diffusion workflows you converse with it and iterate together toward results, and it can also generate infographics; a surprising pivot from Luma's video focus.
Mistral Small 4: 119B MoE with 6B active unifies vision, coding, reasoning
Mistral returned to open source with Small 4, a 119B-parameter MoE with 128 experts and only 6B active per token, released under Apache 2.0. It unifies the previous Pixtral (vision), Devstral (coding), and Magistral (reasoning) lines into one model and can fit on a single H100 when compressed. Early WolfBench results are sobering at ~17% on OpenClaw agent tasks, roughly on par with similarly sized Nemotron.
Xiaomi MiMo revealed as the 1T-param stealth model topping OpenRouter
Xiaomi revealed MiMo, a 1-trillion-parameter family with omni-modal and language-only variants, unmasked as the stealth model that had been sitting at #1 on OpenRouter. The reveal surprised the panel, marking Xiaomi's entry into the frontier-model conversation.
Google launches Gemini Embedding 2, a natively multimodal embedder
Google launched Gemini Embedding 2, a natively multimodal embedding model that supports text, image, video, and audio in a single unified embedding space. It is available through the Gemini Embeddings API.
Google DeepMind launches Nano Banana 2 image model mid-show
Google DeepMind announced Nano Banana 2 during the show, a Flash-quality tier of its image model line. Alex broke in mid-TLDR to describe near-Pro image quality at roughly half the price, plus a new image search capability.
ByteDance Seed 2.0: frontier multimodal family at 73-84% lower pricing
ByteDance released Seed 2.0, a frontier multimodal LLM family with Pro, Lite, Mini, and Code variants that rivals GPT-5.2 and Claude Opus 4.5 at 73-84% lower pricing. Its video understanding surpasses the human benchmark at 77% vs 73%. At 84% cheaper than Opus 4.5 with near-comparable quality, the panel called it a compelling option for price-conscious developers.
ByteDance Seedance 2.0 shatters video generation reality
ByteDance launched Seedance 2.0, a unified multimodal video generation model that accepts up to 9 images, 3 videos, and 3 audio clips as references and produces 15-second multi-shot clips with native stereo audio and strong character consistency (a 45-second internal test mode also exists). The panel compared the quality jump to seeing Sora for the first time. Available on the BytePlus platform.
MiniCPM-o 4.5: first open-source full-duplex omni model
OpenBMB released MiniCPM-o 4.5, the first open-source full-duplex omni-modal LLM that can see, listen, and speak simultaneously. It can listen while speaking and even interrupt the user, bringing real-time conversational behavior to open weights.
Allen AI adds video-input multimodal OLMO models in 4B/7B/8B sizes
Allen AI extended its OLMO family with multimodal models that accept video input, released in 4B, 7B, and 8B sizes. It continues Allen AI's fully open approach to model development alongside the BOLMO byte-level work.
Gemini 3 Pro launches with record ARC-AGI-2 scores
Google's new frontier multimodal model with a 1M-token context window and huge reasoning gains, scoring 31.11% on ARC-AGI-2 (45.14% with Deep Think mode) — roughly double the previous SOTA — plus 81% on MMLU-Pro and major coding improvements. Amp switched to it as their default model on launch day, the first time they have ever switched defaults. Also rolling out across Gmail, Calendar, and AI Mode in Google Search.
Meituan releases LongCat Flash Omni, a 560B (27B active) omni model
Meituan's LongCat team released LongCat Flash Omni, a 560B-parameter mixture-of-experts model with roughly 27B active parameters that accepts text, audio, and video input. It extends the open LongCat Flash line into omni-modal territory from a lab better known for food delivery than frontier models.
Ming-flash-omni Preview: sparse MoE omni-modal open model
Ant Group's InclusionAI team released Ming-flash-omni Preview, a sparse mixture-of-experts omni-modal model on Hugging Face. It handles multiple input and output modalities in a single open-weights model, adding to the wave of Chinese open omni-modal releases.
Qwen3-VL adds compact 2B and 32B multimodal models
Alibaba's Qwen team extended the Qwen3-VL family with newly updated 2B and 32B checkpoints. The 2B is a generic VLM (OCR-capable) that holds up against its 4B and 8B siblings from prior weeks, while the 32B reportedly outperforms GPT-5 mini and Claude 4 Sonnet on benchmarks.
Qwen3-VL adds compact 3B and 8B open vision-language models
Alibaba's Qwen team released smaller Qwen3-VL vision-language models in 3B and 8B sizes, bringing the flagship VL capabilities down to edge- and laptop-friendly scales. Weights are open on Hugging Face as part of the Qwen3-VL collection.
Qwen3-Omni ships open-weights any-to-any audio, vision, and text
Alongside Qwen3-VL, Alibaba released Qwen3-Omni, an end-to-end omni-modal open-weights model that takes text, image, audio, and video input and can respond with streaming speech. The show treated it as direct evidence of how fast open multimodal systems are improving, with weights on Hugging Face, a GitHub repo, demos, and availability in Qwen Chat and the Model Studio API.
Alibaba releases Qwen3-VL open-weights vision-language flagship
Alibaba's Qwen team shipped Qwen3-VL, its new flagship open-weights vision-language family, headlining the episode's 'Qwen-mas' barrage. The panel discussed it as a practical workflow tool for visual understanding and agentic GUI tasks, not just another model card, with weights, a blog post, and a Hugging Face demo all available at launch.
Meta Connect: new AI glasses with a display and neural control interface
At Meta Connect, Meta unveiled new AI glasses featuring a built-in display, a neural wristband control interface, and a new AI mode. The panel treats the glasses as an interface milestone, arguing the product surface for AI is shifting from apps to display-equipped wearables.
Perceptron AI introduces Isaac 0.1, a 2B perceptive-language model
Perceptron AI released Isaac 0.1, a 2B parameter perceptive-language model with open weights on Hugging Face. Despite its small size, the show notes highlight that it 'points better than GPT', excelling at visual grounding and pointing tasks relative to much larger models.
Baidu open-sources ERNIE 4.5, a 10-model multimodal family
Baidu open-sourced the ERNIE 4.5 series, a family of 10 models ranging from 424B down to 0.3B parameters with multimodal capabilities, reportedly beating o1 on DocVQA. The release marks a sharp reversal from Baidu's previous anti-open-source posture and another sign that Chinese labs are setting the pace in open source.
ByteDance publishes Seed1.5-VL, a 20B vision-language thinking model
ByteDance's Seed team published the technical report for Seed1.5-VL, a 20B-parameter vision-language model with thinking capabilities. It was covered among the big-company releases of the week, with the tech report shared on GitHub.
Qwen 2.5 Omni gets an update
Alongside the Qwen 3 launch, Alibaba updated its Qwen 2.5 Omni multimodal model line. Mentioned briefly in the open-source roundup as part of the week's Qwen ecosystem push.
NVIDIA releases DAM-3B for region-based image and video captioning
NVIDIA dropped the Describe Anything Model (DAM-3B), a 3 billion parameter multimodal model for region-based image and video captioning. You can point it at a specific region of an image or video and it generates a detailed description of just that area. NVIDIA also published an accompanying DescribeAnything dataset and a Hugging Face demo.
OpenAI launches o3 and o4-mini, SOTA reasoning models with tool use
OpenAI shipped o3 and o4-mini in ChatGPT and the API, with o3 setting new SOTA records on Codeforces, SWE-bench, MMMU and more. For the first time the models can use tools (web search, Python, image generation) during the reasoning process, and they can think visually by cropping, zooming and rotating images. o3 scored $65k on the Freelancer eval versus o1's $28k, and o4-mini hits 99.5% on AIME with a Python interpreter.
Jina Reranker M0: SOTA multilingual, multimodal document reranker
Jina AI released Jina Reranker M0, a state-of-the-art multimodal and multilingual document reranker model. It reranks documents that include both text and images, targeting retrieval and RAG pipelines, with weights available on Hugging Face.
Meta drops Llama 4 Scout (109B) and Maverick (400B) open-weights MoE models
Meta released the long-awaited Llama 4 family in a chaotic Saturday drop: Scout (17B active / ~109B total, 16 experts) and Maverick (17B active / ~400B total, 128 experts), with a 2T-parameter Behemoth still in training. The models are multimodal, multilingual MoE architectures trained on ~30T tokens with FP8 and interleaved attention (iRoPE), claiming 10M context for Scout and 1M for Maverick. The release was marred by drama: the LMArena version differed from the released model, and the community criticized the lack of small local-friendly sizes.
Nomic Embed Multimodal: SOTA embeddings for visual documents
Nomic AI released Nomic Embed Multimodal, new 3B and 7B parameter embedding models built on Alibaba's Qwen2.5-VL. They achieve SOTA on visual document retrieval by embedding interleaved text-image sequences, ideal for PDFs and complex webpages. The 7B model ships under Apache 2.0 with open weights, code, and data; guest Zach Nussbaum discussed the release on the show.
Qwen launches Omni 7B: sees, hears, reads, and talks back
Qwen released Qwen2.5-Omni-7B, an open-weights omni-modal model that perceives text, images, audio, and video, and generates both text and speech. It packs end-to-end multimodal perception and spoken output into a 7B parameter model available on Hugging Face.
OpenAI enables native image generation in GPT-4o, internet goes Ghibli
OpenAI finally enabled GPT-4o's native auto-regressive image generation in ChatGPT, sparking the biggest mainstream AI buzz of the week as the internet ghiblified itself. Launched right after Gemini 2.5, it excels at instruction following, text rendering, and multi-turn editing, with viral demos ranging from ad mockups to a full Lord of the Rings trailer.
Mistral Small 3.1 24B: open-weights multimodal model
Mistral released Mistral Small 3.1, a 24B-parameter open-weights model that adds multimodal (vision) capabilities to the Small line. Both instruct and base checkpoints were published on Hugging Face, making it a strong local multimodal option at the 24B size class.
Google AI Studio adds native YouTube video understanding via link dropping
Google AI Studio now lets you drop a YouTube link and have Gemini natively understand the video. This unlocks video analysis, summarization, and support use cases without downloading or preprocessing the content.
Gemini Flash gains native image generation and conversational editing
Google enabled native image generation in Gemini Flash Experimental, letting users generate and iteratively edit images conversationally inside the same multimodal model. The crew demoed it live on stream, editing photos of themselves with natural-language instructions, and saw it as a preview of how creative tools like Photoshop will work.
Google open sources Gemma 3, 1B-27B multimodal family with 128K context
Google released Gemma 3, an open-weights model family spanning 1B to 27B parameters with multimodal (text, image, video) capabilities, support for over 140 languages, and a 128K context window. The 27B model runs on a single GPU, with Sundar Pichai claiming competitors need roughly 10x the compute for similar performance. It shipped with day-one open source ecosystem support (Hugging Face, Ollama, Kaggle) plus ShieldGemma 2 for content moderation.
Microsoft releases Phi-4-multimodal and Phi-4-mini open weights
Microsoft expanded the Phi family with Phi-4-multimodal-instruct, a small open-weights model that handles text, vision, and audio in a single model, alongside a compact Phi-4-mini. The weights shipped on Hugging Face, continuing Microsoft's push for capable small models that can run on-device.
Alibaba ships Qwen2.5-VL open vision-language model family
Alibaba's Qwen team released Qwen2.5-VL, open-weights vision-language models up to 72B that handle images, documents, video understanding, and on-screen agentic grounding. The 72B Instruct model was immediately available on Hugging Face and in Qwen Chat.
DeepSeek Janus Pro: open multimodal models in 1.5B and 7B
Amid the R1 frenzy, DeepSeek also released Janus Pro, unified multimodal models at 1.5B and 7B parameters that handle both image understanding and image generation. The open release added to DeepSeek's week of dominating AI news headlines.
NVIDIA releases Eagle 2 open vision-language models
NVIDIA published Eagle 2, a family of open vision-language models with an accompanying paper, model weights on Hugging Face, and a live demo. It is a fully transparent VLM release covering training data strategy and recipes, competitive with much larger vision models.
Follow Multimodal Models and everything else in AI — live every Thursday.