New ModelsOpen weights
Inkling-Small
Thinking Machines releases Inkling-Small: 276B/12B open MoE that beats its 975B sibling on agentic coding
The efficient sibling previewed alongside Inkling ships as open weights: 276B total with 12B active, natively multimodal with an encoder-free architecture (images via hierarchical patch encoding, audio via dMel spectrograms straight into the decoder), and variable thinking effort. On-policy distillation from Inkling plus two extra weeks of agentic-coding RL let it beat the 975B teacher on SWE-Bench Verified (80.2% vs 77.6%) and ARC-AGI-2 (40.1% vs 36.5%), though factual recall regressed hard (SimpleQA 20.6% vs 43.9%). Priced at $0.30/$1.20 per million tokens, roughly 3-4x cheaper than Inkling, with day-zero SGLang, Unsloth GGUF, and Baseten support. Dropped just after the July 30 show aired.
276B / 12B total / active parameters80.2% SWE-Bench Verified, beating the 975B Inkling's 77.6%$0.30 / $1.20 per 1M tokens in/out
New Models
Lyria 3.5
Lyria 3.5 generates full 3-minute songs with BPM and key control inside Flow Music
Google's flagship music model now produces cohesive three-minute songs with tempo and key-signature control in the prompt, more expressive multilingual vocals, style-transfer covers that preserve a track's structure, and lip-synced music videos via Gemini Omni Flash — plus a new Flow Music iOS app. All output is SynthID-watermarked. Google published no benchmarks against Suno or Udio, and early testers say paid Suno 5.5 still edges it, but as a free end-to-end create-to-publish stack it's a real move.
3 min full cohesive songs, up from short clips
New Models
Qwen3.8-Max Preview
Alibaba previews Qwen3.8-Max at a claimed 2.4T parameters, 'second only to Fable 5'
Alibaba's Qwen team previewed Qwen3.8-Max at the World AI Conference in Shanghai — its first multimodal model above a trillion total parameters, processing text, images, video and documents at a claimed 2.4T. Alibaba shares rose as much as 5.4% in Hong Kong on the news. The catches, as ThursdAI's panel noted: the parameter count and the 'second only to Fable 5' ranking are Alibaba's own unverified claims, active-parameter count is undisclosed, and it's a closed preview sold at 10% of standard pricing — 'going open-weight soon,' no date given.
2.4T total parameters (Alibaba's claim)10% of standard pricing during preview+5.4% Alibaba HK share move on announcement day
New ModelsOpen weights
MOSS-VL-Realtime
OpenMOSS open-sources MOSS-VL-Realtime, an 11B VLM that decides when to speak — and when to stay silent
OpenMOSS released MOSS-VL-Realtime, an open-source 11B-parameter vision-language model (~22.7GB) built for real-time streaming video that proactively speaks up or deliberately stays silent instead of only answering prompts. ThursdAI reported it as state-of-the-art on all three open proactivity benchmarks, with the base model included in the release.
11B parameters22.7GB model download size
New ModelsOpen weights
Inkling
Thinking Machines releases Inkling, a 975B open-weights MoE trained on 45T multimodal tokens
Mira Murati's Thinking Machines shipped Inkling, a 975B-total/41B-active Mixture-of-Experts transformer pretrained from scratch on 45 trillion tokens of text, images, audio and video, released under Apache 2.0. The ThursdAI panel called it the top US open-weights model right now — 41 on the Artificial Analysis Index — with encoder-free native reasoning over text, image and audio and a 1M-token context window. A leaner Inkling-Small (276B/12B active) was previewed alongside, and both run on the Tinker platform at a limited-time 50% discount.
975B / 41B total / active parameters45T multimodal training tokens41 Artificial Analysis Index — top US open-weights model
New ModelsOpen weights
Bonsai 27B
PrismML compresses a full 27B model to 3.9GB so it runs on a phone
PrismML released Bonsai 27B, extreme quantizations of Qwen 3.6 27B under Apache 2.0: a 1-bit build at 3.9GB keeping ~90% of full-precision quality — small enough for an iPhone 17 Pro's memory budget — and a ternary build at 5.9GB keeping ~95%. Both stay multimodal with the full 262K-token context window. Nisten demoed it live on the show running on a phone and on a 6GB GTX 1660 Ti.
3.9GB / 90% 1-bit variant size / quality retained5.9GB / 95% ternary variant size / quality retained262K context window (tokens)
New Models
OmniFlash
Google DeepMind debuts OmniFlash, first of the any-to-any Omni family
OmniFlash — first of Google's any-to-any Omni family — generates videos up to 10 seconds with precise conversational multi-turn editing via the Interactions API: say 'make it daytime' and it redoes light, sky and shadows. Editing Elo 1087 at $0.10 per second of output.
1087 editing Elo$0.10 per second of video, up to 10s