New Models
Gemini 3.8 Live & Live Extended Thinking
Gemini 3.8 Live claims #1 on the speech-to-speech index with 97 languages and async tool calls
Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, real-time speech-to-speech models scoring 82.6 to take #1 on the speech-to-speech quality index across 97 languages with asynchronous tool calls. Extended Thinking brings longer reasoning into a live model without breaking real-time interaction, and the release reaches everyone on Android and Google Search rather than just API users.
82.6 #1 on the speech-to-speech quality index97 languages
New Models
Union Alpha
Union Alpha: anonymous stealth model free on OpenRouter with 262K context
A new anonymous stealth model called Union Alpha appeared for free on OpenRouter with a 262K context window, processing over 100B tokens within hours of listing. The community's leading guess for the lab behind it is Z.ai, but the provider remains unconfirmed.
262K context window100B+ tokens processed within hours
New Models
StepAudio 3
StepFun's StepAudio 3 family takes #1 on Artificial Analysis real-time voice
StepFun launched StepAudio 3, a five-model audio family spanning Real-Time Preview, ASR Max, and TTS — a full suite for building end-to-end voice assistants in code. Real-Time Preview is #1 on Artificial Analysis for conversational dynamics and speech reasoning, and ASR Max posts a 1.7% word error rate, significantly below Whisper. API only, no open weights.
#1 Artificial Analysis real-time voice1.7% ASR Max word error rate5 models in the family
New Models
Jev
TypeSafe AI launches Jev, the first public non-LLM System 1 decision model
TypeSafe AI, the stealth lab led by RLHF and ChatGPT co-creator Diogo Almeida, launched Jev — the first public System 1 model, a new class of AI that returns calibrated probabilities instead of generating text. Developers define questions with three primitives (choice, score, null) in natural language and get typed, machine-readable probability outputs back in 70-500ms with a 32K context window, trained with what TypeSafe calls RLCD (reinforcement learning for calibrated decisions). Pricing is $42 per billion input tokens with free output tokens, and TypeSafe cites 133x faster and 444x cheaper than competitive-level LLMs on tested decision tasks. Access is via waitlist.
$42 / 1B input token pricing; output tokens free70-500ms decision latency133x / 444x faster / cheaper vs competitive-level models (TypeSafe)
New ModelsOpen weights
Ling-3.0-flash-VL
InclusionAI open-sources Ling-3.0-flash-VL: 124B vision-language MoE, 5.5B active, MIT
InclusionAI (Ant Group) open-sourced Ling-3.0-flash-VL, a 124B-parameter sparse MoE vision-language model that activates 5.5B parameters per token, adding a ViT encoder and VideoRoPE to the Ling-3.0-flash backbone for image, video and GUI-agent tasks with up to 1M tokens of context. Weights ship in BF16 and FP8 on Hugging Face under MIT. It got a brief mention on the show as a solid workhorse VL release.
124B total parameters5.5B active parameters per token
New Models
SWE-2
Cognition SWE-2: near-frontier coding at up to 70% lower cost, free for a month on Devin
Cognition shipped SWE-2, its closest model yet to the frontier, as breaking news during the show. On the company's charts it scores 50% on Frontier Code 1.1 (Claude Fable 5.1: 50.9, GPT-6 Astra: 53), 73% on DeepSWE (Astra: 74) and 92.8 on Terminal-Bench 2.1, while matching Fable at 64% lower cost and headline pricing up to 70% below frontier models. All numbers are company-reported. It ships in Devin and the Devin CLI, free for a month on paid tiers.
50% Frontier Code 1.1 (Fable 5.1: 50.9, Astra: 53)92.8 Terminal-Bench 2.1, top measured score70% lower cost than frontier, per Cognition
New ModelsOpen weights
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash: 552B MoE, 8B/16B active, KV cache 400x smaller than V1, MIT
DeepSeek's V4.1 Flash is a 552B-parameter multimodal MoE that activates only 8B parameters on prefill and 16B on decode, with a 1M-token context window, trained from scratch on 45 trillion multimodal tokens and released under MIT. It returns to an encoder-decoder architecture at half-trillion scale and cuts the KV cache to about 890 bytes per token, down from 389,000 bytes in the first DeepSeek release (over 400x smaller), which is why it is so cheap to serve. DeepSeek's own evals put it at 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE 1.1, above Opus 5 and GPT-5.6 Sol, and second behind GPT-6 Astra on an Open Design leaderboard at about two cents per task. TokenJuice.ai serves it free in the US for a limited time in exchange for training data.
552B total parameters (8B active prefill / 16B decode)890 bytes KV cache per token, over 400x below DeepSeek V190.6 Terminal-Bench 2.1 (DeepSeek evals)
New Models
On-device model suite (Voz, Redact, Clips, Clear)
Desert Ant Labs debuts 18 on-device models with native SDKs, led by Voz speech-to-text
Desert Ant Labs, a European lab spun out of the Apple Design Award-winning video app Detail, launched 18 small on-device models for audio, vision and text with Swift, Kotlin and JavaScript SDKs. Voz transcribes 10 minutes of speech in about 2 seconds on an iPhone from a 467 MB model (Whisper large-v3-turbo is 1.6 GB); the suite also includes Redact for PII detection, Clips for highlight picking and Clear for speech enhancement. Weights are on Hugging Face under a source-available license and are free up to 100k monthly active devices per SDK.
18 on-device models~2s to transcribe 10 minutes of speech on an iPhone (Voz)467 MB Voz model size vs 1.6 GB Whisper large-v3-turbo
New ModelsOpen weights
YuE 2
YuE 2 open music model lands on Hugging Face
YuE 2, the successor to the open-source YuE music model, can generate vocals and accompaniment and is already available on Hugging Face.
New Models
ChatGPT Images 2.5 (Flare & Sunburst)
OpenAI launches ChatGPT Images 2.5 with Flare and Sunburst API models
OpenAI released ChatGPT Images 2.5 with two new API models: GPT-Image-2.5 Flare for speed and volume and GPT-Image-2.5 Sunburst for precise editing with native transparent backgrounds, both at $30 per million image output tokens and up to 50% lower latency than Images 2.0. ChatGPT gains Sketch (type @Sketch and draw), comment-on-image editing and templates. Alex made this week's thumbnails with Sunburst via Fal; after fixing prompts written for GPT-image-2 and dropping the AI-generated reference photo, the second round beat Nano Banana Pro on likeness and tile text.
$30/M image output tokens, both models50% lower latency than Images 2.0 (OpenAI claim)
New Models
Suno 6
Suno launches Suno 6
Suno released Suno 6, its next music generation model, the same week Google shipped Lyria 3.5. The show noted it in passing without a deep dive.
New ModelsOpen weights
AUK
Tencent AUK: open 1.5B speech model for TTS, cloning and edits from one prompt
Tencent released AUK, an open-source 1.5B-parameter speech model with MIT weights. It handles text-to-speech, voice cloning, content/emotion/accent edits, cleanup and separation from a single natural-language prompt.
1.5B parameters, MIT weights
New Models
Qwen3.8-Max-0902
Qwen3.8-Max-0902: 2.4T API-only refresh claims #1 on Code Arena
Alibaba refreshed its API-only frontier model with a September 2 snapshot: 2.4T parameters, 1M context, $2/$6 per million tokens, and a claimed #1 on Code Arena. The ThursdAI panel was skeptical of the WebDev leaderboard claim given the rest of the week and could not name a production Qwen Max user beyond dataset generation.
2.4T parameters1M context window$2 / $6 per million input / output tokens
New Models
Claude Fable 5.1 & Mythos 5.1
Claude Fable 5.1 & Mythos 5.1: Terminal-Bench 4.0 jumps to 55.8, cache reads 75% cheaper
Anthropic shipped Fable 5.1 and Mythos 5.1 as the same weights with different guardrails. Fable 5.1 scores 55.8 on Terminal-Bench 4.0 (up from 42.0 for Fable 5, versus about 37 for GPT-5.6 Sol) and 52% on the new Terminal-Bench Science, more than double Fable 5. Cache reads drop 75% to $0.25/M while input/output stay at $10/$50 per million tokens. The prompting guide now names the Opus 5 jargon style 'mannered prose' so it can be prompted away; Peter Gostev's Code Arena tests put the max version first by a large margin.
55.8 Terminal-Bench 4.0 (Fable 5: 42.0)52% Terminal-Bench Science$0.25/M cache reads, down 75%
New Models
MiniMax H3 Max Turbo
fal H3 Max Turbo: ~97% of H3 Max quality in 1.4 seconds at $0.01 per second
fal's Turbo variant of MiniMax H3 Max keeps about 97% of H3 Max quality, generates a clip in 1.4 seconds, and costs one cent per second at 768p, about 50% cheaper. Alex's live test found it no longer renders copyrighted characters the way H3 Max did.
$0.01/sec price at 768p1.4s generation time~97% of H3 Max quality
New Models
Gemini 3.8 Flash
Gemini 3.8 Flash: third Flash in three weeks, HLE-Verified 54.9, 1M context
Google's third Flash iteration in as many weeks, billed as the reasoning and coding workhorse for Googlers who all build on Antigravity internally. It scores 54.9 on HLE-Verified, keeps a 1M-token context, is about 3x faster, and is priced at $0.75/$3.75 per million tokens until a doubling on January 1, 2027. The Wall Street Journal reported Google scrapped the 3.5 Pro checkpoints because Flash overtook them; Gemini 4 is still in post-training.
54.9 HLE-Verified$0.75 / $3.75 per million tokens until Jan 1, 20271M context window
New Models
Gemini 3.8 Flash Cyber
Gemini 3.8 Flash Cyber: Fairwind-only cybersecurity variant, CWE-Bench 47.2%
A dedicated cybersecurity variant of Gemini 3.8 Flash, available only to trusted defenders through Google's Fairwind program. The newsletter lists CWE-Bench at 47.2%; no benchmark numbers surfaced during the show, and Alex questioned who will actually get to use it.
47.2% CWE-Bench
New Models
Realtime TTS-2
Inworld Realtime TTS-2 goes GA: sub-100ms, $25/M characters, #1 on AA Controlled Voice Arena
Inworld's real-time text-to-speech model is generally available with sub-100ms latency at $25 per million characters. It ranks #1 on Artificial Analysis's Controlled Voice Arena and #4 on the Provider arena.
<100ms latency$25/M per million characters#1 AA Controlled Voice Arena
New Models
Muse Spark 1.3
Meta Muse Spark 1.3 ties GPT-5.6 Sol and Grok 4.6 on the AA index, max mode ties Fable 5
Spark 1.3 xhigh scores 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and Grok 4.6, and a limited-preview max reasoning mode scores 62, tying Claude Fable 5, the first time Meta has jumped over OpenAI, Microsoft MAI, and Google on that index. Pricing is unchanged at $1.25/$4.25 per million; it runs about 3x faster than Fable and Sol at roughly 4x lower cost than Sol. MRCR long-context jumped from 66% to 98.5% at 1M tokens, and a contributor tier charges $0.10/$0.20 per million if Meta can train on your prompts. Open weights and a mystery 'Watermelon' model are teased as coming soon.
61 / 62 AA Intelligence Index, xhigh / max98.5% MRCR long context at 1M tokens$1.25 / $4.25 per million input / output tokens
New Models
Muse Voice Transcribe
Meta Muse Voice Transcribe: streaming ASR with diarization and endpointing in one model
Meta's streaming speech-to-text model handles transcription, speaker diarization, and endpointing in a single model, claims 3.1% streaming WER, and is API only. It costs about 18 cents per hour, supports a custom dictionary (it transcribes 'ThursdAI' correctly), and powers the live diarized transcript on thursdai.news/live, where Alex says it beats Descript on names and terms.
3.1% streaming WER (claimed)$0.18/hr price per hour of audio
New Models
MAI-Transcribe-2
Microsoft MAI-Transcribe-2: 2.0% WER, 400x real time, $1.67 per 1,000 minutes
Microsoft AI's second transcription model ranks #2 on the Artificial Analysis WER leaderboard at 2.0%, runs about 400x real time, and costs $1.67 per 1,000 minutes, less than half the price of its peers. It ships real-time ASR with diarization the same week as Meta's Muse Voice Transcribe.
2.0% WER, #2 on Artificial Analysis400x real time$1.67 per 1,000 minutes
New Models
GPT-6 Astra
OpenAI launches GPT-6 Astra, calls it its most intelligent and aligned model
OpenAI released GPT-6 Astra on September 3, 2026, live during the ThursdAI stream, with president Greg Brockman telling reporters 'welcome to the AGI era.' OpenAI reports 99.9 on ARC-AGI-3, 97.6% on FrontierMath Tier 4, 72.6% on OSWorld 2.0, 92.7 on ScreenSpot Pro and 64.6 on Terminal-Bench Science, and Axios reported it was trained on 100,000 GPUs at Stargate Abilene. API pricing is $10/$50 per million tokens (same as Fable 5.1, 2.5x Sol's promotional price) with a Fast mode at 2x the price, available in the OpenAI API and Amazon Bedrock, rolling out from a limited set of organizations to Plus, Pro, Business and Enterprise. Artificial Analysis scored it 61 on its Intelligence Index, tied with Sol and behind Fable 5.1, but about 70% more token efficient than Sol.
99.9 ARC-AGI-3 (Sol was 7%)97.6% FrontierMath Tier 472.6% OSWorld 2.0
New Models
GWM Worlds 2
Runway GWM Worlds 2: real-time 720p, 24 fps world model with 48 kHz audio
Runway's second generative world model runs in real time at 720p and 24 fps with open-ended sessions, 48 kHz audio, and generated speech, a big jump over the low-fidelity audio in most world models. LDJ broke it live on the show two days after Runway's previous world model. Research preview.
24 fps @ 720p real-time generation48 kHz audio
New Models
Solaris
Runway Solaris: an Interface World Model that generates clickable UIs frame by frame
Solaris generates interfaces as video, so every element in a scene is clickable and draggable: click a lamp and the room lights up, drag shoes onto a person and he wears them. In Runway's own study the generated behavior was preferred 61 to 24 over Opus 5 (71% preferred it over coded UI). Research preview, no public access yet.
61 to 24 preferred over Opus 5 in Runway's study
New ModelsOpen weights
Hy4 preview
Tencent Hy4 preview: 770B/49B Apache 2.0 MoE, Sherry quant shrinks 1.5 TB to 214 GB
Tencent open-sourced a 770B-parameter MoE with 49B active, 1M context, and an Apache 2.0 license. Its Sherry quantization takes the weights from 1.5 TB to 214 GB at 2.38 bits per weight. Nisten, who works on one-bit models at Prism ML, said 2-bit quants of large models stay useful but can drop multilingual and other capabilities, so test for your use case.
770B / 49B total / active parameters214 GB Sherry quant at 2.38 bpw (from 1.5 TB)1M context window
New Models
Atlas
World Labs Atlas: omnimodel turns 1 to 6 images into walkable 3D and bullet-time video
Atlas is pretrained from scratch to natively operate in text, images, video, and 3D as an autoregressive diffusion transformer. From 1 to 6 images it generates up to a minute of 1440p camera-controlled video and a 3D reconstruction with no Gaussian splats, and it can reframe real video from new angles, producing bullet-time from three ordinary tripods. LDJ noted quality scales with more camera angles. Partner access only for now.
1 to 6 input images1 min @ 1440p max generation3 phones for bullet time