New Models
Qwen3.8-Max-0902
Qwen3.8-Max-0902: 2.4T API-only refresh claims #1 on Code Arena
Alibaba refreshed its API-only frontier model with a September 2 snapshot: 2.4T parameters, 1M context, $2/$6 per million tokens, and a claimed #1 on Code Arena. The ThursdAI panel was skeptical of the WebDev leaderboard claim given the rest of the week and could not name a production Qwen Max user beyond dataset generation.
2.4T parameters1M context window$2 / $6 per million input / output tokens
New Models
Claude Fable 5.1 & Mythos 5.1
Claude Fable 5.1 & Mythos 5.1: Terminal-Bench 4.0 jumps to 55.8, cache reads 75% cheaper
Anthropic shipped Fable 5.1 and Mythos 5.1 as the same weights with different guardrails. Fable 5.1 scores 55.8 on Terminal-Bench 4.0 (up from 42.0 for Fable 5, versus about 37 for GPT-5.6 Sol) and 52% on the new Terminal-Bench Science, more than double Fable 5. Cache reads drop 75% to $0.25/M while input/output stay at $10/$50 per million tokens. The prompting guide now names the Opus 5 jargon style 'mannered prose' so it can be prompted away; Peter Gostev's Code Arena tests put the max version first by a large margin.
55.8 Terminal-Bench 4.0 (Fable 5: 42.0)52% Terminal-Bench Science$0.25/M cache reads, down 75%
New Models
MiniMax H3 Max Turbo
fal H3 Max Turbo: ~97% of H3 Max quality in 1.4 seconds at $0.01 per second
fal's Turbo variant of MiniMax H3 Max keeps about 97% of H3 Max quality, generates a clip in 1.4 seconds, and costs one cent per second at 768p, about 50% cheaper. Alex's live test found it no longer renders copyrighted characters the way H3 Max did.
$0.01/sec price at 768p1.4s generation time~97% of H3 Max quality
New Models
Gemini 3.8 Flash
Gemini 3.8 Flash: third Flash in three weeks, HLE-Verified 54.9, 1M context
Google's third Flash iteration in as many weeks, billed as the reasoning and coding workhorse for Googlers who all build on Antigravity internally. It scores 54.9 on HLE-Verified, keeps a 1M-token context, is about 3x faster, and is priced at $0.75/$3.75 per million tokens until a doubling on January 1, 2027. The Wall Street Journal reported Google scrapped the 3.5 Pro checkpoints because Flash overtook them; Gemini 4 is still in post-training.
54.9 HLE-Verified$0.75 / $3.75 per million tokens until Jan 1, 20271M context window
New Models
Gemini 3.8 Flash Cyber
Gemini 3.8 Flash Cyber: Fairwind-only cybersecurity variant, CWE-Bench 47.2%
A dedicated cybersecurity variant of Gemini 3.8 Flash, available only to trusted defenders through Google's Fairwind program. The newsletter lists CWE-Bench at 47.2%; no benchmark numbers surfaced during the show, and Alex questioned who will actually get to use it.
47.2% CWE-Bench
New Models
Realtime TTS-2
Inworld Realtime TTS-2 goes GA: sub-100ms, $25/M characters, #1 on AA Controlled Voice Arena
Inworld's real-time text-to-speech model is generally available with sub-100ms latency at $25 per million characters. It ranks #1 on Artificial Analysis's Controlled Voice Arena and #4 on the Provider arena.
<100ms latency$25/M per million characters#1 AA Controlled Voice Arena
New Models
Muse Spark 1.3
Meta Muse Spark 1.3 ties GPT-5.6 Sol and Grok 4.6 on the AA index, max mode ties Fable 5
Spark 1.3 xhigh scores 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and Grok 4.6, and a limited-preview max reasoning mode scores 62, tying Claude Fable 5, the first time Meta has jumped over OpenAI, Microsoft MAI, and Google on that index. Pricing is unchanged at $1.25/$4.25 per million; it runs about 3x faster than Fable and Sol at roughly 4x lower cost than Sol. MRCR long-context jumped from 66% to 98.5% at 1M tokens, and a contributor tier charges $0.10/$0.20 per million if Meta can train on your prompts. Open weights and a mystery 'Watermelon' model are teased as coming soon.
61 / 62 AA Intelligence Index, xhigh / max98.5% MRCR long context at 1M tokens$1.25 / $4.25 per million input / output tokens
New Models
Muse Voice Transcribe
Meta Muse Voice Transcribe: streaming ASR with diarization and endpointing in one model
Meta's streaming speech-to-text model handles transcription, speaker diarization, and endpointing in a single model, claims 3.1% streaming WER, and is API only. It costs about 18 cents per hour, supports a custom dictionary (it transcribes 'ThursdAI' correctly), and powers the live diarized transcript on thursdai.news/live, where Alex says it beats Descript on names and terms.
3.1% streaming WER (claimed)$0.18/hr price per hour of audio
New Models
MAI-Transcribe-2
Microsoft MAI-Transcribe-2: 2.0% WER, 400x real time, $1.67 per 1,000 minutes
Microsoft AI's second transcription model ranks #2 on the Artificial Analysis WER leaderboard at 2.0%, runs about 400x real time, and costs $1.67 per 1,000 minutes, less than half the price of its peers. It ships real-time ASR with diarization the same week as Meta's Muse Voice Transcribe.
2.0% WER, #2 on Artificial Analysis400x real time$1.67 per 1,000 minutes
New Models
GPT-6 Astra
OpenAI launches GPT-6 Astra, calls it its most intelligent and aligned model
OpenAI released GPT-6 Astra on September 3, 2026, live during the ThursdAI stream, with president Greg Brockman telling reporters 'welcome to the AGI era.' OpenAI reports 99.9 on ARC-AGI-3, 97.6% on FrontierMath Tier 4, 72.6% on OSWorld 2.0, 92.7 on ScreenSpot Pro and 64.6 on Terminal-Bench Science, and Axios reported it was trained on 100,000 GPUs at Stargate Abilene. API pricing is $10/$50 per million tokens (same as Fable 5.1, 2.5x Sol's promotional price) with a Fast mode at 2x the price, available in the OpenAI API and Amazon Bedrock, rolling out from a limited set of organizations to Plus, Pro, Business and Enterprise. Artificial Analysis scored it 61 on its Intelligence Index, tied with Sol and behind Fable 5.1, but about 70% more token efficient than Sol.
99.9 ARC-AGI-3 (Sol was 7%)97.6% FrontierMath Tier 472.6% OSWorld 2.0
New Models
GWM Worlds 2
Runway GWM Worlds 2: real-time 720p, 24 fps world model with 48 kHz audio
Runway's second generative world model runs in real time at 720p and 24 fps with open-ended sessions, 48 kHz audio, and generated speech, a big jump over the low-fidelity audio in most world models. LDJ broke it live on the show two days after Runway's previous world model. Research preview.
24 fps @ 720p real-time generation48 kHz audio
New Models
Solaris
Runway Solaris: an Interface World Model that generates clickable UIs frame by frame
Solaris generates interfaces as video, so every element in a scene is clickable and draggable: click a lamp and the room lights up, drag shoes onto a person and he wears them. In Runway's own study the generated behavior was preferred 61 to 24 over Opus 5 (71% preferred it over coded UI). Research preview, no public access yet.
61 to 24 preferred over Opus 5 in Runway's study
New ModelsOpen weights
Hy4 preview
Tencent Hy4 preview: 770B/49B Apache 2.0 MoE, Sherry quant shrinks 1.5 TB to 214 GB
Tencent open-sourced a 770B-parameter MoE with 49B active, 1M context, and an Apache 2.0 license. Its Sherry quantization takes the weights from 1.5 TB to 214 GB at 2.38 bits per weight. Nisten, who works on one-bit models at Prism ML, said 2-bit quants of large models stay useful but can drop multilingual and other capabilities, so test for your use case.
770B / 49B total / active parameters214 GB Sherry quant at 2.38 bpw (from 1.5 TB)1M context window
New Models
Atlas
World Labs Atlas: omnimodel turns 1 to 6 images into walkable 3D and bullet-time video
Atlas is pretrained from scratch to natively operate in text, images, video, and 3D as an autoregressive diffusion transformer. From 1 to 6 images it generates up to a minute of 1440p camera-controlled video and a 3D reconstruction with no Gaussian splats, and it can reframe real video from new angles, producing bullet-time from three ordinary tripods. LDJ noted quality scales with more camera angles. Partner access only for now.
1 to 6 input images1 min @ 1440p max generation3 phones for bullet time