Hosts & Guests
By The Numbers
🔥 Breaking During The Show
👋 Opening
Alex records a cold open from the editing floor: this is a special two-part podcast because GPT-6 Astra dropped about two and a half hours into a five-hour livestream. Part 1 is the traditional ThursdAI show (Fable 5.1, Muse Spark 1.3, the Abliteration AI interview, World Labs Atlas); Part 2 is Astra with Peter Gostev and Ryan Carson.
- Two-part episode: the regular show here, GPT-6 Astra in its own episode
- Five hours of continuous livestream, edited down with Fable 5.1's help
🎙️ Show Start
Live from CoreWeave, Alex welcomes Wolfram and LDJ to what he calls possibly the most insane week of frontier-lab updates the show has seen. Summer is officially over, and the GPT-6 rumor (same jump as 3 to 4, already shown to government officials) hangs over everything.
- Summer break is over: every frontier lab shipped something this week
- Rumor mill: GPT-6 is the same jump as GPT-3 to GPT-4
🔥 Banter — Highlights of the Week
Everyone picks a highlight before Astra: Wolfram calls fal's open MiniMax H3 the stable-diffusion moment for video (fal got banned from Twitch and Kick for an infinite Rick and Morty stream, then built fal.live over a weekend). Yam and Alex pick Fable 5.1, and Alex shows the thursdai.news/live page Fable built in one sitting, complete with live Muse Voice Transcribe diarization and an agentic producer throwing chyrons.
- fal's infinite Rick and Morty stream: banned from Twitch, banned from Kick, so they built fal.live
- Fable 5.1 one-shot a live-streaming platform for thursdai.news/live in one sitting
- Anthropic gave 'jargon douche' an official name: mannered prose
- Peter: Meta's Spark 1 → 1.1 → 1.2 → 1.3 shipping speed is the real story
📰 TL;DR
The rundown of everything deemed cover-worthy: Fable 5.1 and Mythos 5.1 (75% cheaper cache reads, a Terminal-Bench jump), OpenAI designating Astra the first model at its critical cyber threshold, Gemini 3.8 Flash and Flash Cyber, Qwen3.8-Max-0902, Muse Spark 1.3 tying GPT-5.6 Sol and Grok 4.6 on Artificial Analysis, Tencent Hy4, GLM-5.3 open weights, the abliterated GLM-5.3, Inworld TTS-2, MAI-Transcribe-2, World Labs Atlas, Runway Solaris, and fal's H3 Max week.
- Frontier: Fable 5.1 / Mythos 5.1, Gemini 3.8 Flash + Cyber, Qwen3.8-Max-0902, Muse Spark 1.3; Grok 4.7 promised next week
- Open source: Tencent Hy4 preview (770B MoE), GLM-5.3 full weights, Abliteration AI's uncensored GLM-5.3
- Voice: Inworld real-time TTS GA, Microsoft MAI-Transcribe-2 at 2% WER and $1.67 per 1,000 minutes
- Vision & video: World Labs Atlas, Runway Solaris, fal H3 Max Turbo at one cent per second
🏢 Frontier AI — Big CO LLMs + APIs
The Frontier AI corner opens. With every big lab shipping a point-one release ahead of OpenAI, Alex starts with the one the whole panel agrees is the biggest: Fable 5.1 and Mythos 5.1.
- Every frontier lab shipped a .1 this week
🏢 Anthropic Claude Fable 5.1 + Mythos 5.1
Fable 5.1 scores 55.8 on Terminal-Bench 4 (up from 42 for Fable 5, versus 37 for GPT-5.6 Sol) and 52% on the new Terminal-Bench Science, more than double Fable 5. Cache reads drop 75% to $0.25/M while input/output stay at $10/$50. Peter tested the max version on Code Arena where it came in first by a large margin, though single generations cost him $40 to $60. Anthropic also officially named the Opus 5 'jargon douche' problem: mannered prose, which you can now prompt away. Nisten demos a two-prompt Mars-launch simulator that built the entire solar system, a NASA mission planner, and a parachute landing on Earth.
- Terminal-Bench 4: 55.8 vs 42.0 for Fable 5 and 37 for GPT-5.6 Sol
- Terminal-Bench Science: 52%, more than double Fable 5
- Cache reads 75% cheaper; $10/$50 per million tokens unchanged
- Peter: #1 on Code Arena by a large margin, but $40 to $60 per generation
- 'Mannered prose' is the official name for load-bearing / control-plane speak, and you can prompt it away
- Nisten's two-prompt Mars maglev launcher sim grew into a full solar system with a parachute landing
🔥 Runway Solaris + GWM Worlds 2 (Breaking News)
LDJ breaks Runway's GWM Worlds 2: a real-time 24 fps world model with 48 kHz audio, including generated speech, two days after Runway's previous world model. Alex then shows Solaris, an Interface World Model where UIs are generated frame by frame: drag shoes onto a person, click a lamp and the room lights up. Runway's own study found 71% preferred the generated behavior over hand-coded UI, and Wolfram calls it the future of interfaces.
- GWM Worlds 2: real-time 24 fps world model with 48 kHz audio and speech
- Solaris: world model for interfaces, every pixel clickable and draggable
- 71% preferred the generated behavior over coded UI in Runway's study
- Alex to Runway: let people actually play with it (CoreWeave has GPUs)
🏢 Google Gemini 3.8 Flash
Google's third Flash iteration in as many weeks: 1M context, 3x faster, billed as the reasoning and coding workhorse for Googlers, who all build on Antigravity internally. Wolfram uses every Flash update in his home assistant but wants a Pro model; the Wall Street Journal reported Google scrapped the 3.5 Pro checkpoints because Flash overtook them, and Gemini 4 is still in post-training.
- Third Flash release in three weeks, 1M context, 3x faster
- WSJ: 3.5 Pro checkpoints scrapped, Gemini 4 still in post-training
- Wolfram: a million tokens matters less when the model burns context thinking
🛡️ Gemini 3.8 Flash Cyber
A dedicated cybersecurity variant of 3.8 Flash, gated behind Google's Fairwind program for trusted defenders. No benchmark numbers surfaced on air, though the newsletter lists CWE-Bench at 47.2%.
- Fairwind-only cyber variant of Gemini 3.8 Flash
🏢 Alibaba Qwen 3.8 Max-0902
Alibaba refreshed its API-only frontier model and claimed #1 on a WebDev leaderboard, a claim Alex is skeptical of given the rest of the week. Nobody on the panel knows anyone using Qwen Max in production, beyond one IT worker running OpenClaw on it and people generating datasets.
- API-only refresh, #1 Code Arena claim
- Panel poll: nobody knows a production Qwen Max user
🏢 Meta Muse Spark 1.3
For the first time, Meta jumps over OpenAI, Microsoft MAI, and Google on the Artificial Analysis index: Spark 1.3 xhigh ties GPT-5.6 Sol and Grok 4.6, and an unreleased max reasoning mode ties Fable 5. LDJ notes it runs about 3x faster than Fable and Sol at about 4x lower cost than Sol. MRCR long-context jumped from 66% to 98.5% at 1M tokens, and the contributor tier is $0.10/$0.20 per million if you let Meta train on your prompts. Zuck and Alex Wang are teasing open weights and a mystery 'Watermelon' model.
- Spark 1.3 xhigh ties GPT-5.6 Sol and Grok 4.6; max mode ties Fable 5 on AA
- About 3x faster than Fable and Sol, about 4x cheaper than Sol
- MRCR long-context 66% → 98.5% at 1M tokens
- Contributor tier: $0.10 in / $0.20 out per million if Meta can train on your data
- Open weights and a 'Watermelon' model teased
🔊 Meta Muse Voice Transcribe
Meta's streaming ASR model does transcription, speaker diarization, and endpointing in one model. Alex shows it powering the live transcript on thursdai.news/live: it labels Yam, Alex, and Wolfram in real time, gets names like Alex Wang and terms like ThursdAI right with a custom dictionary, and costs 18 cents per hour. Its diarization only fails about 17% of the time when someone talks over Alex, which he calls better than Descript.
- Streaming ASR + diarization + endpointing in one model
- 18 cents per hour: a three-hour show for under half a dollar
- Custom dictionary support: transcribes 'ThursdAI' correctly
- Real-time speaker labels power the new thursdai.news/live page
⚡ This Week's Buzz — CoreWeave, Fully Connected, CoreWeave Hacks
Fully Connected 26 runs September 29 to October 1 at Moscone South in San Francisco with 32 sessions, Fei-Fei Li keynoting, battle bots, and a ThursdAI live show from the floor; a code shared on the livestream gets you a 100% free ticket (regular price $1,299). CoreWeave Hacks: Agent Loops is September 12 to 13 in SF with W&B, AGI House, and TypeSafe AI, over $20,000 in prizes, a robot dog, and an F1 ticket for the most production-ready hack. Kimi K3 now runs on CoreWeave Dedicated Inference on GB300s.
- Fully Connected 26: Sept 29 to Oct 1, Moscone South, Fei-Fei Li keynote, free ticket code on the live show
- CoreWeave Hacks Agent Loops: Sept 12 to 13, $20k+ prizes, robot dog, F1 tickets
- Kimi K3 on CoreWeave Dedicated Inference on GB300 NVL72
🔓 Abliteration AI — Obliterated GLM 5.3 Interview
Abliteration AI's founder joins anonymously to explain abliterated-model-large-v2, a GLM-5.3 with refusals removed, hosted as an API at $5 per million tokens ($0.50 cached) on US servers. He stresses they did not invent abliteration (the refusal-direction technique is public and dozens of abliterated models already sit on Hugging Face); the paying market so far is red-teaming enterprise agent fleets, cybersecurity, and trust-and-safety tooling, with only self-harm and CSAM still blocked. Nisten raises social-worker and Metasploit-style use cases, Alex pushes back on vetting, and the founder argues that gated frontier cyber models leave small security startups without tools. Wolfram's worry: that this becomes an excuse to regulate open source.
- abliterated-model-large-v2: refusal-removed GLM-5.3 as a hosted API, $5/M tokens, $0.50 cached
- Only two guardrails left: self-harm and CSAM
- Early customers: agent red-teaming for banks, cyber startups, trust-and-safety
- No KYC: 'how would we be the arbiters' of legitimate use
- Reaction: mostly positive from cybersec, outrage from AI safety
🔓 Open Source LLMs — Tencent HY4, Z.ai GLM 5.3
Tencent's Hy4 preview is a 770B MoE whose Sherry quantization shrinks the weights from 1.5 TB to about 200 GB; Nisten, who works on one-bit models at Prism ML, warns that 2-bit quants of big models stay useful but can drop multilingual and other capabilities. Z.ai shipped the full GLM-5.3 weights a day after last week's Flash reveal. The segment turns into geopolitics: OpenAI pulled its models from Cursor after the SpaceX acquisition (Wolfram calls it a bad precedent, like Anthropic yanking Windsurf), and NVIDIA made the Hugging Face acquisition official.
- Tencent Hy4 preview: 770B MoE, Sherry quant takes 1.5 TB to ~200 GB
- Nisten: 2-bit quants of large models work, but expect losses elsewhere
- GLM-5.3 full weights out a day after the Flash reveal
- OpenAI pulls its models from Cursor post-SpaceX deal; NVIDIA's Hugging Face buy is official
🤖 Agents & Tools — Muse Code, OpenClaw 2.0
Muse Code is out of beta with $5, $20, and $50 plans and a contributor tier at $0.10/$0.20 per million. OpenClaw 2.0 adds native computer use through Cua's driver (Francesco's team, previously on the show), cloud fleets, and platform support down to Wear OS, though the panel and chat have mostly moved on. Anthropic finally shipped background computer use in Claude, Claude Code, and Cowork, with the odd restriction that it refuses to type into IDEs. Yam flags a Codex remote-voice update.
- Muse Code GA: $5 / $20 / $50 plans, contributor tier $0.10/$0.20
- OpenClaw 2.0: Cua computer-use driver, cloud fleets, everything from Mac to Wear OS
- Anthropic background computer use ships, but will not type into your IDE
- Chat consensus: Codex Mobile and Claude Mobile replaced OpenClaw for many
🎥 World Labs Atlas
Wolfram calls it the real big breakthrough of the week. Atlas is an omnimodel pretrained from scratch on text, images, video, and 3D: one image yields a camera-controlled flyover, a few images a full pixel-accurate volumetric scene with no Gaussian splats, and real videos can be reframed from new angles. The showstopper is bullet-time from three tripods: a watermelon splat viewable from every angle, which LDJ (noting Fei-Fei Li was Karpathy's advisor) says scales with more cameras, and which Alex and LDJ immediately want for sports replays. Partner access only for now.
- Omnimodel pretrained on text, images, video, and 3D; autoregressive diffusion transformer
- 1 to 6 images → walkable 3D reconstruction, up to a minute at 1440p
- Bullet-time 4D video from three ordinary tripods
- Fei-Fei Li keynotes Fully Connected 26
🎥 fal MiniMax H3 Max Turbo + fal.live
fal's H3 Max Turbo keeps about 97% of H3 Max quality, generates a clip in 1.4 seconds, and costs one cent per second at 768p. Alex's live Big Bang Theory test reveals the catch: Turbo no longer renders copyrighted characters the way H3 Max did. The segment closes on fal.live, the interactive stream fal built after Twitch and Kick bans, and Alex's take that world models are now advancing faster than computer graphics: we may get GTA 6 before GPT-6.
- H3 Max Turbo: ~97% of H3 Max quality, 1.4 s generation, $0.01/sec
- Turbo strips copyrighted characters that H3 Max happily rendered
- fal.live: viewer-steered infinite generation after Twitch and Kick bans
- World models are advancing faster than computer graphics
Frequently Asked Questions
What is new in Claude Fable 5.1 and Mythos 5.1?
Anthropic shipped Fable 5.1 and Mythos 5.1 as the same weights with different guardrails. Fable 5.1 scores 55.8 on Terminal-Bench 4.0, up from 42.0 for Fable 5 and versus about 37 for GPT-5.6 Sol, and 52% on the new Terminal-Bench Science, more than double Fable 5. Cache reads dropped 75% to $0.25 per million while input and output stay at $10 and $50 per million tokens. Peter Gostev's Code Arena tests put the max version first by a large margin, and the prompting guide now names the Opus 5 'jargon' style 'mannered prose' so you can prompt it away.
Did Meta Muse Spark 1.3 really catch up to Claude Fable 5?
On the Artificial Analysis index, yes. Spark 1.3 xhigh scores 61, tying GPT-5.6 Sol and Grok 4.6, and the limited-preview max reasoning mode scores 62, tying Fable 5. It is the first time Meta has jumped over OpenAI, Microsoft MAI, and Google on that index. Pricing is unchanged at $1.25/$4.25 per million, LDJ noted it runs about 3x faster than Fable and Sol at about 4x lower cost than Sol, and its MRCR long-context score went from 66% to 98.5% at 1M tokens. Meta is teasing open weights and a mystery 'Watermelon' model.
What is the Muse Spark contributor tier?
A pricing tier for people who let Meta train on their prompts and completions. It costs $0.10 per million input tokens and $0.20 per million output tokens, with cached input at two tenths of a cent, for a 1M-context model that matches GPT-5.6 on several benchmarks. Alex called it intelligence too cheap to meter if you are willing to share your data; Nisten said he would use it as a cheap verifier model for medical datasets.
What is abliterated-model-large-v2 from Abliteration AI?
It is Z.ai's GLM-5.3 with refusals removed via abliteration, offered as a hosted API rather than open weights. Pricing is $5 per million tokens, or $0.50 for cached tokens, on US-based servers with reasoning models available. The only remaining guardrails are self-harm and CSAM; everything else, including offensive cybersecurity and exploit writing, is allowed, and the lab reported completed exploit tasks rising from 29 to 105 versus the refusing model. The founder, who appeared anonymously, said the early paying market is red-teaming enterprise agent fleets, cyber startups, and trust-and-safety tooling, and that they do not do KYC.
Is GLM-5.3 open weights now?
Yes. A day after unveiling GLM-5.3 Flash (the model previously seen as OX Alpha), Z.ai released the full GLM-5.3 weights on Hugging Face: 753B total parameters with 40B active, under a custom glm-5.3 license. Z.ai claims 84.5% on CyberGym and 2,436 vulnerabilities found. Abliteration AI's uncensored model is built on these weights.
What is World Labs Atlas?
Atlas is World Labs' omnimodel, pretrained from scratch to natively operate in text, images, video, and 3D as an autoregressive diffusion transformer. From 1 to 6 images it produces up to a minute of 1440p camera-controlled video and a walkable 3D reconstruction, with no Gaussian splats. Its headline demo reframes real video from new camera angles, producing bullet-time from three ordinary tripods, and LDJ noted quality scales with more camera angles. Access is partner-only for now; Fei-Fei Li keynotes Fully Connected 26 on September 29 to October 1.
What is Runway Solaris?
Solaris is Runway's Interface World Model: instead of coding a UI, the model generates the interface frame by frame, so every element in a scene is clickable and draggable (click a lamp and the room lights up, drag shoes onto a person and he wears them). In Runway's own study, 71% preferred the generated behavior over hand-coded UI, reported as 61 to 24 over Opus 5. Runway also released GWM Worlds 2 the same day, a real-time 720p, 24 fps world model with 48 kHz audio and generated speech, both as research previews without public access.
What did Meta Muse Voice Transcribe and Microsoft MAI-Transcribe-2 launch this week?
Both are real-time speech-to-text models with speaker diarization. Meta Muse Voice Transcribe does streaming ASR, diarization, and endpointing in one model, claims 3.1% streaming WER, is API only, costs about 18 cents per hour, and powers the live transcript on thursdai.news/live, where it correctly labels speakers and terms like 'ThursdAI' via a custom dictionary. Microsoft MAI-Transcribe-2 ranks #2 on the Artificial Analysis WER leaderboard at 2.0%, runs about 400x real time, and costs $1.67 per 1,000 minutes, less than half the price of its peers.
TL;DR and show notes
Hosts and Guests
Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
Co-hosts: @WolframRvnwlf, @nisten, @ldjconfirmed, @yampeleg, @petergostev
The founder of Abliteration AI, who joined anonymously
GPT-6 Astra
Launched mid-show, covered in full in its own episode (thursdai.news/astra)
Frontier AI
Anthropic Claude Fable 5.1 and Mythos 5.1, same weights: Terminal-Bench 4.0 55.8 vs 42.0, cache reads down 75% to $0.25/M, and “mannered prose” gets a name in the prompting guide (X, Blog, System card, EFS, Writing density)
Meta Muse Spark 1.3: xhigh scores 61 on the AA index (ties Sol and Grok 4.6), limited-preview max 62 (ties Fable 5), unchanged $1.25/$4.25, open weights and a watermelon model “coming soon” (Zuck, AA analysis, AA model page)
Google Gemini 3.8 Flash and 3.8 Flash Cyber: HLE-Verified 54.9, $0.75/$3.75 until a doubling on Jan 1, 2027, Cyber is Fairwind-only with CWE-Bench 47.2% (X, Cyber, Fairwind, Pricing)
Alibaba Qwen3.8-Max-0902: 2.4T, 1M ctx, #1 on Code Arena, $2/$6, API only (X, Arena, QwenCloud)
Grok 4.7 lands next week, per Elon
Open Source LLMs
Z.ai releases the full GLM-5.3 weights: 753B/40B active, custom glm-5.3 license, CyberGym 84.5% claimed, 2,436 vulnerabilities found (X, HF, Blog)
Tencent Hy4 preview: 770B/49B active, 1M ctx, Apache 2.0, Sherry quant takes 1.5 TB to 214 GB at 2.38 bpw (X, Sherry, HF, Blog)
Abliteration AI abliterated-model-large-v2: refusal-removed GLM-5.3 as a hosted API, $5/M, only CSAM and self-harm blocked (X, Docs, Pricing)
OpenAI pulls its models from Cursor after the SpaceX acquisition (OpenAI)
NVIDIA makes the Hugging Face acquisition official at $12,930,300,000 (Clem, last week’s issue)
This Week’s Buzz
Kimi K3 (2.8T) on CoreWeave Dedicated Inference on GB300 NVL72 (X, Docs, Dedicated Inference)
Fully Connected 26, Sept 29 to Oct 1, Moscone South SF, Fei-Fei Li keynotes, ThursdAI live from the floor, free ticket code on the show (X, Keynote teaser, Register)
CoreWeave Hacks: Agent Loops, Sept 12 to 13 SF, $20k+ prizes, robot dog, F1 tickets (X, Luma)
Voice & Audio
Meta Muse Voice Transcribe: streaming ASR, diarization and endpointing in one model, 3.1% streaming WER claimed, API only, powers thursdai.news/live (X, Architecture, Zuck)
Microsoft MAI-Transcribe-2: #2 on AA WER at 2.0%, about 400x real time, $1.67 per 1,000 minutes (Launch, AA thread, Leaderboard)
Inworld Realtime TTS-2 GA: sub-100ms, $25/M chars, #1 on AA’s Controlled Voice Arena, #4 on the Provider arena (X)
AI Coding & Agents
Vision & Video
World Labs Atlas: up to 1 minute at 1440p from 1 to 6 images, 3D reconstruction, bullet time from three phones, partner access only (X, Blog)
Runway Solaris: Interface World Model, UIs generated frame by frame, preferred 61 to 24 over Opus 5 in Runway’s own study (X, Cristóbal, Blog)
Runway GWM Worlds 2: real-time 720p, 24 fps world model with 48 kHz audio and open-ended sessions, research preview (X, Research)
fal: infinite Rick and Morty stream banned from Twitch and Kick, fal.live built in a weekend, H3 Max Turbo at $0.01/sec, open FastH3 (Turbo, fal.live, Rehan, FastH3, HF)
H3 World: an open LoRA that turns MiniMax H3 into a walkable world model (Github)
Guest
Links & Resources
Frontier AI
- Claude Fable 5.1 & Mythos 5.1 announcement (X) ↗
- Anthropic: Claude Fable and Mythos 5.1 ↗
- Fable 5.1 system card thread ↗
- Anthropic Enterprise Frontier Safeguards ↗
- Prompting Fable 5.1: writing density (mannered prose) ↗
- Alex on the Opus 5 jargon problem ↗
- Zuck announces Muse Spark 1.3 ↗
- Artificial Analysis: Muse Spark 1.3 analysis ↗
- Artificial Analysis: Muse Spark model page ↗
- Gemini 3.8 Flash announcement ↗
- Gemini 3.8 Flash Cyber thread ↗
- Google DeepMind Fairwind program ↗
- Gemini API pricing ↗
- Qwen3.8-Max-0902 announcement ↗
- Qwen3.8-Max on Code Arena ↗
- Qwen3.8-Max-0902 on QwenCloud ↗
Open Source
- Z.ai releases GLM-5.3 weights (X) ↗
- GLM-5.3 on Hugging Face ↗
- Z.ai GLM-5.3 blog ↗
- Tencent Hy4 preview (X) ↗
- Tencent Sherry quantization ↗
- Hy4-preview on Hugging Face ↗
- Tencent Hy4 preview research page ↗
- OpenAI pulls models from Cursor ↗
- Clem on the NVIDIA acquisition closing ↗
- Last week's issue: NVIDIA buys Hugging Face ↗
Guest
This Week's Buzz
Voice & Audio
AI Coding & Agents
Vision & Video
- World Labs Atlas announcement ↗
- World Labs Atlas blog ↗
- Runway Solaris announcement ↗
- Cristóbal Valenzuela on Solaris ↗
- Runway: Introducing Solaris ↗
- Runway GWM Worlds 2 ↗
- Runway research ↗
- fal H3 Max Turbo ↗
- fal.live launch ↗
- Rehan on the infinite stream ↗
- FastH3 open release ↗
- FastH3 on Hugging Face ↗
- H3 World: MiniMax H3 as a walkable world model ↗
- fal set Alex straight on Turbo ↗
GPT-6 Astra