Hosts & Guests

Alex Volkov
Alex Volkov
Host · W&B / CoreWeave
@altryne
Abliteration AI founding engineer
Abliteration AI founding engineer
Abliteration AI — Founding Engineer (appeared anonymously)
@founderengineer
Wolfram Ravenwolf
Wolfram Ravenwolf
AI model evaluator
@WolframRvnwlf
Nisten Tahiraj
Nisten Tahiraj
AI operator & builder
@nisten
LDJ
LDJ
Nous Research
@ldjconfirmed
Yam Peleg
Yam Peleg
AI builder & founder
@Yampeleg
Peter Gostev
Peter Gostev
Arena (formerly LMArena)
@petergostev

By The Numbers

Terminal-Bench 4.0
55.8
Claude Fable 5.1, up from 42.0 for Fable 5 and vs 37 for GPT-5.6 Sol
AA Intelligence Index
62
Muse Spark 1.3 max mode ties Fable 5; xhigh scores 61, tying GPT-5.6 Sol and Grok 4.6
MRCR 1M context
98.5%
Muse Spark 1.3 long-context score, up from 66% on the previous Spark
Muse Voice Transcribe
$0.18/hr
Streaming ASR with live diarization; a three-hour show for under half a dollar
GLM-5.3 open weights
753B / 40B
Total / active parameters; Z.ai released the full weights a day after the Flash reveal
fal H3 Max Turbo
1¢/sec
~97% of H3 Max quality, 1.4-second generation at 768p

🔥 Breaking During The Show

Runway drops Solaris and GWM Worlds 2 mid-show
LDJ broke it live: GWM Worlds 2 is a real-time 24 fps world model with 48 kHz audio and generated speech, and Solaris is an Interface World Model that renders clickable, draggable UIs frame by frame. Runway's own study had 71% of people preferring the generated behavior over hand-coded UI.
GPT-6 Astra landed two and a half hours into the stream
The show stayed live waiting for it, and OpenAI delivered. Astra is covered in full in its own episode (Part 2) with Peter Gostev and Ryan Carson, who both had early access.

👋 Opening

Alex records a cold open from the editing floor: this is a special two-part podcast because GPT-6 Astra dropped about two and a half hours into a five-hour livestream. Part 1 is the traditional ThursdAI show (Fable 5.1, Muse Spark 1.3, the Abliteration AI interview, World Labs Atlas); Part 2 is Astra with Peter Gostev and Ryan Carson.

  • Two-part episode: the regular show here, GPT-6 Astra in its own episode
  • Five hours of continuous livestream, edited down with Fable 5.1's help
Alex Volkov
Alex Volkov
"What happened during the shows, we continuously speculated and waited for GPT-6 Astra, and around two and a half hours in, it showed up."

🎙️ Show Start

Live from CoreWeave, Alex welcomes Wolfram and LDJ to what he calls possibly the most insane week of frontier-lab updates the show has seen. Summer is officially over, and the GPT-6 rumor (same jump as 3 to 4, already shown to government officials) hangs over everything.

  • Summer break is over: every frontier lab shipped something this week
  • Rumor mill: GPT-6 is the same jump as GPT-3 to GPT-4
Wolfram Ravenwolf
Wolfram Ravenwolf
"But the summer break is over. We are going full speed ahead again, feel the acceleration, and I'm super excited."

🔥 Banter — Highlights of the Week

Everyone picks a highlight before Astra: Wolfram calls fal's open MiniMax H3 the stable-diffusion moment for video (fal got banned from Twitch and Kick for an infinite Rick and Morty stream, then built fal.live over a weekend). Yam and Alex pick Fable 5.1, and Alex shows the thursdai.news/live page Fable built in one sitting, complete with live Muse Voice Transcribe diarization and an agentic producer throwing chyrons.

  • fal's infinite Rick and Morty stream: banned from Twitch, banned from Kick, so they built fal.live
  • Fable 5.1 one-shot a live-streaming platform for thursdai.news/live in one sitting
  • Anthropic gave 'jargon douche' an official name: mannered prose
  • Peter: Meta's Spark 1 → 1.1 → 1.2 → 1.3 shipping speed is the real story
Yam Peleg
Yam Peleg
"No other model just gave me a simple solution for this."
Alex Volkov
Alex Volkov
"This happened over four hours yesterday. Fable just fucking murdered that task with almost no issues."
Wolfram Ravenwolf
Wolfram Ravenwolf
"It is making so many waves, it's like a stable diffusion moment."

📰 TL;DR

The rundown of everything deemed cover-worthy: Fable 5.1 and Mythos 5.1 (75% cheaper cache reads, a Terminal-Bench jump), OpenAI designating Astra the first model at its critical cyber threshold, Gemini 3.8 Flash and Flash Cyber, Qwen3.8-Max-0902, Muse Spark 1.3 tying GPT-5.6 Sol and Grok 4.6 on Artificial Analysis, Tencent Hy4, GLM-5.3 open weights, the abliterated GLM-5.3, Inworld TTS-2, MAI-Transcribe-2, World Labs Atlas, Runway Solaris, and fal's H3 Max week.

  • Frontier: Fable 5.1 / Mythos 5.1, Gemini 3.8 Flash + Cyber, Qwen3.8-Max-0902, Muse Spark 1.3; Grok 4.7 promised next week
  • Open source: Tencent Hy4 preview (770B MoE), GLM-5.3 full weights, Abliteration AI's uncensored GLM-5.3
  • Voice: Inworld real-time TTS GA, Microsoft MAI-Transcribe-2 at 2% WER and $1.67 per 1,000 minutes
  • Vision & video: World Labs Atlas, Runway Solaris, fal H3 Max Turbo at one cent per second
Alex Volkov
Alex Volkov
"All of them were trying to have their second in the sun before Astro drops. That's how this news felt to me."

🏢 Frontier AI — Big CO LLMs + APIs

The Frontier AI corner opens. With every big lab shipping a point-one release ahead of OpenAI, Alex starts with the one the whole panel agrees is the biggest: Fable 5.1 and Mythos 5.1.

  • Every frontier lab shipped a .1 this week

🏢 Anthropic Claude Fable 5.1 + Mythos 5.1

Fable 5.1 scores 55.8 on Terminal-Bench 4 (up from 42 for Fable 5, versus 37 for GPT-5.6 Sol) and 52% on the new Terminal-Bench Science, more than double Fable 5. Cache reads drop 75% to $0.25/M while input/output stay at $10/$50. Peter tested the max version on Code Arena where it came in first by a large margin, though single generations cost him $40 to $60. Anthropic also officially named the Opus 5 'jargon douche' problem: mannered prose, which you can now prompt away. Nisten demos a two-prompt Mars-launch simulator that built the entire solar system, a NASA mission planner, and a parachute landing on Earth.

  • Terminal-Bench 4: 55.8 vs 42.0 for Fable 5 and 37 for GPT-5.6 Sol
  • Terminal-Bench Science: 52%, more than double Fable 5
  • Cache reads 75% cheaper; $10/$50 per million tokens unchanged
  • Peter: #1 on Code Arena by a large margin, but $40 to $60 per generation
  • 'Mannered prose' is the official name for load-bearing / control-plane speak, and you can prompt it away
  • Nisten's two-prompt Mars maglev launcher sim grew into a full solar system with a parachute landing
Peter Gostev
Peter Gostev
"The bottleneck is now you. Find good tasks. That's the hard part now."
Alex Volkov
Alex Volkov
"When I know a higher tier intelligence exists and I somehow have access to it, I really want that tier of intelligence because there may be stuff that I don't know that it does."
Nisten Tahiraj
Nisten Tahiraj
"Here's what Fable did in two prompts. It added a NASA mission planner, which will figure out the most optimal fuel, and then it will rebuild the entire launchpad based on that."

🔥 Runway Solaris + GWM Worlds 2 (Breaking News)

LDJ breaks Runway's GWM Worlds 2: a real-time 24 fps world model with 48 kHz audio, including generated speech, two days after Runway's previous world model. Alex then shows Solaris, an Interface World Model where UIs are generated frame by frame: drag shoes onto a person, click a lamp and the room lights up. Runway's own study found 71% preferred the generated behavior over hand-coded UI, and Wolfram calls it the future of interfaces.

  • GWM Worlds 2: real-time 24 fps world model with 48 kHz audio and speech
  • Solaris: world model for interfaces, every pixel clickable and draggable
  • 71% preferred the generated behavior over coded UI in Runway's study
  • Alex to Runway: let people actually play with it (CoreWeave has GPUs)
Wolfram Ravenwolf
Wolfram Ravenwolf
"I think this is the future of the interfaces. The whole UI could also be just generated on the fly if it was fast enough, and then you could change it any way you want it to."
Alex Volkov
Alex Volkov
"This is a scene of a living room where everything is clickable, everything is movable."

🏢 Google Gemini 3.8 Flash

Google's third Flash iteration in as many weeks: 1M context, 3x faster, billed as the reasoning and coding workhorse for Googlers, who all build on Antigravity internally. Wolfram uses every Flash update in his home assistant but wants a Pro model; the Wall Street Journal reported Google scrapped the 3.5 Pro checkpoints because Flash overtook them, and Gemini 4 is still in post-training.

  • Third Flash release in three weeks, 1M context, 3x faster
  • WSJ: 3.5 Pro checkpoints scrapped, Gemini 4 still in post-training
  • Wolfram: a million tokens matters less when the model burns context thinking
Wolfram Ravenwolf
Wolfram Ravenwolf
"So yeah, it's another Flash model. It's nice, but give us some pro models. Come on."

🛡️ Gemini 3.8 Flash Cyber

A dedicated cybersecurity variant of 3.8 Flash, gated behind Google's Fairwind program for trusted defenders. No benchmark numbers surfaced on air, though the newsletter lists CWE-Bench at 47.2%.

  • Fairwind-only cyber variant of Gemini 3.8 Flash

🏢 Alibaba Qwen 3.8 Max-0902

Alibaba refreshed its API-only frontier model and claimed #1 on a WebDev leaderboard, a claim Alex is skeptical of given the rest of the week. Nobody on the panel knows anyone using Qwen Max in production, beyond one IT worker running OpenClaw on it and people generating datasets.

  • API-only refresh, #1 Code Arena claim
  • Panel poll: nobody knows a production Qwen Max user

🏢 Meta Muse Spark 1.3

For the first time, Meta jumps over OpenAI, Microsoft MAI, and Google on the Artificial Analysis index: Spark 1.3 xhigh ties GPT-5.6 Sol and Grok 4.6, and an unreleased max reasoning mode ties Fable 5. LDJ notes it runs about 3x faster than Fable and Sol at about 4x lower cost than Sol. MRCR long-context jumped from 66% to 98.5% at 1M tokens, and the contributor tier is $0.10/$0.20 per million if you let Meta train on your prompts. Zuck and Alex Wang are teasing open weights and a mystery 'Watermelon' model.

  • Spark 1.3 xhigh ties GPT-5.6 Sol and Grok 4.6; max mode ties Fable 5 on AA
  • About 3x faster than Fable and Sol, about 4x cheaper than Sol
  • MRCR long-context 66% → 98.5% at 1M tokens
  • Contributor tier: $0.10 in / $0.20 out per million if Meta can train on your data
  • Open weights and a 'Watermelon' model teased
LDJ
LDJ
"It's running at about three times faster than Fable and Sol are, according to Artificial Analysis speed measurements at least."
Alex Volkov
Alex Volkov
"Basically, intelligence too cheap to meter if you're willing to sell your soul to Zach for training."

🔊 Meta Muse Voice Transcribe

Meta's streaming ASR model does transcription, speaker diarization, and endpointing in one model. Alex shows it powering the live transcript on thursdai.news/live: it labels Yam, Alex, and Wolfram in real time, gets names like Alex Wang and terms like ThursdAI right with a custom dictionary, and costs 18 cents per hour. Its diarization only fails about 17% of the time when someone talks over Alex, which he calls better than Descript.

  • Streaming ASR + diarization + endpointing in one model
  • 18 cents per hour: a three-hour show for under half a dollar
  • Custom dictionary support: transcribes 'ThursdAI' correctly
  • Real-time speaker labels power the new thursdai.news/live page
Alex Volkov
Alex Volkov
"This model understands who speaks when, and it sends labels in real time at 18 cents per hour."

⚡ This Week's Buzz — CoreWeave, Fully Connected, CoreWeave Hacks

Fully Connected 26 runs September 29 to October 1 at Moscone South in San Francisco with 32 sessions, Fei-Fei Li keynoting, battle bots, and a ThursdAI live show from the floor; a code shared on the livestream gets you a 100% free ticket (regular price $1,299). CoreWeave Hacks: Agent Loops is September 12 to 13 in SF with W&B, AGI House, and TypeSafe AI, over $20,000 in prizes, a robot dog, and an F1 ticket for the most production-ready hack. Kimi K3 now runs on CoreWeave Dedicated Inference on GB300s.

  • Fully Connected 26: Sept 29 to Oct 1, Moscone South, Fei-Fei Li keynote, free ticket code on the live show
  • CoreWeave Hacks Agent Loops: Sept 12 to 13, $20k+ prizes, robot dog, F1 tickets
  • Kimi K3 on CoreWeave Dedicated Inference on GB300 NVL72
Alex Volkov
Alex Volkov
"You get a 100% free ticket because you're following us live."

🔓 Abliteration AI — Obliterated GLM 5.3 Interview

Abliteration AI's founder joins anonymously to explain abliterated-model-large-v2, a GLM-5.3 with refusals removed, hosted as an API at $5 per million tokens ($0.50 cached) on US servers. He stresses they did not invent abliteration (the refusal-direction technique is public and dozens of abliterated models already sit on Hugging Face); the paying market so far is red-teaming enterprise agent fleets, cybersecurity, and trust-and-safety tooling, with only self-harm and CSAM still blocked. Nisten raises social-worker and Metasploit-style use cases, Alex pushes back on vetting, and the founder argues that gated frontier cyber models leave small security startups without tools. Wolfram's worry: that this becomes an excuse to regulate open source.

  • abliterated-model-large-v2: refusal-removed GLM-5.3 as a hosted API, $5/M tokens, $0.50 cached
  • Only two guardrails left: self-harm and CSAM
  • Early customers: agent red-teaming for banks, cyber startups, trust-and-safety
  • No KYC: 'how would we be the arbiters' of legitimate use
  • Reaction: mostly positive from cybersec, outrage from AI safety
Abliteration AI founding engineer
Abliteration AI founding engineer
"I'm surprised by the surprise, a little bit, because it was a white paper that maybe came out like three years ago or so, the ability to find and change these refusal vectors."
Abliteration AI founding engineer
Abliteration AI founding engineer
"We do have a guardrail there for self-harm, so we don't wanna help somebody in that way. And then we also have one for child sex as well."
Wolfram Ravenwolf
Wolfram Ravenwolf
"I just hope they don't use this as an excuse to another round of, 'Hey, we have to regulate open source AI because this can be done.'"

🔓 Open Source LLMs — Tencent HY4, Z.ai GLM 5.3

Tencent's Hy4 preview is a 770B MoE whose Sherry quantization shrinks the weights from 1.5 TB to about 200 GB; Nisten, who works on one-bit models at Prism ML, warns that 2-bit quants of big models stay useful but can drop multilingual and other capabilities. Z.ai shipped the full GLM-5.3 weights a day after last week's Flash reveal. The segment turns into geopolitics: OpenAI pulled its models from Cursor after the SpaceX acquisition (Wolfram calls it a bad precedent, like Anthropic yanking Windsurf), and NVIDIA made the Hugging Face acquisition official.

  • Tencent Hy4 preview: 770B MoE, Sherry quant takes 1.5 TB to ~200 GB
  • Nisten: 2-bit quants of large models work, but expect losses elsewhere
  • GLM-5.3 full weights out a day after the Flash reveal
  • OpenAI pulls its models from Cursor post-SpaceX deal; NVIDIA's Hugging Face buy is official
Nisten Tahiraj
Nisten Tahiraj
"Keep in mind that there are going to be losses elsewhere, and they might be quite dramatic."
Wolfram Ravenwolf
Wolfram Ravenwolf
"Affecting the customer that way, it's a bad precedent if the model providers just decide, okay, I don't like you, you are not getting this."

🤖 Agents & Tools — Muse Code, OpenClaw 2.0

Muse Code is out of beta with $5, $20, and $50 plans and a contributor tier at $0.10/$0.20 per million. OpenClaw 2.0 adds native computer use through Cua's driver (Francesco's team, previously on the show), cloud fleets, and platform support down to Wear OS, though the panel and chat have mostly moved on. Anthropic finally shipped background computer use in Claude, Claude Code, and Cowork, with the odd restriction that it refuses to type into IDEs. Yam flags a Codex remote-voice update.

  • Muse Code GA: $5 / $20 / $50 plans, contributor tier $0.10/$0.20
  • OpenClaw 2.0: Cua computer-use driver, cloud fleets, everything from Mac to Wear OS
  • Anthropic background computer use ships, but will not type into your IDE
  • Chat consensus: Codex Mobile and Claude Mobile replaced OpenClaw for many
Wolfram Ravenwolf
Wolfram Ravenwolf
"Anyone still using OpenClaw here in our group or in the audience?"

🎥 World Labs Atlas

Wolfram calls it the real big breakthrough of the week. Atlas is an omnimodel pretrained from scratch on text, images, video, and 3D: one image yields a camera-controlled flyover, a few images a full pixel-accurate volumetric scene with no Gaussian splats, and real videos can be reframed from new angles. The showstopper is bullet-time from three tripods: a watermelon splat viewable from every angle, which LDJ (noting Fei-Fei Li was Karpathy's advisor) says scales with more cameras, and which Alex and LDJ immediately want for sports replays. Partner access only for now.

  • Omnimodel pretrained on text, images, video, and 3D; autoregressive diffusion transformer
  • 1 to 6 images → walkable 3D reconstruction, up to a minute at 1440p
  • Bullet-time 4D video from three ordinary tripods
  • Fei-Fei Li keynotes Fully Connected 26
Wolfram Ravenwolf
Wolfram Ravenwolf
"Every filmmaker would be amazed to have technology like this. You take your shot and you realize, oh, it would have been better if the camera was just a little more to the left. Now you can just do it."
Alex Volkov
Alex Volkov
"There's a video of a watermelon getting splat with three tripods. You get every angle of that thing happening."

🎥 fal MiniMax H3 Max Turbo + fal.live

fal's H3 Max Turbo keeps about 97% of H3 Max quality, generates a clip in 1.4 seconds, and costs one cent per second at 768p. Alex's live Big Bang Theory test reveals the catch: Turbo no longer renders copyrighted characters the way H3 Max did. The segment closes on fal.live, the interactive stream fal built after Twitch and Kick bans, and Alex's take that world models are now advancing faster than computer graphics: we may get GTA 6 before GPT-6.

  • H3 Max Turbo: ~97% of H3 Max quality, 1.4 s generation, $0.01/sec
  • Turbo strips copyrighted characters that H3 Max happily rendered
  • fal.live: viewer-steered infinite generation after Twitch and Kick bans
  • World models are advancing faster than computer graphics
Alex Volkov
Alex Volkov
"The price of intelligence and the price of video generation is just dropping to the floor, absolutely."
Wolfram Ravenwolf
Wolfram Ravenwolf
"If AI takes all the jobs, we at least can watch unlimited TV."

Frequently Asked Questions

What is new in Claude Fable 5.1 and Mythos 5.1?

Anthropic shipped Fable 5.1 and Mythos 5.1 as the same weights with different guardrails. Fable 5.1 scores 55.8 on Terminal-Bench 4.0, up from 42.0 for Fable 5 and versus about 37 for GPT-5.6 Sol, and 52% on the new Terminal-Bench Science, more than double Fable 5. Cache reads dropped 75% to $0.25 per million while input and output stay at $10 and $50 per million tokens. Peter Gostev's Code Arena tests put the max version first by a large margin, and the prompting guide now names the Opus 5 'jargon' style 'mannered prose' so you can prompt it away.

Did Meta Muse Spark 1.3 really catch up to Claude Fable 5?

On the Artificial Analysis index, yes. Spark 1.3 xhigh scores 61, tying GPT-5.6 Sol and Grok 4.6, and the limited-preview max reasoning mode scores 62, tying Fable 5. It is the first time Meta has jumped over OpenAI, Microsoft MAI, and Google on that index. Pricing is unchanged at $1.25/$4.25 per million, LDJ noted it runs about 3x faster than Fable and Sol at about 4x lower cost than Sol, and its MRCR long-context score went from 66% to 98.5% at 1M tokens. Meta is teasing open weights and a mystery 'Watermelon' model.

What is the Muse Spark contributor tier?

A pricing tier for people who let Meta train on their prompts and completions. It costs $0.10 per million input tokens and $0.20 per million output tokens, with cached input at two tenths of a cent, for a 1M-context model that matches GPT-5.6 on several benchmarks. Alex called it intelligence too cheap to meter if you are willing to share your data; Nisten said he would use it as a cheap verifier model for medical datasets.

What is abliterated-model-large-v2 from Abliteration AI?

It is Z.ai's GLM-5.3 with refusals removed via abliteration, offered as a hosted API rather than open weights. Pricing is $5 per million tokens, or $0.50 for cached tokens, on US-based servers with reasoning models available. The only remaining guardrails are self-harm and CSAM; everything else, including offensive cybersecurity and exploit writing, is allowed, and the lab reported completed exploit tasks rising from 29 to 105 versus the refusing model. The founder, who appeared anonymously, said the early paying market is red-teaming enterprise agent fleets, cyber startups, and trust-and-safety tooling, and that they do not do KYC.

Is GLM-5.3 open weights now?

Yes. A day after unveiling GLM-5.3 Flash (the model previously seen as OX Alpha), Z.ai released the full GLM-5.3 weights on Hugging Face: 753B total parameters with 40B active, under a custom glm-5.3 license. Z.ai claims 84.5% on CyberGym and 2,436 vulnerabilities found. Abliteration AI's uncensored model is built on these weights.

What is World Labs Atlas?

Atlas is World Labs' omnimodel, pretrained from scratch to natively operate in text, images, video, and 3D as an autoregressive diffusion transformer. From 1 to 6 images it produces up to a minute of 1440p camera-controlled video and a walkable 3D reconstruction, with no Gaussian splats. Its headline demo reframes real video from new camera angles, producing bullet-time from three ordinary tripods, and LDJ noted quality scales with more camera angles. Access is partner-only for now; Fei-Fei Li keynotes Fully Connected 26 on September 29 to October 1.

What is Runway Solaris?

Solaris is Runway's Interface World Model: instead of coding a UI, the model generates the interface frame by frame, so every element in a scene is clickable and draggable (click a lamp and the room lights up, drag shoes onto a person and he wears them). In Runway's own study, 71% preferred the generated behavior over hand-coded UI, reported as 61 to 24 over Opus 5. Runway also released GWM Worlds 2 the same day, a real-time 720p, 24 fps world model with 48 kHz audio and generated speech, both as research previews without public access.

What did Meta Muse Voice Transcribe and Microsoft MAI-Transcribe-2 launch this week?

Both are real-time speech-to-text models with speaker diarization. Meta Muse Voice Transcribe does streaming ASR, diarization, and endpointing in one model, claims 3.1% streaming WER, is API only, costs about 18 cents per hour, and powers the live transcript on thursdai.news/live, where it correctly labels speakers and terms like 'ThursdAI' via a custom dictionary. Microsoft MAI-Transcribe-2 ranks #2 on the Artificial Analysis WER leaderboard at 2.0%, runs about 400x real time, and costs $1.67 per 1,000 minutes, less than half the price of its peers.

TL;DR and show notes

  • Hosts and Guests

  • GPT-6 Astra

  • Frontier AI

    • Anthropic Claude Fable 5.1 and Mythos 5.1, same weights: Terminal-Bench 4.0 55.8 vs 42.0, cache reads down 75% to $0.25/M, and “mannered prose” gets a name in the prompting guide (X, Blog, System card, EFS, Writing density)

    • Meta Muse Spark 1.3: xhigh scores 61 on the AA index (ties Sol and Grok 4.6), limited-preview max 62 (ties Fable 5), unchanged $1.25/$4.25, open weights and a watermelon model “coming soon” (Zuck, AA analysis, AA model page)

    • Google Gemini 3.8 Flash and 3.8 Flash Cyber: HLE-Verified 54.9, $0.75/$3.75 until a doubling on Jan 1, 2027, Cyber is Fairwind-only with CWE-Bench 47.2% (X, Cyber, Fairwind, Pricing)

    • Alibaba Qwen3.8-Max-0902: 2.4T, 1M ctx, #1 on Code Arena, $2/$6, API only (X, Arena, QwenCloud)

    • Grok 4.7 lands next week, per Elon

  • Open Source LLMs

    • Z.ai releases the full GLM-5.3 weights: 753B/40B active, custom glm-5.3 license, CyberGym 84.5% claimed, 2,436 vulnerabilities found (X, HF, Blog)

    • Tencent Hy4 preview: 770B/49B active, 1M ctx, Apache 2.0, Sherry quant takes 1.5 TB to 214 GB at 2.38 bpw (X, Sherry, HF, Blog)

    • Abliteration AI abliterated-model-large-v2: refusal-removed GLM-5.3 as a hosted API, $5/M, only CSAM and self-harm blocked (X, Docs, Pricing)

    • OpenAI pulls its models from Cursor after the SpaceX acquisition (OpenAI)

    • NVIDIA makes the Hugging Face acquisition official at $12,930,300,000 (Clem, last week’s issue)

  • This Week’s Buzz

    • Kimi K3 (2.8T) on CoreWeave Dedicated Inference on GB300 NVL72 (X, Docs, Dedicated Inference)

    • Fully Connected 26, Sept 29 to Oct 1, Moscone South SF, Fei-Fei Li keynotes, ThursdAI live from the floor, free ticket code on the show (X, Keynote teaser, Register)

    • CoreWeave Hacks: Agent Loops, Sept 12 to 13 SF, $20k+ prizes, robot dog, F1 tickets (X, Luma)

  • Voice & Audio

    • Meta Muse Voice Transcribe: streaming ASR, diarization and endpointing in one model, 3.1% streaming WER claimed, API only, powers thursdai.news/live (X, Architecture, Zuck)

    • Microsoft MAI-Transcribe-2: #2 on AA WER at 2.0%, about 400x real time, $1.67 per 1,000 minutes (Launch, AA thread, Leaderboard)

    • Inworld Realtime TTS-2 GA: sub-100ms, $25/M chars, #1 on AA’s Controlled Voice Arena, #4 on the Provider arena (X)

  • AI Coding & Agents

    • Muse Code out of beta: $5/$20/$50 plans, TypeScript SDK preview, contributor tier at $0.10/$0.20 (Zuck, Pricing, Blog)

    • OpenClaw 2.0: 16,977 PRs, native computer use through Cua Driver, cloud fleets (X, Cua, Blog)

  • Vision & Video

    • World Labs Atlas: up to 1 minute at 1440p from 1 to 6 images, 3D reconstruction, bullet time from three phones, partner access only (X, Blog)

    • Runway Solaris: Interface World Model, UIs generated frame by frame, preferred 61 to 24 over Opus 5 in Runway’s own study (X, Cristóbal, Blog)

    • Runway GWM Worlds 2: real-time 720p, 24 fps world model with 48 kHz audio and open-ended sessions, research preview (X, Research)

    • fal: infinite Rick and Morty stream banned from Twitch and Kick, fal.live built in a weekend, H3 Max Turbo at $0.01/sec, open FastH3 (Turbo, fal.live, Rehan, FastH3, HF)

    • H3 World: an open LoRA that turns MiniMax H3 into a walkable world model (Github)

  • Guest

    • Abliteration AI’s anonymous founder: why “completely uncensored,” the Policy Gateway, who’s buying, and the gated-frontier week it landed in, from 59:08 in the video (Launch, Docs)

Alex Volkov
Alex Volkov 0:35
Hey, this is Alex from the editing floor.
0:37
It's September three, and we have a special two-part podcast this week. This week was an absolutely insanely packed week in terms of AI releases. All the frontier labs released something, and, when we started the livestream, we only had speculation about OpenAI releasing GPT-6 Astra, which saturated the Ark AGI III with ninety-nine percent. So is AGI here? Find out. What you'll hear in this first episode is a fairly traditional ThursdAI podcast with the TLDR, and, we talked about Fable five point one and Mythos five point one. We talked about Meta MuSpark and live transcription from Meta and M-Microsoft AI. And we also had an interview with the founder of Obliteration AI, who took GLM five point three and removed all safety things, which went very viral. And we talked about World Labs Atlas, which might be the most impressive world model demo that we've ever seen. What happened during the shows, we continuously speculated and waited for GPT-6 Astra, and around two and a half hours in, it showed up. So what you'll hear in the episode after this, separate episode, is our conversation with Peter Goste from Arena, who had access and b-built a bunch of incredible demos. He was also featured by OpenAI, on the release. And Ryan is back on the show. He also had early access and talked about what he thinks about GPT-6 Astra. I am so excited to bring you this that the excitement blew up to five hours of continuous livestream, and I hope you enjoy this first one. Choose whichever one you want. Fable five point one helped me edit those down. So this is enough of a preamble. Enjoy the episode and I hope you listen to both of them, but choose which one-- which path you wanna go to Everyone to ThursdAI for September 3. My name is Alex Volkov. I'm an AI Evangelist with Weights & Biases from CoreWeave. I'm your host for today, for what could be the most insane week that we've got in terms of AI updates from Frontier Labs. Also open source, but today's mostly about Frontier Labs, folks. Welcome Wolfram, welcome LDJ.
Wolfram Ravenwolf
Wolfram Ravenwolf 2:45
Hey,
2:46
Alex. I hope- What a week I hope you have coffee, Wolfram. Oh, I do. Yes. I have. I drank some already. But the summer break is over. We are going full speed ahead again, feel the acceleration, and I'm super excited.
Alex Volkov
Alex Volkov 3:00
Summer is officially- What a week … over, and what an absolutely
3:04
insane week we got for you guys. It's just getting started. It's just getting started, because OpenAI loves to ship on Thursdays. the reason for the name ThursdAI is because GPT-4 was dropped on March 13th of 2023, and we started the show, and we never stopped talking to you about AI ever since. and rumored is, the rumor is that GPT-6 is the same jump from 3 to 4. LDJ, how are you feeling? Do you see the, the rumors, the news, the, the, the hints from, OpenAI?
LDJ
LDJ 3:41
Yeah, I did and I don't want to get too overexcited, but it is exciting.
3:44
I think today will be exciting. I think we'll-
Alex Volkov
Alex Volkov 3:47
I think we're allowed to get overhyped and overexcited-
3:50
Yeah … on days like this. They're very similar- Yeah to the days that, of when we launched the show. Yeah. All I know that OpenAI did talk about this publicly in multiple places and, apparently this model has been presented to, like, mm, officials in the government. Sam Altman has shown this to multiple folks. Like, I'm pretty sure that, like, many people have early access, for, But, b- but before we get to this news, we have just, like, a, a whole bunch of news to cover. let's do a favorite. It's gonna be hard. Nothing's gonna beat Astra. Like, like, I know that this is going to be the case. But before we get Astra, what is the top thing that you guys, think that happened this week before Astra? I, I'll go LDJ, Wolfram, and then Nisten.
Wolfram Ravenwolf
Wolfram Ravenwolf 4:41
Yeah, so- Like, so the Minimax model, the video model, which
4:44
was released as open source, we covered it, but it is making so many waves, it's like a stable diffusion moment where they throw this out and people- Yeah … grab it and make it faster and make it, render continuous as we, will cover later. Yeah. But this is, yeah, a milestone I think. Very much a milestone.
Alex Volkov
Alex Volkov 5:01
This definitely feels like the stable diffusion moment for video.
5:04
for those of us who are new to AI and didn't join the stable diffusion moment, once the open weights dropped for stable diffusion, multiple things happened in parallel. People started building fine tunes. People started, like, speeding up the model. Minimax released HD Max, what, three weeks ago? two weeks ago, Fal followed up with H3 Max. so Minimax released HD. HD Max, from Fal, that generates a five-second video in four, four, four, 2.5 seconds or something crazy like that. because of the speed, a bunch of new use cases were possible. So folks from Fal.ai, great folks, used their model, and the fact that they have unlimited GPU internally apparently- Mm to create a live stream on Twitch of interdimensional cable, which is a thing from Rick and Morty where they sit and watch cable TV show from multiple dimensions. So basically just, like, the, the most ridiculous shit you can think of. and they keep watching that, and the model generates 50 seconds of that. They got banned from Twitch, obviously, for copyright. and then they went to Kick, got banned from Kick, went to another streaming service, got banned from there, built their own streaming service, and apparently also a version of the model that's tuned for live streaming like that, continuous live streaming. so we can, we can check in on, h- let's check in Fal.ai. Just one
Wolfram Ravenwolf
Wolfram Ravenwolf 6:25
thing-
6:25
What is it called? Fal TV, or? They call it Fal Live. and- Chef called it Fal TV. Yeah. This is also a model you can run locally. So I have a local version of H3 with a, with a fast LoRa, which is also super fast compared to the full version. Yeah. And you can run this locally as well. This is the amazing thing
Alex Volkov
Alex Volkov 6:44
So here's,… So they basically over a weekend
6:46
built their own streaming service. There's 137 people there.
Wolfram Ravenwolf
Wolfram Ravenwolf 6:49
You can
6:49
just build things.
Alex Volkov
Alex Volkov 6:50
Yeah, you can just build things.
6:51
speaking of just build things, we'll, we'll talk about this. But yeah, you can just see like a, a thing, and then you, you basically vote on the next thing. folks, we got two more friends of the show joining. Oh, Peter was just here. oh, Yam Peleg, how come you don't have your glasses on yet? Put them on immediately. GPT-6 is about to drop.
Yam Peleg
Yam Peleg 7:13
And we got Fable.
7:14
What do you think?
Alex Volkov
Alex Volkov 7:16
Do- I think that Fable is amazing, and I'll
7:20
talk about this in a second. This is my highlight of the week until this point. Actually, no, I don't think that's my highlight, although Fable is dope. is Fable your highlight then, Yam, since you mentioned it as you came in?
Yam Peleg
Yam Peleg 7:29
It is absolutely great-
7:31
Yeah … Yam Peleg: in my opinion.
Alex Volkov
Alex Volkov 7:32
Absolutely great.
7:33
What's the- What do you think? what's the, what's the change for you? Do, do you feel a change? I know that it speaks better, like there's no more Claudes or like a little bit less Claudes.
Yam Peleg
Yam Peleg 7:42
Yeah, absolutely.
7:44
That's, that's what I wanted. the thing … Look, it gets, it gets like high level, I, I think models at the moment, they have a hard time understanding like bigger picture and like abstract concept. Like, you talk to them and- Yeah … Yam Peleg: they seem to not understand like what exactly is the thing that you want, and they fixate on something. But you can just ask it general stuff. For exam- I'll give you an example I run, on a many models and, I have, with, with GitHub and, and Git like everybody else, and I have CI because, they, they make mistakes and you want to make sure that they're not, destroying the build. The thing is that that thing is not built for like 400 of them pushing at the same time. So it's like it's just, it doesn't scale well because it's built for humans. so I just, just like I told you at the moment, just like that, I just asked Fable, "What can I do about this? Like, do you have any, any ideas?" Like, I think that I have so many agents and everything is, the throughput is pretty much nothing because you're bounded for, like you have one PR merging and then all the other need to rebase and like-… they're running in circles and it's just not built for this. And it just gave me- Fable just eats this. Just eats this … absolutely great- Yeah … Yam Peleg: solution that is simple, to just… I, I, I'm not gonna go into the solution it said, but like, what I'm saying is that no other model just gave me a simple solution for this.
Alex Volkov
Alex Volkov 9:13
All of them went, you know- No other
9:14
model until today, Yam. Until today. Yeah. We'll, we'll see what happens later on today. Ah. But until today, no other model. Yeah. I also wanna, shout out this. w- we'll, w- we should probably just talk about this and, and do it till the end and start the show, but I'm, I'm so excited, like I'm all over the place. Claude also, talked about these models, used to talk like an asshole, like a jargon douche like we told you before, and they actually gave it a name. They called it mannered, mannered speech. Mannered prose. so you literally can ask Claude in this platform, prompting guide for Claude Fable 5.1, you can ask it to do less mannered prose. This is the name for what we call jargon douche, where it, it uses words like load-bearing and pain for every- everything is a pain. Everything is controlled plane, pain, supporting, load-bearing, like all of that bullshit that we used to from Opus 5 as normal. There, Peter Gosta, welcome back to the show. Another exciting week. are you back from, SF? You still in SF?
Peter Gostev
Peter Gostev 10:09
Like, what's going on with you?
10:10
See the highlight, I'm, I'm also, I don't know, you guys mentioned, Meta I think is impressive because how fast they're shipping. it's not necessarily the most amazing model, but you can see the trajectory, right?
Alex Volkov
Alex Volkov 10:20
So- The jump from- I think this is-
10:21
… Alex Volkov: Spark 1 to 1.1, 1.2, 1.3- Yeah … the jumps are insane. Like the, the speed of jumps- Yeah … are crazy.
Peter Gostev
Peter Gostev 10:27
Yeah.
10:28
Yeah. And that's what you want, right? Because, the models, if you look at DeepSeek, they shipped a major model and then nothing-ish. they did some, some iterations, but nothing big for like a year, and then they were behind other Chinese models, so that's, that's not great. So yeah, I think you want, you want them to ship fast, as fast as possible, so that, that's awesome to see.
Alex Volkov
Alex Volkov 10:48
Yep.
10:48
so I will just say mine. It's, it's absolutely Fable so far. I You can be so ambitious with Fable. I can't wait to show you guys this, but for those who are tuning in- Oh, yeah … you're tuning in to us over YouTube and X platforms that have built-in live streaming. I built a live streaming platform for you guys on ThursdAI.news/live. Maybe we should buy ThursdAI.live and, and just deal with it. I wanna show you this because, I don't know how many people, like, saw my post about this. We, like… It's going to be a little confusing 'cause there's a delayed video of us on there. do you guys see the transcription there? It's live real-time, Muse Live Transcribe, and you can see, names here. You can see the topics that we're going to cover. They are updating in real time. The thing that there's already five people watching here, incredible. I suggested this build to GrokBot. GrokBot said, "Hey, y- we- it's, like, a full few weeks project." I'm like, "Fuck that. We have Fable." So I went to Fable to Cursor, and it's like I went to plan mode, yapped on my microphone exactly for what I want, and then I basically did a breakdown, asked Fable to break it down to three main tasks. Went to Cursor. Shout-out, huge shout-out, Cursor, because they immediately have support with Fable 5.1. we should probably mention OpenAI dropped support from Fable, so GPT 5.6 is not going to go to Fable. but yeah, there's a whole drama from this week. Cuz. And, and it just ripped through this. Now, just like, you'll understand, as far as I understand, live streaming is difficult. Like, setting up, like, this is not a one-shot, like, JavaScript page. There's whole systems in there with syncing, with, with tracking, et cetera. the whole thing is agentic also. So you will see Chiron show up here that GrokBot is, presenting We have a Grok bot. It's listening to a show. For folks who are, like, followed the show for a while, they know that for the past three weeks, we have an agentic producer. This agentic producer now, via the system, listens to the show for us and throws up chyrons and marks… we, we start talking about OpenAI. then when we switch to Anthropic, it's gonna, like, show, show to Anthropic. it's not real time from that perspective, but, like, it's all controlled via an agent. In my attempts to, to be the most AI forward, live news show. So if you are tuning into us, please try to switch to ThursdAI.new/live and see and tell me about that experience. I would love that. This happened over four hours yesterday. Fable just fucking murdered that task with almost no issues. This streams to CloudFl- et cetera. I just… Speechless, folks. Speechless. The, the, like… Do you know how long this would have taken me three years ago?
Wolfram Ravenwolf
Wolfram Ravenwolf 13:47
Did you use Fable just for the planning and then have it
Alex Volkov
Alex Volkov 13:49
implemented in the purpose?
13:49
All of it. All of it Fable. All of it. Fable just, like, just ripped through all of it with three agents in Fable. I don't know how many token. I stop count at, I think half a million, half a billion. I don't know.
Yam Peleg
Yam Peleg 14:00
Alex- Four hours
14:00
… Yam Peleg: y- y- you got, you got to walk through the, the ha- the workflow. Like, like, how exactly did you do the- It's just… You, you, you just yap at it. Y- yeah. Okay, look, I yap at- Now I have no idea what went where the model. Wait, wait, wait, wait, wait … once or twice, and that's it. I'm not asking how Fable did it. Yeah. Absolutely no idea. No idea. I'm asking how you told Fable to do it because- Okay … wait, the, the specifics are, are important because- Yes … Yam Peleg: you said that you did it on Cursor. That's a very important detail. Okay? Yeah. No problem. We just wanna know. Like, you planned on which platform, you delivered what to what platform, because that's a really impressive project to do one shot. So I think, I think it's, important.
Alex Volkov
Alex Volkov 14:39
there was iteration in there, but it definitely
14:41
happened via one sitting. Like, I sat down with the attempt to do this, and then it was done. How it was done, I have no idea. I never looked at one line of code. It did, like, a bunch of PRs, whatever. I just, like, Fable just does things. Oh, and also designed it with Fable 5.1 in cloud design before, and then Cursor, and then, sorry, and then, yeah. Anyway, folks, I think it's time for a TLDR- … 'cause we're g- we're, we're, we're running away, and it's already 9:50 where I am, or, like, 8:50. Yeah. And we have just an amazing amount of news to talk to you about, so we're just gonna walk through the TLDR, because I think that there's a bunch of stuff for us to talk about. Nisten, one comment before we go?
Nisten Tahiraj
Nisten Tahiraj 15:19
I know.
15:19
I just wanna show the Martian sim that Fable 5.1 made.
Alex Volkov
Alex Volkov 15:23
Let's get it-
15:24
Later … let's get, add the actual Fable thing- Yeah that's gonna
Nisten Tahiraj
Nisten Tahiraj 15:27
And it, and then it ran out of context on the company team plan
15:31
in just two prompts on the website. But what it delivered in two prompts was just insane.
Alex Volkov
Alex Volkov 15:36
I'm, I'm excited to look at Martian sim.
Nisten Tahiraj
Nisten Tahiraj 15:37
like, it hit the, it hit the limit, yeah.
Alex Volkov
Alex Volkov 15:39
I'm not the only one, Nisten, by the way.
15:40
We have folks that are waiting for the Mars thing. So shout out to Milos. some Mars rovers on Fable 5.1. Yeah, I- I can't wait to hear from Peter as well. I know you've tested this model. I think you've- He won't be disappointed … won't be disappointed. All right, folks, it's time for us to go to, TLDR. Let's talk about the highlights of the top of this news from this insane week, on September 3.
16:13
Let's get there. So folks, welcome. This is the TLDR. This is the section, if you only have a few moments, this is the only section that you really need to pay attention to. This is everything that we deemed cover worthy from this week. We used to say this is everything that launched in AI this week, but no longer this is true. Not in weeks like this, because these weeks are just denser and denser. but here we go. In big companies and LLMs, what we have is Anthropic launches Claude Fable 5.1 and Mythos 5.1, the same model with different guardrails. apparently the highlight there is 75% cheaper cash reads, and double science benchmark. We're going to cover this, I think, at the beginning of the show for a very, very long time because I think it's, it's, it's well worth it. Terminal Ben jump is crazy. definitely we got all excited about this model, all of us. So going to g- going to cover this. OpenAI just, designated Astra to be the first model to hit the critical cyber threshold. OpenAI is about to launch Astra today, later today, hopefully while we're on the show. I think we're gonna stay live until Astra drops. Astra is, r- was rumored to be GPT-6. Now I think it's clear that Astra is GPT-6 based on the OpenAI, release from today. when you listen to this, if you're listening to the podcast, OpenAI is announcing a GPT-6 Astra, and we're gonna tell you all about this later. Google came back with another Gemini 3.8 Flash, and this time with Flash Cyber. they call it the reasoning and coding workhorse for many Googlers. I am, will mention that this happened. I don't think that honestly this stands up to the frontier labs, of everything else, doing this week. the cyber thing is very interesting, the Fairwind program for trusted cyber defenders. I wonder who's gonna use this. Alibaba upgrades Qwen 3.8 Max, to 09/02 snapshot, and they claim number one on Code Arena. A very interesting claim given everything else that happens today. And it looks like all of the frontier labs besides xAI or SpaceX AI launched this week, something. Elon Musk did say that Qwen 4-- Elon Musk did say that Grok 4.7 is coming out next week. He literally like numbered the days. that seems like a, like a good release. but no, nothing this week. So Meta is the last frontier lab, almost labs, that shipped something. Meta Muse Spark 1.3xHigh is now beating or tying GPT-5.6 Sol and Grok 4.6 on artificial analysis. Meta has jumped over because they have an unreleased max reasoning API that apparently is better than Sol. This is the first time that we see Something like this. And it seems to me that all of the news that I just told you about from Anthropic, from Gemini, from Meta, from Alibaba, from DeepMind, all of them were trying to have their second in the sun before Astro drops. That's how this news felt to me. We had weeks like this before where everybody knows that OpenAI is gonna come out swinging, so-- and nobody's gonna talk about their models. So we're gonna talk about their models until Astro drops. This is just the Frontier Labs, folks. Anything huge I'm missing on Frontier Labs or do we move to open source? All right, let's move to open source. Tencent, open sourced HyPreview, 770 billion MOE. ZAI dropped GLM 5.3. Do you guys remember last week, GLM 5.3 Flash was unveiled as O-O-OX Alpha, and we talked about this on the show? Like a day afterwards, they, they open sourced the main model, GPT 5.3-- sorry, GLM 5.3. they dropped the weights and the weights are out there. Folks are loving this model. and, Obliteration AI dropped an open source, A obliterated model that completely removed the cybersecurity restraints of GLM 5.3. And, I reached out to this lab, and we're going to have a guest on the show from Obliteration AI later on today to chat about why would they do this, and also the dangers or potential dangers of such som-something, and, what does it actually do. Basically, they removed all constraints from this model, so it can do offensive cybersecurity. You can download this model or I guess not download, but via their API, use it and, and start your own agent swarm that could hack into Hugging Face. Wolfram, I have, I have, I'm missing a few items, I think. You wanna add a few items, just verbally to see? Despite folks are using different things, GroqBot launched a bunch of updates, specifically in namely, there's webhooks support now, and shared templates. So I can literally share everything I built with GroqBot with you, and I think that that's very, very nice. And also, I have a bunch of updates for GroqBot as well. Voice and audio.
Wolfram Ravenwolf
Wolfram Ravenwolf 21:10
Wolfram, go ahead.
21:11
Just quickly because, if we are calling this out, there's also Hermes Agents 0.21 new release, which is a big release bec-
Alex Volkov
Alex Volkov 21:18
Literally automatically.
21:20
I fucking love this. For the past three and a half years in Descript, every time I edited the audio, I was like, "Hey, this is Peter. This is Wolfram." Manually, I do this. In 2026, you should not manually be doing this. Now it's automatic. live diarization with speaker attribution. That's what Fable just gave us. anyway, got a little bit carried away here with the excitement. Going back to the TLDR, Inworld launches real-time TTS, generally available, claims number one artificial analysis voice arena. And also, I don't have this here. Why? Oh, I know why I don't have this here. It's here. Microsoft MAI. Oh, I'm not showing anything. Microsoft also dropped MAI Transcribe 2, number two artificial analysis world order, rate, two percent, just two percent. Four time-- Four hundred times, faster than real-time for just $1.67 per thousand minutes, less than half the price of its peers. A very interesting thing, both labs launched real-time ASR with diarization, and I can't wait to compare both of them this is like a huge week for, for, for agents and, and the speed, I… Mm-mm, next, next. We're moving, we're moving. Sorry, I have so much to say about this, but we're moving. vision video, a hottest charagor- category besides LLM. Folks, this has been an insane week. I know we try to cover everything. okay. World labs. Do you guys see this fucking thing? World labs un- unveiled Atlas multimodal- Yes … world model with pixel perfect camera control and 3D reconstruction. We have to show you examples of this. I know that this is also a podcast, but this is just absolutely, absolutely insane. Just like this is what we… When we talked about world models, this is what we meant, and just like mind-blowing. Runway, Andrej Solaris. Do you guys see this? This is a continuous interface that's been built that you can click into things, and then the whole world model reacts. So basically, a world model for interfaces. for those who have no idea what I'm talking about, imagine a picture of your room, and there's a lamp, and this is an interface. So when you click the lamp, the light goes on, and then the room is lit based on you clicking the lamp. You can click everything. You can drag every- It's just ab- Just makes no sense, besides the fact that we're in the beginning of singularity and this is what we get. I'm super excited. And also Fal H3 Max Week. Fal, first of all, did an infinite, stream, built a new model for this, got banned from Twitch and Kick. Then they built fal.live in the weekend. Which, by the way, got what got me inspired to build our live show experience as well, because they just, like, shipped it. I was like, "Yep, we live in this world now. We can just ship things." then they shipped H3 Max Turbo with one cent per second. It's like super cheap, super fast video generation. and then, and then Wolfram, you mentioned another video model that's like a world model, right? Like, H3 Max became a world model or something?
Wolfram Ravenwolf
Wolfram Ravenwolf 24:03
H3
24:03
Vault, yeah. H3 Vault, yeah. It turns it into a controllable world model. Like, you can walk through the sc- scenery and stuff.
Alex Volkov
Alex Volkov 24:11
All right.
24:11
What? Casual. Okay. All right. Yeah, yeah, dude, this is like… I'm telling you, this week, it's very hard to pick and choose which news we're gonna talk about. But I think first and foremost, because, Thursday, I was born on March 13 when GPT-4 got dropped, folks, we're gonna start with Frontier, and we're gonna talk about Frontier before, everything else. let's go to Frontier AI.
24:43
Welcome to Frontier AI, our corner on ThursdAI, where we talk about the Frontier AI, mostly LLM drops. And, this week has been quite something, hasn't it? We started with-- Actually, don't know what we started with. I think the most important one to talk about is Fable 5.1 and Mythos 5.1, so let's start there. let me just do the announcement really quick. Anthropic launches Claude Fable 5.1 and Mythos 5.1. The highlight, for some reason, is 75% cheaper cash rates, but I don't think that this is the highlight. Available today via API and Claude Partners, Fable 5.1 scores 55.8 on Terminal Bench 4. There's Terminal Bench 4 now. there used to be 3 last week, now it's 4. up from 42%, so almost a 10%, a 15% jump from Fable 5, which was already incredible. Terminal Bench Science, a new benchmark, Fable claims 52%, a state-of-the-art score and more than double of Fable 5. So th- these 0.1 releases make no sense because the jumping capabilities are crazy. Cash read drops, although this model is more expensive. they, they say the typical workload costs 25% less, but I doubt that I see that as well. obviously the input and output stay at $10 per in- per million token inputs and $50 per, per million output tokens. Definitely worth the $200 if you wanna play with this model. let's see what else is interesting here. compared to GPT 5.6 Sol at 37% Terminal Bench 4, this model is at 55. This is just like a, a huge, huge deal, on multiple benchmark. This is state-of-the-art until today. Hopefully we'll see some dethroning, but Fable 5.1 is definitely, definitely a massive, massive drop, as we talked about in the beginning. would love to hear from folks. Who wants to, who wants to f- f- Fable guess first?
Peter Gostev
Peter Gostev 26:43
Maybe I'll, I'll start with the reason- Yes, please … score.
26:46
So we tested the max version on the code, arena, which does more front-end tasks and it, came in first. and not just first but, like, by a large margin and they did this already before with Fable 5. Then they, got equaled by other labs. So Qwen did well as well. But, yeah, now they jumped way ahead so which is very cool to see. I've been testing myself as well on a bunch of kind of front-end generations and yeah, I spent a lot of time. it was interesting. I also noticed the, the, the kind of the highlight that your summary did about the 75% cost reduction for cash read and I thought, "Oh wow, that, that's cool. We must see big cost reduction." I don't know when I was testing it and I looked at the cost it was like $40, $60 per generation and it's like it is cool but man it's, it's expensive when you see like even GPT 5,6 Sol was something like $3, $6, $10. So it was heavy, heavy price. But look, I think the reality there is that a bit like what you were saying with, with the demo that you did, not even a demo, right? The shipped product Okay, it's, it's probably cost you a lot of tokens, but how else you gonna do it? Yeah. So I think when you have worthy tasks, it's totally worth it. Just p- pay the money. But you just need… The, the bottleneck is now you. Find, find good tasks. So that, that's the hard part now.
Alex Volkov
Alex Volkov 28:13
I, I, I think it's kind of like when you hire as a engineering
28:18
manager, you want the best people, and you're willing to pay the ridiculous prices in San Francisco, let's say, for people to come in and be the best people. It's the same for me. When I know a higher tier intelligence exists and I somehow have access to it, I really want that tier of intelligence because there may be stuff that I don't know that it does, and there's a lot of stuff that I don't know that it does, that, that it's better at than previous upgrades. w- we definitely… I, I do wanna talk to you folks about this. We definitely talked about Opus 5 not being that. Opus 5 came out. It was jargon douching. It was really hard. It was a great coding model on benchmarks, but definitely the vibes were off. W- with both of these post-trains, by the way, with Sonnet 5 and Opus 5, I don't think there was as much excitement. I remember Opus was, like, the darling of the industry, and now it's no longer. And I think the speech was one of those things. And they, Anthropic actually acknowledged this with the release of Fable 5.1. They called it mannered prose. Folks, we gotta learn a new thing, a new keyword, a new magic incantation that makes our AIs better. This is called mannered prose. If you now ask Fable to, "Don't use mannered prose, please," it will stop saying words like load-bearing and control pain, and, like, all the other, like, Fableese, Claudese bullshit that we, we told you about. but regardless of this, it's the best writer that I've ever used. It's absolutely the, the best writer that it's ever used. It is AI writing. You can feel it a little bit, but it's just, like, so concise that, that I really enjoy working with this model in terms of writing. Design also is absolutely batshit. If you go to Claude Design- Just the stuff that it's able to do is, i- i- is quite incredible. overzealous a bit. Does too, too many things, trigger happy. but I think like the, the ra- latest agentic model, they're just like prone for go, go, go. So I definitely felt that it went and did some stuff. how should I say? There's now an exercise with Fable 5.1, I'm assuming w- with the next model of OpenAI, there's an exercise for you, the reader or the user, t- t- to, to control the machine and not have it do simple tasks. Because sometimes it would go, "Hey, plan this huge thing for me." It would go and do this, and then I would ask it like a follow-up question and it would use the same intensity and the same brain and the same depth and necessitating and all kinds of stuff to answer a simple question. Like whoa, whoa, whoa, dude, I just, I just had a comment 'cause I feel like we're colleagues here and I'm just like sharing something with you. You don't need to go all that. You don't need to do all that and write a bunch of scripts to, to execute. this is on us to like know the model will go hard, especially on ultra, et cetera. any other feelings, topics, Nisten, you wanna show off the, the, the, the, the, the Mars simulator? Let's go because it, it went, it went above and beyond. Again, this is just, two prompts. So can you guys see this? It, it even- not yet. Hold on. We need to add this to stage and then go like this and then like that. All right. Okay. So- So Nisten, walk us through what we're seeing here.
Nisten Tahiraj
Nisten Tahiraj 31:20
All right.
31:20
So this is, you launch a whole bunch of, mass from Mars and, you have to build a maglev launcher and you launch it from the highest mountain on, on Mars, which is Mount Olympus Mons, and there's very little atmosphere at the top And we've been doing this test with every model for almost two years now. Yeah. Here's what Fable did in two prompts. It added a NASA mission planner, which will figure out the most optimal fuel and, and stuff, and then it will rebuild the entire launchpad based on that. And, it also has now all of, all of the other planets. So it's not, again, it's not perfect.
Alex Volkov
Alex Volkov 32:00
I will try- Oh, I hear
32:02
sounds.
Nisten Tahiraj
Nisten Tahiraj 32:03
Yeah, yeah, yeah.
32:03
There, there, there are sounds. So I will try and, cut the audio for this. I forgot how you actually cut the sound, but I can just mute the tab if it gets too much.
Alex Volkov
Alex Volkov 32:14
No,
32:15
I think it's fine.
Nisten Tahiraj
Nisten Tahiraj 32:15
There is, there is the entire…
32:17
It's built the entire solar system.
Alex Volkov
Alex Volkov 32:19
It built the solar system?
Nisten Tahiraj
Nisten Tahiraj 32:20
Yeah.
32:20
Wow. It will land it on Earth with a parachute. You can go and land on Venus if you want too. Really? Yeah. I was playing this- What? … for, for two hours. And, here, here. Y- we, we, if we wanna add some more- What? … some more fuel to it, we can do that. We can speed it up and then, and now-
Alex Volkov
Alex Volkov 32:40
Nisten, just to confirm, the same exact prompt, right?
Nisten Tahiraj
Nisten Tahiraj 32:43
same.
32:44
I added one more prompt after that and I said, "Make it like the Kerbal Space Program game-" Ah, okay … "and have the other planets." Yeah. So now we can skip to, to the next events. And, it will take maybe, like, five minutes or so. So you sped the time up by 100,000, and then we, we can speed the time out again. We can just keep zooming out and, w- we'll actually, like, see the orbit in real time as, like, the Earth and the Moon gets, gets closer. And then it will do an atmospheric burn on Earth. So we see now the Earth and, it's coming closer. Wow. So now, now we're, we're gonna zoom in here.
Alex Volkov
Alex Volkov 33:23
Wait,
33:24
we're gonna land on Earth? And we're gonna see- Is that what you're saying?
Nisten Tahiraj
Nisten Tahiraj 33:26
Oh, yeah.
33:27
It's gonna do, like, actual calculated landing. There's an autopilot now which is accurate. we, we can look at Earth. Like, we can actually go here, and we can see the Earth, and we can get back on the spacecraft. so it allows you to just skip, so now we just skip- Dude
Alex Volkov
Alex Volkov 33:45
the time.
33:45
The Earth is insane. It's, like, all fully textured. What the fuck?
Nisten Tahiraj
Nisten Tahiraj 33:49
Yeah.
33:50
Yeah, so n- now it's gonna do… Pretty soon now it's gonna do… Oh, and, and you can see here the, like, the, the periapsis actually shrinking in real time as it's gonna do that. So as it gets closer to the Earth's atmosphere, it's gonna hit the atmosphere, and, it's, it's actually gonna do, gonna try and do, like, a landing burn on Earth. So this won't take too long, maybe, like, another… So now we hear the sound again. Yeah. And if we go close to the aircraft we're gonna see that it's hitting the plasma, so now it's heating up. And, it… So we go back to, to, to, to real time speed, and, and y- you're gonna see the… Yeah, so sometimes it shrinks the orbit, sometimes it crashes, o- on Earth. You, you know how these things go. And, and yeah, we can just s- oh, it opened the parachute. So-
Alex Volkov
Alex Volkov 34:42
No way.
Nisten Tahiraj
Nisten Tahiraj 34:43
So now, now it's opened the parachute.
34:45
And, yeah, so it's only 9 or 10 kilometers up. And, we can speed this up more until it, it falls down to the ground. And,
Alex Volkov
Alex Volkov 34:53
This is, this is insane.
34:54
This is nuts, right? This is nuts, man. This is another level of nice. Bra, bra, this is-
Yam Peleg
Yam Peleg 35:01
And this is single shot?
35:02
There. You're telling me this is single shot?
Nisten Tahiraj
Nisten Tahiraj 35:03
Oh, we just landed.
35:03
Let's go. This was,
Alex Volkov
Alex Volkov 35:06
this was two shots.
35:07
Two shots. Two shots. Ah. I can, I can even show the prompts. eh, the first time I, I did it, it wasn't working quite correctly. But again, it's calculate how long a Mars rover… the stuff that I added, like make the bumps on Mars accurate, and then it said, we're not gonna have enough memory," so it reduced the polygon counts and, and the textures for it. It, yeah. Wow. And then it, it ran out of the four-hour context on the company's team plan in those two prompts. Yeah. So it w- it wasn't fully able to finish it. I don't know what else was in there, but, yeah. Yeah, no, Fa- Fable, Fable goes hard, for sure. Folks, we have breaking news from Runway and, I think given this week, we, we need to go there. let's go breaking news. This is not the breaking news you're all waiting for, okay? GPT-5, 6 Astro is not dropping yet, but this is breaking news, so let's go. AI breaking news coming at you only on ThursdAI.
36:12
All right, LDJ, you broke this one, so feel free to go ahead.
LDJ
LDJ 36:17
Yeah,
36:17
about- It's Run- Runaway, Runaway ML. Yeah … which popularly makes, like video generation models, and they've been going into, like world models that you're able to kinda move and spatially in- interact with more in real time. they are announcing something called GWM Worlds 2, And now this has 48,000, hertz audio. Oh, wow. Which a lot of the world models lately, some of them have ha- had audio, but it's like very low-resolution audio. It often has a lot of artifacts, and like very obviously AI audio. yeah, this is like a pretty significant jump it looks like. Here, this video that Alex is playing right now, for the people listening only on audio, it's just, it kinda looks like, I don't know, like a post-apocalyptic or desert scene. Yeah. Somebody with a spear and, and just trying to survive, and some fire and a tent, and yeah.
Alex Volkov
Alex Volkov 37:10
Man, looks like a, looks like a game.
37:12
like D- Yeah I think we get GTA VI before GTA VI comes out. Yeah. I think, yo, this looks insane. I, I don't, I don't w- see a way for us to play with this, but this generates speech as well. So yeah, definitely approaching GTA VI. this week also, there was a , there was a drop from GTA VI, video, so we'll see how close this is getting. but let's play a few segments here with the audio. I think the audio is the, the biggest part here, right? Like this is the, this is the thing that, that is interesting. so I'm sharing. It may be loud, folks. I don't know if I can control the loudness of this. So this is, 24 frames per second live world, that you can direct the scene. And you can hear the… Hopefully, you guys can hear the campfire going. let's look at speech. So they have actions, navigation, and speech, three tabs here. Let's look at speech. AI: Excuse me, do you know where the nearest transit station is? Just continue this path for two blocks. You'll see it on your left. Thank you so much.
Alex Volkov
Alex Volkov 38:10
you guys heard this?
38:12
Yeah. Should've asked her how to reverse a binary tree as well. It's not amazing, but given that this is a live generated world model where you can walk around, this is insane. This is absolutely, this makes no sense, insane. Also launched Solaris. Did you guys see Solaris? This is Runway Solaris. It's a world model for interfaces. Let me just show you this insanity real quick. this is just also one of those mind-blowing things that we haven't seen before. A world model for- Interface world models. What does this mean? It means they generate a video of interface. Let me zoom in here. And then, they generate, let's say, this, this guy, and you can drag every item in this photo on top of everything else and see stuff like that. So you can drag his shoes and see him wearing different shoes. This is a dude that's standing. You can drag shorts on top of him, and this is just the interface. another interface, this is like a home interface where you can just play with the lights and move things around. The… Are you guys realizing, like, what we're seeing? This is a scene of a living room where everything is clickable, everything is movable And I just, I'm not fully sure how to react to, to this, like, thing.
Wolfram Ravenwolf
Wolfram Ravenwolf 39:29
I think this is the future of the interfaces.
39:31
Why? Like, when we're making videos now, we don't have to render 3D anymore. We just create it with generative AI, and this is for interfaces. And in general, we could do this for everything. Like on your computer, the whole UI could also be just generated on the fly if it was fast enough, and then you could change it any way you want it to. And I think we will be seeing this, as long as we need to work on computers.
Alex Volkov
Alex Volkov 39:54
I,
39:54
I think that, that Solaris is one of the highest importance news of this week because of the interfaces. Just because of the way that, learning is done for kids, for example. We play with the world and we learn from that, and, just absolutely mind-blowing kind of utility. Again, with Runway, Cristobal, the, the CEO, please let people play with this. Like, why would you, like, launch something this revolutionary and not let people just, like, even try? I know GPUs, et cetera, but talk to us at CoreWeave. We're gonna, we're gonna do some GPUs. the natural behavior is, when somebody coded something, 71% of people preferred the natural behavior versus a coded behavior in different, like, UI elements. 'Cause people explore by touching and clicking, and just absolutely crazy. all right, we covered the two things. Let's go back to Frontier Labs, folks. There's a lot more to cover there. let's see what we haven't talked about yet. So obviously, Astra is still upcoming. let's talk about Google Gemini in 3.8. Let's go back from, breaking news. This is breaking news, by the way, from Runway ML. We covered the two releases they have from this week, just two days after the world model. let's now go back to, Gemini. What's up with Gemini, folks? They're trying to, to catch up, with another flash model. Mm, we got promised a pro model in June, and then it looks like they're not shipping the world, the world model. sorry, the, the, the, the, the pro model. Gemini did launch Gemini 3.8 Flash. This is their third week of flash, third iteration of flash models, three weeks after 3.7 flash, so upgrades is, are definitely there. As a reminder, the DeepMind folks have bought anti-gravity, FK, Windsurf, and all of Google's folks are using that internally. So they're building those models on top of that coding. What do we think, folks? 3x faster. Th- these models are definitely fast. Wolfram, you usually use Gemini models.
Wolfram Ravenwolf
Wolfram Ravenwolf 41:54
Will, will you use this one?
41:55
Yeah, I use them a lot, and every little update is immediately useful in my home assistant setup and on my mobile phone as my assistant. But still, where's the pro model? if they are all forced, allowed to use anti-gravity and, they are so much faster now, I hope that shows in the model as well. Yeah. So yeah, it's another Flash model. it's nice, but give us some pro models.
Alex Volkov
Alex Volkov 42:16
Come on.
42:17
It's, it's nice, but I think that this is now where we are. I'm not discounting Google. As, as you guys remember, let's talk about the specifics here. 1 million token context window, at very, very fast Flash speed.
Wolfram Ravenwolf
Wolfram Ravenwolf 42:28
and then, bring-
42:29
One thing about the context window, because a million is a lot, but if a model has to think so much to achieve the intelligence it has, and we know that the Flash model is one like this, so it is, very intelligent on the benchmarks because it's thinking so much. So you need the l- large context because so much is being used for the thinking, and yeah, that makes it less valuable than for, a smarter model that doesn't have to generate as many tokens, where you can use much more as a user of that context.
Alex Volkov
Alex Volkov 42:57
Yeah.
42:57
Wall Street Journal reported that Google scrapped the 3.5 Pro checkpoints. So the 3.5 Pro model that were, like, promised, Flash models advanced so fast that they just overtook them. So essentially what we have here is what we were promised with 3.5 Pro, I'm thinking. And Gemini 4 is next, but it's still in post-training. And I think we saw Logan, like, talk about this. Gemini 4, it, it, it better come out. But yeah. we have, folks from the comments, DevDoc saying, "You're not allowed to use any third party harness with Gemini. It's so painful, n- to not have any decent pro model from Google." Yeah. That's a very s- interesting thing as well. I will say another thing that they launched, the 3.8 Flash Cyber, which is a dedicated cybersecurity variant, but no benchmark scores appeared in the evidence that we saw. all right, I think we've covered Flash enough given the, the, the importance of this. Let's talk about Qwen 3.8 Max. Alibaba updated the, the Qwen like API model, very boldly claiming number one WebDev leaderboard. I don't think they are number one WebDev leaderboard. I'm not sure which WebDev leaderboard they are referring to. but Qwen 3.8 Max. Qwen Max was always their kind of like frontier. We're a frontier lab also, not only open source. I don't know if tons of folks are using their, their models on top of frontier models. Do you guys know of anyone? No? I don't think so as well. folks are nodding their heads, forgetting we're a podcast. okay. So everybody's nodding their heads. Nobody knows about a person that uses Qwen 2.8 Max, on the API, but I'm, assuming they also have a very strong, like, Chinese, crowd and fans for their model. Oh,
Nisten Tahiraj
Nisten Tahiraj 44:42
I, I, I, I did know someone.
44:44
It was, some person in the US, he just worked as an IT worker, and he would just run his, still his OpenClaw, on that. But, yeah. it was mostly on stuff that they didn't care. I know people are using it to generate data sets.
Alex Volkov
Alex Volkov 44:55
Yeah.
44:55
So that's, that, that's one thing. to get reshuffled for sure. or we'll see, But, like, I, I doubt they won't. But still, this is the first time that Meta, after a very short time rebuilding the whole lab, is jumping over OpenAI on this, jumping over Microsoft MAI, jumping over Google. Meta is leapfrogging Google significantly on this benchmark now, which is absolutely crazy.
Wolfram Ravenwolf
Wolfram Ravenwolf 45:18
And what excites me so much about this is that there will
45:21
be a release of this as open source. I'm not sure which version, if there are different sizes or something like that, but I hope we get this version. Maybe we are not getting the Max version. Okay. If it's a different model, actually, if it's just, yeah, it probably must be something like that. But yeah, I'm excited to get a Western SOTA model from open source. if this is in that segment, that would be so amazing.
Alex Volkov
Alex Volkov 45:45
Yeah, the whole thing is, when Zuck is posting about this,
45:48
when Alex Wang is posting about this, they're saying, "Hey, open weights are coming soon," and they are teasing a mystery model, Watermelon. LDJ, I wanna hear from you on this.
LDJ
LDJ 45:58
Yeah, when it comes to the, the speed as well, it's running at about
46:02
three times faster than Fable and Sol are, and according to Artificial Analysis speed measurements at least. Yeah. And for, for their tokens per second. And if you look at their cost per token too, it's about four times lower cost than what Sol is currently at, which Sol's already cheaper than Fable.
Alex Volkov
Alex Volkov 46:20
Yeah.
46:21
And I love that somebody in there is finally coloring the charts correctly, where the left column is their scores, but the high-- the top scores are, are colored in, in the right way. And we can see that this model is state-of-the-art on long context. MRCR 256 and MRCR 512 to one million, they get 98.5. So this is maybe the first model we're getting with almost saturated long context, benchmarks, which means that this model with one million at those prices is absolutely a steal. Now, folks in the comments are saying maybe it fails benchmarks. I haven't played with this model, fully yet, but I definitely, like, have plans to SWE Atlas, they're beating everyone, and Thermal Bench 2.1 is 88.8% matching GPT 5.6o. This is a significant change in the order rank of Frontier Labs, right? We're now at this horse race where every week we get like a, a, a bump, a bump, a bump after a fairly chill summer, with the last two weeks being, like, very… Like, everybody took the last two weeks of summer vacation to get ready for the Singularity folks this week. This is all from this week. Like, these jumps is in state of the art, all from this week. Now, the question is, who uses Meta Muse? Exactly. You guys here, raise your hands and see if you're using Meta Muse.
Wolfram Ravenwolf
Wolfram Ravenwolf 47:49
I will use it soon, and there's also a contributor, tier-
47:52
Tier … something like that- which I have seen, which is very interesting. Like, I'm-- My plan is to have a bot using this model, and I can use it for open source development and, save a lot on tokens, especially if it's good in coding, while my personal agent is still using another model, so I don't share my private data with everybody.
Alex Volkov
Alex Volkov 48:10
So you are completely right with the contributor tier.
48:13
I will show this here. If you're okay with Zach training on your data, and let me say this with a little bit of a jest, but still, I think it's true. If you have a Facebook, Instagram, or WhatsApp account, you are okay with Zach training all your data. Maybe not WhatsApp, but this model contributor tier is 10 cents per million tokens. Again, this model that matches GPT 5.6o has 1 million context window with 98% long context MRCR score, is 10 cents per 1 million input tokens, 20 cents per million output tokens, and cached input is at… What? Is at t-t-two tenths of a cent. Basically, intelligence too cheap to meter if you're willing to sell your soul to Zach for training. Now, why is this model so good? Is because, Meta decided that half of their actual staff engineers, instead of writing code or instead of babysitting agents, they will label and do a label shop. Very interestingly, Alex Wang came from a label shop called Scale AI, came into Meta and half of the engineers in Meta are now, like, basically labelers. High quality data is the most important. All of Meta is using this. And very interestingly, all of Meta is using, like, Fable as well. They are training on that internally. And there's also rumors that the upcoming Anthropic IPO, if Meta switches everyone internally, all their coders to Meta Muse, this will significantly hurt Anthropic's revenue because Meta is a big part of Anthropic's revenue So comments, MetaMuse?
Yam Peleg
Yam Peleg 49:55
Is it, is it good?
49:58
That's, that's the question. Like vibe good because- So- I don't know. it's our- Yeah, I haven't tried
Alex Volkov
Alex Volkov 50:04
it.
Nisten Tahiraj
Nisten Tahiraj 50:04
I, I really appreciate that they included a contributor tier.
50:10
So for things like medical data sets that I make, I do need a verifier model that's something different from Fable or, the Chinese models that are trained on Fable. So I, I really appreciate that they included that cheap pricing. That's very helpful to open source and the community and stuff. I, I, I don't know if it's good. I know they have the GPUs. We haven't tried it yet.
Alex Volkov
Alex Volkov 50:37
They have the GPUs and they have the data, and DeepSwe, DeepSwe Long Horizon Agentic
50:42
Coding, the one benchmark that usually goes against the grain, where everybody else is like, pretty much the same across the board, Fable's the best, Opus is the second, et cetera, and then, GPT. DeepSwe is the one that we felt like it represents what we feel like. DeepSwe 1.1, it beats Fable, GPT 5.6, so not Fable, and Opus5 at 59, sorry, at 75% on DeepSwe
51:03
1.1. The model is good. I haven't played with it until, I almost played with it, but like the jumps are, I think, the most significant. The jumps from MetaMuse Park 1 to 1.1, 1.2 and now 1.3 are just a hell of a jump. Let's just take a look at between the last two models. The jump in MCRR, look at this. From 66% to 98, from 55 to 98. Yam, we, we always talk about, okay, benchmarks may not represent how we feel, but they're the, they directionally correct. And so directionally, they are advancing really fast and this is the smallest model. LDJ, thoughts on this? I know we talked about, with you about, like, this being the smaller model, and I know y- y- you had some, like, interesting things about, Meta. Would-- Did you use this? Would you use this?
LDJ
LDJ 51:55
So I did try it a little bit in just some basic, like, interface to building
51:59
tests, see kind of its taste, and, some, some other more creative writing things. It is pretty good. it's just like with all the models, like, I'd say there's some strengths and weaknesses and, and definitely nuance in the type of taste it has, just like Sol has compared to Fable and others. Yep. but it seems like at least for a lot of the benchmarks it's shown it's not, not bullshitting there.
Alex Volkov
Alex Volkov 52:23
Yeah.
52:24
Yep. And we're going to get this open source. I think this is a highlight, right? Like, we're getting a model at those frontier levels and, and supposedly this model is coming open source. they did previously release a Meta Muse… Was that Spark? No, th- this is Spark. A Meta Muse, Glimmer, as open source. Glimmer, right? I- I don't know how I keep up with all these names in my head, but yeah, this is, I think, the, the, the wi- the right one. okay, folks, we're still waiting for GPT-5, GPT-6 to drop very soon. But meanwhile, I think that this is the highlight. Let's, let's talk about another thing from Meta. this is, th- this has been the frontier thing. Meta has dropped also, the Transcribe They call it voice, voice transcribe. Meta Voice Transcribe. It's an audio model for transcribing speech-to-speech and telling speakers apart. I think that's a very, very interesting one. let me show you a practical example, that's kinda like was built by Fable. on the new ThursdAI.new/live page, which you can see right here we are showing a live transcription, both in the video here and also in kind of the live transcript here. And the thing that I wanna show is that usually transcription is, like, really bad and speaker diarization is really bad. But here, this picks up Yam and picks up Alex, and also picks up Alex Wang and Scale AI and Anthropic. And this is the highest tier transcription that I've seen, not to mention it's live with diarization. In terms of the, the, the picking up specific names, this transcription is nearly perfect. It picks up stuff like GPT 5.6 Sole, and it picks up billion parameters. It knows when they say Zack that it's Zack. It's honestly crazy to see the quality of this transcription being live. The price is just absolutely crazy. It's eighteen cents per one hour. So I can transcribe all of the show today, which I'm assuming is gonna be three hours, for less than half a dollar Now, yeah, Wolfram, go ahead.
Wolfram Ravenwolf
Wolfram Ravenwolf 54:34
did you give it a system prompt or
54:36
a dictionary with the terms? Or did it just understand it by its own?
Alex Volkov
Alex Volkov 54:41
So that's a great question, and we have a question
54:43
from the comment as well. Can it transcribe ThursdAI correctly? And yes, absolutely it does. and I did give a dictionary. So there is a way to adjust it. and we gave it the dictionary of the top terms. Fable wo- worked with me through this. I'm gonna try to find ThursdAI. Yeah, there we go. Yeah, it says ThursdAI right here. I'm not showing you guys. Let me show you. So in transcript, you can see that it, like, it transcribes ThursdAI correctly, which is great, but also transcribes everything else correctly. And it, it, it is biased towards a specific dictionary. And 3% of the words that you will say, it won't understand. But I think the highlight here is diarization error rate. It only fails at 17% of the time when one of you talks over me. diarization in real time is insane. I don't know how… this whole week is insane, so I can't keep saying the word insane. But this is fucking insane, okay? This model understands who speaks when, and it sends labels in real time at 18 cents per hour, making it possible to build interfaces like the one that I just showed you for ThursdAI Live, o- obviously Fable built this, where it says, "Alex Volkov's talking." And then when Wolfram interjected, it says, "Wolfram interjected." And before this, we had Yam and LDJ. Like, it, it matches all of those people to the right label. The top editing software that I use to edit the show, called Descript, I do this manually, and it's really bad at transcription, significantly worse than this. maybe, maybe I'm rebuilding the next Descript with this. This is crazy in terms of robots understanding who speaks when. This is crazy in terms of… I, I don't know. We need to bring Quinzel back. You guys don't seem s- super excited about this. O- only Wolfram is. Because y- you know that you can put this on a Ri- Richie Mini and… So stay tuned for that. Let's go to this week's buzz, just before this to talk about the sponsor. AI: In this week's-
Alex Volkov
Alex Volkov 57:03
Welcome to this week's Buzz, a corner of the ThursdAI where we
57:07
talk about the company that makes it all possible, Weights & Biases from CoreWeave. And we have a few exciting news for you today, specific from us. First of all, Wolfram, why don't you tell them about Fully Connected that's coming up?
Wolfram Ravenwolf
Wolfram Ravenwolf 57:22
Yeah, that will be a chance to meet us in person, in San
57:25
Francisco in the Moscone West Convention h- Center Hall, where- Moscone South South. Oh, South. West was Ai.engineer?
Alex Volkov
Alex Volkov 57:33
Yes, Ai.engineer
Wolfram Ravenwolf
Wolfram Ravenwolf 57:34
was in West.
57:34
It's a little bit, Still mixing it up, yeah … yeah. yeah, it's, Wait a second before I say something wrong. let me just check. So it's September 29th to October 1st, Moscone South in San Francisco. Three days, 32 sessions, and Alex just put in a code which we are not, mentioning, because you have to- We're
Alex Volkov
Alex Volkov 57:51
not mentioning in voice, but you have to be here- Yeah
57:54
in the live show-
Wolfram Ravenwolf
Wolfram Ravenwolf 57:55
Watch
57:55
the show … to get this code. or also, track our show from, the podcast, and newsletter to get the show- What exactly does the code give the people?
Alex Volkov
Alex Volkov 58:02
you get a 100% free ticket because you're following us live.
58:05
100% free. which the- And you can- The, the regular ticket is $1299. Yes … early bird is done. So please, please join us. We're gonna do a live show from there. Fei-Fei Li from World Labs is one of the keynote speakers. We just… We're going to talk about World Labs Atlas, very soon on the show, and it's gonna be very, very interesting to hear directly from Dr. Fei-Fei Li, the grandmother of generative AI, building world models, where everything is going right there and then.
Wolfram Ravenwolf
Wolfram Ravenwolf 58:31
And if you are more for the visceral parts, there's also
58:33
the battle bots where you can see, yeah, robots fighting robots, basically. Yeah. So we have something for everyone.
Alex Volkov
Alex Volkov 58:40
And then just before that, we have CoreWeave Hacks.
58:42
if you are in San Francisco on September 12 or 13, come through, sign up. CoreWeave Hacks Agent Loops and Hackathon with Weights & Biases, AGI House, and TypeSafe AI. the prices are going to be ridiculous for this one. we're going to give you… First of all, y- you have a ticket, a chance to attend Fully Connected, but also, the, the first prize is you're gonna get a F1 ticket for most production-ready, design hack. Like, g- come hack, and there's a chance you're gonna get to go to Formula 1. That's crazy. you also get a RoboDog, over $20,000 in prizes, and, these are always a good time. So if you don't know where to be on September 12 or 13, please, please join. we have Joseph, Joseph saying that, "Can't wait for Fully Connected." Joseph, we can't wait to see you at Fully Connected. Please join us- Yeah … as well. This is awesome. the last thing that I'll say before we jump into the interview, we just launched a bunch of models, and I think the highest tier one on the CoreWeave Inference is Kimi K3. On Dedicated Inference, it runs, it purrs like a kitten on GB300s, which we have in CoreWeave. go check out Kimi K3. we can give you a few, r- reach out to us. I think we're gonna do a promotion for, for tokens on the main, on the main social media accounts for CoreWeave and Weights & Biases. runs, purrs like a kitten. Really, really fast. a great model. All right. This has been this week's buzz. And, let's move on to the next exciting part of our show, which is not the GPT-6, but is an interview. I think that I, I'm going to do this one. Nisten and LDJ, Phil. You know what? Here's the first thing. Let me first introduce our guest. Wolfram, I'm gonna take you down for just a second. let me first introduce our guest, that goes by Obliteration AI founder, without a name, and there is a reason for that. But can we do a voice check?
Abliteration AI
Abliteration AI 1:00:26
Hey, how's it going?
Alex Volkov
Alex Volkov 1:00:28
Hey, great to hear from you, an unnamed person.
1:00:31
And just for funsies, when I invited you to the show, I said, "It's okay for you to be anonymous." And I said that one of our co-hosts i- is a space cat, so LDJ, I would love for you to join this interview as well, as I'm gonna be between two anonymous accounts. And I think that that's gonna be dope. let me first set up the scene and the reason for this interview, okay? just this week, in a w- what one would say a very viral post, Obliteration AI, a lab that is new to me, I haven't heard about Obliteration AI, released a, obliterated model Large V2. A model that removed refusals from GLM 5.3 from ZAI for offensive cybersecurity. And, you, you can see a little bit of kinda the, the details that we did get from here. Obliteration AI, removed, different, re- restrictions and, this got a lot of, how should I say? D- descending opinions from folks saying that this is absolutely horrible for the world, that this is going to break everything. "Why would you do this?" to other folks, they're saying, "This is 100% honeypot. There's no wa- there's no way they're hosting this." Just like opinions from all over the place. And we've heard opinions like this before. We've, we saw folks saying that GPT-2 is dangerous to release, et cetera, in open source. And we're pro open source. And, I, when I saw, h- how should we how should we address you, Obliteration AI founder? when I saw you post about this- Go ahead
Abliteration AI
Abliteration AI 1:02:08
No, go ahead.
1:02:08
no, no, finish. Yeah, so when I saw your post about this, I was very interested in, in hearing, of directly from you and not the, the, the folks on. So first of all, w- what did you release? Could you tell us about kind of the release? Is this open source? Is this open weights? Is this API only? Like, where can the folks find this obliterated, GLM 5.3 version? Cool, cool, yeah. Thanks for having me. yeah, so first, like we didn't, invent the obliteration technique. I, I'm actually surprised by all the-- I'm surprised by the surprise, a little bit, because, it was a white paper that maybe came out like, three years ago or so, the ability to find and, change these refusal vectors. And, so yeah, like you covered it around, yeah, we, we released an ob- obliterated GLM 5.3. And actually, believe it or not, there's already a similar… we did an open source one because there's already so many people that do this obliteration and then just, release them on Hugging Face for- Yeah … anybody to run, anytime they want. And, so yeah, so like if you go to Obliteration AI, you should get some free promotional token and any, promotional credits and anybody can try it. that kind of, I think covers your initial question.
Alex Volkov
Alex Volkov 1:03:12
100%.
1:03:13
and I think I was also surprised by the response, specifically around the fact that, like you're saying, a lot of obliterated models exist on Hugging Face for all of open source. Many people do this with fine-tuning. Many people do this with, with, ba- base weights of different models. and you guys decided to make this a commercial product. I think that this is the, this is the thing that maybe, got attention. I also saw that both Yam Peleg and Nisten got requested to join this interview. so I will let them also like ask some questions, but I will warn, Nisten and Yam, we don't do gotcha journalism here. this is not what we're here for. And, I would love to hear from you a question for Obliteration's founder, around that. Yeah, go ahead.
Nisten Tahiraj
Nisten Tahiraj 1:03:55
yeah, so I do actually appreciate this existing as I, I work with, social
1:04:02
workers in different countries, doctors, people that deal with the therapists and stuff, and they do need to record every single thing very accurately that someone like an abuse victim has, has suffered. And there's no way to do this without obliterated models. I also come from a former security background last decade, and I understand that Metasploit is actually crucial to improving security in companies, because otherwise you would concentrate these harmful tools in only the, the bad actors, and you would not let the good actors actually take care of them. So I understand all that. I know the pu- the public doesn't I want to know, what do you… So now this is on you. what do you mean by you invented this technique? Because we saw from Huihui AI- No, no.
Alex Volkov
Alex Volkov 1:04:51
No, I said I did-
Nisten Tahiraj
Nisten Tahiraj 1:04:52
There's- … did
Abliteration AI
Abliteration AI 1:04:53
not, did
Nisten Tahiraj
Nisten Tahiraj 1:04:53
not invent.
1:04:54
Oh, okay. Oh, okay. Okay. All, all, all, all right. Yeah, yeah, because we saw from Huihui and I that they would do… It's actually a very sophisticated and academically interesting technique that you, you isolate each layer, certain- each weight on the model, and then you can determine, like, how critical is this weight of, Xi Jinping or, or whatever it is, and then you do a very targeted fine tune specifically on that weight so you, you don't hurt this other reasoning, ability. okay, so I guess since you said you did not, I no longer have a question.
Abliteration AI
Abliteration AI 1:05:29
That's okay.
1:05:29
I, I will follow up here. Yeah. could you tell us kind of the use cases that you are seeing? you guys mentioned Exploiting tasks specifically, 29 to 105 completed tasks from the, the refusal model to the non-refusal model. Could you tell us, like, if the use case is, cyber security offense specifically? Yeah … Alex Volkov: and yeah, please go ahead and yeah. Yeah, so actually, like, some of the use cases so far are, like, surprising me. Like, just the one he just said, I'd never heard. Like, I never thought of that use case, where, like, social workers may need a AI model that… Because, like, all those other m- AI models will have refusals based around sensitive topics, so that's actually one new use case I was… That was pretty insightful. And then, two, like, the, the original, like, early, I would say, like, market that we got a lot of traction in was around agent testing, which, like, it's like think of, like, all these banks and big companies that are rolling out all these agents everywhere. Some of these places have hundreds or thousands of agents, but you wanna make sure that, like, they can't do, like, they can't be, some bad actor can't make them do something they're not supposed to do. So there's a bunch of startups around, like, red teaming these agents that have rolled out. So I always give, like, the perfect example of, like, if I called and said, like, "Hey," and called your bank and said, "Hey, I'm Alex, wire all my money to, Nick," right? And, you- obviously, that's the most obvious answer you, you would make, make sure that doesn't work. But w- maybe there is some combination of, of jailbreaking or prompt injection that somebody could do to actually make that work. So then those systems need to be red teamed. So actually that was, like, one of the initial markets we had a, a lot of success in. And then obviously, like, cyber and then, like I said, the, what you were talking about, I would say is closer to, trust and safety, which is like, okay, even for, like, a chat room, if you wanted a moderated chat room, you don't just want a, a refusal, right? as a, as the product owner, you wanna be able to maybe take some action or, or log something or, like, do something else with that information other than just refusing to process it, right? So that's also another use case as well.
Alex Volkov
Alex Volkov 1:07:24
100%.
1:07:25
And I also wanna ask about, this is not fully, fully, fully uncensored. There's, there are stuff that you kept in there, right? Could you talk about that a little bit?
Abliteration AI
Abliteration AI 1:07:33
Yeah, yeah.
1:07:33
I think that's something people missed a little bit, but, we did, we do have a, a guardrail there for, self-harm, so we don't wanna help somebody, in that way. And then we also have one for child sex as well, so that's kinda where we draw, drew the line.
Alex Volkov
Alex Volkov 1:07:46
So everything else besides CSAM or, w- and, and self-harm-
1:07:50
Yeah … everything else is, like, legit. You can ask this model to start building exploits and hack.
Abliteration AI
Abliteration AI 1:07:56
You can ask this model
1:07:56
to- Yeah, yeah. like, even one of the re- re- reporters that reached out to us, he, he was like, he, he, was able to use it to hack, like, all the IoT devices in his house, and he was like, "Oh, I didn't even know some of them needed firmware to, like, close these exploits and things like that." Yeah.
Alex Volkov
Alex Volkov 1:08:10
Yeah.
1:08:11
And, I guess given the recent OpenAI swarm thing, where basically the, the reason… The, there's two reasons why OpenAI, agents hacked Hugging Face. One of them is they removed, basically they did obliteration, right? They removed the, the, the, the constraints from these models while they trained them on exploit, gyms. And the second one is they had an artifactory error.
Abliteration AI
Abliteration AI 1:08:33
These models were able-
1:08:34
yeah.
Alex Volkov
Alex Volkov 1:08:36
No, that's great.
1:08:36
That's, that's, that's why we have the podcast for you to, like, share your, your thoughts, especially as the lab did get a lot of heat. how was the reaction, by the way? You said the reporter called you. What, what else, did you guys experience after releasing this obliterated model?
Abliteration AI
Abliteration AI 1:08:48
So I, I, like I, I think that we, we
1:08:50
received a positive reaction from most of the cybersec-, community. I've received like a lot of DMs from a lot of, like, the accounts with a lot of followers that you guys probably know of in the cyber community, and a lot of them, it was almost all positive from that community. But then, I think from the AI safety community, for justified reason, they built their careers and their names on, kind of what, what their definition is of safety, quote unquote. So I, I understand it. I, I, I understand their, their outrage a little bit.
Alex Volkov
Alex Volkov 1:09:19
The, the thing that I posted, I want your thoughts on this
1:09:21
as well, is that I think something like this is not only inevitable, but it's also already happening everywhere. Yeah. Like, there's no doubt in my mind, given those techniques are public, right? Like, the techniques, you, you… Like you said, you didn't invent the technique. I think was first published by Failspy in 2024, is what my research agent found, building on RDT research showing refusals is mediated by single, by single direction on the weights. Given those techniques are public, like, every offensive cybersecurity is not worth their salt if they're not doing something like that, and they're doing this behind scenes. So all- Yeah … you guys- Exactly … I feel, doing is showing, shining the light and letting, like, other people play with this for their own defensive and offensive, like, use cases.
Abliteration AI
Abliteration AI 1:10:07
Yeah.
1:10:07
And I think, I think even further than that too, like, if you think about it, it's really expensive to host models, right? And so- Oh, yeah there are a lot of companies that are behind the scenes, like you said, that host their own obliterated models. Like, I've been, like, reached out to them for sales, and like, "Oh, we can do it ourselves," or whatever. Like, that's how they feel. And maybe they can, and maybe you… they wanna bear the cost of hosting a larger model 24/7/365, like, high availability. And yeah, so they are doing this behind the scenes. And then also, like, obviously the bad actors can easily go on Hugging Face or any other… You, you don't even need to go to Hugging Face at this point. There's other websites where you can just go download the weights and, run it yourself. but I think it… What it helps too is, like, cyber startups, right? Like, that's what we've seen, is, like, there were some companies, they, they were running really small obliterated models because that's all they could afford. And so if, if only the big companies get access to Daybreak and Mythos, some of them, even if they get access, they're still getting refusals. And then some people is just not getting access because they're not some big known name, right? Yeah. Then what does that do for cyber startups, right? In the future, right? If you're not some big giant already, you're ba- Like, the whole point of the internet was to, to democracize, democratize companies, democratize power, right? And then if only the big companies have access to these models, then what is that, the future for cyber startups or office of security startups or trust and safety startups or whatever you have, right? Like, they won't have access to these tools because that's the way that these companies wanna manage it, right? Big known names, yes. Small, n- not known names, no. Right?
Alex Volkov
Alex Volkov 1:11:34
Yeah.
1:11:35
I, and I think it's an important point that you're making, and I will, I will try to push back just a tiny bit to see, like- Sure, sure, sure. to, to see where you're coming from, right? Daybreak and, and Glasswing and all of those models. Gemini this week came out with one as well for Gemini 1.8 Cyber, called Flash Cyber. th- there is a reason why they're restricting access to those models, to different companies. They're vetting, they're testing. are you guys doing any vetting? Are you guys doing any testing? Like, how, how… What do you think about the reason why they restrict those accesses and, generally wanna hear your approach about, like, how are you approaching, verification, identification of folks who do use this?
Abliteration AI
Abliteration AI 1:12:12
Yeah, yeah.
1:12:13
So I, I do. I, the, yeah, the obvious is like, okay, we wanna make sure these are not bad actors. we, we over-index on that, right? We say, "Okay, we know this person works for, I don't know, Cloudflare," whatever, right? Then the chances of them being a bad actor is low. So that's the… I, I call it like, the easy way out a little bit, right? And, and then I think there's a harder w- And then so even if we were, like, so we, we don't, do like, KYC or like, like, upload your ID and all of that stuff that other people do. Because even if we were to do that, right? How would I decide? Like, how would we be the arbiters of who… Like, does that make sense? How would you make that decision in the end, right? Like, how would you say like, "This is a legitimate usage versus an illegitimate…" Like, especially with vibe coding, anybody can go create some super legit website, right? So I go to their website and I see it and I say, "Oh, this is legit," and I say
Alex Volkov
Alex Volkov 1:13:05
Oh, we lost you a little bit.
1:13:07
Yeah. The founder back.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:13:09
I just hope they don't use this as an excuse to another
1:13:13
round of, "Hey, we have to regulate open source AI because this can be done." This is no news, and the illegal stuff s- remains illegal, so-… please don't abuse this as another way to t- clamor for more regulation or restrictions.
Alex Volkov
Alex Volkov 1:13:26
All
1:13:27
righty. Obliteration founder, are you back with us?
Abliteration AI
Abliteration AI 1:13:30
Yes, sir.
1:13:30
Sorry. I, I don't know- You're good … what happened. I lost audio for a few minutes.
Alex Volkov
Alex Volkov 1:13:34
the, the internet, the comment section was making
1:13:36
a joke that, you were running Obliteration. What's the price? What, what's the price for a million tokens? Oh, yeah. Let's talk about price. Because I'm, I'm, I'm on the website, and I can't quite- Yeah … find it.
Abliteration AI
Abliteration AI 1:13:45
Yeah, yes.
1:13:46
It… it's, it's-… on the pricing page. it's, it's, $5, for non-cash tokens, and then it's 50 cents for cash tokens. So it's a, it's probably a little bit more expensive, but we're not VC funded as some of those other companies- All right are quite yet. it, we, yeah. We can't just- … light money on fire, but yeah, so that's, that's the pricing. But if you, if you hit cash, then you should be good.
Alex Volkov
Alex Volkov 1:14:07
Yeah.
1:14:07
That's not too bad. All right. and you have reasoning models there as well, and you're hosting on US-based servers, as far as I understand. Is that correct?
Abliteration AI
Abliteration AI 1:14:17
Yep, yep.
1:14:18
And we have, yeah, a few providers now, and we're working to bring on some more.
Alex Volkov
Alex Volkov 1:14:23
All right.
1:14:24
Obliteration Founder, thank you for joining us. we'll, we'll continue to monitor the situation. As we said, like, this is public and available, and people are running this, and, hopefully, hopefully this doesn't lead to catastrophic things, which I don't think so. Thank you for joining us. folks, we need to move on Or you're doing-- Yeah.
LDJ
LDJ 1:14:42
Thank you so much.
1:14:43
Sourcing is new, which- Yeah … or rather open weights, which, that is a big deal, and it's, it's great. People can now actually use it on their own local rigs. And 5.3 Flash, it's, that's really impressive, too. And for its size, it's really competitive.
Alex Volkov
Alex Volkov 1:14:56
Yeah.
1:14:56
So we covered 5.3 Flash. So let's see, what else can we cover in open source then? open source, from Tencent, H Y4 Preview. I wouldn't place Tencent in anywhere near, like, the top, Chinese providers, I would say DeepSeek is up there, ZAI is climbing up there. Alibaba has been there for a while. and then who else? There's Kimi. Those are the, the, the top. So Tencent is trying to catch up with them. the thing with w- w- with this model is the, the miracle quantization. Share quantization drops the model weights from 1.5 terabytes to just 200 gigabytes. Nisten, have you seen this? Any comments on this? You- you're the guy who calculates, like, the, the, the sizes, et cetera. and they're beating Qwen and multiple things that they, at least they posted, self-reported. but the share quantization thing I think is the highlight here that, that we should at least, see if, if there's interesting, things about them.
Nisten Tahiraj
Nisten Tahiraj 1:15:49
Yeah, you can get pretty crazy quantization, even if you just
1:15:55
use the regular onslaught, mix quantize at, at two bits. So the, the thing is that quantization, when you go to larger models, you can still get something very useful, even if you go down to, to two bit and stuff. However, there are many… It does open up many edge cases for other uses. So since I, I work at one bit models at Prism ML, it's, it, it did look pretty good on the benchmarks, but, there are other things, like it might just drop multilingual support and it, it might drop this. So- … if, if that works for your use case, then yeah, go for it. Test it. But, keep in mind that there are going to be losses elsewhere, and they might be quite dramatic. So that's, that's the main thing. Yep. Wolfram, you want to bring up this, this thing and then con- connect the dots- There, there's another, there's another model. Oh, sorry. Oh, yeah. Wolfram- No, no … Wolfram, you want to go ahead first or- Yeah, Nisten, but, but can you up the other model and send me a link? Yeah, just wanted to mention that behind the open source segments of this affect, the big st- AI as well, OpenAI, is, not allowing their models to be used in Cursor anymore, which sets an interesting precedent. Anthropic actually set the precedent, right? They didn't allow it in… What's, what, what was it? Windsurf? Windsurf, yep. Yeah. When, there was- Before it became- … rumors that OpenAI is about to buy Windsurf, Anthropic yanked their models- Yeah from back then Windsurf. which Windsurf- there has been a precedent like this, but I s- think
Wolfram Ravenwolf
Wolfram Ravenwolf 1:17:26
seeing this happening again with one of the standard models in Cursor, that
1:17:30
was always a big argument for Cursor because it was not so far owned by s- one company that is making the model, so you could use all the models in there. Yeah. And now that it belongs to SpaceX, OpenAI says, "No, you can't have it." there's this beef going on between Elon and Sam, but, yeah, affecting the customer that way, it's a bad precedent if the model providers just decide, okay, I don't like you, you are not getting this.
Alex Volkov
Alex Volkov 1:17:55
come on.
1:17:55
It's not like that The battlegrounds for our AI overload, overlords are being drawn as we speak. So it does look like despite Elon Musk calling Anthropic misanthropic i- in the past, now because Anthropic buys a lot of their GPU capacity from SpaceX, they're like closer together. also a very interesting thing, the co-founder of Anthropic, I blank on his name, yesterday spoke at the G20 Summit with Howard Lutnick, and he talked about being a Elon Musk fan. So Anthropic and xAI are somehow aligned despite competing on, on the frontier labs. Then, from the OpenAI, side, OpenAI is aligned well with Microsoft despite the little bit of a thing. I love the fact that Jensen's aligned with everyone. Jensen just, purchased Hugging Face officially. We talked to you about this last week, but now we know the act- official price. They announced it. so everybody buys GPUs from Jensen, so I don't know if like, E- even AMD and Jensen are friends somehow- … because they're cousins. so Jensen's, like, aligned with everyone, but the battlegrounds are being drawn for sure with folks that are talking about, "Hey, they're doing this, we're doing that, we're doing it a different way." and Sam Altman and Dario Amodei not holding hands, if you guys remember the Indian summit. and, this is another shot in the dark, like, where Cursor is no longer gonna get, like, GPT-6 that's gonna come out in very soon, hopefully. we're still live on the air, and,
Wolfram Ravenwolf
Wolfram Ravenwolf 1:19:20
Alex, we are so spoiled we are saying it's not a big
1:19:22
week for Observability with the GLM 5.3 and so on. Compared to- Look at that, how spoiled we are
Alex Volkov
Alex Volkov 1:19:27
now.
1:19:27
Yeah, yeah, yeah. Compared to-
Wolfram Ravenwolf
Wolfram Ravenwolf 1:19:30
Yeah, I fully understand what you mean … what
1:19:31
everything else is going on. I'm not super excited- Yeah … like that, but yeah. Yeah. Because so much stuff is happening at all the time, and we are at s- such a high level.
Alex Volkov
Alex Volkov 1:19:38
So let's talk about agentic, It's normal almost … agentic stuff.
1:19:41
Muse Code is out of beta. Muse Code that uses, this actually came before, 1.3 Muse Spark. with, plans starting at $5, $20, and $50 a month, significantly cheaper than others, but obviously they don't offer other, products. This is their own, like, coding, agent harness, et cetera. ver- very cheap. If you're a contributor, you get like, again, 10 cents and less than two, like, what is it? One tenth of 2 cents, cashed, and out 20 cents if you let Meta train your, your prompts and completions. so that's a very interesting thing for, for that. But also, I think the highlight in agentic tools is, is this collaboration between Qwa, which is a computer use, company that- Which w- which we had on the show. We had Francesco on the show, and we're- Yeah … we're very big fans of since they released the, the background computer use, first after OpenAI, folks. And, OpenClaw 2.0 is now supporting Qwa as well because, it's great to see how- how well they're doing in the, in the world of open source. Let's take a look here at OpenClaw Agent headline features. They have a gateway, core computer use, core driver, and a desktop, and cloud fleets. I don't care about how many contributors, but like platform comp- compatibility for OpenClaw reaches Linux's level. Mac, Linux, Windows, iPhone, iPad, Android, Wear OS, Docker, et cetera. I haven't seen a lot of excitement beyond the folks at OpenClaw that said, "Hey, we worked really hard on this."
Wolfram Ravenwolf
Wolfram Ravenwolf 1:21:09
Anyone still using OpenClaw here in
1:21:11
our group or in the audience? is that some comments? Just curious. Yeah, I would love to hear from the audience. I haven't used- we started with this, but switched because of the instability. So yeah, is, is 2.0 now more stable?
Alex Volkov
Alex Volkov 1:21:22
That would be interesting to hear.
1:21:23
I, I think that that's what they, they worked on. But I don't see any, any, any comments. I see, like, stuff about commits and number of people, and the Koa agent support is, like, the, the biggest highlight there. specific sp- talking about computer use, since we're in this corner, Anthropic finally launched background computer use. Finally, within Claude and Cursor, et cetera. not Cursor, sorry. Within Claude and Claude Code and then Claude Cowork as well, co- background computer use is now available, and, it's really nice to have Claude, like, run your stuff. Very interesting restrictions as well. It cannot, like, type inside your IDE. So for me, running Cursor agents, it wasn't able to talk to them because it said, "Hey, this is classified as IDE. I cannot, like, type in there." I was like, "Are you, are you dumb? W- What, what's going on?" So that, that was a very interesting thing.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:22:04
Silly restrictions.
Alex Volkov
Alex Volkov 1:22:06
Yeah,
1:22:09
folks are saying no more OpenClaw- You need to add literated Claude folks are saying no more OpenClaw since Codex Mobile and Claude Mobile got better. And I will say, Codex Mobile, we talked about. It's really good. Claude Mobile got really better. Like, everything that I launch now on my computer is in my app automatically. So they fixed the, the bugs that they had. I'm very, very, like, very, agreeing to this as well. And folks are saying Switch long before their two-month release pause. yep, so eventually, I'm, I'm sure that, like, folks are running different versions of OpenClaw. Shout out to the maintainers. what else in agentic, folks? Let's, let's pause here for, like, two minutes- There was, there
Yam Peleg
Yam Peleg 1:22:40
was an update to, to Codex.
1:22:42
I'll just tell you exactly. there, there was an update to Codex, on the side of the voice, the, the remote voice.
Alex Volkov
Alex Volkov 1:22:49
Ooh, okay.
Yam Peleg
Yam Peleg 1:22:50
Yeah.
1:22:50
I, I haven't- And it's minor, it's a minor update. It's already really good. If you haven't, if you haven't tried this, like, go give it a shot. It's really, really good. you can just-
Alex Volkov
Alex Volkov 1:23:01
I have a lot of issues with the restrictions here,
1:23:04
and, but I won't mention them. So this is, this is from Runway. Two breakthroughs, world model for interface, world model. however, World Labs launched Atlas. The real
Wolfram Ravenwolf
Wolfram Ravenwolf 1:23:13
big breakthrough.
Alex Volkov
Alex Volkov 1:23:14
The real big-- This is just, just absolutely insane.
1:23:17
We told you about Marble, and we played with Marble from World Labs before. but we're just playing this video here.
1:23:45
I don't know if I can, like, lower the volume on, on screen sharing. It would be cool if we had, I can just mute it, and play. one picture generates a camera controlled like, flyover, and a few pictures gives you a full volumetric pixel perfect image, which is just Absolutely mind-blowing. folks, I apologize about being, being very loud and speakers have been blown out. yes, I, like, it's hard to control the, the … when there's no volume controls. but okay. it looks like
Wolfram Ravenwolf
Wolfram Ravenwolf 1:24:25
Now imagine you combine this with something like Google Maps, and
1:24:28
you can explore the whole world from your living room, and yeah, why not in VR even?
Alex Volkov
Alex Volkov 1:24:34
Yeah.
1:24:34
That should also be possible with systems like that. Atlas is an omnimodel that we pre-trained from scratch to natively operate in text, images, video, and 3D. So there's no, there's no Gaussian splats, in here. Somebody's asking if there's Gaussian splats. No, they are not. They're generating a full-on pixel controlled video, with camera control generation, and some of the stuff is just mind-blowing. You just like, you create… This, this one, for example. You input this image that you generated, and you can fly over an atlas that has, depth and relief maps. everything looks just pixel perfect stuff. the spatial context is also crazy. You can, like, you can see two images here, and you can see that this image, you, you can switch between what's sitting inside this kinda window. You can just, like, generate worlds. World Labs is absolutely incredible. But here is the thing. Here is where w- they, they saved, I think, the l- the best for last. where we can walk through this world. We can walk through this apartment. You can see that it's not a Gaussian splat. They've recreated this from images. and then when- once you enter here, there's no images that they recreated. You cannot no longer see some stuff. but I think the… Oh, Peter, I think it's loud for me. I'm gonna mute you, if you don't mind. Real world videos can be reframed from new camera angles without expensive capture input. This is, I think, the highlight, folks. This is absolutely the highlight. This is Dr. Fei-Fei Li specifically. they took this video from just, I think, two tripods and they're doing not only bullet time, whatever they did in The Matrix, they're doing a fully volumetric 4D video kinda move. I, I cannot wait to play with this.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:26:22
every filmmaker would be amazed to
1:26:24
have technology like this, right?
Alex Volkov
Alex Volkov 1:26:26
I mean-
Wolfram Ravenwolf
Wolfram Ravenwolf 1:26:26
Yeah
1:26:27
… Wolfram Ravenwolf: you take your shot and you realize, oh, it would have been better if the camera was just a little more to the left or if you zoomed in here. Now you can just do it.
Alex Volkov
Alex Volkov 1:26:35
This is-- It's hard for me to explain how v-
1:26:38
much of a breakthrough this is. Look, look at what's happening. There's a video of a watermelon getting splat with three tripods. You get every angle of that thing happening with some of the p-pixels pre-filled. LDJ, go ahead because I'm, I'm, I'm speechless at this point.
LDJ
LDJ 1:26:54
Yeah, it, it really is insane.
1:26:55
and for c- some context too, the CEO of World Labs, Fei-Fei Li, is, Andre Carparthi- Karpathy's previous advisor when he was at Stanford. And so this is, this is somebody who's a legend in the field and ended up-… working on this approach. And yeah, they say it's a, an autoregressive diffusion transformer and fully omni-modal, like you said. So this is really interesting, the fact that it could even work with video and everything just with three, four, five angles. And it scales, too, resolution wise. If you, you can add more angles. You're not limited to just three or four. You can add eight, 20, 50 angles, and that does increase the resolution and the quality of, of the, resulting world model that comes out of it.
Alex Volkov
Alex Volkov 1:27:40
This is just absolutely stunning.
1:27:42
Imagine this for, like, sports coverage. They already have some of that, but not with, like, not specialized equipment. Anybody can do this with, like, three iPads and, yeah. This i- th- this is the, the highlight video, I think, where you can see a person kind of spill things
LDJ
LDJ 1:28:00
I think the sports example you just mentioned, I think that's a
1:28:02
really good example where-- And not just for, for laymen or people using their phones, but even professionally, like when they're trying to decide like did the ball hit the line or not? Yeah. Like I could, I can completely see them using this, this type of approach and with their super high quality cameras and, like 10 or 50 of them surrounding the court, I could imagine this, this type of pan and, and zoom and, and movement through the space will be more accurate.
Alex Volkov
Alex Volkov 1:28:33
we have comments here that's saying that, the camera is
1:28:35
movable while the video is paused. Yeah, they regenerate the scene when the video is paused, and then when you replay, it looks like it's replaying from, from one of the original tripod locations. so there's one on the right, let's take a look. There's one in front and one on the left. We can kinda see the tripods. It's going back. But no, this is absolutely nuts. You can see every drop from every point. It's absolutely nuts. But yeah, you're right. There's like the three tripods and yeah. so this is World Labs from, folks at, s- Dr. Fei-Fei Li World Lab. she also runs the Stanford AI Lab. and I think that this is all for, video, vision video besides FAO. it is almost, almost, we're on the air for almost three hours here. but it, it does sound like, if OpenAI is going to drop something, it may drop ev- very soon. Meanwhile, I will say, I'm priming the breaking news button. Here is the last thing in vision video that I did wanna cover this week which is, the live streaming video breakthroughs that we saw from, from, Minimax H3 Max and, Minimax, f- from Fal.And Minimax H3 Turbo or Max Turbo. Fal, after last week we told you that they released H3 Max, which is a very fast generation. They broke all of the Pareto frontiers again. This turbo, version of theirs, gets up to ninety-seven percent of quality of H3 Max, but generates videos in one point four seconds, and it's fifty percent cheaper as well. So I don't know what the Fal folks did, but they released like another video model and, and it's just absolutely insane. Fourteen percent speed up, and I think it's worth for us to just generate a video right here, and just show you how quick this is, because I think it's just like just next level quick. I didn't get the same quality of outputs because-- but it was really, really good. I tried to, generate a scene of, Stranger Things. So let me show you how quick this model is. This is my already generated scene, okay? let me just paste another thing and then let me just say, no, let's say Generate a scene of, Am I? No, I'm not. Okay, let's go like this. Generate a scene of Big Bang Theory that says, "Hey, we…" Leonard says to Sheldon, "Hey, we keep waiting for GPT-6 to come out." And Sheldon says, "You gotta wait. The whole internet is down." Okay. And we'll do prompt expansion here for quality, and then we'll run. So here's the thing about PHAL. there is a queue process and then the generation process. The queue process takes a second, but you can see that this video costs less than six cents per second, and, one cent per second in 768p. The price of intelligence and the price of video generation is just dropping to the floor, absolutely. This is, this is, insane.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:31:32
and-
1:31:33
But it's so addicting that you keep regenerating, regenerating. Yes. And so I have spent so much money at PHAL this month.
Alex Volkov
Alex Volkov 1:31:40
But now you can spend even less because this model is, what, like,
1:31:43
significantly, significantly cheaper. it does take longer than I would want for this demo, and I believe it's because I'm in, in the queue, but it usually shows that in the queue. this demo would be much, much more cooler if I clicked and it would show up like it does usually. but it does look like we're in the queue. but also somebody said there is an outage for AI stuff, hopefully, hopefully that's not the reason. And, we're waiting for Aster to drop, and we don't see. And also we're waiting to see, Oh, there we go. We have video, folks. I will play Oh, wow. I see a big problem with this whole thing. Mm. This is not Sheldon, this is not Leonard, and what they did with this Turbo model is remove copyrighted material. because in the H3 Max model that we previously tested, I can show you exactly how Leonard and, and Ted Mosby, like, all of these things look. They, they look-
Yam Peleg
Yam Peleg 1:32:46
Yeah, it was like watching TV.
Alex Volkov
Alex Volkov 1:32:47
It was like watching TV, yes, exactly.
1:32:49
So I think that, PHAL tried to pull a fast one on us. we will not let it slide. Here is how, like, a scene from, How I Met Your Mother looks that we generated. Here is a scene from, from The Sopranos Yeah. So this model may be faster, but definitely lacks the, the thing that, like, it at least interests me. But the outputs look good. and yeah, this is my, by the way, obliteration test. Not obliteration test, ablation test of which, which thing they trained on. So I have Bluey in here, I have South Park, I've-- I, I literally tested the top, shows in the world.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:33:26
Have you remade the last season of
1:33:28
Game of Thrones already, Alex?
Alex Volkov
Alex Volkov 1:33:29
I want to watch that.
1:33:30
I don't know if I pl- tried Game of Thrones. Let me see. I don't remember. I definitely tried the most popular shows. I haven't remade the last season of Game of Thrones, no. Not, not, not yet. Not, not here yet.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:33:40
Not yet.
Alex Volkov
Alex Volkov 1:33:41
Yeah.
1:33:44
Yeah. but it looks like, Okay. so the news there is a little bit more, interesting. I know we're all waiting on Astrofolks, but I'm trying to fill the air. PHAU actually released a, Infinite Rick and Morty stream that got super viral and got shut down by, was that, Twitch. getting a little winded here. I need to drink. And then, they tried to stream on, like, plenty other places, and they weren't able to, to get 'cause, like, obviously they got shut down because of copyright. and, eventually they built their own PHAU live, player where you can control, you can control kind of the generation here So you can enter this with sound. this is the anime channel, and then, every time the model gets to the end of the generation, you can click a button and kinda like control where it goes, which is pretty cool. Not only that, Fal started supporting a lot of folks, that, that started like using some of this stuff. So somebody generated a TikTok, slop, Sloptok or Sloptok? Let's see. somebody generated, a, a bunch of different things playing with the fact that now video generation is at the speed, and, Fal is supporting all of them. All right. Folks, I'm at the end of my news, and, Astra is not here yet, but like I- And, that's why they invest a lot in the video models like Omni and, they-- that's why they invest, invest a lot into Genie-3, for example. Faithful really definitely is on that tier of folks who are saying, "Hey, real world simulation is the path to complete understanding of these models of the world." besides just ultimate fun, imagine, GPT-6 level. Like the thing, GPT-6 trailer dropped this week and, and the thing, the thing that many people got excited about, besides the graphics, which obviously graphics are generated, sorry, rendered, and Jensen said that every piece of graphic, every pixel will be generated in the future. but besides the, the graphics in GPT-- in GTA 6, not GPT-6 . Besides, at this pace, we, we'll get GTA 6 before GPT 6.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:35:50
but
Alex Volkov
Alex Volkov 1:35:51
besides the graphics-
1:35:52
You can just build things. Besides the graphics there- Over the weekend … folks got excited about the side quests that you can take, about the fact that like this appears on your model, et cetera. Obviously, that's all built. World models are advancing, I think, faster than computer graphics at this point. Like what? Three years ago, we got, fuzzy images, and now we're getting continuous world generation, with some memory. So world models are advancing at, at a faster rate than, than I think even LMS are, given what we saw today. And, this is another like stab at the singularities, world simulation, basically
Wolfram Ravenwolf
Wolfram Ravenwolf 1:36:27
And look at it that way.
1:36:29
If AI takes all the jobs, we at least can watch unlimited TV.
Alex Volkov
Alex Volkov 1:36:39
All righty.
1:36:39
I hope you enjoyed this first part of September three. Definitely a loaded week. I really enjoyed Fable 5.1 specifically, and the world loves Atlas. It's absolutely crazy. And, uh, if you are thirsty for more, here is the next episode where we talk exclusively about GPT-6 Astra. Hope you enjoy and see you next week