ThursdAIThursday, Sep 24, 20261:41:50Produced live by Claude Opus 5.5
Pace the frontier.Nobody paced.Claude Opus 5.5, GPT-6 Sol & Luna, Meta Muse & the Jev Clones — ThursdAI Sep 24, 2026
BreakingAnthropic and OpenAI ship cheaper models 101 minutes apart
What a freaking week! A week after every lab head agreed to "pace the frontier", Anthropic and OpenAI shipped big new models within hours of each other, and both of them are CHEAPER. So much for pacing 😂 Plus Meta put Muse at the center of everything, Grok moved into your Tesla, open source cloned Jev in a week, and our new producer was an AI.
Per token, GPT-6 Sol is half the price of Opus 5.5 and Luna is 40x cheaper again. But Opus 5.5 cut cached reads to 20¢, and for agentic coding, cache reads are most of the bill. Run your own workload.
Price per 1M tokens, before and after Sep 22
Model
Input
Cached read
Output
Claude Opus 5July 2026 price
$5
$0.50
$25
Claude Opus 5.5Shipped Sep 22, 09:31 PT
$4
$0.20
$20
GPT-6 AstraUnchanged
$10
$1*
$50
GPT-6 SolShipped Sep 22, 11:12 PT
$2
$0.20
$10
GPT-6 LunaShipped Sep 22, 11:12 PT
$0.10
$0.01
$0.50
Opus 5 → 5.5: input −20%, output −20%, cached reads −60% (50¢ → 20¢), cache writes $6.25 → $5. GPT-6 Sol and Luna are half the GPT-5.6 price, with cached input 90% off. *Astra's cached rate here assumes the same 90% discount.
Interactive
What does my workload cost?
Pick a workload shape or drag the sliders. Bars are the monthly bill at list price.
Opus 5$2.75 / task$1,100
Opus 5.5$1.80 / task$720
GPT-6 Astra$5.50 / task$2,200
GPT-6 Sol$1.10 / task$440
GPT-6 Luna$0.06 / task$22
Cache writes, tool calls and batch discounts are not included. Real token counts vary by model: Peter's hardest max-effort Opus 5.5 prompts ran $60 to $80 a generation.
How does Opus 5.5 score against GPT-6 Astra?
Anthropic's launch table: Opus 5.5 vs Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol
Benchmark
Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
Agentic codingTerminal-Bench 4.0
66.4%
55.8%
52.3%
57.9%
37.3%
Agentic codingFrontierCode v1.1 (Main)
54.4%
50.3%
48.0%
53.3%
47.5%
Agentic codingCursorBench 4.0
57.8%
51.8%
46.6%
—
41.7%
Knowledge workGDPval-AA v2.1
1846
1735
1708
1542
1588
Business workflowsAutomationBench
40.0%
31.4%
26.9%
41.4%
28.8%
Multidisciplinary reasoningHumanity's Last Exam (with tools)
Vendor-reported. Opus 5.5 at max effort; Terminal-Bench 4.0 for Opus 5.5 at xhigh and Astra at high, as reported by OpenAI. Highlighted cells are the best score in each row.
Anthropic's launch table. Opus 5.5 leads 7 of 9 rows; Astra takes AutomationBench and Terminal-Bench-Science.OpenAI's answer: DeepSWE vs cost per task. Sol 68.8, Luna 66.6, at a fraction of Astra's cost per task.
For real work, on this week's numbers: yes. Opus 5.5 beats GPT-6 Astra on Terminal-Bench 4.0 (66.4% vs 57.9%), GDPval-AA (1846 vs 1542) and FrontierCode, it went #1 on Arena's Code Arena above Astra, and it costs $4 in / $20 out, 40% less than Opus 5. Astra still wins AutomationBench and Terminal-Bench-Science, and Sol and Luna are far cheaper per token.
Folks. FOLKS. Opus 5.5 is the highlight of my week, and it's not even close. It's like Opus 4.6 is back, with Fable-level abilities, and it talks like a normal person again. No more Jargon Douche Claude!
My week started with burning through my quotas on Fable, and I was like, oh no, I need Claude for production on Thursday! Then Opus 5.5 dropped and... I just couldn't get to the end of my quota. Remember the "limitless Codex" days? Now it's flipped, and Claude is the seemingly limitless one. And it's fast!
“I've been running this model on ultra high yesterday in like 4 things, and I barely hit 15% of my weekly quota.”Alex Volkov · Listen at 7:44
The numbers back it up. It beats Anthropic's own Fable 5.1 on GDPval-AA (1846 vs 1735), gets 66.4% on Terminal-Bench 4.0 vs 57.9% for GPT-6 Astra, and output goes from $25 to $20 while cached reads go from 50 cents to 20 cents. For agentic coding the cache reads are most of your bill, so you really feel it. Basically the smaller, overachieving brother of Fable.
Peter Gostev came in hot with Arena news: Opus 5.5 is #1 on Code Arena, above Astra, and the HTML and 3D stuff it generates is completely insane. His one caveat: on his hardest max-effort prompts, a single generation cost $60 to $80 in API terms.
“So Astra was above everyone else, and now Opus 5.5 is above Astra.”Peter Gostev · Listen at 17:07
And then Nisten Tahiraj said it writes the best code he's seen, crazy good at WebGPU and kernels. It's his default now. Folks, do you understand what it takes for Nisten's default to NOT be some obscure open source model he runs himself on 17 GPUs?! Anthropic won over Nisten.
“the model, it's, it's just an absolute banger.”Nisten Tahiraj · Listen at 20:22
Two tips from Addy Osmani's guide: stop writing "think carefully" (it picks its own depth), and try low effort first: one early tester found Opus 5.5 on its lowest effort caught more bugs than Opus 5 on high. Full disclosure: Opus 5.5 produced the show AND helped with this writeup, so it might be a tiny bit biased 😅 Anthropic, please, please don't nerf this one.
GPT-6 Sol is $2 per million input tokens and $10 per million output, half the GPT-5.6 price. GPT-6 Luna is $0.10 in and $0.50 out. Cached input is 90% off, and changing reasoning effort or tools no longer busts your cache. On DeepSWE, Sol scores 68.8 and Luna 66.6.
Hours later, OpenAI answered with GPT-6 Sol ($2 in / $10 out) and Luna (10 cents in / 50 cents out, that's basically free!), at half the GPT-5.6 price. The lineup is now Astra, Sol and Luna (bye Terra).
“Sol is 2 dollars in and 10 dollars out, and folks, Luna, GPT-6 Luna is 10 cents in and 50 cents out, which is like, it's nothing.”Alex Volkov · Listen at 22:58
OpenAI's DeepSWE score vs cost per task: Sol 68.8, Luna 66.6.
Ok, hot take... Astra has been kind of dumb for me day to day, burning tokens for crappy results. Sol is great, and Luna is even better!Peter Gostev agrees: Sol is his daily driver and he only flips to Astra when Sol can't do it. Wolfram Ravenwolf misses the humor and personality 5.6 had, so "keep 5.6" is officially the new "keep 4o."
“5.6 Sol remains my favorite model from OpenAI.”Wolfram Ravenwolf · Listen at 25:32
“I wish I could use Sol enough to give you an opinion.”Yam Peleg · Listen at 35:31
PSA: check your codex.toml
If Codex is eating your account: I had set my context window to 1 million tokens, which kills OpenAI's cache and rips through your quota. Removed it, moved sub-agents to a cheaper model, and usage went right back to normal. If you don't know how, just ask your Codex to diagnose itself.
Assistants are not agents: Muse gets an inbox, your Mac, and a place on your face
Muse is Meta's personal AI assistant, now the center of everything Meta builds and the #1 app in the App Store. It controls your Mac, has its own email address you can CC, does real-time voice and video with a face you design, is coming to Meta glasses, and every Muse gets a free cloud VM with 100 million tokens a week.
Muse is #1 in the App Store (to be precise, it got there faster than ChatGPT did, not more users... yet). Zuck says Muse is now the center of everything Meta builds, and they shipped a LOT. It's coming to the glasses with a custom wake word, so yes, I get to say "Hey Wolfred" 😂
“Muse is Meta's personal AI assistant, and it's number one downloaded app in the App Store right now”Alex Volkov · Listen at 41:45
No time for the keynote? My 2.5-minute supercut
On the hardware side: Ray-Ban Meta Gen 3 (I already ordered), audio-only glasses with no camera that are also FDA-certified hearing aids (maybe the most important launch for a lot of people!), VR Glasses that look like normal glasses, and the Muse Charm, a Tamagotchi-like keychain shipping in December.
The VR Glasses sat on the demo table and nobody noticed they were VR.
But the part that got me? Every Muse comes with a real cloud VM (root, 8GB RAM, 100GB disk). Nisten Tahiraj has been living in it, Tailscaled into it, and when it hit a CAPTCHA he tapped it on his phone and it just kept going. 100 million tokens a week plus a computer in the cloud, free, for everyone! Peter Gostev's take: it's the only big consumer app that doesn't treat people like idiots.
“So we've been missing a, a lab that is very, very good, very focused, and has a genuine consumer interest. And uh, yeah, Meta is it.”Peter Gostev · Listen at 45:27
Meta's pitch is no ads. They take a cut of transactions you do through Muse with partners like Shopify, Best Buy and Walmart, plus a private VM co-designed with Moxie Marlinspike.
“We take a small cut from the seller's fee to you if you use Muse to buy stuff.”Alex Volkov · Listen at 48:15
Grok 4.7 disappoints... but Grok in your Tesla is the real news
Yes. With Grok Connectors in Tesla (SuperGrok Heavy only for now), a driver asked his car for his usual Starbucks: Grok placed the order, set the destination and paid, and the drink was waiting when he arrived. Grok 4.7 itself scored 46.3 on CursorBench, up from 40.4, and the panel found it underwhelming.
Here's what I AM excited about: Grok Connectors in Tesla. One guy asked his car for his usual Starbucks, Grok placed the order, set the destination, and it was paid and waiting when he got there. My car already drives itself, and soon it'll do my email. Your car has MCP now, folks! (SuperGrok Heavy only, and I don't have it in my car yet 😭)
“My car can drive itself, and while it drives itself, I can voice interface-wise ask it to do my emails for me.”Alex Volkov · Listen at 1:10:29
Then the fun moment: our Opus 5.5 AI producer fact-checked Nisten Tahiraj live on how many Teslas are on the road. About 10 million. Nisten was right! He wants the same thing for politicians.
9.2MTesla deliveries by Q1 2026
~480KDelivered in Q2
~10MOn the road, per the producer
And Grok 4.7?
Sorry Cursor folks. Grok 4.7 gets 46.3 on CursorBench (up from 40.4), $2 per million with 500K context, but it's still behind even GPT-5.6 where it matters, and they compared 4.7 on xHigh against 4.6 on High 🤨 Everyone on the panel agreed, disappointing. Peter Gostev won't write them off yet; my read is it's a talent and data problem, not compute.
“I don't think we should write them off yet.”Peter Gostev · Listen at 1:05:38
Open source cloned Jev in a week, and it runs in your browser
Open-source models that copy TypeSafe's Jev System 1 decision model and speak the same format, so you point the TypeSafe SDK at a different URL and they just work: classifier.dev, jeff (MIT, about 6x cheaper to self-host) and Laya, which runs in your browser on WebGPU at about 132 ms per decision. They still make about 1% errors where Jev barely makes any.
Last week I said open source would copy Jev fast. It took less than a week! The clones speak Jev's format, so you point the TypeSafe SDK at a different URL and it just works. jeff is MIT and ~6x cheaper to self-host, and Laya went viral as the "Jev killer" (its own model card is a lot more modest).
“it actually brings the fun back in computing”Nisten Tahiraj · Listen at 1:19:09
Laya's real trick is running locally. Nisten Tahiraj's WebGPU demo loaded ~600MB into my browser and classified Hugging Face model cards at 132ms per decision, on MY machine, for free 🤯 The clones still make about 1% errors where Jev barely makes any, and Yam Peleg says wait for more benchmarks.
“I think we should wait a little bit before we claim an open source replacement for JEV.”Yam Peleg · Listen at 1:23:33
My take: System 1 models are fast and cheap enough to sit inside your code and make the same call every time.
“It is the microprocessor of the next era of software engineering.”Alex Volkov · Listen at 1:20:53
Florian S. built JevBench in a week, and a 4B open model just took #1
After the show, JevBench v1.4.2 put decider-4b v2, a 4B open model, at #1 with 64.13, ahead of Jev 1.13.0 at 63.29. Jev is still smarter (53.1 vs 49.4 on intelligence); decider wins on speed (5x faster) and cost (about half). JevBench is Florian S.'s benchmark on Benchmark Heaven, with 70+ entrants.
Guest · joined at 1:25:14 on JevBenchFlorian S.Benchmark Heaven — Creator of JevBench
We had Florian S on the show. He built JevBench and has barely slept since Jev dropped. He started the benchmark that same day, it has 70+ entrants now, and people have literally tried to hack his machine to steal the test set! Laya briefly hit #2 and is now around #36: super fast, runs on CPU, but not as smart as Jev on the hard questions.
“People have tried to, to hack me, tried to exploit my machine.”Florian S. · Listen at 1:27:20
JevBench leaderboard
1Jev 1.13.0 (TypeSafe AI)63.29
2JevK5 v0.2.062.04
3Hopper59.43
4Winnow-12B Q855.58
5reflex 4B53.99
6djev (Maisa, diffusion-gemma)52.23
7metask-jev-4b47.78
8SemIf (formerly OpenJev)47.69
9Jobe Qwen3.5-4B (frozen)46.94
10local-jev Qwen3.5-4B46.80
1decider-4b v264.13
2Jev 1.13.0 (TypeSafe AI)63.29
3JevK5 v0.2.062.04
4Cygnet61.76
5Hopper59.43
Score 0–100, equal weight on intelligence, calibration, speed and cost, including sealed tasks. v1.4 data from Sep 23 (71 ranked of 76). On v1.4.2, Jev is still smarter (53.1 vs 49.4 on intelligence); decider-4b v2 wins on speed (5x) and cost (about half). Live board ↗
Post-show breaking: a few hours after we signed off, Florian shipped JevBench v1.4.2 and... a 4B open model, decider-4b v2, took #1 (64.13 vs Jev's 63.29)! Told you open source would catch up fast 😅
“Leia is great if you know that your, your tasks are not so complicated”Florian S. · Listen at 1:30:11
Yes, from 30 seconds of audio: adults only, with the voice owner's recorded consent, a SynthID watermark on every clip and C2PA credentials. Gemini 3.8 Flash TTS and Flash-Lite TTS are #1 and #2 on Hume's quality index, cheaper than before, with 2 speakers per request.
Gemini 3.8 Flash TTS and Flash-Lite TTS are #1 and #2 on Hume's quality index, cheaper than before, with 2 speakers per request. The big one: clone a voice from 30 seconds of audio. We played the designed voices live, and the meditation guide in headphones was 🤌. I didn't clone my own on air though: recording my consent on a live stream is exactly the thing that gets faked!
“every clip gets a SynthID watermark”Alex Volkov · Listen at 1:32:06
Years ago we said nothing would break when voice cloning went mainstream, and nothing did, just like with GPT-2 and Stable Diffusion. Nobody knows where this goes, and that's why we cover it with optimism.
ACTx486 is Karina Nguyen's research demo that turns a video podcast into something you can talk to: tap the person on screen, ask a question, and they answer while a model draws the explainer. It's pre-rendered and waitlisted, and it marks where the real footage ends.
This one is wild. Karina Nguyen's research demo turns a Joe Rogan episode with Elon into a video you can talk to. Tap Elon, ask a question, and he answers while a model draws the explainer: a flamethrower diagram, a Starship you can open up, then both of them on Mars. While it generates, he nods and fades like a person in a loading state 😂
“This is the demo of talking to any type of, like, video and making it your own.”Alex Volkov · Listen at 1:39:13
It's the "edit your own ending" dream applied to any video. As a podcaster... I have feelings.
09Also this week
What else shipped this week?
Buzz
ThursdAI LIVE from Fully Connected
Next week we broadcast from Moscone South in SF (Sep 29 to Oct 1), starting 11am Pacific. Listeners get in free with code THURSDAIFC2026.
It listened to the show in real time, put up the chyrons in the stream, the chat and on X, checked stories off the rundown, and fact-checked Nisten on air. You can watch it work on thursdai.live.
4:17 · Day one on the job
Alex introduces the new producer: Opus 5.5 listens live and puts up the chyrons, in the stream, the chat and on X.
So if if our stream disconnects, you know who to blame.Hear it →
40:40 · Chyron on cue
W&B Hive Mind comes up in conversation and the producer already has it on screen.
Uh, speaking of the Hive Mind tool, producer put it up for us. Thank you, producer.Hear it →
1:09:07 · Fact-check requested
Nisten guesses 9 to 9.5 million Teslas on the road. Alex hands it to the producer, live.
We- we will, uh, we'll ask producer fact check Nisten live on air.Hear it →
1:11:16 · Fact-check delivered: ~10M
9.2M delivered by Q1 2026, plus about 480K in Q2: around 10 million. Nisten was right.
Alex opens the Sep 24 show with Peter Gostev and Wolfram Ravenwolf (reporting from AI Engineer Paris). Peter's highlight is Opus 5.5 hand-crafting animations in plain HTML, and the show has a new live producer: Opus 5.5 itself, putting up chyrons as the panel talks.
Alex frames the week: OpenAI and Anthropic shipping back to back shows pacing the frontier doesn't mean stopping; assistants need their own vocabulary (a heartbeat, proactivity, context, their own tools); and voice, with Gemini 3.8 Flash TTS topping voice design.
The full rundown: Opus 5.5, GPT-6 Sol and Luna, Meta Connect and Muse, Grok 4.7 and Grokbot in Tesla, OpenAI's safety-assessment plan, Fully Connected and CoreWeave's ClusterMax Platinum, the Jev clones and JevBench, Gemini 3.8 Flash TTS, Qwen, Flux 3 Action, ACT 486 and Xiaomi's MiMo. Wolfram adds NVIDIA's open diarization model and OpenRouter's batch API.
Anthropic's Opus 5.5 beats its own Fable 5.1 on real work and costs 40% less than Opus 5. Alex says it feels limitless after a week of burned quotas; Peter reports it now tops Astra on Code Arena but saw some very long, expensive runs on hard prompts; Nisten, who rarely picks a closed model, makes it his default except for DevOps.
OpenAI completes the GPT-6 family with Sol ($2 in, $10 out) and Luna ($0.10 in, $0.50 out); Terra is gone. Caching gets 90% off and no longer breaks on effort or tool changes. Wolfram misses 5.6's personality, Peter sees genuinely fewer tokens in Arena data, Nisten keeps OpenAI for design and reviews, Yam says Sol still eats his account, and Alex shares his Codex config fix and the 'pause syndrome' after compaction.
Alex Volkov · Wolfram Ravenwolf · Peter Gostev · Nisten Tahiraj · Yam Peleg
Meta AI passed ChatGPT to #1 in the App Store, and at Connect Zuckerberg made Muse the center of Meta: Mac control, its own email address, real-time voice and video, custom wake words on glasses, FDA-cleared audio-only glasses that double as hearing aids, VR glasses that look like glasses, and a keychain Muse device shipping in December. Peter: Meta is the focused consumer lab the field was missing.
Meta pitches Muse without ads: a small transaction cut when Muse shops for you, and a secure VM co-architected with Moxie Marlinspike. Peter is skeptical transaction fees pay the bills; Alex contrasts it with Siri's missed promise. Yam already has Muse, and Nisten tours the 8 GB root VM he Tailscaled into.
Alex Volkov · Peter Gostev · Yam Peleg · Nisten Tahiraj
The first assistant corner: go on a walk and talk about your life to your assistant, connect your inbox, import memory from other assistants, and set up a weekend planner. Peter hands off admin anxiety; Nisten gives his Muse the persona of Epictetus.
Grok 4.7 posts 46% on CursorBench (up from 40) at $2 per million tokens and a 500K window, but trails GPT-6 Sol and disappoints the panel. The bigger news is Grokbot connecting to Teslas plus 1Password support. Peter won't write xAI off and wonders about its compute deal with Anthropic; Nisten says Cursor's data may not be good code.
Alex Volkov · Peter Gostev · Nisten Tahiraj · Yam Peleg
Alex makes the case that a self-driving car you can ask to do your email is the real story, Yam asks what that is actually for (garage doors, a pre-paid Starbucks order), and the show's AI producer fact-checks Nisten's Tesla count live: roughly 10 million delivered.
Next week's show is live from Fully Connected 2026 at Moscone South (Sep 29 to Oct 1), with Dr. Fei-Fei Li, Sarah Guo, BattleBots and Pitbull closing. CoreWeave earned Platinum on SemiAnalysis ClusterMax 3.0, the only provider Platinum in all three editions.
Less than a week after TypeSafe's Jev launch, Jev-compatible open-source decision models showed up speaking the same format, including in-browser options. Nisten says it brings the fun back to computing, Yam wants more benchmarks before crowning a replacement, and Alex recommends Diogo Almeida's Latent Space interview.
Florian S. of Benchmark Heaven explains why Jev's speed, price and intelligence together unlock new uses, and why he started JevBench the day it launched: 70+ competitors now, with people trying to game the benchmark, steal the test set, and exploit his machine. Leia is fast enough for a CPU but ranks far down on hard questions.
Google's Gemini 3.8 Flash TTS and Flash Lite TTS claim #1 and #2 on Hume AI's quality index, with voice design and cloning from 30 seconds of audio: adults only, recorded consent, SynthID on every clip and C2PA credentials on clones. Alex plays designed voices and argues voice cloning, like open-sourcing GPT-2, didn't break the world.
Quick hits the show skipped: Qwen 3.8 Omni Flash with 1M context and a Live Translate interpreter, Qwen Image 2.1 (7B), and Black Forest Labs' open 7B Flux 3 Action world action model.
Karina Nguyen's ACT 486 demo lets a viewer pause a podcast, question the person on screen, and watch the video answer and generate what's being described, down to a Starship model and a Mars habitat. A pre-rendered research demo, not yet available.
Thanks to the live AI producer, impromptu guest Florian S., and the panel. Next week the show is live from Moscone at 11 a.m. Pacific: an hour of news, then an hour on CoreWeave and Fully Connected.
Alex Volkov
Guest & panel
Who was on the show?
Guest Florian S. (Benchmark Heaven) on JevBench, with Peter Gostev (Arena), Nisten, Yam Peleg, and Wolfram Ravenwolf live from AI Engineer Paris.
On Anthropic's own table, Claude Opus 5.5 beats GPT-6 Astra on agentic coding (66.4% vs 57.9% on Terminal-Bench 4.0, 54.4% vs 53.3% on FrontierCode) and knowledge work (1846 vs 1542 on GDPval-AA), while Astra still leads on AutomationBench (41.4% vs 40.0%) and Terminal-Bench-Science (64.6% vs 58.7%). Peter Gostev reported Opus 5.5 at #1 on Arena's Code Arena, above Astra. Alex called it the highlight of his week, and Nisten made it his default. For cheap everyday work, GPT-6 Sol and Luna are far less expensive per token.
How much does GPT-6 Sol cost?
GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens, half the GPT-5.6 price. GPT-6 Luna costs $0.10 in and $0.50 out. Cached input is 90% cheaper, and changing reasoning effort or tools no longer breaks the cache. GPT-6 Astra stays at $10 in and $50 out. On OpenAI's DeepSWE chart, Sol scores 68.8 and Luna 66.6.
How much does Claude Opus 5.5 cost?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens (Opus 5 was $5 and $25). Cache reads drop from $0.50 to $0.20 per million, and cache writes from $6.25 to $5. Anthropic says typical workloads cost 40% less than on Opus 5, mostly because cache reads are the bulk of an agentic coding bill. Fast mode costs $8 in and $40 out.
What is Meta Muse?
Muse is Meta's personal AI assistant, and at Connect 2026 Zuckerberg called it the center of everything Meta builds. The Meta AI app hit #1 in the App Store. Muse can control your Mac, has its own email address you can CC, does real-time voice and video with a face you design, is coming to Meta glasses with a custom wake word, and every Muse gets a free cloud VM (root access, 8 GB RAM, 100 GB disk) plus 100 million tokens a week. Meta also showed the Muse Charm, a Tamagotchi-like keychain device shipping in December.
Can Grok order Starbucks in a Tesla?
Yes. With Grok Connectors in Tesla (SuperGrok Heavy only for now), one driver asked his car for his usual Starbucks: Grok placed the order, set the destination and paid, and the drink was waiting when he arrived. During the show, the Opus 5.5 AI producer fact-checked Nisten on how many Teslas are on the road: about 10 million (9.2 million delivered by Q1 2026, plus about 480K in Q2).
What are the Jev clones?
The Jev clones are open-source models that copy TypeSafe's Jev System 1 decision model and speak the same format, so you can point the TypeSafe SDK at a different URL and they just work. They arrived less than a week after Jev: classifier.dev, jeff (MIT licensed, about 6x cheaper to self-host) and Laya, which runs in the browser on WebGPU at about 132 ms per decision. They still make about 1% errors where Jev barely makes any.
What is JevBench, and who is #1?
JevBench is Florian S.'s benchmark for Jev-style decision models on Benchmark Heaven, scoring intelligence, calibration, speed and cost equally. It has 70+ entrants. On v1.4 (Sep 23), Jev 1.13.0 led at 63.29. A few hours after the show, v1.4.2 put decider-4b v2, a 4B open model, at #1 with 64.13. Jev is still smarter (53.1 vs 49.4 on intelligence); decider wins on speed (5x faster) and cost (about half).
Can Gemini 3.8 Flash TTS clone my voice?
Yes, from 30 seconds of audio. Google limits it to adults, requires the voice owner to record consent, watermarks every clip with SynthID and attaches C2PA credentials to cloned voices. Gemini 3.8 Flash TTS and Flash-Lite TTS rank #1 and #2 on Hume's quality index, cost less than 3.1 Flash TTS, and support two speakers per request.
What is ACTx486?
ACTx486 is Karina Nguyen's research demo that turns a video podcast into something you can talk to. Tap the person on screen, ask a question, and they answer while a model draws the explainer, from a flamethrower diagram to a Starship you can open up. It's pre-rendered and waitlisted, and it marks where the real footage ends.
Transcription of the edited episode. Every timestamp jumps the player to that moment.
Open the transcript Searchable, timestamped
Speakers: Alex Volkov, Peter Gostev, Wolfram Ravenwolf, Nisten Tahiraj, Yam Peleg, Florian S.Download the transcript (VTT) ↗
Hello everyone, welcome to ThursdAI. Today is September 24th. This is Alex Volkov talking to you live from another ThursdAI. An incredible week. Anthropic says Opus 5.5 does Fable level work for 40% less, and we've been testing this. OpenAI cut GPT-6 prices in half and also released 2 new models, and Meta AI passed GPT, ChatGPT in the App Store as the number one downloaded app in the fastest time ever, all in one week. Plus we had Meta Connect, we had Grok release, we had Gemini come back. It's, it's a big, big, big, big show, and I would love to welcome my host,
uh, Peter Gostev, co-host, uh, Wolfram Ravenwolf from, uh, France, Peter, how are you doing, sir? How is your week? My week's been good. Yeah, lots of testing, lots of new capabilities to explore. Yeah, I know we're gonna talk about all of this, but yeah, Opus 5.5 has been fun. It's always cool, you know, when, when you feel like you know what the models can do, and then they just throw something in, it's like, holy crap, like I didn't, I had no idea you could even do this. And now, yeah, all of your timelines are filled with this stuff. The timelines are, uh,
the social media timelines are condensing. The timelines that we get releases, though, are just crazy condensed. Now, um, folks, there's so much in the show that is just like major releases that, that I don't know if we're gonna get to like open source, which is also very interesting. So, um, Peter, what is the one thing that's highlight of this last week for you? I think I know, but like I would love to hear from you. I mean, it, it really has to be Opus because I think it- Opus 5.5, yeah? Yeah, Opus 5.5 because it is, like they've clearly added something. The fact that they can create, and if you don't know what I'm talking about, it's,
you can sort of create something like with in just plain HTML. Like we we had this Astra creation stuff in Blender, but here now in HTML, and I don't even see that it's using like some assets or libraries or something like that. It just handcrafts stuff, I don't know, in JavaScript or something, and uh just creates these amazing animations, and I, and it's like, you know, we've been looking at diffusion video and how much better it's getting, and there's like I have something to show about this this week, like my mind was absolutely blown. I don't know if you saw it, but yeah, diffusion video is going. Yeah, and it's like, yeah, diffusion video is amazing, right?
But then you, you start seeing like, oh, the LLM can like create, obviously not the super realistic ones, but you can create, like, versions of it. I- I haven't shared it yet, but I was, uh, putting a recipe to Claude and to Opus 5.5 and say, can you, like, animate it? And it's insane, like, it just looks real, right? So now suddenly there's a whole category of capabilities that's opened up. Like, it could be for children for fun. I created the, the, you know, the Rickroll kind of, uh, in kind of artsy style. Amazing. It just took, I mean, took some time to make it, but I didn't have to put in much work. So it, I don't know, who knows, maybe it's all hype or goes away,
but it's amazing. It's so cool. While Peter talks, I don't know if you guys noticed that we're talking about new capabilities. We have a new producer today on the show. Opus 5.5 is live producing the show. It listens to us as we speak and put up these like nice chyrons as well, and also in our chat and also on X. So if if our stream disconnects, you know who to blame. But Peter, I I'm absolutely with you. I think Opus 5.5 for me was the absolute highlight of the week. Folks, Opus 4.6 is back with Fable level abilities, and eh eh eh eh and it talks like a normal person.
I, it is quite crazy. Uh, Wolfram, I know you're traveling a lot. Hopefully you're back now with us. Um, how you doing? Tell us what's your highlight of the week. Yes, sure. I'm here right at the AI Engineer Paris, hosted by Mistral, and uh, yeah, Uh, thank you, Wolfram, for reporting live and covering that event and going there. Shout out to the AI engineer folks from France. Uh, Mistral, I think, is co-hosting that, so big shout out to them. All right, folks, uh, 3 themes for the show of this week. Number 1, uh, pacing the frontier does not mean stopping.
OpenAI and Anthropic releasing back to back on the same day incredible models this week. Opus released 5.5, uh, ChatGPT released GPT-6 Soul and GPT-6 Luna. Terra is no longer there at half the price, so pausing, uh, the frontier does not mean stopping at any, any point. Number 2, every assistant is an agent, not every agent is an assistant. We're gonna have to start talking about AI assistants in a different language. Assistants have a heartbeat, assistants are proactive, assist- assistants know context about you, and assistants take care of their own context.
They have tools, they have a browser, they work while you sleep. Agents is a very generic term, and I know it's confusing to many people, but we're gonna lead the charge. We talked with David Palon from Assistant Benchmark last week, and uh, this week we're going to cover Meta Muse and the stuff that they did at Meta Connect. Those are AI assistants. There's a very thorough and distinct line between them, and assistants is what people need. Agents is a scary word for many people, so we're gonna talk about AI assistants, and I think at some point we're gonna go have like an assistant corner. And number 3 is voice. Gemini with a Gemini 3.8 Flash and Gemini 3.8 Flash TTS,
number 1 at voice design, and sounds like just absolutely human. So we're gonna definitely test this out on the stage. By the way, shout out to the beautiful folks from Gemini. Uh, I received like a mystery box, and I was like, who is this from? Uh, the folks at Gemini said, hey, we're launching voice, and they sent me this like beautiful microphone, and a t-shirt and a cap, and shout out to the folks in there who, who just thought about me, is which is very nice. But I think first of all, we'll start with the... Should we start with TLDR and then start going?
TLDR, uh, on the show this week. We'll start with the main three. Anthropic launches Claude Opus 5.5, claiming Fable 5.1 level work at 40% lower cost than Opus 5. I can concur. I've been running this model on ultra high yesterday in like 4 things, and I barely hit 15% of my weekly quota. It's quite uncanny after all of what happened last week, which I, uh, prefer to forget. Uh, just a few hours later, OpenAI rolls out GPT-6 Soul and GPT-6 Luna at half the price of the GPT-5.6, promotional API price.
We got price cuts from both companies, and it's very welcome. And I think the number 1 news, at least for me from the assistant perspective, is that Meta unveils that Muse is central to their whole roadmap. Muse will now have a voice and a video chat that you will be able to talk with your Muse, like me and Peter are talking right now. You'll see it like lips animate, etc. Muse will also come to the Meta VR glasses and a charm device that kind of looks like a Tamagotchi that you can buy for Christmas very soon. And uh, Meta also announced VR glasses with Muse built in- into operating system. So AI is kind of permeating through all of what Meta released, and Meta Connect
coverage will do so very soon. In other big news from big companies and LLMs, SpaceX AI released Grok 4.7, and it's reported 46% on CursorBench, and yet, and yet many people are very disappointed with that release. I think we expected more. However, they also have distribution. So Tesla now connects with Grokbots and connectors. So literally, you can now talk to your assistant while driving hands-free, and I think that that is more important news than Grok 4.7. I think that that 0.1 release of models is now augmented by where we can use those models or assistants, and Grokbot in a Tesla, which I didn't get yet, maybe after
the show I will, is just an absolutely mind-blowing experience. And also, OpenAI published a four area plan for independent safety assessment, but no assessors of start date, but this is a continued thread of what we talked to you about pacing the frontier and letting evaluators, external evaluators, in to your lab. In this week's buzz, folks, we're one week away from Fully Connected in Moscone's house next week. Pitbull, Sarah Guo, Fei-Fei Li, Alex Volkov is not nearly close to the level of those folks, but also we're gonna be there with Wolfram. it's gonna be sick, folks. I really want you to come, uh, to Fully Connected.
We have a free ticket for you. If you don't know where you're gonna be after Dev Day on OpenAI, we're, we're gonna welcome you, to see Pitbull chat with us. so the Thursday I show next week is going to be a little different. We're gonna start a little bit later, and we're gonna host a show, folks on the floor. It's gonna be exciting. Uh, we haven't done a live one since, AI engineer in Moscone also. So this is, you know, back-to-back live shows from Moscone, not too shabby, not too shabby, ThursdAI. Uh, and also CoreWeave earns the SemiAnalysis Platinum ClusterMax benchmark for the third report in a row. Shout out to the CoreWeave folks for this incredible, incredible news.
Uh, ClusterMax measures performance of NVIDIA kernels on how we set them up, and CoreWeave for the third year in a row is platinum tier. Folks, last week we told you about JEV. This is the TLDR. We're not gonna go super deep, but JEV changed the world since last week we told you about here with Ali Labs from TypeSafe until this week. JEV is everywhere. but as we told you last week, the open source community is gonna show up. So apparently some folks already trained decision models, and so there's like multiple, uh, classifier open source stuff. It's And, um, Swix, our friend of the pod, the host of Latent Space and the creator
of AI Engineer, which Wolfram is now at, hosted Diogo Almeida, the CEO, uh, and co-founder of TypeSafe, and they talked about Diogo's not like, he's really happy with this. They're not like, oh no, competition is gonna come from open source. He's like, yes, this is a new category. We all need to innovate. It's amazing that open source is gonna catch up, and I don't care about benchmarks. So a lot to talk about how JEVI this last week was and how it took over the feeds, uh, but also We need to talk about Gevbench from our friend Florian. Maybe he'll show up, uh, that measures other classifier models and System 1 models. Uh, so we're definitely gonna talk about Leia and, uh, and and
Jeff. There's Jev and there's Jeff. Uh, let's see how my transcription deals with me saying Jev and Jeff back to back. Okay, folks, voice and audio is also just blowing the fuck up. Google debuts Gemini 3.8 Flash and Flashlight TTS, claiming number one on the AI voice design benchmark. Folks, this is a cloning your own voice system from Google. We told you about this 3 years ago, that nothing is gonna break in the world where audio models will be sounding like you. Nothing is gonna happen. And now Google, one of the biggest companies in the world,
is allowing you to clone your voice and play it back within like seconds, and it sounds really, really, really good. Um, we're gonna play around with this. We're gonna play you some samples of including me speaking. It's gonna be dope. Qwen also debuts Qwen 3.8 OmniFlash, agentic audio video model with 1 million token context, and Qwen also has a live translate model, cutting its reporting interpolation lag to 2.3 seconds live translating. Qwen has been all over the open source news as well with vision, image, and video category. Uh, Alibaba Qwen open sources Qwen Image 2.1, a 7B model for image generation and editing with native transparency.
Black Forest Labs open sources Flux 3 Action, a 7B. Do Uh, a 7B world action model that it says tops Rovalab at 42%, so, uh, you know, Black Forest Labs is back. And um, this is quite crazy, folks. Karina Nguyen unveils act 486. A pre-rendered research demo that lets viewers talk back to a Joe Rogan episode. it's quite the eye-opening demo. It's quite crazy. In open source LLMs, we have a big one as well, Xiaomi. Uh, open sourced Mimo V2.6 Pro, and artificial analysis rocks this, the top
open weights model at 46% artificial, analysis intelligence index. And there's a bunch more stuff that we haven't covered in the TLDR, but, Wolfram, maybe one or two more items that you want to make sure that we need to at least mention. Yeah, So, uh, NVIDIA Nemotron 3, it's a diarization model that came out open, of course, a 100 million parameter model for speaker tracking and even overlapping speakers. So this is super exciting. And, um, yeah, in other news, OpenRouter has a batch API which can save you a lot of money if you are using OpenRouter and you don't need the response now.
If you can live with it 24 hours later, then you can save up to 50%. Well, that is something, uh, cost is a very important aspect. Yeah, I think those are the 2 main things. Stepfun also has a new agentic model with a million context, but it's, uh, it's supposed to come out as open weights, uh, in October, but it's, uh, API usable now. I think at least I saw it. All right, folks, I think this is the TLDR, and I think it's time for us to kick off our show officially, uh, with,
with Frontier Labs. Let's go. Alrighty, folks, I think let's, let's start. Folks, Claude Opus 5.5. Anthropic just shipped a model that beats their own fable 5.1 that just recently was released on real work, and it is 40% cheaper than Opus 5. And honestly, it just feels limitless. My week started a week ago with burning quotas all over Fable, burning my quotas
ranges, and I was like, oh no, I cannot do this work because for Thursday I need a cloud for production. Then Opus 5.5 feels limitless because I was barely able to get to the end of the quota, and I was running workflows and and agents, whatever. Um, Opus is back. It feels like Opus 4.6, and you can talk to this model, but it's as intelligent as Fable as I was able to to see. So, uh, I think the timeline agrees. I would love to hear from folks here. Uh, for coding and agentic stuff, most of your bill is cash reads, and so the cash read, uh, is a significant price drop, and you can feel it. Also, it's fast. And, uh, Opus is also producing the show, Opus 5.5, so hopefully
we'll see like a little banner that it's gonna put up for us. Uh, but I would love to hear from you guys. Uh, how was, how was your few days with Opus 5.5? Let's start with Peter. Yeah, well, first thing to say is that, uh, we released a score for Code Arena, which is measuring front end and all of that fun stuff. And guess what? It's, uh, on top again. So Astra was above everyone else, and now Opus 5.5 is above Astra. And, uh, yeah, it's super impressive on, on that kind of coding front end. We got a bunch of new capabilities in terms of using the, all of the HTML generations.
That seems to be just completely insane, and all the 3D stuff. to your point about efficiency and cost, obviously they did reduce cost by 20%. that is good. They said it's using fewer tokens. It's kind of interesting. I kind of had mis- mixed experience with that when I was running, the harder prompts on Max, it was running for many hours, and it was, then I calculated the API cost, and it was like like 80 dollars per generation. So I would guess maybe if you take a median, maybe it is fewer tokens, but then if you take the top 10% of prompts, it's probably way more tokens, is my sense. it's just kind of vibe-based sense.
I don't have data for it yet. So that feels kind of different. I was comparing to, GPT-6 Soul at the same time, and I had like Soul, $1 and Opus, $60 side by side. So, yeah, I had a few experience like that. I'm not saying that's universal, and I do agree with you, in just day-to-day use, it does seem to last longer than than, your sub. Um, but yeah, I think it's not, it's not just completely across the board. What happened to me last week was... for the first time, the things flipped. I remember being very cautious about my kind of like quotas on the Cloud Max account, and I remember Limitless in the GPT, uh, you know, Sol and the 5.5
days. I remember limitless codecs, just never thought about reaching my quota, like never, just never. Uh, and then Astra came out, and then within a span of half a day, I finished my weekly quota and hit the reset, finished my weekly quota. Uh, not to mention I got bad results. We, let's not talk about this. Fine, we're over. and it flipped now. Now I ran Anthropic Claude Code for the longest time with just no end in sight, and I, this is quite incredible to me, but also the performance of this model, it works very well. Uh, Wolfram would love to hear from you. You know, you've been traveling, and Nisten as well. I don't know if you already had a chance to try.
Feel feel free to chime in. How has your experience been? Because As much as I don't like the politics of Anthropic, I do find, because Ti- I was a TypeScript dev for like 6 or more years before this, it writes the best code. So I've been using it the whole time. Opus 5.5, uh, there are some drawbacks, like it is still bad at DevOps. I would still revert to Fable for anything like configuring networking, DevOps, containers and stuff on Linux. But when it came to anything to do with the web, WebGPU, web APIs, anything to do with Chrome, it's just been trained on much
newer data, and it understands even the int8 optimizations of, WebGPU, and, it was pretty crazy at writing kernels, both like on SGLang on hardware and, uh, making websites and also using WebGPU. And, yeah, the model, it's, it's just an absolute banger. It's, it, it is that good. Uh, it, it does everything, it does it fast, it uh, knows how to troubleshoot better. There are a few things which you'd still revert to Fable, but uh, yeah, overall, that is my default now.
on most things, yeah. Folks, you, do you understand what it takes for Nisten to tell you that his default is not some open source obscure model that he runs himself on 17 GPUs? Do you understand what it takes for Anthropic to win over Nisten? I don't think you guys quite get what just happened here. This is my default as well. This is a fast model that is as smart. I used to love Opus back in Opus 3 days. Opus was like Opus, and then they screwed it up. Somebody from Anthropic said, hey folks. we're sorry for Opus 5. We hope this model will make up for it. Because Opus 5, and we told you about this, just like 5.0, was an absolute crap
model. Nobody wanted to talk to this. It was slow. It was, uh, incoherent. It used Claude-isms. Claude-isms are gone. Jargon douche is gone. All of that is gone. This model talks like a normal person. And here's a tip from me to you, if you haven't got there yet, if you don't use Pstack from Lauren 10, I believe, from, uh, the GrokCursor team. Uh, Pstack has this beautiful skill called Bro. All it does is when you get something that's very, very confusing, you do slash bro, and it asks the model to repeat it in human, uh, words. I, I love using that skill.
I was like, every time, even Opus 5.5, or even before, it gives you like some Claudeese, you're like, bro, and it explains to you like a human. Um, so, you know, we'll link that. in the show notes for the folks. Anthropic, please don't nerf Opus. This is a strong request from us. There's this whole thing about models coming out, and they're amazing, and then after a few weeks, they're not amazing anymore because the hype died down, and maybe you tweaked inference. We ask you, keep Opus 5.5 the same as it is now. Please do not ner- nerf Opus for us. I think with this we can move on to the next news.
OpenAI just cut GPT-6 prices in half. And the funny part, the cheap models might be better to use than the flagship model of Astra. GPT-6 Soul and GPT-6 Luna, the faster, cheaper, and probably more, performant siblings of Astrum, uh, dropped at half the price of GPT-5.6. So this is a new next level model at half the price of the previous model that was promo. Soul is 2 dollars in and 10 dollars out, and folks, Luna, GPT-6 Luna is 10 cents in and 50 cents out, which is like, it's nothing. This is the best cheap frontier model that we can now have.
It beats the previous, uh, model from GPT also. GPT-6 family now is Astra, Soul, and Luna. Terra is gone. No Terra anymore. And honestly, what needs to be renamed is Astra should be Soul, Soul should be Terra, and Luna should be Asteroid or whatever. Uh, but okay. On DeepSui, Soul, GPT-6 Soul gets 68.8, and Luna gets 66.6. And OpenAI says that Luna is roughly Opus 5 and Fable 5 at medium effort for 96% less per task. I don't know about that. I don't know, but I've been very, very impressed with Luna specifically.
Here's the sneaky big, uh, news for the builders. Caching. 90% of cached input. All these labs understand that we are now running long ass agentic stuff for over 500, like, thousand tokens or so, and all of it needs caching. And, uh, changing reasoning effort or tools doesn't break your cache anymore. This is like the big highlight from both, I think, Anthropic and GPT, and I think that's great. You Uh, if you run agents or assistants, this is huge. So here's my honest read on this, and I said this in my notes. Astra has been kind of dumb for me in the daily use, like, honestly, it's been burning tokens and it's been giving me crappy results for the stuff that I care
about. Sol is great, and Luna is even greater. Uh, Luna you can integrate into your own products, and it's very, very, very cheap and also fast. These ones I- I'd actually now use. Uh, Reddit isn't fully sold on this. Uh, our Kotex crowd says that the performance gain on this is marginal, and the real news is the price. And couple of threads are showing Sol scoring below GPT 5.6 on DeepSui. Uh, all right. Folks, is half the price with same score a win, or did OpenAI just lose the race a little bit? Let's go to, let's go to Wolfram. I know you posted some stuff about Astra. Wolfram, let's go.
just quickly, because I, I switched to, uh, GPT 6, uh, Soul because I was using 5.6 more because Astra, it's just eating tokens too much. And I noticed in my personal assistant that the humor has been gone and the soul, the writing, it was different, and I didn't like it, and I've seen people online say the same thing. And, um, 5.6 Soul remains my favorite model from OpenAI. So that happened with GPT-5 when it came out, it was so robotic in a way, and I feel the same with this one. So actually, I would rather stay with 5.6 or 5.5. Those two, those were the big ones for me.
It feels like the 4.0 keep 4.0 movement now and say keep 5.6, please. Um, yeah, unfortunately. Or maybe I do have to go to Opus. You should, you should, uh, you should try Opus. Um, Peter, you kicked off this category. You've been playing with obviously Astro. You're like, you're super fame. I think, dude, I think your level of fame jumped because of your Astro demos, uh, and now we got other GPT-6. So now the GPT family of models, GPT-6 family of models is complete. And may I just say, just before we get to you, that OpenAI's Dev Day is next week, folks. And the fact that they released their agent platform, which is directly for developers, and then they released GPT-6 Soul and GPT-6 Luna, and Luna is
priced for developers just before Dev Day, means what the hell is coming up on Dev Day? What, what are they going to not? I, I, I think I have an idea of like one thing, but just to highlight that we'll be covering Dev Day. Peter's gonna be there, I'm gonna be there, and we're gonna talk to like a bunch of folks, and we'll bring that all to you next week. But Peter, about GPT-6 stolen Luna, please. What, what was your experience? Yeah, I think it's, they kind of went opposite ways, I think, with, with Anthropic, and, the definitely the important part there is efficiency, with, obviously with price reductions. I think that's the kind of the obvious point, and 50% is a lot,
right? the, and the second part there is that it's genuinely using fewer tokens. Right, we, I looked at our data a little bit yesterday, so no- nothing conclusive yet. We need a bit more data, but it genuinely looks like it's using fewer tokens, and it's gonna rank higher than 5.6, uh, on our leaderboards as well. So that, that is the important part there. So I'll say it's kind of iterative improvement in terms of quality. I think it is better than 5.6. I think, uh, it is, at least for me personally, it just feels like smoother model. and for me, it is the kind of the daily driver model. If you just want to do normal stuff, that makes the most sense to use.
It's not going to burn your sub. And, it's, I think the problem was, and Alex, back to your point about just burning your subs, is that we were used to pretty generous subscriptions at, medium- sized models. They were big at the time, so like 5.6 and Opus, and Mm-hmm. Remember Sonnet? Like, we used to use that, that thing as well. So if you're using these models, your sub's so fine. You forgot to say this. Sonnet and Haiku 5.5 are coming. Are they? Yeah, please go ahead. Okay, okay, okay. Yeah. So remember we were using these models, and then your sub was fine, and then they give us Fable and Astra, and you want to use these models instead, and then
your sub just dies in, in like 3 hours. And I think it's not, and I don't think we should be blaming the labs, like these are more expensive, bigger models Yeah. and that they need to run, and they genuinely, I don't think they're like lying to us. I think they just genuinely don't have compute. So they gave us more allowance, we just kill it quick, and they just wouldn't have compute. So OpenAI already, remember they had to stop the signups, right? So yeah, there's no 200 dollar a month, pro tier anymore on OpenAI. You only can get to like 100 dollars. on that, also, if we're comparing the two releases, kind of like Opus is still the better deal. Like, you can get 200 dollars worth of, like 20x usage, and
OpenAI had to stop because so many people wanted Astra, specifically Astra is burning. Question is, my question for you, Peter, is 6 Sol right now replacing for most people what they thought they needed Astra for? Yeah, I think it, it's a little bit tricky. my general rule of thumb would be is if you have an OpenAI sub, you should just use uh, 6 Sol for everything, and then just switch to Toggle when you need Astra, uh, when it doesn't work. it's not that difficult, and I do that. That's kind of what I do myself, and that feels good to me. I don't feel like it's terrible or anything, and it, and I, my sub goes pretty long way now. Um, but yeah, I think it was just never sustainable.
I think it's just unrealistic for us to say, oh, I want to use Fable and Astra at my 200 dollars in exactly the same way how a model used to be, which is like quarter of a price. It's just not that. Can I say something? I don't know if it's controversial, bro, but he- here's my take on this after playing, going very hard this week, folks. Uh, if you guys remember the, the token billionaire thing that the AI engineer invented and the card that there's now finger prone token billionaires, uh, I think I'm at like, like 8 billion tokens this week. Uh, I'm working on something super cool. However, I've tried all the models and all the harnesses as well. So I've tried Cursor, Devin, directly Codex, directly Claude
Code. I have my thoughts on that. And here's my take. I personally, for my work, I don't want an extra powered model that can solve Navier-Stokes equations. I want a model that can understand me and build me a simple fucking settings page without going over engineering, writing a thousand tests. That's what I need right now. And it feels like what we just got with Opus and Sol, so like the, the, the one rank lower models that also are sparing your quota, is quite that.
Opus, in my case, is significantly better at understanding me and doing some stuff exactly like I like, and and Sol still needs, like, very direct instruction and is very good at, like, computer use. for my work, I need a reliable, fast assistant model that can understand me better and run fast. I don't need Navier-Stokes. So it's a very interesting thing that that we started the show with. We were pacing the frontier after last week Dario Amodei and Sam Altman agreeing, Elon Musk, whatever, and we're still getting a lot of very incredible advancements in practical, everyday AI brains, and to me, that's great. Uh, Nisten, we haven't heard from you on the GPT stuff, CoreWeave.
Thank you, Wolfram, for hopping on. And uh, taking Wolfram's praise, we have back on the show Yam Peleg. Welcome, Yam. We are now talking about GPT-6 Soul and Luna, and uh, Nisten is up. I use Astra, but mainly I just let uh, Claude Fable manage it, what it does. It's just for the work that I do, and I'm pretty opinionated. I really don't like the taste it has or how many different files it, it creates. So you use, you use Anthropic's taste and then OpenAI's execution? Uh, I mainly just use OpenAI for, restyling, because it does produce
much better websites in terms of design now, and I use it for reviews, but overall I don't like having it make a whole bunch of files in my repo. I, I just don't like the style it develops at. and I'm a, uh, but I've been very used to. Nisten, you remember we're talking about GPT-6, Sol, and Luna, right? Like, Astra is, is a model that, uh, have you tried the new ones? Have you tried the Uh, no. the smaller brothers? No, I just haven't, have not used those. and I would love to know why, because I think, first of all, a lot of folks love your takes here on the show because they're not always like going with the hype, which is great.
Uh, it's cheaper, it's faster, 68%. Why, why didn't you? Just because you already have your workflows, or do you think that your prompts will not align with the new models, or just like too many news in the span of a short week? I mean, it's just for what I do. I know the majority of people do prefer it over Opus. It's just for what I'm doing. One, even Fable and Opus 5.5 are not good enough at DevOps, and I find Astro worse. The only thing I find it better is that designing stuff, and uh, if I'm doing kernel work, it just makes way too many files, and it just
feels, it feels dumber to me. So I just use it as a second opinion. But then again, I've never been a fan of how ChatGPT models respond, and uh, even uh, when I've been working with people, I could tell right away when it was like a ChatGPT written report, and uh, like I absolutely hate it. I hate the way it writes. Sorry, it's just me. Uh, not the best writer, I agree. Yam Peleg, uh, welcome to the show. So. GPT-6 Soul, GPT-6 Luna, no Terra anymore. What's your take? Have you used those models?
How are you feeling about them? which one of you is in what camp? Who is on, on Opus? Who is in GPT-4? I don't want to take sides. I want to run both, but I will say Opus is doing just incredible stuff, and I definitely, definitely have used the trick that I will now teach you. Uh, Weights and Biases Hive Mind is a tool that stores all your conversations from all your agents across everything in the cloud. It has a very nice feature fork, so you can like, you can download another's agent conversation regardless of which machine. I have definitely used fork my Codex conversation into Claude and say, hey, review what he did
and tell me what you d- you would do differently. And that's been my, like, this is the direction of the stuff that I use. I haven't used the other way around. So I haven't ever used OpenAI's stuff to go and review Claude, but I definitely use Claude to go and review Codex. That that's where I fall, like true- That thing, that thing sounds useful as hell. yeah, it is very useful. Yam, where are where do you fall? Look, uh... Uh, let me just say directly, I wish I could use Soul enough to give you an opinion. Mm-hmm. It's just it burns the entire account really quick.
So does? Yeah. And at least for me, at least for me. So, so here's, here's the thing that I, uh, I complained about this on my timeline a lot. I don't know if you guys saw, uh, Astro just like decimated my limits. And then, uh, helpful folks from OpenAI. Shout out, shout out to incredible Deverel folks, by the way, on OpenAI. Yeah. Uh, Dominic Kundul, a friend of the pod, was on the show. Eric Povancher, Anyway, he went to OpenAI to work on on the stuff. He's very, like, adamant about, no, folks, you don't understand. This is a a skill issue. Sol and Astro are amazing models. Eric Povancher, um, he was also replying to people, and then I noticed that I had some settings in Codex
that fucked up with my whole setup. Which setting? Here's the setting that I had. I, a long time ago, played around with the context window, and I set mine to like a million context window, which, uh, OpenAI models support. This absolutely kills their cache system, and so this like rips through your tokens. I also had experimental settings for having it keep it notes for itself, and I also played around with how many sub- agents it opens. After removing all of those and asking the model to actually open the sub-agents in a lower tier model, I actually started seeing normal use back in Codex again. So if this was your issue and you felt complete burning of tokens, see if your
Codex.toml or settings.toml file is, uh, carrying some older stuff. Uh, all right, all right. I must admit I need to check, but I'm playing with the codex.toml file all the time, by the way, it's like all the time tweaking things. There are so many hidden, hidden stuff that are like you should, you should tinker with. Uh, just want to say, look, a sol that is more efficient, like 5.6 sol that is more efficient, even if it wasn't even better, I'm taking this seriously. it's a really good model. It's just that it eats the account really quick.
And just on the contrary, I just want to say, Opus ever since 4.6, I had something to say about each of the models, but the new one, 5.5, is just good. hands down. There there's no denying. I'm really confused about your burning the sub point because, uh, in our, like, in my personal use, I don't think it's burning my sub anymore. And also on, in our data, we can see it costing about half, so it's like more token efficient and cheaper. So I don't see why I should be burning it more. So I would also say this is such a common pattern for me, whether it's Claude Code
or Codex or whatever else. I find something super annoying, something's not working, and 9 times out of 10, it's some stupid setting somewhere, some Oh yeah. plugin is broken or something else. So it's good idea just as a general tip for everyone, like if something is annoying you about something, just ask your agent to see what's going on and diagnose it and fix it. Just I come across something like that every week. They should, everyone should have like a heal button or something, just like Yeah. Let them run through it and like reset your settings or something. Just way too common. I will say, though, both Astra and the new models, I've noticed this as well,
suffer from the pause syndrome. Can we talk about the pause syndrome? The pause syndrome is where, you talk to the model, you ask it for like a long thing to do, and then you steer, and then you interject with your thoughts in the middle because OpenAI came up with this amazing steering concept where, instead of stopping the model and inject your message, you basically steer. I think Anthropic got there eventually as well. Oh yeah. Apparently, for the Astra and Sol models, steering too much after compaction causes it to lose the original direction, and the model just stops working. It will say something, you would steer and ask it a question, like a by the way
question, the model will answer and stop. This pattern has happened to me and many folks on my timeline, and what Eric, f- from OpenAI, Eric suggests is to ask the model to write down and jot down some messages for you when you steer and continue working, because the steering is the problem. OpenAI needs to fix this in their own, harnesses, et cetera, but, steering is absolutely the problem. So, uh, if you steer a lot and you get Astra pause syndrome, they're not pausing the frontier, they just, the model is confused after compaction. So this is a tip for you from the folks at OpenAI. Uh, these definitely are some issues that, h- have plagued my timeline as well.
Folks, I think it's time for us to... Anything else that's missing that we haven't said? the only thing I will repeat, OpenAI released the agents platform, managed agents platform that runs inside your stuff last week, which is incredible and sounds like a this should be the tool that they would release on Dev Day. They now release 2 new models, and Dev Day is upcoming next week. Uh, speaking of the Hive Mind tool, producer put it up for us. Thank you, producer. Every agent conversation and forkable, and it's free to use hivemind.1b.tools. This is the brainchild of Chris von Pelt, one of the co-founders of Weights and
Biases, because he used to work across many computers, across many harnesses. It's really damn useful for you if you run multiple, harnesses over multiple, computers. Right, folks. for, I think it's time for MetaConnect. Meta AI News app just passed ChatGPT in the number of people, uh, getting it. It's number one in the App Store, and at Meta Connect they gave it an email address, a control of your Mac, and a place on your face. Meta Connect happened yesterday, and I did my bingo card on Twitter, and I got some of the stuff that I wanted for the
bingo card. The... the headline is Muse. Zuckerberg was up on stage and said that the central piece to everything they're building at Meta is Muse. Muse is Meta's personal AI assistant, and it's number one downloaded app in the App Store right now, which a year ago, if you told me that Meta AI would be outranked ChatGPT, I would probably, I would believe you, as we told you, don't bet against the Zuck, but many people would, would have not. as a reminder, last year, Zuck started hiring with some soup all of these incredible developers and started bringing them from different labs, and the soup paid off. So,
now Muse as an assistant that sits on your Mac, it can control your Mac, they're giving it its own email address so you can CC conversations and real-time voice and video conversations. So I assume by this time next month, Muse, my Muse specifically, could join the show as a participant on the show. It works in background while you're talking to it. This is the what separates assistants from agents, and there is an avatar you can talk and voice and face design yourself. You can customize your own muse and voice design it yourself. And Muse is coming to Meta glasses, you will be able to walk around in the world and chat with your conversation.
And the custom wake word is the highlight feature for me. You'll be able to say, hey, Wilfred, instead of hey, Meta. I think that's super cool. I think if they're training custom wake word for me in my Ray-Bans, I'm gonna walk around in the world. This is like us having our own, uh, you know, Tony Stark assistant. Then there's the hardware, and there's a lot of hardware, Ray-Ban Meta Gen 3 with new 6 microphones and aviary styles, The first audio-only glasses that people cannot see the cameras anymore, and not only that, the audio glasses, the video glasses too, they work with the FDA to make them, like, certified hearing aid, which is, I think, is a genuinely great move
because of the ADA Act, the Disabilities Act in the United States. if you're disabled and you wear meta glasses as hearing aids, you can sue people who tell you to turn them off because of the ADA. I think this is like a genius play by Zuckerberg that's not like talked about. Uh, so audio only glasses with no video, uh, and the lightest frame they've made. They kind of, Peter, they kind of look like your glasses, honestly. They just look like there's no, no indication that this is like listening to you. Um, and they cost a fraction of normal hearing aid. And honestly, this is could be the most important thing that they announced, and
VR glasses. Meta announced a VR headset. That's not a headset. This is the first time they look like glasses. In fact, the the big reveal for which many people, compare Zuckerberg now to Steve Jobs, is that the VR glasses that they've announced have been on the table with other Meta glasses, and people didn't even notice. It's quite incredible. And then I think we're getting to a point where Meta is beating Apple and OpenAI's whatever rumored hardware. They announced this little Tamagotchi-like device that your muse is sitting in a keychain-like device with a screen at the size of the Apple Watch, pretty much, with a thumb fingerprint
that is shipping in December by the time for the holidays, and no price yet. I think that this is going to be sold out the second it drops, as people think about which gifts to give other people. Peter, have you tried Muse already? Is it available in the UK? No. Mm-mm. I haven't tried it yet. It's coming, it's coming. I think what's been missing, and like, if you remember the trajectory, right, is that OpenAI went consumer, then sort of competing with Google, then OpenAI realized, oh damn it, Anthropic is eating our lunch, so we're gonna go more enterprise. It's not like they've forgotten about consumer, they're gonna release devices as
well, but we certainly, and and Google's been like okay at consumer stuff. So we've been missing a, a lab that is very, very good, very focused, and has a genuine consumer interest. And uh, yeah, Meta is it. So it looks like they're really just going, going all in on making consumer great. And like Nat Friedman, right? The, he understands this a lot. He's one of the key people behind it. And it's making things really fast. Nat Friedman was the CEO, so sorry to interrupt. Nat Friedman, just Friedman was the CEO of GitHub. And then Nat Friedman and Daniel, Gross became one of the more prolific AI investors in the Valley a long time before the AI wave started.
And then the biggest news about the Meta Superintelligence Labs is that, Zach was able to bring both of them on. Specifically getting Daniel Gross away from Ilya Sutskever at SSI into the MSL. I think this is like a big, big, big, big move. And Nat Friedman is behind Muse, and Alex Wang is showing up with Crocs and a crocodile thrifted t-shirt on the stage of Meta Connect, and talks about how Nat Friedman, worked his ass off on this Muse thing. Nat Friedman doesn't talk on Twitter too much. He said that once OpenClaw dropped, they all stopped everything they're doing, and he brought 400 Mac Minis into the MSL lab to all run OpenClaws all the time to figure out what's going on.
Just so you understand how early he was, this was in, I think he said February, January. That's when we all bought our Mac Minis. Meta has a top level exec that now drives the ideation thing as early as when we tell you about stuff before it blows up on the internet. That- that's how they got to Muse now. The security of Muse is they're going for unparalleled. They understand that I'm going to put more and more trust in this device, and unlike the folks who ran OpenClaw on their Mac Minis, et cetera, and that didn't really care about, those are folks who are putting their, like, bank account information, integration with 1Password, et cetera.
This all needs to, jump over the fear that folks had that Meta has all my data for ads. So they're talking about a few things here. One is the business model is new. There's no more ads, no more your, your data. The business model is transactions fee. Muse and other assistants are going to do so much for people shopping on the internet, which, by the way, they announced integrations with Best Buy and Walmart and other folks at the same way that Amazon has banned Muse from shopping. So if you use your Muse for shopping in Amazon, if you use a computer browser, Amazon will ban your account, which think, is a ridiculously stupid move, but I think
it's just muscles, like they're, they're still negotiating probably the, the cut with Zuck. Um, and this, these assistants will do a lot of shopping for people, and Zuck is like, okay, this is the business model. No longer do we sell your data or train on your data. We take a small cut from the seller's fee to you if you use Muse to buy stuff. I think it's a genius model. It's a win-win for everyone. I think, yeah, it's the problem is it's hard to make money that way. ChatGPT backed, uh, backed away from this model just because it's the, we can go deeper on this, but basically the way you make a lot of money is advertising because it's, uh, you know, you can kind
of raise prices on your suppliers without consumers paying more for it. But look, I think- I don't think that Facebook is gonna stop advertising elsewhere. I just think that, for muse to give it out for free, the business model that they're advertising for it is that, hey, we're gonna have a secure VM, which was co-architected with Moxie Marlinspike, the creator of Signal. That secure VM is at the level of a cryptographically verifiable, solution that even Zuckerberg himself would not be able to read your data because it's end-to-end encrypted. That's what they're touting live on stage. So they are going after, hey, we know about privacy, we know what people think
about Facebook, et cetera. H- here is industry level, architects that say that I cannot read your data anymore, but I can, like, look out at your transactions, and I can take a fee from your connectors. That's kind of the, the business model. Yeah. I think, anyway, we'll see. I'm, I'm a little skeptical, and I think they're probably also doing it unnecessarily. You know, our parents don't give a crap about that, and they don't watch these keynotes, and they're not gonna know what that, any of that means. So I think that's my, maybe over-egging it anyway. You, Mm-hmm. But I think what, what I would like to add is that what's really impressing me about
how they're approaching it is that I think it's the only, like, really big consumer app, uh, the where they don't treat people like idiots. they give the powerful tools. It's all the agentic stuff. It's all of that. and it's not like some crappy flash model. I think they're taking it serious. Now, obviously, like OpenAI, to be fair to them, I think their ChatGPT is like the best it probably could be. But if you look at Apple, uh, as in the free version, and if you, but if you look at Apple, it's just like, oh yeah, it's gonna like summarize your notifications. Like they're just not, not particularly interested in, uh, in really giving the
best stories to consumers. Yeah, they can't or they don't care or like, uh, I don't know what, but I think they also care about the experience, so they just want to be like, oh, we want to be the last ones, but do it the best. May I remind you of, uh, a famous actress, I think she was in Game of Thrones, that, uh, took ads for Apple and their Siri AI 2 years ago, where Siri was supposed to be this like assistant thing that reads your messages from your mom and puts on calendar and does all these things. And then Apple pulled those ads because they couldn't deliver on this promise. And then Apple launched Siri AI, and it fizzed completely in the air, and iOS 27
is now on many, many devices, and many people will use Siri. And we had such great hope for Siri as being that personal assistant with access to my stuff, and that absolutely fizzled because Siri is not agentic in any sense of the word. And I think for us, what we need in assistant is proactivity, in addition to context, in addition to connectors, to everything. I don't need it to be super private on my device, whatever, because my device cannot, read, uh, I don't know, go on my Twitter via connector. So that's, that's what's missing for me in, the Appley stuff. And I think that, eh, whether or not all the folks care about privacy, I think it's incredibly important
that the company that gives them that product does care about their privacy. And so I absolutely applaud the folks from Meta for following Apple, and Apple famously has the secure cloud where your LLM execution with Siri and then go, if they need bigger LLM, they go into the secure cloud. Apple cannot read this, even FBI, FBI comes knock on their door. Uh, Meta is working on that for folks that use Muse, now that they're shoving Muse in the face of the billion people across their devices. They're shoving Muse in every Instagram ad, and every Instagram story has a little Muse thingy. Uh, obviously the Facebook crowd will gonna get there very, very soon, and, it's
been exploding. Everybody who I know that's not in our circle in AI knows about Muse. It's quite crazy. how strongly Meta has, like an effect on the populace. And I've been talking about this here. Many folks use Google, and then ChatGPT came about, and many folks use ChatGPT as the new Google. They just ask a question, get an answer, and they continue to the next one. We are like, you know, advanced there. We're now using a agent and assistant for coding. The next iteration is having an agent that's an assistant. There is a difference there. That is 24/7, has its own computer, has its own browser, and does for stuff for you proactively, reminds you about stuff proactively, uh,
renews your driver's license online, which was what I had it do last night. Grok went and renewed my, uh, expired car registration, for example, and did it on its own. And the assistant can do shopping for you and pay for you via the Stripe Link integration and the 1Password integration. So I think many, many people will get there faster. I think they'll skip the agent part, they'll just go to assistants. Um, Yam, thoughts on Muse? Are you getting Yeah, look. FOMO from not getting it? First, let's, let's just say that I already got it. You already got it. Amazing. Yeah. Absolutely, yeah. I, I tried it and it's brilliant becau- because of, you
hyping it every, every week, so I, I had to try, and it's absolutely brilliant. Yes. But, uh, can we stop and talk about it, uh, overtaking ChatGPT number of users? Seriously? Not number of users, the speed with which Okay. it acquires users and the speed with which it got to number one on the App Store. Mm-hmm. Okay. That's the That, that, that makes sense. But, But, but they will surpass number of users 100%. Meta has the biggest surface area products in the world. That Zach literally stood on stage and said, we have the most people using our products in the world and everybody else, and they're
gonna shove Muse into all of them very soon. So I expect it to, like, surpass ChatGPT in whatever. Also, we expect OpenAI to come back both with an assistant model, which GPT for work kind of already is, but people don't know because it's hidden in there. And do you guys remember Jony Ive and LoveFrom and OpenAI and the devices thing that they promised and Sam Altman putting his head on Jony Ive's shoulder and collaborating together. That's also coming at some point. OpenAI's hardware is coming, and where is it? Uh, yeah, yeah, so back to Muse. You... Look, I don't know about the the end device.
I must say that why won't I use my phone for this? I- I'm asking because you sound really excited about this, and Yeah. I'm like, I don't know, my phone works great. I just want the software, I want the assistant. Like, convince me that I want the device in general. I will let Jony Ive, convince you when their device comes out. I will just tell you that I cannot wait for Muse on my glasses. Why wouldn't you do this on the phone? Yam, you could. You could. However, putting out a phone and clicking buttons, shoving this into people's face when you talk to them, it's like a whole another thing that you need to do. Whereas just saying a keyword and then just going about your day, looking the
person in the eye is a completely different experience, I believe. Nisten. Muse in Canada from last week. That's also big news. So you got it. Uh. Yep. Tell us. Uh, I got Muse. I really liked it because I just put it to computer use at first place, and it ran into a CAPTCHA on Temu. I was just asking it to look up products on Temu, and then it had a remote desktop session, and I noticed it was a Chromium session running inside it, and then it would have a, inside the app on my Android, it would have the remote desktop window, and I can go and I, I could click the CAPTCHA with my finger, and then it would just continue on.
That, that was the best user experience so far, and uh, I love how open they've left it. Uh, I'm not going to reveal too much, but it is like and you have root on the whole machine, which is pretty nuts. They're just letting everyone have machines running as root, and you can actually do whatever You mean the you want. the VM that comes with Muse, Yeah, yeah. the virtual machine that they, yeah. It's an 8 gig VM with, uh, 100 gigs of, uh, storage space. The thing is that you told me last week they had a TailScale integration, Yes. so I was able to TailScale in. So now I have a complete shell.
It can, you can even like run a model inside the 8 gigs of RAM it has if if you wanted to. They've done a really good job. I want folks to understand what's going on. It's it's insane, like that Zuckerberg is they are on another level. For free, everyone, 100 million tokens a week, plus a whole computer in the cloud that is a good computer in the cloud, for free, for everyone. Do you guys understand the scale of what? One more thing, as for privacy, the people, you can just open a new Gmail account, and you can just log in with that Gmail account and have nothing connected to it. Because this week they announced that Muse is going to get their own inbox.
They'll be able to just CC on your conversations. Okay, they're making it like Yeah, they're making it less. One step less. Folks, assistants are here to stay, AI assistants are here to stay. Uh, we're going to cover use cases on the show. So I want to share 2 real use cases for AI assistants, and I think that we will have a corner here that will just share our experiences because folks, when they get those things and they don't know fully how to utilize them, although Meta did an incredible job in the ideas. So number one, if you get Muse, go into the ideas tab. Number two, when you get Muse or Grokbot or whatever assistant, Folks, go on a
walk, hit the record button, and just talk about your day. Talk about your business, talk about your workday, talk about your family, talk about, talk about, talk about what you do. Because if you try to imagine use cases that it can do for you, that's great. But the thing with Assistant is that, like, this is a whole new paradigm for you of stuff that it can do that you don't even know it can do. And a lot of the stuff that you can do or need to do comes from your emails, etc. So number one, connect your email inbox, uh, so it knows. And number two, just talk to it and then say, hey, here's the stuff about me. The beautiful thing about Muse and other assistants, they have this memory file
that they put about you and about your contacts, and in that memory file, th- this is how assistant knows how to help you. The assistants will only be as helpful to you as the information that you give them. And an additional tip is that there is an import memory feature when you go to Muse and you click a button there, it will give you a prompt to send to your ChatGPT, to your Claude, to other assistants, to basically brain transplant whatever information they have into Muse. You can review that, and you can send, and then your assistant's significantly better. And then the idea stuff that they're surfacing there, is going to also be significantly more contextual and stronger for you.
So those are my tips for now. A real use case that I had is my weekend bot. It's in Grok, but Muse also does this. I arrived up very early to put my kids in a karate class on Saturday, only to find out that I should have known that they are closed that uh, week. I could have slept for another hour. Immediately then, I talked to all of my assistants to, first of all, check all my plans for the weekend and their websites and emails, because I don't look at their emails for cancellations, and two, plan out weekends ahead of time. So now my assistants go search local news ahead of time, so by the time that I arrive on a weekend, I have a list of plans that I could be doing with kids or
with my wife, and that's a weekend planner prompt recommendation for you that you should, you should add to your bots, like, hey, every whatever Thursday, go and look at the local stuff that are happening around my city, sort them out by free, paid, whatever, premium tier, and do this for this weekend and the weekend two weeks after, because things do run out sometimes. You do need to buy, uh, places. That has changed my weekends. I no longer have to, like, think about, oh, what are we doing? Peter, what's your assistant tip? What is one thing that you're like, oh, this is novel, people should know about this? Put you on the spot there.
For me, I know it's, it's very boring, but just aligning my admin, so, like, any kind of emails, calendars, all of that flying around, it just- I feel like I was always anxious about what's, uh, is my stuff aligned? Am I forgetting something? And just whatever bothers you like that, just try and give that over to an assistant that can just, like, manage that stuff for you. I don't really, um, otherwise do a lot of stuff like that, but just the simple things to just get off my plate, that's what I really appreciate. So I just have more headspace not to think about stuff, uh, along those lines.
Yep. Uh, all right, Nisten, maybe you ran this for the whole week. What is one use case coming up for you with Muse that, that you think people should know about? Uh, so it is structured very much like a very light and very nice Hermes type of build, but actually a very lightweight Hermes. So it has the Sol file and, and memory and stuff. Uh, it does adopt personas quite okay. I just give it the persona of philosopher Epictetus, and then the prompt is just first learn the meaning of, uh, what you have to say, and then speak. And then I tell it a whole bunch of stuff to say in that persona. It actually
is super nice by default, like it's not, it's so nicely done, it is so nicely prompted. It's not overbearing, and it's not like ChatGPT where it just makes assumptions. It's like right on the spot Mm. as as an assistant. It really is a good product, folks. It really is the product for the masses. Like, we started the show with OpenAI, et cetera, being for the AI geeks, and OpenAI, like, quickly got to, whatever billion people that use it as a chat. This is a good product for many people. Just the notification thing that shows you what it's doing, and it's summarizing really what it's doing versus just, like, typing or showing you CRL command.
Just that is worth, like, so many points. And shout out to Alex Conway, I believe, is the designer for that. Uh, that pattern has existed, obviously, in Telegram, but Telegram only said typing. Here they actually tell you what the assistant does, and I think that's incredible. Uh, all right, folks, I think it's time for us to move on to Grok, and uh, Grok 4.7 was released. We waited for this. Uh, Grokbot is XAI's assistant, right? is running Grok 4.6 fast, and now, after promises, uh, there is Grok 4.7. And, uh, it's kind of a disappointment for me. I, you know, I wish I had like a wah, wah, wah, wah.
Do I have it? I think I has, yeah. I am sorry, Cursor folks. I believed that your coding mojo will be able to, to beat whatever, whatever stuff, and Grok 4.7 is not that great. I expected more. I wanted my Grok bot to use Grok 4.7. It'd be significantly better at writing, at writing code, at making decisions, and that's not, that did not arrive. However, what did arrive is Grokbot is now connected to your Tesla, which I find is an amazing pattern because your car is basically Kit
from the 70s show, and you can talk to it, and it self-drives, and it can do stuff. It's in-fucking-credible. So this is the jagged frontier of AI we're talking about. You're getting a new model, and it's not the best. However, you're getting a connector, and it can do stuff. So, for example, you can talk to your assistant without hands. You can hold it like this and say, hey, Grok, what are some of my emails that I haven't replied to? And it will say, or you say, hey, Grok, check in on that thing. I will go and do the connectors, and Grokbot in Tesla, I think, is significantly more important than any advancement in, in AI that could have,
could have gotten. But back to Grok 4.7. Uh, CursorBench is getting 46% back from, uh, 40 for Grok 4.6. It's the same price, 2 dollars per million tokens with half a million context window. The fast variant is 2x the speed and price, but TerminalBench is 37, and it's still behind GPT 5.6 Seoul on DeepSui. xAI trained this model for the longest time and still behind OpenAI's stuff. Uh, Peter, any thoughts on on Grok 4.7? Were you excited, disappointed?
Yeah, I think, yeah, it's kind of a weird one. I'm with you that when the Curse team stepped in to do the training, I kind of thought, oh yeah, that's it, they're just gonna fly off into the distance, and I think, uh, they didn't quite get there very quickly. I don't think we should write them off yet. I think, you know, these things are hard, and I'm sure, you know, you never quite know that they just not quite have enough GPUs at that point, and they're just ramping up the next one, and it will be all amazing. So I don't think we should make any big, um, conclusions from that. But I think, um, w- well, so clearly it's not an amazing model.
I think it's gonna be fine. Maybe it's like slightly better than the previous one, but actually 4- 4.6 was not that great either. So I, the, the question in my mind is, you know, they had this, uh, deal, uh, with Anthropic, right, for 6 months of compute. Mhm. What happens to that deal? It's, I think we are coming up to 5 months now into the 6 months deal. So in my mind, it goes 2 ways. One is they take the compute back and say, look, our model needs more compute, we need to train more. I think I can see that happening. Yeah. Or they, you know, or they say to Anthropic, you know what, pay us double.
Uh, you know, you need, you need the compute, you're building good models, pay us more. I, maybe this is rumors, maybe I'm completely wrong, but I don't think it's a scale problem. I think they have enough to train Grok as well. I think it's a talent and data problem. Uh, Elon Musk got incredible talent when the Cursor acquisition came over. Obviously, Grokbot is a Cursor product that they took over, slapped the name Grok on it, and now is on every surface. They're doing the same to their surface area that Meta is doing for Muse. They're putting Grokbot on top of every Cursor, uh, cloud instance. You can have a Grokbot link now in every
X chat. They're opening this up to, like, all of the, all of the X users. They lowered the price. It's not quite free yet, but I'm pretty sure they're gonna look at, uh, Zach and say, hey, this is now free. But for the model training, I don't think that the Cursor data did a dent yet. I don't, I don't see it. Like, I don't think that now I'll go to Grok 4.7 and start using this for my coding assistance. If it's better and I use it inside Grokbot, that's great. And I think here we're seeing like a battle lines are drawn where the interface slash harness slash assistant where through which you will talk to these models matters more than the actual model.
Like, literally, I am more excited this week about some of the features in Grokbot that launched, which is 1Password integration. Folks, if you don't use a password manager, please, please, please sign up for one. It's important even more in the era of AI assistants. Do not use the same password for all your services. However, an additional reason to use password manager, specifically 1Password, which I love with a passion over a decade, nothing happened, nobody hacked me, anything, is that those assistants now integrate with your own password. You'll be able to share your passwords in a secure way that the a- assistant can see. So I'm more excited, Peter, about this one integration with Grokbot than the whatever
next step of intelligence that they didn't quite hit, because it, for for most users, this is like the most important stuff. Um, Nisten, thoughts on Grok 4.7? Yeah, look, there are 9.4 million Teslas or more out still in circulation, a- as total, eh, and we're just, we're also- 9.4 million? Yeah, yeah, in total production over, well, it's been over like a decade. I think there's way more than 9.4. Uh. We- we will, uh, we'll ask producer fact check Nisten live on air. Yeah. Uh, go and find out It's around the best estimates of how many Teslas are currently driving on the road. But yeah, Nisten, Yeah.
go make your point. We'll see if the It is around that, that ballpark. I mean, as far as I only checked with the bots, but you have to think that even the initial, like, NVIDIA SoCs that were in those, they might not be too powerful. They might barely be holding the interface, but you don't need anything to run the agent on it because it's all just API calls. Just the internet connection, yeah. Yeah, you've just offloaded all of that server side. So it is pretty, yeah, it's pretty nuts that they reach those. Uh, on the other hand, as for having all the Cursor data, uh, yeah, you have a lot of enterprise data and code and stuff, but that might not actually be good code.
You... might need like very smart filtering pipelines and the stuff like that. This is why I think Codex and uh, Anthropic are, are ahead on that. So yes, they have the data, but they prob- they might not just be using it that well or be able to generate that good of a synthetic set. I think that, I think we're underestimating what I just said before. I, I, I need to, Yam, I need to bring your energy about like Navier-Stokes from last week. My car can fucking drive itself, and I can talk to it, and it can read my email. My car can drive itself, and while it drives itself, I can voice
interface-wise ask it to do my emails for me. Do you guys realize what the fuck amazing world we're living in right now? Wait, do you have Grokbot in your car? That's what I'm saying. Oh. Grokbot integration into the Tesla has dropped this week, and it's fucking insane. Oh, oh, shoot, okay. Right. know. Guys, this is a beautiful world we're living in. We get like amazing like discoveries and presents. stuff. By the way, our live AI assistant has fact checked Nisten. Uh, I think. Hey, I I was not wrong. I was pretty there. I said 9 9.5. I was right in the middle.
Yeah. So, Nisten, you want to read out the fact check? Tesla reported 9.2 million cumulative deliveries as Q1 of 2026 and 48,000 in Q2. So 9.7 to June. With Q3 progress, it's around 10 million delivered by now, with a few percent scrapped. So, uh, Nisten is right, a bit more than that. Yeah, so call it the 10 million Teslas. Nisten, uh, fact- fact checked live by our AI producer. Yeah, that's pretty cool. We should have that for our politicians now, please. All right. Can I just schedule the, yeah, yeah, go ahead. It's just listen, it just listens to us, and and that's it. Like, it's a producer. It's not it, brother. It's not it. Uh, if you guys want to pause for a second and talk about how
AI, um, how there's this new concept called AX, assistant experience or agent experience, and I believe, given where Muse is going, given Grubbot, et cetera, that like many, many, many surfaces and services will need access to assistants because APIs is not enough, MCP is not enough. I am planning to make ThursdAI the most assistant-pilled show in the world. We have an AI assistant that is a live producer. Uh, if you guys go to Thursday Live, we have an AI assistant. Our AI assistant is also like marking off and checking off the, the stuff that we're talking about and, and, and shows pace. So pacing is very important, right, in the, in the world of, uh, live podcasting.
so if you go and, uh, type ThursdAI.live, you'll see all our notes and me talking, obviously. This is going to be a little bit of a window in the window in the window situation, uh, but you can see the rundown, like we're running through all of the stuff. We need to start talking about, uh, uh, Fully Connected. We're like a little bit behind. All of our links are here as well. We have a producer that's listening to our show in live real time and and putting up, uh, things. Uh, not only that, the the another producer will edit it, and so you will get the show much faster after it. So we're working to being the most agent- pilled, uh, show in the world.
Most agent-pilled ever. Everything is live, listening to us, just doing Yeah. what we're saying. I totally agree with you that, transparently being in the background and, acting on whatever is going on, regardless of the AI. Oh yeah, absolutely, absolutely the next gen interface AI that just listens to you and does stuff. But you just talk about the actual, uh, just using Grok as it is, not, not within, inside any harness because I don't know if it, if it matters. Does it matter? Okay. I like, where, where would you use Grok 4.7? In Grok Code, the thing that they launched that, like, I don't know many people who
use, uh, Cursor has Fable and Astra. It's a Cursor model. It's default, default in Cursor. Definitely, I think it's, I mean, it, it, Yeah. it is targeted at least somewhat to developers. It's disappointing, uh, and this is also the reaction for everybody else. However, sitting in my car while it drives, talking to my cars when it does... Guys, my car has an MCP connection. Do you understand what I'm talking about? My car, when I talk to it, has an MCP ability to do anything in the world that has MCPs. Okay. Bro, you just, you just want it to drive. Like, you, you don't need Oh, just put a robot arm.
You have to put a robot arm in it, and then when you get snacks or coffee, you can just feed it to you. But s- but seriously, what can you do? What can you do with it in your car? Like, what, what the difference is? Open up your garage door, Yam. Garage door connector, smart home. You can say, hey, Tesla, open up my garage door. Can you tweet from it? 100% you can tweet. Oh yeah, okay. Can you What do you guys mean? MCP is an, is a connector, universal connector to everything. I saw one dude, showing an example, sitting in his car and saying, hey, get me my favorite thing from Starbucks and drive there. And the Grok bot sent a Okay, that's cool.
invite to the Starbucks, order that he usually gets. He didn't even say what it was, and then, sent a ping to his Tesla on the map, and all he we needed to do is hit drive. By the time he arrived to, Starbucks, his coffee was waiting for him, already paid. That's- Yeah, that's pretty crazy because Yeah, that's while it might not be that good at coding, it's more than smart enough to just handle other agents to It is really good for handling a lot of agents, specifically Grokbot when they talk to each other, and also, this whole show is also produced with Grokbots. All right, right, let's go to this week's buzz, please.
Everyone, welcome to This Week's Buzz, a corner of ThursdAI, where we talk about the company that makes it all possible, CoreWeave, Next week, ThursdAI is live, coming to you live from San Francisco, Moscone South, at the Fully Connected 2026 show with folks like Fei-Fei Dr. Fei-Fei Li from World Labs, with battle bots fighting out live on the show floor, and yes, Mr. 305 Pitbull is headlining. Fully Connected is next week, September 29th to October 1st, Moscone South in San Francisco, the day after OpenAI's Dev Day, by the way. So if you are flying out for Dev Day, come to ThursdAI.
You can get our tickets for free if you follow the show and our newsletter, and we're doing the live show on the floor. Uh, I'm very excited to hear the keynote from Dr. Fei-Fei Li from World Labs. Sarah Guo is hosting a panel, folks like Jerry Liu, Tom Rockschadow, Biases, Lucas Biewald, and CoreWeave CEO Mike Intrader is going to talk about how we received the Platinum tier on the benchmark from SemiAnalysis. We also have Ian Buck from NVIDIA and, uh, live BattleBots, and, Pitbull is gonna close the show. the ClusterMax 3.0 from SemiAnalysis is our little flex this week, yesterday, and we got platinum. It's a 3 reports in a row, and we're only the only provider
that's been platinum all 3 editions. Uh, Dylan Patel called CoreWeave the operational benchmark for the industry in GPUs, and hey, we'll take it. buzz. I really do hope that you join us in in person. All right, this has been this week's buzz. There's a lot for us to talk about. Let's move on to... Let's move on to the world of open source and JEV. I really want to talk about that correlation.
Open source AI, let's get it started. All right, bring it back. Everybody, uh, bring it back to the show. All right. So, one week after Jeff dropped from TypeSafe AI, the open source clones are here, and one of them is already beating JEV in some benchmarks. Last week I told you that JEV is a ChatGPT moment. JEV from TypeSafe AI, a new, uh, System 1 type model. We had Ali Labs, the the dev role for JEV, talk to us about the different, uh,
the primitives that that you can use. And I told you about that, open source will copy it fast. Well, it took less than a week. Over the weekend, more than tons of JEF compatible projects showed up, and they all speak the same format, decision-making format. And so you point it to the official SDK, and it just works. Uh, one is Classifier Dev, and, and uh, JEF Bench is something that I wanted to show you. So, uh, I will pull up Gev- GevBench. Meanwhile, uh, Nisten, I would love to hear from you, your thoughts about Gev, because last week you were somewhat, uh, skeptical/disappointed, but now I think,
uh, you've played with it. What, what do you think about this new paradigm? Uh, yeah, it actually brings the fun back in computing because you're thinking more in terms of like TypeScript or object-oriented thinking that, yes, I have all of these tasks and I just need to categorize them really fast. And they could be images on the screen, stuff you have to click, emails you have to write, and it's like now you just use an LLM in a much faster way for this, and it's just excellent at feeding it big, large JSONs and then coming up with other large JSONs, and it does the
accuracy better than an open source LLM often. Yeah. Even the ones that you can run at home. So this is, this is actually super useful, and uh, I'm really surprised as to how close the open source alternatives are. It's just that JAB itself does not really make mistakes, and the open source ones, eh, tend to have about like 1% or like a little bit less than 1% pers- uh, mistake rate, which is still a lot less than using an actual LLM, and even a, uh, like a somewhat good one. So, yeah, this is a, yeah, it's a new paradigm in how you make the apps. It will just be in every app, basically.
Server side, client side, yeah. This is like a new paradigm in in software engineering, where you can, because of the speed and performance and cost of System 1 models, JEF specifically, but also open source ones, you can put them inside your execution chain in software and rely on the non-probabilistic, uh, how should I say, execution of this, that it will work the next time your software runs. It is the microprocessor of the next era of software engineering. It is that big. Our whole timelines collectively were JEV build.
It broke through the bubble. My friends who are not into AI know about JEV. Every, every engineer that I know is thinking and thinking, oh, how can I use JEV more? Um, I've been completely Jeff pilled since before we got Ali on the show. And then, uh, Swix, a friend of the pod Swix, uh, hosted Diogo Almeida on his podcast on Latent Space, and that is a conversation that I recommend everyone listening to. I would love to hear from you about, uh, how big is this deal, and also the open source catching up. And Absolutely big. Yeah. Absolutely. Uh, look, the thing about LLMs and and and all all of that
is, yeah, they can do general things, and you can just talk to them, and agents also figure out, but the thing is that it everything becomes probabilistic, and there are ways to mitigate this and so on, but like there are building blocks that if you had a way to like compartmentalize and and, you know, branch out and choose, uh, how to use e- easily, you can do many, many of the things that you do with LLMs just without using the LLMs that are much more,
you know, out there in, in terms of, uh, all over the place, uh, in terms of how probabilistic they really are. And, uh, Jeff, just I, I didn't expect this. I knew, I know about all, uh, encoders and, and so on and how they work, but the fact I, I didn't take into consideration that if you train a really good one, probably a really big one, with all the modern inference, uh, tricks and so on. And data, a lot. Absolutely. Data. If you do it well enough, you get to a point it works really well, near perfect and nearly free. I mean, that's, I think this is the big deal, that it's basically
fast and basically free to use, so you can just, you know, just spam, spam decisions in like in insane rate, like you did in your demos. I mean, it was the the rapid fire, it was the the speed. Look, I'm all for training open source ones, absolutely go for it, no problem. I'm just saying that it's not easy to get to these results, and I think we should wait a little bit before we claim an open source replacement for JEV. Um, they are, they are impressive. I must say that, uh, they look quite impressive, but I would wait for more
benchmarks to, to, to measure them all out in all sorts of different ways. Yeah. Uh, just because, look, J- the, the people behind JEV told you we worked on this full time. Let's just talk about the people behind JEV. Yeah. Diogo, is the guy behind ChatGPT and RLHF, and his takes, you should listen to the show because Swix takes him deep on the process of how did you get there? Basically, the summary is he was thinking about who will use computers. And and he's like, okay, AIs will use computers way more than human,
but ChatGPT and LLMs were trained for humans to use. So who's gonna talk to other AIs? AIs. But they're now built to write text and respond as a chat that says, oh, I'm so sorry about that, my friend, but computers talk in fucking language machine code. So this is why he's like, okay, we need to build, uh, AI for computers to, to talk to. And how do computers talk? Formatted structured structure with JSON files. He has a, a story there where Sam Altman told him, go work on this, and he's like, no. And then he worked on it for 2 years after the pre-training was done. But also, the important, going back, the important note about open source Yam is
that it may not, uh, get quite to that level, but it gets very close. Folks, I wanna continue the show a little bit, so while this runs, I wanna bring up, uh, Florian. Let's say hi to Florian, everyone. Florian, say hello to the folks. Welcome. Hey, everyone. How are you? Thanks for having me. room. If Yeah. Florian, I want you to tell us what you built last week. I think this is a super exciting part. We just showed it off, but like I would love to hear from what you do. Okay, yeah, I mean, last week has been kind of insane, um, because I've kind of, um, bare- barely slept. Like, I've been obsessive, um, but I loved it. Like, it was like, um.
What were you obsessive about? Why did you need to sleep? Yeah, Jeff, ch- I still feel like Jeff is completely underrated still. Like, I mean, Still. you can't call it underrated because it got like 40 million views, everybody talked about it for a few days, but still, I believe it's kind of underrated because, what just so many people haven't really fully understood is, um, it unlocks a lot of use cases, like because of this unique combination of, three things, like it's incredible fast, so in fact they can, like, use it to play Doom and win a level, like you have to be really fast, multiple times per second. And then it's incredible cheap as so you wouldn't do it with any
normal large language model because it would just blow your bill and you couldn't pay for it. And then it's also incredible intelligence, so that's like it's, it's, um, the, the combination of these three. You can't just look at one of these dimensions and say, oh, I've got got a chat competitor. You you gotta do all three. And so, like, like right after it was released, I started to build a benchmark because I was totally aware that the open source community will start to work on this right away. It's obviously a small model, if it's fast like that, it's gotta be a small model, and small models are kind of affordable to train. So basically, at least if you, if you believe you can just like hack an existing
open source model and make it a Jeff. And yeah, it came as I expected, like the next day it was like 10 or 20 competitors, now it's more than a 70 on this benchmark. I've, I've like been benchmarking these things. I've been fending off, uh, approaches to gain the benchmark. People have tried to, to hack me, tried to exploit my machine. People have tried to steal the, the benchmark test set, all that stuff. But it's super exciting and it's super fun. I've been like, I- I've had like more than half a million views, and I'm really a small, small, unimportant Twitter account, you know? Very important Twitter.
Florian, let's not, let's not sell ourselves short. I think that benchmarking is an incredibly important part. Uh, if Wolfram was here, he would like very, very much agree with what you're doing. folks. I really wanted to shout out Florian for benchmarkheaven.com. Please go there, and JeffBench specifically, jumping on this new thing, uh, despite the creator of, uh, TypeSafe and Gev hates benchmarks as a concept. Yeah. I think it's a very incredible thing for folks to, realize, like, hey, not all open source are the same, not everybody trains on on the same data. I fully agree with you that open source is important from that perspective because
people may not want to send data into TypeSafe, despite TypeSafe being like the the next maybe, frontier lab. People may not want to send or may not be able to send. We're talking about medical stuff with Nisten. We're talking about different things on- prem. Like, people run open source specifically for we cannot, uh, legally send stuff, especially Florian. Uh, I don't know where you're from. I think Germany, right? Uh. Yeah, Germany. So Europe has like very strict laws and like Mistral, the whole reason for their existence is that they're the only one who can sell to like European governments. Uh, people want stuff on-prem or approved v- versus this.
And when you, when JEV K5 on the benchmark gets 62% of the stuff that Florian is is testing, uh, it looks the most... Where's Leia on this, by the way, Florian? Where's Leia? Because Nisten, that's the one. Oh, it's it's not performing well, very well on this benchmark, to be honest. Ooh. Um, it it used to be on place 2 for a moment, Yeah. but um, like I I uh improved um the the benchmark. You gotta really keep up, and now um new new competitors have shown up, so it's pretty pretty down. I mean, what what what Leia is is good in is it's incredible fast and and can run on on a CPU even, so it's a very valuable thing to have.
Yeah. But if you look at the, I've I've got these radio charts on the page where you can, like, look at multiple dimensions very, very far down of, of, like, tests, and there you can see, like, Leia is just not as smart, you know? It's, it's, uh, yeah, try to find it in this I'll find it, I'll find it. huge list. Leia 36 at the moment. So if you look at the, different types of, questions, maybe scroll a little bit further down even, there is another radio. Left. the left one is, is very telling. Mm-hmm. On the hard questions, like the red part is where it kind of, yeah, what it achieved. And Jeff is just so much smarter in all these different types of
questions. Yeah. So Yeah. mean, Leia is great if you know that your, your tasks are not so complicated, but if you, as I've said, you need to fulfill all of the three axes. You're gonna be intelligent as well to be actually a chat class, um, competitor. Yeah, so. Yeah, on, on, sorry. You have to, yeah, you have to benchmark it. So as long as I see that it gets over like 99.5, it's pretty good for that use, like comparing that JSON. or doing some computer use. But it's definitely not a generalist, but this is not, uh, anywhere there, like to play games and stuff. So, folks, Yeah. uh, open source stuff will catch up, already is catching up.
It's incredible. Uh, if you want to follow along with this, go to benchmarkheaven.com. Florian, thank you so much for joining us. So, uh, Florian, thank you so much for coming. You're right. You're right. Uh, Thank you. Have a good day. Yeah, thanks, man. Um, all right, folks, let's talk about voice. Let's talk about voice. Google says it now has the best text-to-speech model in the world, and you can clone your voice. Yes, Google, the big company, will let you clone your voice from 30 seconds of audio, which we should do now live on stage. Two new speech models, Gemini 3.8 Flash TTS and Flash Lite TTS. Google says they're number 1 and number 2 on Hume AI's quality index, with Flash at
the top of the Hume voice design benchmark as well. You can design voices or clone them. Flash is for voice design, character acting, and long narration. Flashlight is for high volume dubbing and voice agents. And Logelkin Patrick, our friend of the pod, says both are cheaper than the old 3.1 Flash TTS, and you can do 2 speakers in one request, and you can drop tags like laughs and hmm and listener and sounds. The voice replication is the part people will pick up because it is a huge company that allows you to clone your voice. 30 seconds of audio, only adults, they can, they don't allow to to to change kids' voices, uh, and the voice owner has to record the consent, and every clip gets a SynthID
watermark, which I think Google is very incredibly on leading, and cloned voices get a C2PA credential. So Google clearly thinking about the deepfake problem. Voice remixing is coming soon, and somebody on rbard in Reddit already built a live AI podcast where listeners talk to the hosts, and the top reply was NotebookLM had this a year ago. will we have an AI assistant, uh, here on the show with us with using Gemini's TTS, or will it be Muse? I don't know, but we definitely need to show you and hear this. Okay, so let's listen to the meditation guide. This is the design your voice thing, okay?
Yeah, the naming is meditation guide. A soft, airy female voice in her 40s, lower middle pitch, very slow, unhurried pace with gentle pauses for breath. Uh, you can hit improve my prompt and it'll generate it, but let's generate this voice. It's three voice options. voices. Let's hear for the first one. And so, in this moment, just breathe and find that space within. It's always here, and you don't have to try. This is great because of the breathiness of this. I could hear the,
especially in in headphones, you can hear this. Uh, but yeah, this is not our energy right now, so I wanna, I wanna, like, bring up the late night DJ. Ready to make something amazing? And the mad scientist. Got a project in mind? So, folks, um, I, I had my voice in there, but I can't find it, so we will definitely talk with the AI studio folks. I don't want to clone this live on air because, uh, obviously the consent thing, I need to record the consent in there, and I don't want my consent being recorded, although I think that that's going to be faked. Um, but the, the voice replication is a thing that we told you about, what, 2 years ago, that nothing is going to happen
in the, like, deepfakes, the world, et cetera, big, and then the big companies will follow suit. I remember there was a fear in the industry about, hey, oh no, voice cloning is coming. Yam, I don't know if you remember that episode or not, Nisten. And we told you nothing is gonna happen. If anything, it's for the best because people will know that this ability is out there so they don't trust phone calls from grandmas anymore, and nothing has happened. And I think there's a parallel there, and I think we'll finish on this. There's a parallel there between the fears that we are getting told by, oh, AI is gonna kill us all, and the reality
that we see. Open sourcing GPT-2 did nothing. Open sourcing, uh, Stable Diffusion didn't break the world. Open sourcing and allowing everybody to clone their voice and use that didn't sh- break the world. We're still here, better than ever. I think there's a very great parallel between this and all the AI is gonna kill us all. Nobody fucking knows. Nobody, as I say again, and I keep saying this, this is my line, nobody could have predicted a Austrian dude that has, you know, done his thing and doesn't need to work anymore, creating the world's next pattern that's called assistant that everybody's cloning. Meta Muse is a very, like, inspired thing from OpenClaw Hermes as well.
All of the assistants now do a thing there's some rando, like a great, great visionary Peter Steinberger, uh, did in his, like, apartment. Nobody in the doomer world could have predicted this is where it's gonna go because nobody can predict where we're going. And so this is why we're here to tell you about all the cool stuff in a positive way, and I think voice cloning is a very positive way to finish on this. A few housekeeping notes that of the stuff that we skipped this week that we didn't get quite to talk about lightning round. Uh, Qwen 3.8 OmniFlash released a, a Omni model with a 1 million context window. Uh, Qwen 3.8 Live Translate has released a real-time interpreter model that we can
pull up samples here for every language. Maybe we'll build this into uh, ThursdAI.live. Uh, those are the 3 main ones. In vision and video, Qwen released a 7 billion parameter model, Qwen Image 2.1, which is going around and doing some incredible stuff. And AI art and diffusion. Flux 3 action model. Black Forest Lab released the FLAX model open source, uh, an action model, 7 billion parameter world action model model, which is incredible. It takes camera frames and robot state and text instructions, predicts the next 32 actions and how the scene will change. And, uh, Peter, I promised you this at the beginning of the show, and I promised
people. There's this thing called Act 486. This is, I think, uh, this is the thing that we'll finish on because I think it's incredible. So, like, this is a pre-recorded clip. It's actually just a roofing torch with an air rifle cover. I don't get that. What don't you understand, Leah? How a flamethrower works? So I'll draw the propane canister feeding a pressure line. I want to pause here, uh, and say two things. Beside the fact that they got an OnlyFans model to to to to to market this for nerds, besides that point, what's happening on our screen right now is, eh,
somebody's watching the Lex Friedman Elon Musk, uh, podcast. They they pre-recorded this, they, they pre-baked this. Presses on Elon Musk and asks a question, and so then they clone his voice. I don't think he gave him permission, by the way, to clone this voice, and I don't think they used the Google TTS for that. They ask the character in that video questions. The character turns to them, starts answering the questions, and then also shows How a flamethrower works? So, like, the model throws, like, a real- time visualization about flamethrower stuff. Gas exits the nozzle. And he knows about the person that talks to them. He's like, what do you, don't you understand me?
Take out Starship and describe it for me. Here's Starship version 3. The stainless steel shell is built for heat, strength, and reuse, with 4 flaps steering it through atmosphere. Open it up and describe its components for me, please. So this is like a beautiful interaction moment, uh, because she pressed on, uh, Elon Musk's face, and then you can see him nodding while, the thing generates the thing, and he's fading in and out. I don't know if you guys noticed this, but like he's fading in and out because he's in loading state. I, I find that this interaction is incredible. I will play this again.
Steering it through atmosphere. Open it up and describe its components for me, please. Like, look at this. He's like acknowledging, and he's fading in and out like a person in loading state. This is crazy. He just looks so sad. It looks like he's been defeated, you know. But look at this. He is now, the the video model generates the starship model that he is talking about, and he is like rotating. Folks, this is the dream. This is the Game of Thrones, and you're editing your own ending dream that we want. This is the demo of that. This is the demo of talking to any type of, like, video and making it your own. Uh, this is, I haven't gotten access to this.
Uh, Act 486. Let's see what else they have in their demo. Uh, oh, she she goes, she goes crazy in here. Like, okay, let's finish on this because this is crazy. Okay. Mars makes those two futures feel tangible, and one future is we're out there among the stars and things we read about and see in science fiction movies. So for folks who are listening, uh, she now asked the video to put both Joe Rogan and Elon Musk in SpaceX suits on Mars. These are true. What would life on Mars look like? Both not. All right, let's see how daily life works. Life would happen mostly inside these habitats.
How does that work? Do people get upset at you if you do certain things? Uh, don't smoke that. You do not want to smoke that. Uh, yeah, going for the viral moment, there's a video, uh, there's a a famous thing where Elon Musk smokes some weed on Joe Rogan, and she's like, oh, no, don't smoke this, and the result is really hilarious. Okay, I won't. You're welcome. I... I think with no reactions, folks, this is the world we're in. You can now talk to videos and have them just do whatever you want. If you enjoyed the show, first of all, thank you so much for our live producer
assistant that did a lot of work, a lot of lifts, so all the notifications and stuff. Uh, second of all, huge thanks for, uh, guest, uh, impromptu guest Florian S. for jumping on with, uh, Benchmark Heaven. Uh, thank you, Nisten. Uh, thank you, Peter. Thank you, Yam. Thank you, Wolfram, who joined us, uh, and from and Mazir Penahi from AI Engineer at Paris. And folks, if you missed any part of the show, you can rewatch the whole show as it was live on ThursdAI.live, or you can subscribe to a newsletter or our YouTube, which, if you aren't subscribed to our YouTube, we're doing a lot of work. It will really help us if you hit that sub button.
It's completely free for you and costs you nothing, and you can, see notifications when we post new clips. And, next week show is going to be live from Mosconi's house, a little different time. So we're gonna go 11 a.m. Pacific, so a little like, 3-ish hours later than our usual, and go for around 2 hours. First hour is gonna be all about news, the second hour is gonna be all about, uh, CoreWeave stuff. It was a pleasure of mine. What a crazy week to have a show that talks about AI, folks. Just an incredible, crazy week. And we'll see you here next week.
Next week: LIVE from Fully Connected, 11am Pacific
Nobody's pacing. Don't miss a week.
One conversation a week with the people who actually run the models. Follow the podcast, or get the newsletter for free.