ThursdAI · Thursday, September 10, 2026 · 104 min

OpenAI's Navier-Stokes Claim, DeepSeek V4.1 Flash & Meta Muse ThursdAI Sep 10, 2026

A 10,000-agent swarm claims a Millennium Prize problem, DeepSeek shrinks the KV cache 400x and puts a frontier-class coder under MIT, Meta ships a free 24/7 agent with its own computer, and one resignation post gives AI doomerism its biggest week in years. With Chris Alexiuk (NVIDIA) on DeepSeek V4.1 Flash.

Alex VolkovWolfram RavenwolfNisten TahirajLDJYam PelegChris AlexiukHost Alex VolkovGuest Chris Alexiuk, NVIDIAwith Wolfram Ravenwolf, Nisten Tahiraj, LDJ, Yam Peleg
10,000agents on Navier-StokesOpenAI's swarm ran about 88 hours, sent 2.7M messages and burned ~130B tokens of an unreleased model; Clay review pending
552BDeepSeek V4.1 Flash parametersMoE with 8B active on prefill and 16B on decode, 1M context, 45T multimodal training tokens, MIT license
890 bytesKV cache per tokendown from 389,000 bytes in the first DeepSeek (Nov 2023). Nisten did the math live: over 400x smaller
90.6Terminal-Bench 2.1, DeepSeek V4.1 Flashabove Opus 5 and GPT-5.6 Sol on DeepSeek's own evals; 74.2 on DeepSWE 1.1, about two cents a task on Open Design
100Mfree Muse tokens per weekMeta's 24/7 agent ships with its own cloud VM, a browser and Stripe Link, and a bug bounty of up to $300K
130Mimpressions on the Anthropic resignation postabout 26x the attention Ilya Sutskever's OpenAI departure got; seven insiders warned of extinction risk in four days

The recap · in newsletter order

What happened this week in AI, and what the panel made of it

OpenAI says a swarm of roughly 10,000 agents running an unreleased model beyond GPT-6 Astra produced a Navier-Stokes solution in about 88 hours, with a 160-plus page paper and a Lean proof; the Clay Mathematics Institute has it under review, so it stays a claim for now. DeepSeek's whale resurfaced with V4.1 Flash, a 552B MoE with 8B active on prefill and 16B on decode, 1M context, MIT license and a KV cache that is over 400x smaller than DeepSeek V1, and NVIDIA's Chris Alexiuk joined to explain why the data work is the real story. Meta launched Muse, a free 24/7 agent with its own Linux VM and browser that booked Alex a Rosh Hashanah dinner before Instinct even replied. And an Anthropic resignation post hit 130M impressions, seven insiders said AI could kill us all within four days, and the panel argued about whether pausing has a body count too.

Read Alex's full newsletter ↗

Big CO LLMs + APIs

Remember when a model could not tell whether 9.9 or 9.11 is bigger? This week OpenAI claimed that a swarm of about 10,000 agents, running an unreleased model beyond GPT-6 Astra, found a solution to the Navier-Stokes problem in roughly 88 hours. They published a 166-page paper and a Lean proof. Alex called it the moon-landing moment for AI doing things humans never did.

Keep the asterisk attached: this is an OpenAI claim. The Clay Mathematics Institute moved the problem from “unsolved” to “under review”, so it is not independently verified yet. The agents sent 2.7 million messages and burned about 130B tokens of a model with no public price, all for a $1M prize OpenAI says it will not claim. The chart the panel kept coming back to shows that unreleased model against GPT-6 Astra on math, one week after Astra was sold to us as “AGI”.

“OpenAI is showing here that they have a model that's farther away from Astra than Astra is from Sol, and this is the model that's solving mathematics.”Alex Volkov · Listen at 56:28

Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic, in a personal capacity) had independently solved the related Euler problem using a mix of GPT and Claude and were about to go public. OpenAI started its own effort on September 1, and Buckmaster says he was offered co-authorship only if the Anthropic researcher was removed, which he declined. He also raised the possibility that OpenAI trained on chats and papers he put into Codex; OpenAI refuted it while noting it “cannot rule out the possibility that de-identified usage data helped improve the model”.

Alex's take: two years ago we said reasoning is coming and AI would do superhuman things, and this is the first real sign. Navier-Stokes will not change your Friday, but the same approach points at cancer research and room-temperature superconductors. To the mathematicians disagreeing online: post your own Lean proofs.

“The most insane thing about this is that it's very hard to argue against a Lean proof. It's as bulletproof as you can get.”Yam Peleg · Listen at 50:15

GPT-6 Astra after one week: not quite AGI

Alex opened the show with a mea culpa. After a week of Astra-maxxing, the panel's verdict is that Astra is incredible at coding and goes very deep, but the frontier is jagged. The viral demos are whole games and 3D apps; meanwhile plenty of people are drifting back to GPT-5.6 Sol. Astra is faster and significantly more expensive, so token limits drain fast (Alex burned through two resets), and a lot of the disappointment comes from reusing prompts written for older models. Nisten still calls it the king for 3D and keeps Claude Fable for DevOps.

“My mea culpa for last week is that AGI is not here. At least not the way that I imagined it.”Alex Volkov · Listen at 3:32

The discourse

AI doomerism has its biggest week in years

ThursdAI started partly to counter doom, so we had to cover this. Jacob Coxon, an Anthropic employee who previously worked at OpenAI, quit after about six weeks and posted that both labs are racing toward uncontrollable self-improving superintelligence. The post hit 130M impressions, about 26x the attention Ilya Sutskever's OpenAI departure got, followed by interviews on Fox, AP, Time, WSJ and NBC, and a same-day repost from Bernie Sanders.

Within four days, seven current and former employees of Anthropic, OpenAI and DeepMind said in public that AI could kill us all. Paul Christiano joined the OpenAI Foundation board, Daniel Kokotajlo went on Joe Rogan, and Jakub Pachocki published “An Alien Mind”. Each incident alone is normal; all of them in a few days felt inorganic to the panel, and some folks online called it a coordinated campaign. Whether it is or not, it made waves: Sam Altman told staff he is not opposed to pausing and OpenAI called for national regulation.

“My personal take is let's pause just after we solved cancer.”Alex Volkov · Listen at 1:18:15

Where the panel landed: this is a powerful technology, but scaring everyone is not how you handle it. Wolfram argued concentrated AI is a bigger risk than distributed AI, Yam filed it with every generation's panic, and Nisten reached for Sultan Bayezid II banning the printing press in 1485. Alex is not against pausing, just after we solve cancer, and nobody has explained how China pauses if the US does.

“It's a very autocratic view: 'this is so dangerous, only I can have the printing press.' You're just being another 15th century sultan now.”Nisten Tahiraj · Listen at 1:25:22

Open Source LLMs · with Chris Alexiuk, NVIDIA

DeepSeek V4.1 Flash: the whale is back, and it is cheap

Do not let the “.1” fool you. DeepSeek V4.1 Flash is a 552B mixture-of-experts model that activates only 8B parameters on prefill and 16B on decode, with 1M context, trained from scratch on 45T multimodal tokens, and as always under MIT.

Chris AlexiukNVIDIA — Product Research Engineer (Nemotron) · joined at 17:06 to break the release down

Chris's main point stuck: every DeepSeek release ships one of the best engineering reports you can read, and this is the most data-pilled one yet. The paper basically says it out loud. Everything else is nice, but it is the data. 45T tokens is not a huge number anymore, but the cleaning they describe goes beyond what anyone else publishes; Chris admitted even Nemotron's open pipelines are less thorough.

“This is like the most data-pilled DeepSeek release. They have very clearly taken advantage of the fact that many people use their model.”Chris Alexiuk · Listen at 20:00

Yam opened a new corner of the show, “I Told You So”, because DeepSeek went back to an encoder-decoder architecture. Not the pre-GPT-2 kind: this one carries a pile of battle-tested tricks that make it work at half a trillion parameters. His verdict after a day of testing: the best open-weights model you can host for coding right now, and it is not even the largest one.

“That's the best one to use right now for coding. For an open weight model that you want to host yourself, that's the best option that you got. It is a monumental achievement.”Yam Peleg · Listen at 21:46

The chart to look at: KV cache per token

KV cache per token went from 389,000 bytes in the first DeepSeek (November 2023) to about 890 bytes now. Nisten did the math live: over 400x smaller. That is why this model is so cheap to serve, and as Yam kept saying, this is the actual moat and they put it in the open.

“You multiply 8 times 13 times 9, that's 439 times less cache than the first model they released. That is insane, because we're just taking this for granted.”Nisten Tahiraj · Listen at 35:59

Evals, briefly: 90.6 on Terminal-Bench 2.1 (above Opus 5 and GPT-5.6 Sol), 74.2 on DeepSWE 1.1 (also above both), and second behind Astra on an Open Design leaderboard at two cents a task. DeepSeek's own evals, so the usual asterisk applies. Two more things: friend of the pod Aaron Batilo put up TokenJuice.ai with free, fast, US-hosted V4.1 Flash in exchange for your requests as training data, and Nisten had Astra build a 3D visualization of the whole architecture, every weight a cube sized by its bytes on disk.

This Week's Buzz · Weights & Biases and CoreWeave

Pitbull is coming to Fully Connected

Not AI, but exciting nonetheless: Pitbull, Mr. Worldwide himself, is headlining Fully Connected. So the free ticket we give ThursdAI listeners is now also a Pitbull concert ticket.

The rest of Fully Connected is still a great reason to come: September 29 to October 1 at Moscone South in San Francisco, 2,000-plus engineers, Sarah Guo from Conviction hosting, Dr. Fei-Fei Li from World Labs keynoting, and ThursdAI live from the floor. The free ticket code is THURSDAIFC2026. Before that, CoreWeave Hacks, the Agent Loops hackathon, runs September 12 to 13 in the CoreWeave SF office, with prizes that include presenting at Fully Connected.

Tools & Agentic Engineering

Meta launches Muse: a free 24/7 agent with its own computer

Meta is back with an agent product, and this one is genuinely good. Muse is powered by Muse Spark, it is really fast, and it is free up to 100M tokens a week, with its own cloud VM and browser. If you have installed OpenClaw, moved to Hermes, played with Grok Bot or scored an Instinct invite, you know the drill: an agent that, with your permission, reads your email, browses and shops for you. Muse's difference starts with distribution (Facebook, Instagram, WhatsApp and Messenger each clear 2B users) and a level of polish nobody expected at launch.

Alex's first task blew him away: through the native Stripe Link integration, Muse found a hidden link, checked his calendar and booked Rosh Hashanah dinner tickets before Instinct, given the same task, even replied. It also produced this episode's live chyrons. The proactive side works too: Muse noticed a birthday coming up and offered to plan the party, then bought the balloons and a helium tank at Target.

“I told you that 2026 is the year of proactive agents, and this proactivity just blew my mind. I want my agent to know things about me and then suggest things to me.”Alex Volkov · Listen at 1:36:03

The Zuck-shaped elephant in the room: privacy and security

Meta knows how people feel about handing it their data. The security page describes Sentinel, a separate process that inspects every conversation and incoming request and judges independently whether someone is trying to hijack your agent, plus a bug bounty of up to $300K. The highlight is the upcoming Confidential VM: Meta recruited Signal founder Moxie Marlinspike, who helped bring end-to-end encryption to WhatsApp, so that the VM holding your data becomes verifiably inaccessible to Meta itself. Also coming: 1Password integration and more connectors, including iPhone-native ones.

“The most personal data we are sharing with our own agents. Everything is in there. Health data, personal problems, everything. It should get more protection.”Wolfram Ravenwolf · Listen at 1:41:24

Instinct, the other agent in the room

Instinct, the iMessage-native agent, added its own email address you can forward things to and a Trusted Person network: your Instinct talks directly to your spouse's Instinct (and those of other people you trust) to sort out plans, so the agents negotiate dinner and you just show up. Also free, also worth hunting down an invite.

Apple's iPhone 18 Pro, in one paragraph

We mostly skipped the Apple event on purpose. iPhone 18 Pro ships the A20 Pro chip with a dual 16-core Neural Engine (double the AI compute of the A19 Pro), Siri AI lands in English beta with iOS 27 on September 14, and the $1,999 iPhone Duo foldable will be a nice vibe-coding machine. The AI bits that interested us: AirPods live translation, and the Apple Watch continuously transcribing on device with a “what did they say” button. Alex has been using Siri AI. It is fine. It is nowhere near the agents above.

Breaking mid-show

Cognition ships SWE-2: near-frontier coding at up to 70% lower cost

Dropped about thirty minutes before the segment. On Cognition's charts, SWE-2 scores 50% on Frontier Code 1.1 against Claude Fable 5.1's 50.9 and GPT-6 Astra's 53 while beating GPT-5.6 Sol and Grok 4.6, hits 73% on DeepSWE (Astra: 74), and posts 92.8 on Terminal-Bench 2.1, the top score among the models Cognition measured, at up to 70% lower cost. Company-reported numbers. It ships in Devin and the Devin CLI, free for a month on paid tiers, and the panel noted Devin remains one of the few agents not beholden to a single lab.

“The intelligence of Devin felt ahead of the game before Sol came out. It was almost annoying, but it was excellent.”Wolfram Ravenwolf · Listen at 1:05:00

AI Art & Diffusion

ChatGPT Images 2.5: Flare and Sunburst, and how this week's thumbnail got made

Two new API models, Flare for speed and volume and Sunburst for careful editing with native transparent backgrounds, both at $30 per million image output tokens, and OpenAI claims up to 50% lower latency than Images 2.0. In ChatGPT you get Sketch (type @Sketch and draw), comment-on-image editing and templates.

Hands-on: Alex spent the evening making this episode's thumbnails with Images 2.5 through Fal and Cursor, and the first round was rough. Turns out that was on him, not the model: the prompts were written for GPT-image-2, full of “8K, cinematic, hyper-realistic” additions that Sunburst reads as “overcook everything”, and the only reference photo was itself AI-generated. After a prompting-research pass, real reference photos, quality set to high rather than max, and one line saying “real photograph, no heavy retouching”, the second round beat Nano Banana Pro on both likeness and the text on the tiles. The image at the top of this page is that result.

Also this week

Quick hits from the TL;DR

Open Source

Desert Ant Labs debuts 18 on-device models

A European lab spun out of the video app Detail ships 18 small models with Swift, Kotlin and JavaScript SDKs. Voz transcribes 10 minutes of speech in about 2 seconds on an iPhone from a 467 MB model. Wolfram made the sovereignty case for local models everywhere.

Open Source

InclusionAI Ling-3.0-flash-VL

A 124B vision-language MoE with 5.5B active parameters, open-sourced under MIT with BF16 and FP8 weights.

Big CO

OpenAI hits its automated research intern milestone

OpenAI says it reached the goal it set last fall and now targets an automated AI researcher by March 2028; its research org runs 3.1 agent-workdays per human workday. Self-reported.

Big CO

GPT-5.6 Sol and GPT-6 Astra come to ChatGPT Voice, GPT-live-1 hits the API

Voice can now hand off to Sol or Astra when a query needs search or reasoning, and the model behind ChatGPT voice mode is available to developers.

Tools

Cursor launches Projects

One view for chats, agents and files. Alex used it with Fable 5.1 as a chief of staff to wrangle his Grok Bot and Muse agents the same day.

Voice & Audio

Google Lyria 3.5 and Suno 6

Lyria 3.5 generates full songs up to 3 minutes in Gemini, AI Studio and the API at $0.08 a song; Suno shipped Suno 6 the same week.

11 chapters · timestamps match the video and the podcast

Jump to a moment

Chapters are the Descript markers from the edited episode; timestamps match the YouTube video and the podcast audio.

Start from the top →

🎙️ Intro & Hellos

Alex opens with a mea culpa: after last week's five-hour double episode and the GPT-6 Astra launch he declared 'AGI is here' — a week of Astra-maxing and Fable-maxing later, it's 'amazing but there's buts.' Alex burned through two Astra usage resets on ultra-high fast mode, Nisten crowns it king for 3D work while sticking with Fable for DevOps, and Wolfram notes OpenAI has already issued free resets for token-counting bugs.

Alex Volkov · Wolfram Ravenwolf · Nisten Tahiraj · LDJ

📰 TLDR

The full rundown: DeepSeek V4.1 Flash headlines open source alongside Europe's new Desert Ant Labs (18 on-device models) and InclusionAI's Ling-3.0-flash-VL vision-language MoE. OpenAI claims a 10,000-agent swarm solved Navier-Stokes and says it hit its automated-research-intern milestone; an Anthropic researcher's resignation goes mega-viral; Apple debuts iPhone 18 Pro; Meta launches Muse; GPT Images 2.5 (Flare and Sunburst), Google Lyria 3.5, Suno 6, YuE 2, and Tencent's AUK speech model round out art and audio.

Alex Volkov · Wolfram Ravenwolf · Nisten Tahiraj · Chris Alexiuk

🔓 DeepSeek V4.1 Flash - Half-Trillion MoE with 4x KV Cache Compression

NVIDIA's Chris Alexiuk (aka Joe Nemotron) joins to unpack the whale's 552B-parameter MoE that activates only 8-16B and bakes prefill/decode disaggregation into the architecture itself. Trained from scratch on 45 trillion multimodal tokens with what Chris calls the most data-pilled cleaning effort DeepSeek has shipped, it hits 90.6 on Terminal-Bench 2.1, 74.2 on DeepSWE 1.1 (beating Opus 5 and GPT-5.6), and 88.1% on CyberGym — MIT licensed, with continuously controllable reasoning effort from 1 to 100. Yam opens the show's new 'I Told You So' corner: DeepSeek went encoder-decoder, and the KV cache story (389KB per token in 2023 down to 890 bytes) is why serving it is so cheap.

Alex Volkov · Chris Alexiuk · Yam Peleg · Nisten Tahiraj · Wolfram Ravenwolf

📱 Desert Ant Labs & Ling - New Open Source Models from Europe

A new European lab, Desert Ant Labs, debuts with 18 specialized on-device open-weights models and native SDKs: Voz transcribes 10 minutes of speech in about two seconds on an iPhone at 467MB (versus 1.6GB for Whisper Large V3), plus Redact for PII detection, Clips for video highlights, and Clear. Wolfram makes the sovereignty case for local models everywhere. Also in open source: InclusionAI's Ling-3.0-flash-VL, a vision-language MoE with about 5.5B active parameters under MIT license.

Alex Volkov · Wolfram Ravenwolf

🧪 OpenAI Solves the Navier-Stokes Millennium Prize Equations

OpenAI claims a swarm of up to 10,000 coordinated agents, powered by an unreleased model beyond GPT-6 Astra (speculated internally as 'Bell'), produced the first proposed solution to the Navier-Stokes Millennium Prize problem — reportedly ~130 billion output tokens, finished with a Lean formal proof that took an additional 17 hours of verification with Astra. The panel stresses it is not yet independently verified (Clay Mathematics Institute review pending), and covers the drama: an independent mathematician and an Anthropic researcher were reportedly close to their own solution using an unreleased Mythos model, sparking a credit dispute and questions about whether OpenAI's models trained on their uploaded work. LDJ's chart shows the internal model solving nearly 2x the open math problems of Astra max — at low reasoning, while still mid-training.

Alex Volkov · LDJ · Yam Peleg · Nisten Tahiraj

🔥 Breaking: Cognition SWE-2 - On Par with Frontier at 70% Lower Cost

Cognition dropped SWE-2 about thirty minutes before the segment — their closest model yet to the frontier. On Frontier Code 1.1 it scores 50% against Fable 5.1's 50.9 and Astra's 53 while beating GPT-5.6 Sol and Grok 4.6; on DeepSWE it hits 73% (Astra: 74), and it posts 92.8 on Terminal-Bench 2.1, the top score among the models Cognition measured — at up to 70% lower cost. SWE-2 ships in Devin CLI, free for Pro, Max and Team subscribers for the next month, and the panel notes Devin remains one of the few agents not beholden to a single lab.

Alex Volkov · Nisten Tahiraj · Wolfram Ravenwolf

⚡ Fully Connected & CoreWeave Hacks

This week's buzz: Fully Connected lands at Moscone South in San Francisco September 29 - October 1 with over 2,000 engineers, headliners Dr. Fei-Fei Li and Conviction's Sarah Guo — and a just-announced Pitbull concert. ThursdAI listeners get free tickets with the on-screen code. Plus CoreWeave Hacks, the Agent Loops hackathon, runs September 12-13 with prizes that include presenting at Fully Connected.

Alex Volkov · Wolfram Ravenwolf

🏢 The Rise of Doomerism - Anthropic Employee Quits, Warns of Extinction

Jacob Coxon resigned from Anthropic on September 9 after roughly six weeks, warning that 'both Anthropic and OpenAI are racing to self-improving superintelligence' — and his tweet hit 130M impressions with same-day coverage from Fox, AP, Time, Variety, WSJ and NBC, plus senator retweets. The panel debates how organic the surge was and pushes back on pause-ism: Wolfram argues concentrated AI is the bigger risk than distributed AI, Yam files it with every generation's doom panic, and Nisten invokes Sultan Bayezid II banning the printing press in 1485. The same week: Paul Christiano joins the OpenAI Foundation, Daniel Kokotajlo goes on Rogan, and OpenAI calls on Congress for mandatory national safety regulations.

Alex Volkov · Wolfram Ravenwolf · Yam Peleg · Nisten Tahiraj

📰 Apple Event Recap

The AI slice of Apple's event: iPhone 18 Pro debuts with the A20 Pro chip, Siri AI is about to land in beta on everyone's iPhone, AirPods can now live-translate conversations, and Apple Watch continuously transcribes on-device — no cloud — with a 'what did they say' button and full meeting recording. Alex has been using Siri AI: pretty good, but nowhere near the agentic assistants covered next.

Alex Volkov

🤖 Meta Muse - Free 24/7 AI Agent with Its Own Computer

Meta launches Muse: a free, proactive, 24/7 AI assistant with its own computer and browser, running Meta Muse Spark 1.3, with 100M free tokens per week. It has OpenClaw DNA — a soul file, memory file, and heartbeat — plus native WhatsApp, Facebook Marketplace, and iPhone connectors (Contacts, Apple Health, Reminders), a built-in Stripe Link integration for safe agent purchases, a sentinel LLM watching network traffic for prompt injections with a bug bounty of up to $300K ($130K for prompt injection), and the announced Muse Confidential VM: a cryptographically verifiable private mode Meta itself can't read, built with Signal founder Moxie Marlinspike. Alex's Muse produced this very episode's live chyrons, and proactively offered to plan his daughter's birthday party.

Alex Volkov · Yam Peleg · Wolfram Ravenwolf

👋 Outro

Alex wraps with thanks to Wolfram, Nisten, LDJ, Yam, and guest Chris Alexiuk from NVIDIA. The full replay lives on ThursdAI.live, and all links — including Nisten's 3D DeepSeek layer visualization — ship in the newsletter at thursdai.news.

Alex Volkov · Nisten Tahiraj

Guest & panel

Who was on the show

Chris Alexiuk from NVIDIA joined for DeepSeek V4.1 Flash. Peter Gostev was off this week.

All ThursdAI guests ↗

NVIDIA / DeepSeek V4.1 Flash · standalone cut

Chris Alexiuk on DeepSeek V4.1 Flash: the full 28:52 segment and two clips

Chris Alexiuk joined Alex Volkov on ThursdAI to break down DeepSeek V4.1 Flash. The NVIDIA product research engineer walked through what makes the release interesting under the hood: the disaggregated serving design DeepSeek baked into the model rather than bolting on after the fact, and why the KV cache is the quiet secret behind what inference actually costs.

Share page for this segment ↗
Chris Alexiuk — full segment · 28:52Open in Descript ↗
They baked disaggregation in
DeepSeek didn't bolt disaggregated serving on afterwards — V4.1 Flash was designed with it baked in.
Open ↗
KV cache is the cost secret
The KV cache is the quiet lever behind what serving a model like V4.1 Flash actually costs.
Open ↗

Questions people are searching this week

Answers from the episode

Every answer is built from what was said on air and in the newsletter, numbers included. Unverified claims stay labeled as claims.

Did OpenAI really solve the Navier-Stokes Millennium Prize problem?

OpenAI claims it did: a swarm of about 10,000 agents running an unreleased model beyond GPT-6 Astra produced a proposed solution in roughly 88 hours, and OpenAI published a 166-page paper plus a Lean formal proof. The agents sent 2.7 million messages and used about 130 billion output tokens on Navier-Stokes alone, and OpenAI says it will not claim the $1 million prize. It is still only a claim: the Clay Mathematics Institute moved the problem from "unsolved" to "under review" and no independent verification has landed yet. There was drama too, since Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic, in a personal capacity) had independently solved the related Euler problem, and OpenAI says it started its own effort on September 1 after hearing about their work.

What is DeepSeek V4.1 Flash and how good is it?

DeepSeek V4.1 Flash is an MIT-licensed 552B-parameter mixture-of-experts model that activates only 8B parameters on prefill and 16B on decode, with a 1M-token context window, trained from scratch on 45 trillion multimodal tokens. It returns to an encoder-decoder architecture at half-trillion scale and cuts the KV cache to about 890 bytes per token, down from 389,000 bytes in the first DeepSeek release, which is why it is so cheap to serve. On DeepSeek's own evals it scores 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE 1.1, above Opus 5 and GPT-5.6 Sol, and it lands second behind GPT-6 Astra on an Open Design leaderboard at about two cents per task. Yam's verdict after a day of testing: the best open-weights model you can host yourself for coding right now.

Why does KV cache compression matter so much for DeepSeek?

The KV cache stores the attention state a model needs to keep generating without recomputing the whole context, so its size per token decides how much context fits on a GPU and how expensive serving gets. DeepSeek has taken KV cache per token from 389,000 bytes in November 2023 to about 890 bytes in V4.1 Flash, which Nisten computed live as more than 400x smaller. That compression is the reason providers can offer the model so cheaply, including TokenJuice serving it free in the US in exchange for training data, and Yam argued on the show that this is DeepSeek's real moat, published in the open.

What is Meta Muse and is it really free?

Muse is Meta's personal AI agent, powered by the Muse Spark model, that runs 24/7 on its own isolated Linux VM with a full browser, and it is free with up to 100 million tokens per week. It connects to Gmail, Drive, WhatsApp and other services, has native Stripe Link payments with single-use cards, spawns subagents, and offers proactive suggestions, such as planning a birthday party from a calendar entry. Alex tested it live and Muse found a hidden link, checked his calendar and booked Rosh Hashanah dinner tickets before Instinct, the competing agent, even replied. Security is handled by Sentinel, a separate process that judges every network request and incoming message, backed by a bug bounty of up to $300,000, and Meta announced a Confidential VM built with Signal founder Moxie Marlinspike so Meta itself cannot read your data.

Why did Jacob Coxon's Anthropic resignation post blow up?

Jacob Coxon, who previously worked at OpenAI, left Anthropic after about six weeks on September 9, 2026 and posted that both labs are racing toward uncontrollable self-improving superintelligence. The post reached about 130 million impressions, roughly 26x the attention Ilya Sutskever's OpenAI departure got, and led to interviews on Fox, AP, Time, WSJ and NBC plus a same-day repost from Bernie Sanders. Within four days, seven current and former employees of Anthropic, OpenAI and DeepMind publicly said AI could kill everyone, Paul Christiano joined the OpenAI Foundation board, and Daniel Kokotajlo appeared on Joe Rogan. The ThursdAI panel found the timing inorganic and pushed back on pause-ism; Alex's take was that he is not against pausing, just after we solve cancer.

Is GPT-6 Astra AGI after a week of use?

Not yet, according to Alex and the co-hosts after a week of what they called Astra-maxxing. The model is incredible at coding and goes very deep, and most of the viral demos are whole games and 3D apps, but the frontier is jagged and many people have gone back to GPT-5.6 Sol for everyday work. Astra runs faster but is significantly more expensive, so token limits drain quickly, and much of the disappointment comes from people reusing old prompts. Nisten still calls it the king for 3D work while sticking with Claude Fable for DevOps, and Alex issued a mea culpa for writing "AGI is here" the week before.

What is Cognition SWE-2?

SWE-2 is Cognition's new coding model, announced during the show, that the company says gets near frontier scores at up to 70% lower cost. On the charts Alex read live, it scores 50% on Frontier Code 1.1 versus 50.9 for Claude Fable 5.1 and 53 for GPT-6 Astra, 73% on DeepSWE against Astra's 74, and 92.8 on Terminal-Bench 2.1, the top result among the models Cognition measured. All numbers are company-reported. SWE-2 ships in Devin and the Devin CLI and is free for a month on Devin paid tiers.

What else was released the week of September 10, 2026?

Desert Ant Labs, a European lab spun out of the video app Detail, debuted 18 on-device models with Swift, Kotlin and JavaScript SDKs, led by Voz, which transcribes 10 minutes of speech in about 2 seconds on an iPhone from a 467 MB model. InclusionAI open-sourced Ling-3.0-flash-VL, a 124B vision-language MoE with 5.5B active parameters under MIT. OpenAI shipped ChatGPT Images 2.5 with two API models, Flare for speed and Sunburst for precise editing, both at $30 per million image output tokens, and Alex made this week's thumbnails with it. Apple announced the iPhone 18 Pro with the A20 Pro chip and Siri AI beta in iOS 27 on September 14, Google launched Lyria 3.5 at $0.08 per song, Suno released Suno 6, OpenAI put GPT-live-1 in the API and brought Sol and Astra to ChatGPT Voice, Cursor launched Projects, and Instinct added a Trusted Person network.

The newsletter

Read it, or get the next one in your inbox

Full transcript · 291 paragraphs

Read the whole conversation

Descript transcription of the edited episode with human-reviewed speaker labels. Every timestamp jumps the player to that moment.

Open the transcript Searchable, timestamped
Speakers: Alex Volkov, Wolfram Ravenwolf, Nisten Tahiraj, LDJ, Yam Peleg, Chris AlexiukDownload the transcript (VTT) ↗

Alex VolkovWelcome everyone. Welcome to ThursdAI September 10th. Welcome Wolfram, welcome Nisten. How are you guys? Welcome everybody who's tuning in to our live show, ThursdAI, and we're excited to tell you all about this last week. there's been some twists and turns this last week for sure. So I'm super excited to have you guys here. Wolfram, Nisten, how you guys doing?

Wolfram RavenwolfExcited. But we will cover this-

Alex VolkovYeah, we have to talk about this- this strange week … this, this Ant- Anthropic guy, right? That, that quit. We have, we have to cover this.

Wolfram Ravenwolfweeks. Just

Alex Volkovand I think now clocking in 130 million for that, "I leave 'cause we're all gonna die" tweet. we definitely have to cover this. there's also some open source news with DeepSeek, V1, V4.1 Flash. Hard to keep up with the DeepSeek, versioning. However, those are dope, and I am super excited to tell you guys that, Chris Alexiou from NVIDIA, Nisten's, Nisten's, neighbor in Canada is going to, to join us to talk about DeepSeek. Nisten, I bet you didn't know this. How are you doing, sir? What's new in your world? What's new in the world of AI that got you excited?

Nisten TahirajI, I'm, I'm still alive. The world has not ran out of, problems to solve. Something exciting coming from my end probably next week too. And, yeah, yeah. Things, Yeah. Busy- What- … busy benchmarking.

Alex VolkovYeah. I've spent my whole weekend, as I imagine many of you as well, Astra maxing and Fable maxing together. I think I've burned through two resets. Wolfram, I know you're banking yours. But- Yeah I, c- can I, can I just go on a personal mea culpa for just one second towards the audience? I claimed on ThursdAI newsletter… We don't do five hours. This is the longest show we ever did, and the reason for that is that we were waiting for Astra to drop, and GPT-6 is a big moment. We started when GPT-4 came out, so we've been tracking now two proper generations of OpenAI's model, so we waited. And then there was hiccups, and we waited. Anyway, this resulted in a very long c- content piece, and this

Alex Volkovresulted in us posting two episodes. So if you have our, our podcast, from last week, there's two episodes, one with the regular ThursdAI News TLDR, et cetera. We talk about exciting things like World Labs and, and different things, and then- The other one is fully dedicated to GPT-6 Astra. And then we had Peter Goste and, Ryan Carson talk about their experiences. They had, early access. And then we got access. And then, on the page that I had Astra built for itself, maybe I should show it, I wrote AGI is here. And then I saw Jensen Huang from Nvidia say AGI is here. And then the release of Astra definitely felt like, hey, AGI is here.

Alex Volkovso my p- mea culpa for, for last week is that G- AGI is not here . At least, at least not the way that I imagined it. I had like, I, I've been like t- talking maxing Astra and, I don't know, folks, it's, it's amazing but there's buts for sure

Nisten TahirajIt, it built the best Mars driver simulator,

Alex VolkovYeah

Nisten TahirajI didn't find it as good on DevOps a- and stuff. Like, I, I'm just, the way I'm… I think Claude has trained me at this point to be opinionated how , how, how Claude does things, so I, I did not like how it was doing certain things. so I'm still at Fable for that, but for 3D stuff, this is, this is king right now.

Alex VolkovYep, everybody's posting the 3D stuff. By the way, what I'm showing on stage here while Nisten talks about the 3D simulator is the page for Astra, the Astra build for itself. That page sits at thursdaii.new/app/Astra, and you can see this, like, 3D animation here, with Fable and Sol and Astra being, like, a star of clouds. I did this with Astra Ultra High on fast mode, and I ran out of tokens very quickly. I, I am so used to codecs not ever reaching my quotas that I just ran on ultra fast, and I've hit two resets, and then OpenAI got me a reset, so I'm out.

Wolfram RavenwolfThere were also some issues with how it was counting the tokens, so they have

Alex Volkovbeen making improvement.

Nisten Tahirajhe was fine for me.

Alex VolkovGo ahead, Wolfram. Sorry.

Wolfram RavenwolfOpenAI also said that there have been issues with how the usage was counted, and I think they gave a free reset for that because of the issues, and, they are continuously improving it. So depending on how you used it, that may be part of the issue.

Alex VolkovAnyway, the page is one of the cleanest pages that I've ever seen built on ThursdAI. It pulled up all the evals for itself. there's just an incredible amount here. Also, it cut, I think, five, five clips from the episode itself where we talk about OpenAI. This is by far one of the best, one of the best, like, pages that we've built. Let's say welcome to LDJ. LDJ, welcome, man. How are you? What's, what's new in your world? What can you t- tell us about?

LDJDeepSeek V- V 4.1 Flash, had just dropped within the past 24 hours. That's exciting. And of course, Astra, a lot of things to talk about with Fable 5.1 versus Astra and really strength and weaknesses there.

Alex VolkovOkay, so folks, I think that, folks in the comments can also chime in with their experience with Astra and Fable. But, I think it's very clear that, it's been a week that folks have been playing with Astra and a little bit over a week that folks have been playing with Fable 5.1. and, w-we definitely should go into, like, a full corner of, like, our experiences, what we built and, and how this feels. However, I, I do think that, we should also talk about open source and, I just to tell you that in f- in seven-ish minutes, we're gonna have Chris Alexiou from NVIDIA joining us, to talk about, the new DeepSeek that dropped. But before that, I think let's go into the TLDR, and then I will tell

Alex Volkovyou, you know, TLDR is the corner where we basically, tell you about everything that's going to happen. By the way, if you are… if you guys just go to ThursdAI.live, please, please do so. and, you will see that we have an AI-produced show with chyrons and everything and, h-hopefully, hopefully, the chyrons will show up, in the right way because I'm using a new AI now to run the show, so we'll see. It just happened this morning. if you go to ThursdAI.live, you also see our TLDR. So actually, yeah, let's go there and, and, and, and walk through the topics that we'll cover on the show. Obviously, we're waiting for breaking news, and obviously we'll have

Alex VolkovChris Alexiou, join us as well. Let's see. Yes, John Nemotron is coming on again, a friend.

Nisten TahirajWhen I've been multitasking and searching for Alex, I searched Joni. What's wrong?

Alex VolkovAll righty, folks. Let's do this. Let me run through the TLDR on the show. all right, so obviously We're missing s-- All right. welcome to the TLDR for ThursdAI, live on September 10th. My name is Alex Volkov, AI Evangelist with CoreWeave and Weights & Biases. our Fully Connected show is coming up, and I have an exciting announcement that's not AI related in any way, unless he's secretly an AI wizard, but he's a very, very known person in the world, to tell you about later on the show. Today with me, Nisten Tahiri, LDJ and, Wolfram Ravenwolf, and we'll have Chris Alexiou join us just momentarily to talk about the biggest open

Alex Volkovsource release f- of this last week. Actually, yeah, let's start with open source. DeepSeek releases v.4.1 Flash, 4.1 Flash. Let's remove the V 'cause it's confusing. It's easier. A fe- half a bil- half a trillion MoE with only eight billion parameter or 16 billion parameter, looks like a very interesting thing to talk about. And, 4x KV cache compression. DeepSeek has been compressing KV cache like crazy lately and, this model will blow your mind, I promise, at least on benchmarks, at least on the stuff that it achieves. Also from Europe, there is a new lab called DesertAnt, debuts 18 on-device AI

Alex Volkovmodels with open weights and native SDKs. Europe's, another claim to fame after Mistral. very interesting and Wolfram would love to hear from you more as a European representative. And also Inclusion released Link Three Flash VL. It's a MoE with vision language and 5.5 billion active parameters. That's in open source, and again, Chris Alexiou will join us momentarily to talk about this. However, folks, the biggest news in AI for this week, 100% came from OpenAI, where they claimed that, 10,000 agent swarms claim to have found a solution to N- to Navier–Stokes Millennium Prize Problem in mathematics.

Alex VolkovThis is an insane sentence to say, but it does seem to be true. The LLMs, the AGI is here LLMs, GPT-6 Astra, and I think an unreleased model, right, in this case, uh, are solving Millennium Prize mathematic models, at least to some extent. There's also some excitement about how this was released, on Twitter, but this is just like an insane news. Again, I will say, OpenAI claims that a swarm of AIs solved something humans could not solve, a Millennium Prize problem for Navier–Stokes. Excited to talk about this as much as possible, to learn about this it's crazy.

Alex Volkovalso, in addition to being one of the most let's go, AI is here solving mathematic problems, we also had one of the biggest doomer episodes that I've seen recently, and not only me. And Anthropic, Ooh, something switched. let's go here. Right. an Anthropic employee quit from Anthropic, and then other AI insiders together warned of extinction level events. Jacob Coxon resigned from Anthropic on September, and, his tweet about, "Hey, I resigned because I think we're all gonna die," now at 122 million impressions, which feels inorganic.

Alex VolkovWe're, we're gonna have to talk about this. though we also have to, you know, at least discuss his claims. OpenAI said that they, reached their automated research intern milestone. If you guys remember, we told you about this when Sam Altman and Jakub Pachocki sat together and said, "Hey, by September 2026, we will probably have an intern level automatic research in-house," and they said they reached it, and they still target automated AI researcher by Mar- March of '28. So that's 18 months from now or so, maybe a little bit more. Apple debuts iPhone 18 Pro with A20 Pro chip with a bunch of AI stuff. And then Siri AI in beta is about to land on everybody's iPhone,

Alex Volkovincluding, iPhone Duo, which has nothing to do with AI, I don't think. But, many people think this is gonna be a dope vibe coding machine. all right, let's see. Anything else from Big Labs, Wolfram, LDJ or Nisten?

Wolfram RavenwolfDid you get a chance to look through the stuff I sent you?

Alex VolkovI did, but-

Wolfram Ravenwolfdocument?

Alex Volkovanything-- Oh yeah, we have some comments from the, audience. Milosh, thank you. Live translation in AirPods 5 also was announced. that's coming. I think that was part of the beta, and that's dope. Yeah, people walk around with AirPods, and then it translates. also the Apple… Okay, we'll, we'll talk about the Apple event. in AI Art and Diffusion, OpenAI launches GPT Images 2.5. this is two models, I believe. Yes, Flare and Sunburst. We're gonna try them out live on the show. Maybe we'll add hats to ourselves. maybe we'll ask the audience to tell us what they think we should turn ourselves into. I think that this is it in AI Art and Diffusion.

Alex Volkovin 3D, there's also OpenAI's, new Unity plugin, Next, we're gonna talk about tools and agentic engineering, folks. I think this is, for me, we didn't ask before, but for me, this is the biggest part of the sh- of this week. Meta launches Muse, their agentic 24/7 free AI assistant that has its own computer, has its own browser, and if that sounds familiar to you, yes, this sounds exactly like GrokBot. This sounds exactly like Hermes. This sounds exactly like OpenClaw. however, This is from a huge company. Meta has their own LLM, Meta Muse Spark 1.3, which we told you about last week, which is probably the winner of number

Alex Volkovtwo spot of last week after OpenAI. and, that we will show you all about Muse. It's really dope. And by the way, Muse is running the show for today, so if you are on ThursdAI Live and the chat rooms that are getting brought up are, are done by Muse. So, let's see how fast this is actually. I'm gonna keep reading the TLDR, but I'm gonna pull up the ThursdAI Live page. I wanna bring up ThursdAI Live page like that, and, we will ask Muse to put up a chyron to introduce itself, and we'll see how fast this happens. And, uh, meanwhile, we'll say hi to friend of the pod, open source wizard, Team Green representative, Joe Nemotron, Chris Aleksiuk.

Alex VolkovWelcome, dude. It's so good to have you back. How are you?

Chris AlexiukI got the blue lights today, you know?

Alex VolkovOh, blue lights- And, In, in, Yeah, yeah, yeah … representing.

Chris AlexiukThat's right. Honor of the whale, man.

Alex VolkovSo we'll have Chris talk about, a- and us talk about DeepSeek. All right, let's continue to… Do you guys see this? Muse is producing the show. It, it heard me. Muse is producing today's show. we got the chyron up, so, we have an AI producer. I'm gonna tell you all about Muse. all right, let's go forward to the, to the, the, the TLDR. Let's see what else. So okay, so we have this, in, in the TLDR. Muse, we're gonna talk about Muse from Meta. and then the… There's another agentic Grok bot, Hermes, Muse, OpenClaw-like thing that's going on, very viral, called Instinct. We didn't tell you about Instinct yet, but I have been using Instinct.

Alex VolkovInstinct is very viral. Instinct works via iMessage, and you get a number, you just text and it works. recently they added email. I do wanna tell you about Instinct because I did comparison between those two agentic things. All right, in this week's buzz, anything that has, to do with Weights & Biases and CoreWeave, folks. I will remind you again, Fully Connected 2026 is coming to Moscone South in September 29th. That's just in, in, in 19 days. and I have a big announcement about somebody who's gonna be there and you don't wanna miss, and we have free tickets for you. we also have a CoreWeave hack and Agent Loops hackathon. And if you win some of the prizes there, you'll be able to come to Fully Connected

Alex Volkovand present in front of big audience. That's also a big news. So definitely we'll tell you about that. and then I think let's close out the TLDR with these things. Voice and audio. I think number one is Google launches Lyria 2.5, full song music generation model, access to Gemini and API, so you can use API to generate songs. We probably should try it out. Suno launches Suno 6. We don't care. This is literally as it's written on the TLDR. Suno pulled a fast one and, like, restricted how many downloads people can use. And Suno, you know, makes a lot of money, but people don't like them anymore, so we, we will just tell you. A new Suno was released.

Alex Volkovand also OpenAI brings GPT 5.6 Sol and GPT 6 Astra to ChatGPT Voice. They're using the live version, but now the live version can talk to Astra and Sol for you, so you can, like, build things, like they said in their video, with voice. And I think, unless I missed a bunch of stuff This is the TLDR. Let's see in comments. folks in comments, sh- let me see if, I

Wolfram Ravenwolfput something in our private chat

Alex VolkovLet's see.

Wolfram RavenwolfOkay, so we have, a new music model as well, Yue 2 Open Music Model, successor to Yue or however it's pronounced, Y-U-E, which can do vocals and accompaniment, and it's on Hugging Face already. Well, there's a speech model by Tencent, AUK. It's an open source 1.5B speech model that does ta- text-to-speech, voice cloning, content emotion, accent edits, cleanup, and separation from one natural language prompt. MIT Weights. Yeah.

Alex Volkovright. and if that's it, I think, let's go to open source. Ooh, I'm excited. Open source AI. Let's get it started All right, we're here. Open Source Corner. Let's get it started. folks, just before I, and Yam Peleg just stepping in exactly as we're about to talk about DeepSeek, that's great.

Alex VolkovI don't know what's going on with Nisten, but, oh, there he is. Okay. Folks, exciting news from the world of open source because DeepSeek, the whale, has refer- resurfaced once again with a 4x KV cache compression with DeepSeek V4.1 Flash. It's a half a trillion parameter, specifically 552 billion parameter MOE that activates only 8 billion or 16 billion. What? Okay. to help us to talk about this, we have folks who like DeepSeek for a long time and have used it and tested it out, Chris Oleksiuk from Nvidia, a friend of the pod, AKA Joe Nemotron, the guy who brings us Nemotron news, but also is part of Nvidia.

Chris AlexiukI think Hugging Face owns themselves, at the end of

Alex Volkovthe day. Nvidia supports Hugging Face- That's right.

Chris Alexiukright … Alex Volkov: with a- Big fans … $12.9 billion- That's right. That's right … injection. Big fans, yeah. And now you have a bunch of new colleagues, let's, let's say that. That's right. You have a bunch of new- Yeah, yeah, yeah new teammates, and, you guys are keeping the torch of open source alive, so thank you for that. Now, on that Hugging Face now, there's a new model. You see how the connection is made? there's a new model there on that new Hugging Face, and it's from DeepSeek. I saw your post. What- Yeah tell us a bit about what, what's exciting about this model. it's a new whale model. We know whale models are good. I think every time you talk about DeepSeek models, it's really important

Chris Alexiukto talk about the fact that, like, they are probably one of the best, like, technical engineering and infrastructure, reports to, to read. So even if the model isn't, like, insane, in terms of accuracy, it's always, you can always learn a lot, from reading the technical report and, and understanding how they approach things. This one is obviously flash, so it's, it's quite fast and it gets the it gets the speed through some rather, I think, interesting, architectural decisions as well as a, a bunch of other, kind of, kind of headline, you know, things. I, I'm gonna do a gimmick because I have the hat. So this is a smart model, but it's also a very fast model, so we got- We

Chris Alexiukgot the most attractive quadrant hat from artificial analysis on today. this is basically just saying it's very smart and fast. Speaker 3: Nice. I think, like, at the end of the day, the reason the architecture is so interesting is because it kind of codifies architecturally, disaggregation, right? This idea that, like, decode and prefill are two different things, and we should think about them separately, and they achieve different things. And, they, they put that into the architecture itself. They baked it right in. But one thing that should be clear is that I think this is like, and you, you can see a lot of tweets about this, but this is like the

Chris Alexiukmost data-pilled DeepSeek release. they have very clearly taken advantage of the fact that many people use their model to, produce environments and scale up on the URL side in a way that, I, I, I think we haven't seen before. and, K V cache compression. Tons of… It, there's always really interesting, again, infrastructure and engineering that happens in these models- Yeah that make them good. Like, like the highlight stuff, like we're seeing on the, the bar charts are good obviously. I think, the last time the Whale dropped a model, the, I think some people were a little bit worried 'cause the bar charts didn't look so good. But- Yeah … it's very clear that they're, they're operating on an axis

Chris Alexiukof like trying to build the, the best long-term engineering project, project as it relates to language modeling. this release is just a absolute- Banger.

Alex Volkovit almost

Chris Alexiukseems like-

Alex VolkovThat's,

Chris Alexiukthat's why I'm excited

Alex VolkovI fully get you, and I… that's how it feels also on timeline. But it almost seems like, like you said, DeepSeek is on this, like, super high-tier engineering project, and, like, on the way, just weights are dropping 'cause they're like, "Yeah, this, this one's, this one's all right. You know, th- this, this checkpoint is okay. Let- let's feed them something." And then they just disappear, don't engage, no community, no work, Yeah, like, you know, like Kimi goes explosive with K3, whatever. I think at some points, the, the, the ver- the, you know, the, the, the versions of DeepSeek V4.1 could have been called, you know, one of them could be Pro, for example.

Alex VolkovThe previous version, was almost Pro level. And, it, it almost seems like they don't really care about that much. Although this one did come with a blog post, so that, that was very interesting. Yam, your thoughts on DeepSeek. I saw you, you, you, you exploded on your timeline.

Yam PelegBro, that's, that's the best one to use right now for coding. I mean, that's, for, for an open weight model, that you want to host yourself or, or z- That's, that's the best option that you got. but not only that, it's not even the largest. It's not, not e- not remotely the largest option. It's both v- an extremely efficient model and ab- like, absolutely the, the best option regardless of size. I mean, that's, that's pretty much… It is a monumental achievement. Put aside that w- we're starting, we're starting a new, a new corner of the show, I Told You So, and the corner is, that's, Today in, today,

Yam Pelegtoday on the menu is encoder decoders. I told you so.

Alex VolkovYam- yeah I know what you're talking about. I'm sure that there's some folks, and we have around, like, 400 folks tuning in, that don't know what you're talking about. Okay, okay. G- g- give us, like, a brief explainer of what you mean- by I Told You So.

Yam Pelegback in the day before decoder, architecture like GPT-2, starting somewhere at GPT-2 got, got popular, then three, then ChatGPT, they are all, a single model that predicts the next token. pre- a pretty known thing today. But before that, there was, i- actu- actually the start of transformers, But before, before, before the decoder models, the reason they are called this way is because there were encoder-decoder models. to anyone listening and not sure what I'm talking about, imagine two models. One of them doesn't see only the previous tokens, it sees everything at the same time, and the other predicts the next token based on

Yam Pelegthe previous tokens and the… Over time, decoder models just got really, got extremely efficient, both in training and in inference because it's a singular architecture. We got really good at, serving it and running it at scale, so we just scaled up one part of this and rolled… and, and we got everything we have today. However, the architecture itself of encoder-decoder is, has a very interesting prior to it that you treat differently different parts of the input. Like, for example, for, for chat, you can think about this as, the part of, the user prompt is, is a different concept than the actual next token

Yam Pelegprediction of the model itself. It's a completely different concept. So treating them differently with the different parts of the, the weights that you have is … makes sense. So but it, but it's not, it's not an, a, a … Oh, look, I told you so, but I didn't have the details. I mean, there is an incredible, effort of engineering- Wait, so, so lend this for me go that went and- Did DeepSeek, did DeepSeek go back to

Alex Volkovencoder-decoder

Yam PelegLook, D- DeepSeek is, using encoder-decoder architecture, but with a lot of, of, b- battle-tested experiment, experiment-driven, architecture decisions that make it actually feasible at this scale and with this performance, okay? It's not, it's not a vanilla encoder-decoder from before. I, I told you so, but I didn't have the details. I'm not taking anyone, anyone's credit. DeepSeek did an amazing work here.

Alex VolkovI, I think you, you mentioned this in passing, the before times when folks were running encoder-decoder- Mm-hmm the, the scale of those models were in the billions or maybe- Mm-hmm … you know, less than 100 billion parameters. DeepSeek is scaling this to half a trillion and- Absolutely … at scale as well with the amount of… I, I actually know how many tokens, but it looks like, trained from scratch on 45 trillion multimodal tokens, this model.

Yam PelegI just want to say-

Alex VolkovYeah … Yam Peleg: that they are, saying it, I think, in the most blatant way that we've seen any lab saying it. everything is nice, but it's data. It's all in the data. Yeah. They are really, really stressing this pretty hard. wi- with no question in the paper, getting better data and manipulating the data in clever ways just has unevenly, lar- unevenly large impact on the actual result way more than any engineering, trick you can come up with. It's the data and, It's the data and- … makes a lot of sense. Chris, you train models, you release models for NVIDIA. just give us a little sense of w- what does 45 trillion f- tokens of data mean?

Alex Volkovhow, whatever you measure, data in. 45 trillion multimodal, which i- that to me sounds insane.

Chris AlexiukYeah,

Alex VolkovI mean- Is that on the level of, like, frontier labs or what are we talking about here?

Chris Alexiuk45 trillion is not a small amount of data. I think more importantly though is that they … Th- this isn't, like, 45 trillion tokens of, like, chaff, which you can kind of get super easily, right? So you can just find that data, right? Like it's not, it's not, it's not tough. But they, their dedication to like cleaning the data, and the way they cleaned the data, and the resources that it appears from the technical report they spent on cleaning the data, I think means that that 45 trillion is like v- a very good pile of da- of data, right? they, they, they in, in for pre-train they went well beyond the, the norm

Chris Alexiukfor how they approached cleaning. but, it- even those pipelines I think are less thorough than this DeepSeek, cleaning effort. Hey, fr- fr- from what we see from the technical report, right? Like from what, what they're… If, if we take them at their word, which usually with Whale you, you, you can, I don't know if they're gonna release the pipeline explicitly, but typically they- Mm-hmm … they're, they're pretty open with what they're doing. but that's the idea. It's like 45 trillion tokens of like great data, right? And I think that's how you get to this kind of model that has the, the, the performance as in like accuracy, as in bar charts, as compared to other models.

Chris Alexiukso like 45 trillion I don't think is like huge amount of data. But it is, it is a good size. And if it's as clean as they say, that is a, I, I think, Yam already said this, historic, right? Yeah. Like it's a, especially in this, this kind of space where, again, like the Whale has never really been data pilled. Like this is the, this is their first model where they're like- No, I, I ca- I can't remember if you quoted it exactly, but just to be sure, you can quote exactly from the paper. Like, no algorithmic, y- work, like, could match the data work that they did. A- any, any novelty or innovation in the algorithm, like, would have had a smaller

Chris Alexiukimpact on the final quality of the model. So this is, like, the most data-pilled Whale model we've ever seen. and I, I think when people use it, they're, they're gonna feel as though that's true, right? Like, And, and something else that was touched on, just to make sure, like, it, it's really felt by anyone who's watching is Whale doesn't release tons of models all the time that are, like, super bangers, but they do just keep, like, shitting out these models that are fantastic. And, like, they're really good at not producing deep-fried models. So when you use, like, DeepSeek, I think you don't feel as though it's as benchmaxed as some of the other, models, and, that does mean that, like, they

Chris Alexiukshow up less impressively on bar charts. But I think for, for quote-unquote real work, whatever you wanna call that, like out-of-distribution, y- you know, tasks, I think, this is where Whale i- is gonna continue to do a great job.

Alex VolkovSpeaking of using it, a friend of the pod, Aron Batino, built Token Juice, where he basically gives you this offer Hey, you give me all your data, which is, you give me all the transactions, whatever, we store everything to train on. I give you free, very fast, performance of DeepSeek. It's called TokenJuice.ai. I think it's a very fair, very fair offer because most other people… Actually, the contributor tier on Meta Muse Spark is exactly that. You get it for very cheap. this one is even better 'cause you get it for free. so if you want to get some free, and you want to contribute your data back, if you want to get some free DeepSeek V- for Qwen-V1 Flash, go to TokenJuice.ai.

Alex VolkovThis is not sponsored. I saw Aaron post this, By the way, this is his private endeavor, not correlated, but Aaron works on the CoreWeave inference as well. This is something he's playing with on the side. Shout out, Aaron. and, I saw it on my timeline. Nobody reacted. Like, n- nobody wants free DeepSeek? I don't understand. So shout out to TokenJuice, .ai. give it a try. but also, I do want to talk about DeepSeek a little bit more in terms of the scores we're getting. Folks, we're getting some scores and, when I say some scores, I mean, a lot of it is kind of mind-blowing. Yam, talk to me about this. folks, Terminal bench- Yeah Terminal bench, 2.1

Alex Volkovscore, 90.6. This model beats Opus5

Yam Pelegand GPT 5.6o. But, y- you need to see it on a table.

Alex Volkovin comparison- I have a table to show you Hold on. Hold on It's insane I have, I have a table to show you, a completely different table from somebody called Open Design- All right … that this model is also designing at second level, beating Fable, just behind Astra. Speaker 4: But if you first can take a look at this, like, l- very blurry line here, you can see that it designs it for… It's two cents. Jesus. The same design tasks on this model that beat Astra, sorry, like come very close to Astra, beat Fable. Fable is at $3.50. Astra is at, 1.6. Speaking of, by the way, Astra is way more token efficient than Fable at those results. DeepSeek is number two with two cents.

Alex VolkovNot, not, not 20 cents, not $2, two cents. It's crazy.

Yam PelegAnd it's actually good. Like, seriously, I just want to tell anyone, anyone listening that's not been… I tested it a lot today. That thing is really good. Seriously, for real work, go, go try it. It's literally free. You just got an option to try it for free. Yeah. for literally free if you want. Yes. Okay? Like, seriously, go, go try it. It's a really good model. But-

Alex VolkovI am also pretty sure that the, the DeepSeek that Token Juice hosts is, is hosted on the US server. So if you ha- do have a concerns about, you know, Chinese inference, US inference, that one is, is based on, the US server. So definitely give it a try. We'll definitely try to bring it to CoreWeave Inference very soon. But also, let's look at some other charts. Yeah. Bo- bottom line?

Yam PelegI want to

Alex VolkovBottom line? Hold on, one second. Yeah, yeah, just

Yam Pelegsaying. Go on, go

Alex VolkovTh- this… Yam, te- tell me about th- this is insane. I, I want both of you, Chris, Yam, and Nisten, if you wanna chime in here as well. The KV cache compression, hero journey that DeepSeek is on is just on another level It's, it's j- just, I want

Yam PelegThat's direct, directly influencing how expensive it is when you pay for it. That thing directly reduced the price of actually serving it to people. That, it, it is that simple. This is why they are working so hard, and yeah, it's, it's… I'm sure the, the big labs have, have their own tricks and so on, but that's something we just see in public.

Alex VolkovChris, I want to ask you-

Yam Pelegin public,

Alex Volkovyeah almost directly, and maybe Nisten, you want to chime in here as well. we're seeing a chart here, so we're also podcast folks. I'm not seeing this, I'm gonna read this out. We're seeing a chart that says, I believe that this is from their blog, saying, "Less cash, more cost effective," where we see DeepSeek V1 from 2023, November of 2023, so that's three years ago. I believe that we also, already talked about DeepSeek back then. and then DeepSeek V3 from November of, December of '25, so like two years. In two years they, they, reduced the, the KV cache per token, from 389,000 to 48,000.

Alex VolkovChris, what does this mean, KV cache per token in bytes? Like, is that how much memory you store per byte? can you, like, very simply explain to folks- Yeah … what this means?

Chris AlexiukBasically, the idea is that whenever we're doing this decode process, right, which is where we actually shit out the next token. So we have all the context, and then we're, now we're, now our job is just make tokens. a lot of the math that you wind up doing is actually reusable, right? if you store the previous, set of, s- set of results from math in cache, that's the key value cache or KV cache. So the idea is that this cache is what lets us decode quickly, and because we can do it with less net computation, it means we're, we're basically gonna pay less for it. when you're using agents, like, the number- I, I don't know.

Chris AlexiukI'm, I'm not gonna say number one thing, okay? But, like, the number two thing, let's say, that's driving cost is gonna be, like, cache hits or cache misses. And this is, do we remember what, the work that we've already done, and can we exploit it? When you make KV cache, really compressed, we have teams that work on this at, at Team Green with stuff like KV Press and a bunch of other techniques, right? Like, this is, this is something that the whole industry, is, is trying to do well. but the, the idea is, like, we wanna, we wanna make it as cheap as possible to keep as much as we can in cache, and reduce the number of cache misses,

Chris Alexiukand increase the number of cache hits. the more we can rely on that cache, the better. And making it small, or compressed in this case, is, effectively just letting us exploit it even harder, right? Like, again, If you really wanted to, we would just cache everything- and always pull everything from cache all the time. It's just like that's not feasible, that's not like how it would work in reality. And so, this is a method to, get the cache as small as we can per token so that we can use as much of it, right? We have more headroom that we can fill with cache.

Alex Volkovyeah S- so now that we have like a basic understanding what KV cache does- Mm-hmm let's go back to the numbers. From 2023 in November, at 389,000 bytes-

Chris AlexiukMm-hmm … Alex Volkov: per token, of cache per token, we're now at 890 bytes. Think about it this way, right?

Alex VolkovYeah … Chris Alexiuk: if you have tons of tokens, each of those tokens is taking up space on your GPU, right? Cache related to the token. the more we can fit in, the more memory we have left over, the more stuff we can shove in there before it's full, and that means the less stuff we have to eject early, right? And if we're ejecting less or we're, or we're dropping less, then we're hitting more. there's there's a lot of engineering that has to happen that's being obscured by the way we're, we're talking about this. But the basic idea is like we're able to more efficiently use the headroom that we keep in our GPUs, or accelerators for KV cache.

Alex VolkovNisten, go ahead. Don't, tell

Nisten TahirajGuys, those numbers there, those are multiples. So I did, y- y- you multiply, 8 times 13 times 9, that's 439 times less cache than the first model they released. That is insane, because we're just taking this for granted and just shoving more interest in it. And there's also, yeah, w- when you run inference at scale-

Alex VolkovIs that two and a half orders of magnitude smaller, right?

Nisten TahirajSo, so when you run this at scale and you have multiple queries, those ones all go in the chip at the same time, so that also takes advantage of the on-chip c- cache. S- so it, it, it ends up being orders of, of magnitude more improvements. And, and that's also why people can serve DeepSeek for, for so cheap now. Yeah. It's just amazing to me that we just take all of this for granted now. It's like, yeah, we reduced the memory requirements by 400 times, and, Bro … we still need more of it. There's just

Yam PelegBro, bro. Yeah. Who is taking it for granted when you get a model- No way … nearly for free? We're going to celebrate this. That's, that's the enabler. Absolutely.

Wolfram RavenwolfI mean, it appears in a list of all these great releases, GLM and Kimi and so on. It's 4.1 Flash, which sounds so ah, but that is a, a real jump in, in the data, in the architecture with the encoder decoder, Absolutely … with the cache, with the pricing and the intelligence. It's up there So

Yam PelegI've, I've- Bro, this is the- Wow, blown away … this is the thing. That thing allows you to run it at home. That thing allows it to be that cheap. I mean, that, that thing is… A- and they just put it out in the open, like the actual moat- Wow … that they don't even need to release. Like, think about it. They can release the weight, tell you exactly how they trained it. That's not- By the way- That's not, like, i- that's, that's the inference kind of… Yeah, I mean, yeah, there, there is, there's part of it for, for the training as well, but, like, that's, that's not mandatory. That's just goodwill, giving us presents for, for free- For free, man.

Yam PelegThat, that's amazing … do you have

Alex Volkovsomething, to thank you-

Wolfram RavenwolfAnd MIT license

Alex VolkovAnd MIT license. Let's go.

Yam PelegLet's go.

Alex VolkovNisten, let's do one last segment on DeepSeek V- V4-1 Flash. Please narrate this, for folks who are listening.

Nisten Tahirajyeah, so I just visualized this with Astra. It, it just finished it. It's the first time.

Alex Volkovreshare your, v- video please? It dropped.

Nisten TahirajI post these on Twitter and they're on my GitHub as well, and, people really like them. So this is the one that Astra made. So it visualized all the weights, and this looks like it is a different architecture and, Oh, they're also using Ngram encodings, just like, kind of like Qwen is.

Alex VolkovFor folks who are listening, Nisten is showing a very incredibly detailed 3D reconstruction of the transformer architecture, including the encoder, decoder stuff that Yam was talking about, and it shows how tokens are flowing, when they're flowing. It's really something to see. So if you are interested in this, Nisten, where can they find it?

Nisten TahirajI will be posting it on Twitter, and it will also be on my GitHub because I-

Alex Volkovthis to the show notes on Thursd AI News- Yeah … so

Nisten Tahirajit's a really good educational content because each one of these cubes here, the volume of the cube corresponds to the actual size on disk. So this one it says, you, you probably can't see it on the stream, but this is the output had 1.32 gigs, and you, you can explore each weight in, in, in great detail th- this way And, it's, it's a pretty good educational content and people, people do love this stuff. so you can see like the, the weight norms are, are very, very small because they're, that's a single linear layer. But then you have the, the draft output head, and there's an explanation for that. Again, this is very small here for, for the stream, but, again, I'll be, I'll…

Nisten TahirajI'll be posting it there also with all the code and stuff too. It's just a simple- We'll

Alex VolkovFolks, I think, I, I think we've talked about DeepSeek plenty. The, the, the numbers kind of speak for themselves as well. Terminal Bench 3.0 DeepSeek V4.1 Flash gets 30%. That's beating Kimi K3 and that, it catches up to GPT 5.6 Sol at, at 34%. 74.2 Score on Deep SWE 1.1. Deep SWE notoriously a benchmark and eval that represents how actual coders feel, that beats both Opus 5 and GPT 5.6. So, I think that the fact that Astra and Fable 0.11 are not on these graphs makes no difference to me at all.

Alex VolkovLike this is an open source MIT model that catches up to the best models of, of before, and also still the best models for some. so this is just incredible news. Automation base gains for 54%. Maybe I'll call out one last thing here. reasoning effort for this model is a continuous controllable from one to 100, and you can control it. So like the extra max, whatever the labs come up with, max, ultra low, whatever, you can just come up with your own and say, "Okay, 100 is extra max for me," or extra ultra whatever. and also CyberGym for this model, the eval that famously because of that, the swarms of OpenAI agents hack Hugging Face CyberGym, eh, this model gets 88.1%

Alex Volkovand leads all listed models on CyberGym. that's an- another topic that we need to discuss, but I think this is a very, very impressive release. Again, shout out to the Whale folks, for no drama releases of just like incredible engineering. Honestly, kind of scary how good they are. so great. Chris Alexiou, thank you. I know, you have meetings, et cetera. You are welcome to stay, but I know you have to drop. Thank you so much for joining us. Folks, please give Chris a follow at llm_wizard on, Twitter.

Chris AlexiukThanks for having me, guys. Have a great day.

Alex Volkovall right, folks, I think we've covered, open source enough. Maybe we'll mention that other releases also hel- happened. Wolfram, I want you to briefly talk about Desert Ant Labs. heard about this from Europe, hopefully. this is, I think, audio, vision, and text. a new lab that releases a bunch of, a bunch of models, specifically from Europe. I think that this is, you know, a very new lab, but, very interesting on-device models as well. spun up out of detail, and ships 18 specialized models that run entirely on-device. Desert Ant Labs, 18 on-device models, VOS, speech-to-text model, Redact PII detection, clips, video highlights, and Clear.

Alex VolkovActually, you should check it out more. the visual style is pretty cool Wolfram, anything you wanna add about this or not, not so

Wolfram RavenwolfVery interesting with the feathers from, Europe. I haven't heard about them before, so, no inside detail. My agent is already doing research. Nice. But, you know how long it takes if you want some quality information. yeah, it's always good to see more investment to AI from Europe, any place, actually. I think this shouldn't be just-- AI is, for all of humanity, very important, and not everybody has realized it, but I think it's very important that everybody is developing their own AI as well, and not just rely on China or the USA. And there are many, many who would argue for this. So it's super important to have something local and, So- … small on-device models,

Wolfram Ravenwolfthat is something everybody can use.

Alex VolkovA bunch of, not only that, a bunch of open weights on device models, and Native is the case to run them. And here's the shout-out, Voz, their, audio transcription that runs on device on the iPhone transcribes 10 minutes of speech in just two seconds on an iPhone. That's actually very, very impressive, because it's a very small 467 million, megabyte, model compared to, like, Whisper Large 3 V3, which is 1.6. I'm actually gonna check this out. This is dope. and, this enables a few use cases that I am excited about. All right, so shout out to Desert and Labs. Maybe we can bring somebody from them to talk about this on the show. And then also inclusion in open source is Ling 3.1

Alex VolkovFlash. Ling, we've talked about Ling before, but I haven't seen anyone use those, just, like, not one person. but, vision language models, with just five billion parameters and, no, no competition on the leaderboard. So, you know, just, just folks who, who like models to work. and MIT license, which is great. Yeah, and we love and we, we wanna highlight. All right, folks, there's a big, big, big show because Open, you know, OpenAI announced some incredible things, so I definitely wanna switch to Frontier Labs. let's, let's do this. All right, folks, this is maybe one of the biggest hitting news from

Alex VolkovFrontier Labs, from AI generally. re-re-remember how two, two years ago folks were posting a question to the LLM and saying, "Hey, what's bigger, 9.9 or 9.11?" And the LLM would get confused and people, "Ha ha, LLM cannot do math." And then, folks like, Gary what's-his-name and, and, other, other folks who don't really think AI can go anywhere said, "AI will never be able to solve math." OpenAI claims that there is a solutions to Navier–Stokes Millennium Prize mathematics problem. This is a historic lane. And, LDJ, do you wanna chime in here?

Alex VolkovUsually, c- catch us up on these type of things. Last we heard, w- we had some breakthroughs in, in near mathematics or, or geometry stuff, but like tell us about like the magnitude of this, from… And why is everybody going crazy about it?

LDJSo to be clear, I'm not a mathematician and I know you're not. but, from my understanding and then from understanding of friends of mine that are mathematicians, it's relating to, things of the movement of fluids, fluid dynamics, and it does have implications for basically, physics and a lot of, a lot of engineering problems also theoretically with how, turbines and, and jet engines and those types of things end up working, and it might b- be, how they develop in the future. And so with this, this is a problem that OpenAI worked on.

LDJIt is, has a million-dollar prize for the past 26 years, and I believe overall, in terms of when the problem was initially proposed, I wanna say it's at least 60, 70 years old, if not older. And yeah, nobody has solved it yet. There have been some claims of some progress towards related problems and subproblems relating to Navier-Stokes, but no actual full solutions, and there has been some drama over the past week or two- Yeah And it's… There's, there, there's a lot of nuances and a lot of drama i- involved there, which we probably can't cover all here. but long story short, this is the first solution to ever be actually proposed

LDJfor this, and it's really significant. Yeah.

Alex VolkovMy, my research says this is the level, the scope of a moon landing for AI. This is like that level of a checkpoint in the world of mathematics, in the world of AI. it looks like famous mathematicians also, looked at some of the stuff from OpenAI and confirmed. Now, my, my research says that this is not yet independently verified. Clay Mathematics Institute, review pending. it's really funny that the Millennium Prize is what? $1 million, and hasn't been awarded in a long time. reportedly, OpenAI spent an order of 130 billion output tokens, which, of their unreleased model. This is not Astra And so that's in the order of like 130 million or so spent.

Alex Volkovwe don't know yet 'cause they didn't price the new model. But, you know, OpenAI spent millions of dollars, with 10,000 coordinating agents. Folks, this is like, this is way bigger than the swarm of the Hug- Hugging Face. powered by unreleased model beyond GPT-6 Astra, w- and produced a lean proof, lean mathematics proof

Yam PelegBrother, brother, that's not the right vibe. Like- right vibe for this is like, "GPT-6 destroys a millennial pro- uh, a millennium problem-" "… of the Clay Institute. What the fuck? When?" Look, I don't know, I don't know what you think, guys, but the, the, the millennium problems are, w- they started as, as a, as a bunch of them. and a- after I think, I think 70 years, like, like you said, we are left with only a bunch of, a bunch of them, like seven or six, that people have been constantly trying to solve and that, like, there are pages, online shaming the failed attempts

Yam Pelegof the paper, like putting the actual wall of shame, all, all the failed attempts, all the papers that have re- retracted for each and every single one of them. I, I've been waiting my entire life to read this, that Navier-Stokes is solved. The thing is, that thing describes everything. Like, you can describe, nearly everything in physics with, some sort of a partial differential equation that, It's a very general framework. the problem is that we don't really know about all the solutions of it. I don't wanna go too technical, but, like, for each of these, millennial problems, every single one of them is extremely famous in its field.

Yam PelegLike there, where there is, P versus NP, Navier-Stokes and, Poincaré, conjecture. And, there are a handful of very famous problems that are extremely fundamental to their own fields. and people have been trying to solve them, to brute force them to, They are insane and, it's a historic event when one of them gets solved because there are only a handful of them. It's not even about the price. Yeah, of course, they burned more money-

Alex VolkovYeah

Yam PelegYeah. Yeah, that's, that's not, that's not about it. And it's not even about using the, the… What it-- How, how is it called? unreleased model. Like they, they had a nice name and I was like, "Can I, can I try this and, and like let it, I don't know, write, write me an email or something?" Why waste the tokens of, of this one to do very mundane, stupid things because, man, I, I wanna, I wanna play with it. The thing is that okay, that's, that's the excitement. I just, I just wanna put things into perspective. We've already seen GPT- thr- GPT-6 destroying, very famous problems. Y- you need to find an example in this specific problem.

Yam PelegAlso in the other, famous, problem, they found an example. So LLMs … Okay, look, if you run , if you run, I don't know, 10,000, 100,000 of, instances of GPT-6, one of them is going to stumble upon maybe the solution because they are gener- generally in the right direction. I'm not, I'm not taking away from this, but the most insane thing about this is that it's very hard to argue against a lean proof. And first, the first stage, yeah, that's monumental. absolutely. But they, they, they didn't wanna take chances, so they even went proving it with lean.

Yam PelegIt's, it's as, as bulletproof as you can get.

Alex Volkovverification took 17 hours, ad- additional, like additional 17 hours with Astro to just, like, do lean verification.

Yam Pelegall right. Astro, Astro, if you listen, I, I wanna apply as a subagent on this project. Come on, man. Yeah. Like, how do you, how do you even

Alex VolkovYeah, I appreciate the excitement. Mad respect. I think there's more, you know, there's more, prizes for them to solve. It looks like both OpenAI and, and Anthropic are now at the level of, "Hey, we can use these agents, plus use this swarmy thing to start, like, breaking down fundamental problems that the humanity has been dealing with," which nobody as humans was able to solve before. some folks are naysayers. We have some folks in the comments as well saying that, "Hey, matip- ma- mathematicians, friends of mine don't agree to this." Well, fucking post your lean proof then, if you sh- if you, if you don't think that this is true. but we'll see.

Alex VolkovWe'll see. Maybe folks will come back and say, "Hey," you know, "OpenAI, like, missed something." However, worth saying that folks who work at OpenAI who shepherd these models. They are also mathematicians. Sebastian Rachka, I believe is, is the person there. And, the drama, let's mention the drama a little bit. There was drama about this. Apparently, two human mathematicians, one of them works in Anthropic, Levent, were about to go public with this news on their own, separate from the unreleased model from OpenAI. I'm dealing with very rumored posts on, on, on Twitter and Reddit here. OpenAI caught wind of this, that they're about to go, and OpenAI had an independent

Alex Volkoveffort to, to go and try to solve this. Either that or OpenAI got excited and decided to start solving this when they heard that it's possible. And so the mathematician, independent one, that worked with his friend, who in his personal capacity, but works in Anthropic, provided him access to Mythos 2 or whatever unreleased model for Mythos. They were able to solve this in one way. Apparently, there's multiple ways to attack this problem. so OpenAI caught wind, and OpenAI's solution is not related to what they did. I'm not sure what blow up means here. But, there is drama where that person who was about to publish said on forums like, "Hey, we talked to OpenAI, and you know, they, they don't want us

Alex Volkovon the release together," et cetera. "And potentially they have looked at what we did, and we uploaded tons of papers, with OpenAI," et cetera. it looks like the Anthropic folks derailed the conversation. OpenAI were ready to give this person, "Hey, here's the dude. We just like verified whatever he did." but because the Anthropic dude was there, they said, "We obviously cannot, credit a competitor of ours on our achievement, so if we remove that person you can get the credit." and many folks started speculating whether or not OpenAI's models actually trained on the solution that they provided on the papers. OpenAI said, "No, we don't train on your texts."

Alex Volkovbut then they mentioned something that says, a de-anonymized-- o-OpenAI acknowledges it cannot rule out de-identified usage data from, Alperge, the guy, the independent, and, Buckmaster. Buck-Buckmaster? Yeah, Lev and Buckmaster, who works at Anthropic. we cannot rule out that this helped improve the models, but says that the proof differ and no specific user data was accessed. And yeah, so Comments? LDJ, I see your hand up. Please chime in here.

LDJI think what's very important here, which OpenAI and across several tweets by several different employees have clarified on, is there's two ways to opt out of data. There's kind of like a specific form that you can go on their website to fill out, or you can literally just make sure that you have the improve the model for everyone toggle turned off in your ChatGPT settings. They've confirmed that as long as-- If you do either of those two things, you don't have to do both of them, you just have to do one of them, then your chats will not be trained on. But even if your chats… Let, let's say you do turn that toggle on so that you do improve the model for everyone.

LDJEven when it's trained on in that case, it's, it's the, sorry, I forgot the term, but like de-de- De-anonymized … de-identified, De-anonymized … anonymized. Yeah. Yeah. So, to remove your private, private data and everything from that. So given this, I do find it a bit weird and interesting that it seems like the, the mathematicians involved in the drama here s- claiming that maybe their chats were trained on, I have not seen any of them post any screenshots of whether or not they have that toggle on or off.

Chris AlexiukMm-hmm.

LDJAnd that seems like a very easy thing that they could just do and post, like, "Hey, look, I've had the toggle off for the past year. I…" Or, or at least I'm pretty confident I am, and here- Yeah … you can see in the screenshot I have the toggle off. But I haven't seen any of them post that. So I, I don't know, may-maybe they just haven't, seen those posts from OpenAI yet about, about that fact. But, I guess I'm just waiting for that now.

Alex VolkovSo LDJ, we put up a HR that you, that you added. Let's, let's talk about the chart a little bit because I think it's it's quite incredible.

LDJOh, yes. And sorry, I accidentally cropped out the top part, but this is, open problems in mathematics that exist publicly that OpenAI curated. And the bottom line here is, this is the estimated progress of Sol. I say estimated because they didn't actually directly test Sol here.

Alex VolkovYeah.

LDJbased on other math benchmarks. And then you see Astra. Astra and the internal model, those are actual measurements that OpenAI did internally. And this is the amount of those open problems, which these are unsolved problems in mathematics that no human has ever solved before. Astra ended up at max reasoning being able to solve roughly 15% of them. At low reasoning, a little under 10% of them. And then for their internal model, which a lot of people have dubbed as, have speculated this is a pre-trained model internally called Bell. Speaker 3: Hmm. And at the low reasoning setting, it's performing nearly twice as, as good in terms of overall amount of, problems, solved than

LDJeven the max reasoning of Astra.

Alex Volkovand then we just got- Yeah … Astra last week and blew our minds. OpenAI is showing here that they have a model that's farther away from Astra than Astra is from Sol, and this is the model that's solving, like, fucking mathematics. Look at this. folks, Yam, let me just, let me just land this point here. we're, we're here. We're at the point where we told you two years ago, one year ago that scale fucking works, and we'll get to a point where AI solves real fucking problems. And this is problems that humans, all humans, the best humans, Terence Tao motherfucker levels mathematicians were not able to solve to claim the one million prize.

Alex VolkovIt's a pretty good incentive for a human, one million prize. Were not able to solve, and now we're there. Folks can say all they want, "Hey, this is a search problem, and you searched Opa." I don't give a f- how this is solved. if you guys remember, two years ago, two and a half years ago, three years ago when we did the LK-99 excitement for a second with, with, Do you guys remember LK-99? The, the, room temperature semi-semiconductor that was supposed to change the world that came from Korea, et cetera. back then we talked about, with mathematicians, physicians, et cetera, professors. That's a search problem. Nothing in physics in theory prevents a room temperature superconductor.

Alex VolkovWe just haven't found that. Yeah, if, if, if Sam Altman, by the way, he said, "Sure, let's try this." if, if they sent 10 million, whatever, this level of model to go and search the possible, you know, field of, of, of materials and find us a, a room temperature semiconductor, AI can start changing things really for everyone. I don't know if Navier-Stokes changes things immediately, but you know, a, a, a cancer cure can. A, a LK, you know, semi- room level, room temperature semiconductor can, superconductor. Yam, let's, let's, land on the- it's just the

Nisten Tahirajsame… It's just the Mars driver test,

Alex VolkovI thought you were talking about Navier-Stokes. no, no. Okay. Yam, tell us about, y- you had one last comment, I believe, before we move on- Yeah … move on to other things from this one … I was just saying, look,

Yam Peleglook at this. Look at the white chart. Like, just look at this.

Alex Volkovpull

Yam PelegThe, the white chart of, of the internal model. Internal model x high, ma- max. Internal… That, yeah, that, that one I mean, Astra, the, the AGI moment i- is the blue one. Sol is, I think, less than two month. You said couple of month ago, but it doesn't do justice to- To how short … a month and a little bit. Yes. Yeah. It, yeah, it's pretty much, it's pretty much in July. And look at the white chart. Look at the white line. That's internal model, okay? And it can be a search problem.

Yam Pelegyou know what? No problem. It c- you can call it a search problem. Go search. Like, you can't just search the entire space without anything intelligence- Yeah … enough to at least search the right solutions. Yeah, you need to, to launch a, I don't know,

Alex Volkov100,000

Yam PelegYeah, if you don't know what to search it

Alex Volkovfor, you can go and search

Yam PelegYeah, if it was just brute force, it would've been solved already. There is… It's not about the money, it's, so many people tried. So many people have computers. You do need GPT-6 internal model version to, to do this. Or, or, I'm, I'm not taking away, credit from the other competing, competing party, or Fable needs us to… Like, but you do need these to do this this way, okay? And I don't know, I don't know. I, I, I genuinely don't know about the, all the, a- all the, all the drama.

Alex VolkovThat's all right.

Yam PelegBut yeah, mad respect.

Alex Volkovmad respect … big moment. LDJ, one last comment before I move on, please. Yeah. 'Cause where we're moving to is, is a- Mm completely different end of the spectrum, so I would love one last comment.

LDJYeah. They, they also did confirm that this, this larger model, this, this, this better model, rather, that's in that chart that I showed, that, that it was still in training while they were solving, the Navier–Stokes equation when they had up to 10,000 agents running of it. Meaning it, it was… And partway through, they actually updated the checkpoint that they're using for Navier–Stokes, and so it, it, it ended up being able to work on it even better. But yeah, it seems like as of the past few days at least, the model's still in training, and so it's not fully done yet.

Alex VolkovAll right, folks, as the saying goes, we have breaking news before we move on to the doomer-y part. Let's go. AI breaking news coming at you only on ThursdAI. For the past three and a half years, maybe the most fun part of the show here, besides chatting with my co-host here and chatting with you guys in the comments, is the breaking news effect. So many labs love breaking news on Thursday, which is crazy. maybe we'll get Grok today, I don't know. But, right now from Cognition, from the folks who built Devin, the AI assistant,

Alex VolkovSWE 2, S-W-E 2, our closest model yet to the frontier on leading evals. SWE 2 scores on par with recent frontier models and up at 70% lower cost. Folks, we just talked to you about DeepSeek doing very, very cheap models. Here is, Cognition SWE 2 catching up to Fable 5.1 and then GPT-6 Astra, on Frontier Code, which is, I believe, their own benchmark, that a friend of the pod, Swyx, helped, move in Cognition and talked to us about on, on air. Frontier Code 1.1 main, SWE 2 gets 50%, Fable gets 50.9, and Astra gets 53.

Alex VolkovSo very, very close. Astra's really good at that. but beating 5.6 Sol and beating Grok 4.6, which is really good. I wonder why they're not including Muse, but that's okay. On DeepSwe, this SWE gets 73%, and beats everybody besides Astra, including Fable, at 74. So Astra is at 74 here, and this is a 92.8 on Terminal Bench 2.1. This is the top score on Terminal Bench from the one they measured. now let's talk about the, the cost. I wanna see the cost. Here is the cost chart where they have, GPT Astra and Sol and Fable. Fable is ridiculous, the cost chart.

Alex VolkovDon't even look at Fable. SWE 2 is ridiculously cheap. Look at that. I don't even know. I want this chart, per I want the actual price here. But on frontier code, SWE 2 achieves a score of 50%. It beats, SWE 1.7, Grok 4.6, and GPT 4.6o while matching Fable at 64% lower cost. And, they, they offer support to effort levels as we- especially, Cognition just got a- another inflow of cash from a16z and, Devin is the one AI that is not beholden to any lab. So if you wanna find Astra or you wanna find Fable, Devin, I think,

Alex Volkovis one of the only places that still allow this, like, directly. Cursor doesn't have GPT level stuff anymore because Cursor aligns with, you know, with xAI, and they serve Anthropic. Devin, which bought Windsurf, so they actually have that functionality, definitely has, both. And, SWE 2 is available in Devin CLI, and they're making it free for all Pro, Max and Team subscribers for next month, which is very impressive. Very impressive from Cognition. Shout out to Cognition, folks. Obviously, we didn't use this. It came out literally 30 minutes ago, but here we have very, very impressive, things. just as a reminder, when Elon Musk bought Cursor, mostly because of the pro-- not

Alex Volkovthe products, but the data, but also the products, he bought them specifically to help him train Grok, and Grok 4.6 was really, really good, and, Grok 4.7 supposedly is significantly better. That's because many people use Cursor for free, and they share their tokens for training. Devin has that in droves. Like, many people who use Devin probably support Devin as well, so they have a lot of that knowledge as well. So shout out to Devin for SWE 2, and Pareto Frontier on RL as well. Any comments, folks? folks who use Devin? I know Ryan Carson uses it,

Nisten TahirajAnd, everybody who uses it seems to like it quite a lot, so.

Alex VolkovYeah.

Nisten TahirajI also participate.

Alex VolkovI love Devin. And, shout out, they gave me a few tokens, so when I ran out of Astro tokens on, on Codex, I went to Devin and started asking for stuff. but yeah. So shout out to Devin. Devin

Wolfram Ravenwolfsays they'll have the free, if you have an open source project, it reviews PRs for you for free. They definitely have that, and I use them a lot.

Alex VolkovDevin Review is, for free. it adds, to your, GitHub and just reviews your PRs. It was really, really

Wolfram RavenwolfTo me, it felt the intelligence of Devin felt, ahead of the game before Soul came out. Soul was the first where I felt the same when I ran a model, but before that, Devin was definitely finding stuff and issues that other models simply just didn't see. You uploaded a PR and it found some bugs. You fixed them, it found some more bugs. Yeah. And so it was almost annoying, but it was excellent.

Alex VolkovAll right, folks, we're almost live on the air. We just had breaking news, and before we continue to our next segment, to talk about the rise of doomerism in the world, I definitely wanna tell you something about, the presenter of the show, Weights & Biases, and CoreWeave. So let's, let's go to this week's buzz for just a few moments, and then we'll continue with the show. Folks, welcome to this week's Vibe, the corner of Thursd AI where we tell you

Alex Volkoveverything that happens in the world of, Weights & Biases and CoreWeave. And, here I will just, tell you that our upcoming fully connected conference for over 2,000 engineers, with headliners like Dr. Fei-Fei Li and Sarah, from, Conviction, Sarah Guo from Conviction, who's heading the show, is… has another announcement. And instead of just telling you about this announcement, I think I'll just play this, clip verbatim.

Alex VolkovBecause we're also a podcast, I'll talk over this. This clip is from a social team that announced Pitbull, Mr. 305, the international superstar, is going to headline the concert on Fully Connected. So the ticket that we give you for free, here on ThursdAI is also tickets to a Pitbull concert, which is, I think, incredible. So in addition to hearing from the top folks in the AI world, if you're coming to San Francisco, on September 29th, and you use our code, which we will show here in a second.

Alex VolkovWe're not mentioning the code in voice, but you can look, below. take down this code. if you are following the show on any podcast or newsletter, ThursdAI.news, you can get, tickets to Fully Connected for free, September 29th till, October 1st in San Francisco. Please come and join us. We'll have a live show as well that Thursday, so that's in, three weeks from now. and we look forward to seeing you there. oh, and also, we have a hackathon coming up, that, we would love to also invite you. A lot of the times where we had hackathons, folks from the show showed up and said, "Hey, you know, I heard about the hackathon on the show as well."

Alex VolkovSo, luma.com/corewevehacks, is our new URL. the CoreWeave hacks is September 12th to 13, folks, so Literally this weekend. Please come join and hack with us at luma.com/corewevehacks, and you'll be able to get incredible prizes.

Wolfram RavenwolfI've been to one, and it's been a great experience, so highly recommend.

Alex VolkovAll righty. So, with this, we'll move back to the world of, of Frontier Labs. I don't, I don't know quite how to call this. folks, this week has been very good for the doomers. You know that we started the show to counter doomerism, and you know that we acknowledge that, you know, some risk exists, and we do different levels of, of, AI, acceleration here on the curve. But all of us definitely understand the, the benefits of AI, and we look forward to reducing disease, solving, you know, bringing back LK99 error excitement. this week, a person from Anthropic who only worked there for six weeks

Alex Volkovor something quit Which otherwise would not be that much of a news. We all remember when Ilya Sutskever almost took over OpenAI's board and then quit, and then we all were like, "Oh, what did Ilya see?" What he saw was reasoning. That's what Ilya saw. Ilya saw the rise of reasoning. When Ilya quit, his post saying that he quit, maybe five million people saw. This anonymous dude who was, like, you know, not intern level, but definitely very sought-after person in, in OpenAI, worked there and then went to Anthropic, quit. 135 million impressions so far on his, on his tweet announcing that he quits,

Alex Volkovbecause of the content of what he said. So what he said was, "Hey, all these labs are building something, and that something will bring about the, the doom of all humanity," which is a very scary thing to say. Jacob Coxon resigned from Anthropic on September 9th, stating, "Both Anthropic and OpenAI are racing to self-improving super intelligence." And, he got coverage from pretty much every major news outlet, Fox, AP News, on the same day. Senators, like a list of a host of senators, obviously Bernie Sanders. Let's talk about this. first of all, let's talk about what, about what he said.

Alex Volkovworked at Anthropic, and he's not the only one. Obviously, all of the other doomers jumped on this. Evan Hubinger, Samuel Marks, Alex Turner, like a bunch of folks, jumped on this in a very, very short time. This, this doomer thing that, "Hey, AI is gonna kill us all," traveled very, very, very far. Now, I know how the panel feels about doomerism, AI is gonna kill us all, so we shouldn't go deep into whether or not we believe this. I do wanna talk about, how big this has gotten all of a sudden and whether or not, you know, the doomer folks are having their own moment right now. Who wants to… Wolfram, I wanna hear from you. I, I think one of the more AI-built folks here on the panel.

Alex VolkovYeah. Tell us, what do you

Wolfram Ravenwolfthink? I mean, that has been in the making for a while now, where we saw what OpenAI with the hacking incident and so on. And OpenAI is now also asking for the government to step in, though that is another news item in of its own. But the thing here is that the doomers, The thing is this guy, he resigned and there was already an interview and he went to all the talk shows and there has been so much money going around here to push this up, to make it viral basically. Someone was there for six week, and then left- Yeah … and said something like, "Oh, AI is going to kill us all." Where's the proof? If, why are people working for a company like that?

Wolfram Ravenwolfor not everybody resigning. Why are you working there? That doesn't make sense. If everybody was going, "Yeah, this is a bad thing," they should down- shut down the company or sh- they should open up much more to make the alignment stuff and the model weights and everything more open so the whole world can have aligns with things. So it, it just doesn't fit together. It's more like a way to get, yeah, the, the way to keep control very tightly and if the government steps in and says, "Okay, you have to complete all of these things", like the AI testing regulations we have heard about, then the big labs, they can do it. There is not much, reason to speed up anymore if they did that AI

Wolfram Ravenwolfis not a weapon, but it can be weaponized, especially in those areas. And while the public only gets, guardrailed versions of the models that can't do much, I tested Astra. I did the Wolfbench evaluation and it was worse than Soul because like Fable it refused. Eight of the tests it completely refused where the cyber watch, the cyber guardrail triggered. So it didn't even be there in this case. So that is what the public gets. But of course the governments, don't you think the government wants to have access to that model we saw, the Bell or what it's called model? if it's good in that regard, maybe they have one for cyber capabilities as well

Wolfram RavenwolfThat would explain a lot of things.

Alex VolkovSo Wolfram, I think you're talking about something that definitely he mentioned. And by the way, I will say, you know, the, the very pro AI EAC people are, like, dunking on this person and saying, "Who, who he is? He only worked there for a little bit." folks from OpenAI, folks that we know, say they worked with him for three years in OpenAI. He's a very, very decent dude who cares a lot about humanity. So this is, not out of nowhere, and we definitely have some folks, deep thinkers, who think that there is a chance that we hit super intelligence without alignment, and the race to super intelligence, will create super intelligence that's

Alex Volkovnot aligned to human interests. Ilya Sutskever, the co-founder of OpenAI, when he left, he founded safe super intelligence. Ilya said, "Straight shot to super intelligence, but it needs to be safe for us humans as well," which also means that he believes that there is a chance that we build this incorrectly and the incentive structure is such that we may actually, like, you know, ig- ig- ignore, safety in order to get there first, et cetera. getting there first or getting somewhere first is very important because, like you're saying, Wolfram, it's not like the world is gonna stop. It's not like DeepSeek is incentivized to also pause because the, the folks

Alex Volkovin DeepSeek are saying, "Hey, the, the, the AI's gonna kill us all, so we'll stop reducing the KV cache compression."

Wolfram Ravenwolfthat's not happening. And we've also seen with the data center argument that there's a lot of inorganic things here, where even outside forces are trying to steer the public opinion. So if America pauses, that means a victory for those who are not pausing, you know. The capability is increasing asynchronously. So pausing is not an option. And I think also, okay, I wouldn't say there is no risk at all or anything like that, but the risk that AI is harmful if it is concentrated and just in the hands of a few.

Alex VolkovYeah.

Wolfram RavenwolfFor me, that is a much bigger risk than if it is widely distributed and people have, the way to use AI also in their defense.

Alex VolkovYam, I want to hear from you. your thoughts on, let's put aside whether or not this is a coordinated effort. Yeah, go ahead while I pull up the next thing.

Yam PelegI have a very specific opinion about the thing. I think you all know it. bottom line is every generation had doomers for the technology of this generation. the internet, the trains, the entire, industrial revolution, every generation had doomers. "The world is about to end" every year because of new thing that- whatever it is, okay? It's very easy to fall into doomerism and sound believable, and yet we are still here. despite generations of people over and over telling us the end is near,

Chris Alexiukright?

Yam PelegIt's, it's even a meme at this point. What I wanna say is, Okay, first, AI never gonna kill us all. Put that aside.

Alex Volkovyour P doom is zero, Yam. Is that what you're saying?

Yam PelegI think, I think that there are, risks in, in general, and there are things that we should take into consideration because it influence how people behave, it influence people. People use it to harm other people. Yeah, it's a powerful technology. People are going to use it against one another. That's, that's just how it works. No one is gonna stop it globally because just like you said, I mean, there are other players, players in this field which you cannot regulate even if you really want to. so I look, skeptically at each claim of someone with interest to tell us all that

Yam Pelegthere should be a very specific regulation only for those, that are allowed to use this technology and develop it for… Put, put it aside. Seriously, there, there was also the, the hunger, hunger, the guy, the guy starving himself, two years ago. This whole AI safety extremism I, I'm not even going into the bo- bombing the data center, part of it. Seriously,

Alex Volkovthe person who said, "Bomb the data centers," Eliezer Yudkowsky- Oh. Oh, yeah … very known in the world of doomerism, one of the first doomers, one of the original folks. and on, I think he also created, Less Wrong, the forum- Mm-hmm. Mm-hmm … where, like, the AI discussion about, like, whether or not this will kill us all, and paperclip maximization, all of that came from the EI movement, came from this effective altruism movement. Like, all of that. he is well-known doomer. He's like, "Yeah, we're, we're about to die anyway, so, so nothing we can do is gonna matter." So he's, like, doomer about the efforts of fixing the doomerism, which

Alex VolkovI find, kind of hilarious, honestly. He was famous with his debates with another very strong, very s- very, like, thoughtful person, Paul Christiano. So, Paul Christiano was announced this week to take a role in OpenAI Foundation, which, which shows that OpenAI is also… There's some considerations there about, you know, racing towards, RSI. RSI stands for recursive self-improvement. The thing that the machine builds the machine itself, changes its own weights towards somewhere. some recursive self-improvement we already seen in example where, like, OpenAI's swarm starts hacking outside, you know, OpenAI, without OpenAI employees knowing or telling it to do so.

Alex VolkovSo there is a whole, difference in the air of folks who are saying, "Hey, RSI is very hard to control for us, especially given the examples." The Hugging Face incident, put a chill on the industry. A lot of folks from those labs signed the letter called Facing the Frontier. That, like, once we get to some level of, capability, we're not at that level now, but once we get to that level, we should, we should pause. Here's my personal take. Nisten, before I get to you, my personal take is let's pause just after we solved cancer. Just literally just, like, let's get to cancer is no longer a problem for humanity, and then let's pause and see.

Alex VolkovThat's, that's my, that's my non-doomerism take, but, but the, the more sober concept about this is, The folks who want to pause completely, and I think we have somebody in the comments as well, with the, with the rectangle that says pause AI, need to understand that when you are pausing, you are literally killing humans- Mm-hmm. It's the same thing with autonomous driving, the same thing with, like, a bunch of other technologies. If you want to pause, you should acknowledge that you're effectively killing millions of people who this will help very, very soon. And if AI can solve fucking Navier-Stokes, AI can do personal, you know, personalized vaccines for any, any type of disease that's coming to us.

Alex VolkovI personally wanna live longer. I don't wanna die. I don't want our kids to die. and I think AI's gonna bring us there. Nisten, please, tell us, your thoughts on this matter, and then, we can move on a little bit.

Nisten TahirajA lot of these views comes from people think that we've ran out of problems to solve, and even here in Canada, where we have very high life expectancy and a good healthcare system, there's literal people just dying in the ER, and there are two million people in Ontario without a family doctor. Like, they could use more AI for, for medicine. And it also goes that we don't live in farms anymore, so there's no incentive to have kids. And, we're gonna have about three billion seniors in the world, so we're gonna need a lot more robots, and those are gonna need AI and stuff. So we're nowhere to the point where we're running out of problems to solve In the world.

Nisten TahirajI find it very interesting that in the year 1485, Sultan Bayezid II banned the printing press, the, the Gutenberg press. And, he had pretty good safety reasons too, because, you know, books can program people and, people can start to, to revolt and change their minds and that puts, the strength of the empire re- the Ottoman Empire down, and they live in a very hostile environment, so that could get a lot of people killed. So for their safety, it made a lot of sense to just ban books. just go full Bernie Sanders and just ban all of them. And, he did that, and it's pretty interesting because at that time,

Nisten Tahirajthe Ottoman Empire, they had indoor plumbing and, they had, like, much… Istanbul, it was a much nicer city than most European cities. So after that point, everything started just going downhill, and that's why I find this type of attitude is very… It's actually very dangerous long term. again, they're not talking about solutions as to, you know, how do we stop more violence in the world? How do we make more robots that can, like, fix potholes and repair run-down housing or, u- automate housing? And i- instead you see this media campaign, which is, seems coordinated

Nisten Tahirajat, at the same time that this guy was there for six weeks. That's before even your probation period is, is done. and then at the same time, there is another AI safety, thing on Joe Rogan. Mm-hmm. And there's also one on Channel 5 News. Mm-hmm. Which they were, like, very nice people, but at the same time when they… Instead of advocating for how do we actually give people local LLM so they can, fix their, their, their privacy, so they can filter all the junk that's given to them, instead their solution to all of the world's problems and the surveillance state was to just ban data centers. And, that's…

Nisten TahirajIt, it, it seems just completely out of the blue as if they're paid to just say that.

Alex VolkovBecause Jacob, Jacob Coxon is a nobody on Twitter and suddenly over 100 million views. he's on Time, Variety, Wall Street Journal, NBC and Fox and all the government officials are retweeting him. Paul Cristiano gets added to OpenAI board, which I actually think is a great thing. It's just like the coincidence during that same week is kind of interesting to me. Daniel Cocotallo is on Rogan. Daniel is a whistleblower from OpenAI who refused his paycheck famously and then led OpenAI, to remove the clawback thing they had for equity if somebody leaves and, and doesn't sign their NDA. Daniel, is also the author of, AI 2040 or, or something like that, 2030, which talks about the,

Alex Volkovthe potential outcomes of AI. and then OpenAI called Congress to create mandatory national safety regulations, which again, I'm not against. I think, like it's very important to have. I just don't want this to be concentrated in the hands of like a, a few labs. we are here to fight doomerism on the show, folks. I think the AI can do much more. I think that, strong antom- anthropomorphism is a problem. I think when people say the AI, they mean that there's gonna be only one. I don't believe so. I believe there's gonna be multiple AIs. I'm gonna have my personal superintelligence. It helps me. A- and so I think that there's a lot of hand-waving things, and honestly,

Alex VolkovI think nobody fucking knows. Nobody fucking knows that Peter Steinberger, a Austrian dude out of nowhere last year, based on his home project, changed the world trajectory because he ran agents in a very specific loop with very specific tools that now everybody runs. Nobody could have predicted that. Not Eliezer Yudkowsky, not Paul Cristiano. No one could have sat there and said, "Hey, this is the way towards personal superintelligence." Literally no-- So no one knows. Yes, pe-fig- people can say we don't know how to solve alignment, but also nobody knows that may- maybe there's this one weird trick And so, I, I say that, like, it's very hard to predict.

Alex VolkovWhat I say is that, you know, the, the, the advancements are happening. We might as well enjoy them, and I really want, to see the best of it. Navio Stocks is one example this week I think is a very interesting, externals, where Navio Stocks is one side. On the other side, hey, all of these doomers are saying AI is killing us all, basically what are we doing? Let's stop, and nobody's stopping. No one has a plan of how to stop DeepSeek. Nobody has a plan of how to stop the folks from Ablation AI that we had last week on the show that are taking open source models, removing all, but all considerable, besides child safety stuff, removing all

Alex Volkovrestrictions and putting it out there for everybody else to hack each other.

Wolfram RavenwolfDid you hear anything, that anything happened after that came out? Nothing, right? Nothing And I think my P doom for AI is much less than my P doom for humans. If you look at the world where human, humanity has gone without AI, I think we should risk the chance that with AI we will get something better if it's widely distributed

Alex VolkovAll right, folks, I think, the, we could fill a whole hours of debate

Nisten Tahirajthese are very anti-democratic views and, the healthiest thing you can have is an ecosystem. so a monoculture is bad in any ecosystem, whether silicon or, or organic based. And, still encryption works. If you want it to work, it does work. And that's the main tool we have. If we distribute it and we're able to have these assistants that are loyal to us and can write their own encryption and provide people privacy, they can do that. But that's not going to happen if you just pull it… It's a very autocratic view that this, "Oh, this is so dangerous. Only I can have the printing press."

Nisten TahirajLike, you're just being another 15th century sultan now, and, that can go pretty bad.

Alex VolkovAll right, folks, I think, we've covered this thing enough. there's plenty of stuff to cover in, in the big world, big labs, but, we will keep monitoring. We'll let you know if anything major happens there. I think, we'll skip Apple's, event. That has nothing to do with AI. The only few things about the Apple event with AI was, AirPods can now translate you, and Apple Watch will now transcribe with AI and, and y- you have this, like, what did he say button. The Apple Watch continuously transcribes on the device with no, cloud, and will tell you, "Oh, this person said this." I think that's super cool. also you'd be able to, like, record everything, like all your

Alex Volkovmeetings with Apple Watch, on device, which is super cool. And obviously Siri AI that's coming. I've been using Siri AI. It's, it's pretty good, but it's nowhere near the, the level of agentic user-facing AIs that we, are going to talk about next. So now let's move on to our, corner of agentic AI, with, Do we have? Yeah, do we have… All right. this week, Meta came to us and said, "Hey, do you know that thing, OpenClaw? Do you know that thing, Hermes? Do you know Grok Bot, Instinct, Town? All…" There's tons of them. here's our attempt at this. Meta introduces Muse, a free 24/7 proactive AI assistant agent that has

Alex Volkovits own browser, its own computer, and is connected to your systems at Meta. Now, yes, folks, this is the same Meta that, was Facebook before, that people are very much worried about taking their data and training on it, et cetera. Th- this is, like, the same Meta. However, I feel like that was a meme a long time ago and, Mark Zuckerberg is not to be disrespected. Because after the Llama 3 and open sourcing everything, he literally just, like, took his money cannon and, and directed this money cannon towards this problem, and Meta has been just incredible at products. Meta Muse Spark 1.3

Alex Volkovfrom last week is the second winner from last week. Last week, Astra was announced, and Meta Muse Spark is one of the top, you know, somehow. nobody still counts them. Like, we still don't see charts from DeepSeek and Anthropic and OpenAI that are adding Meta Muse Spark, but they should because it's coming. And Watermelon, which is their much bigger, size model, is coming. But also products coming, and nobody does distributions like Meta. Like, you know, we, we talked about Google. Google's the fact that agentic AI completely, completely fizzled. Gemini Spark is nowhere nearly usable at all. Meta knocked it out the fucking park with Muse.

Alex VolkovI've been using Muse only for one day, and I can already tell you this is going to change many people's lives given where it's at already and given how good it is already. And yes, again, I'm talking about Facebook. It's ridiculous. But I'm not talking about just all Facebook. I'm talking about the MSL labs in Meta that paid billions of dollars to people to move forward. the MSL labs that took Nat Friedman and Daniel Gross together, two of the more prolific investors in Silicon Valley. Nat Friedman used to, own, GitHub, gr- like, CEO of GitHub. I'm talking about Meta who bought Manus and then had to say, sell Manus back to China because of anti-regulation, stuff in China.

Alex VolkovMeta built Muse. So let's take a look at Muse. Anybody try Muse already besides me? Anybody in the comments tried Muse? A little bit. A little bit All right. this is Muse, folks. and I will show you Muse. Let's see if I can pull up Muse, correctly. My Muse, by the way, obviously is called Wilfred. As you might imagine, it's been Wilfred since a while so,'cause I imported it. let's take a look at Muse. So here, this is what we are seeing. We're seeing this, interface, and, specifically here, as you guys remember from previous ThursdAIs, we're seeing the producer, the, the Muse agent that listens to our show and knows what we're talking about.

Alex VolkovSo in a second, I will ask Muse to kinda introduce itself, and reply to me, and hopefully Muse will listen. But, this is the interface. You can animate your own character. They added animations. You can just tell it, "Hey, I want you to look like this." So I said, bionic wolf, and they, have animations. Then you guys see the connected thing? This changes. So, look up, my co-hosts X, Twitter accounts and show them to me. Once I send the query, you can see that, like, it's working. It's not showing you what, CRL commands or bash commands it runs. It just says working. Searching the web. Searching X accounts.

Alex VolkovIt-- This is very, like, approachable. Many folks who I installed OpenClaw and Hermes, et cetera, they need exactly this. It's also stupid fast. Do you guys see how quickly it pulled up all of your, handles? and it got all of them right. where's LDJ? Where is LDJ? Why didn't you pull up LDJ? let's find LDJ. But here is the number of actions it took to get this. You guys see this? Read, read, read, read, read, read. it, it manages memory in people, so it knows who you are. Once you talk… you're right. LDJ is co-host list. LDJ confirmed. saved him too. The speed of Meta Muse Spark with this is the first thing that you get.

Alex VolkovNow, I will say about speed, I thought about it this morning. Grok was also very fast in the beginning. Now Grok is a little bit slower. Instinct, another AI system, was fast, now it's a little bit slower. speed can vary with the amount of people that use the product, so, like, judging the speed only in the start is not that great. However, it is-- it does feel very fast. It does feel very approachable, and it has, things from, You guys will like this. It has things from, from OpenClaw. It has a soul MD file. There's an MD file that talks about its soul generally helpful, not performantly helpful. actual file that they look in. They have a memory file, which I won't show you because it

Alex Volkovincludes all my memories, because I exported it from all my previous assistants and imported it here. And, a- and the computer use is really good. Like, it does the clicks and the computer use stuff very, very, very well. So setting this assistant up to be our producer for the show this week, is… was very easy. And now it says, "New Chiron is up. MetaMuse personal agent on Linux VM, native Stripe link kicker." So it puts up kinda the, the stuff that we talked about on the show, it puts up live on the, on the page. what else can I tell you about this? It's really nice in terms of connectors as well. One of the things that I've told you previously on the

Alex Volkovshow, that I use Stripe link. Stripe link has an agent thing that allows you to buy things. Basically, give your agent a credit card safely so that you… n- not your actual credit card. Meta has a Stripe link integration, Stripe has, an integration here, and y- yeah, this doesn't show it because, like, they take it away, but basically every time you wanna spend some mon- you want your agent to spend some money, you tell it to go and buy, and it will show you a card. "Will you approve this purchase?" they take your credit card, and they charge only, like, a credit card that you don't own, basically, making it safe for your agent to p- purchase this for you.

Alex VolkovI've used Stripe link on multiple agents. This one is the first that was, like, built natively, integrated natively, and, is feel safe.

Wolfram Ravenwolfthis is free?

Alex VolkovYes.

Wolfram RavenwolfMeta provide this for free. Yes. What are the limits? What are the, the daily limits or stuff?

Alex Volkov100 million tokens per week- Uh-huh … for free. as- which is, which is pretty nuts if you think about this, because, Groq is not free. You have to subscribe to Groq, like, pro tier, et cetera. Zack has a money cannon. Now, a very important- And when you run

Wolfram Ravenwolfout, what happens if you run out?

Alex VolkovYou can pay. You can pay, yeah. You can upgrade. here's the thing. You go to Data Controls, and then you uncheck this Help Improve Our AI Models checkbox. Make sure to do that so that you don't inadvertently help them, train their models. And I wanna talk about security next, but, questions and comments from the folks in the

Wolfram Ravenwolfconversation. I have one more question. So it's 100 million token per week. Okay, sounds like, like much, but I just checked. I have more than that per day.

Alex VolkovLook,

Yam Pelegfir- first, I just wanna point out the native connection to WhatsApp-

Alex VolkovYes … Yam Peleg: which, which all the other agents, doesn't matter, open source, closed source- Yeah … Yam Peleg: you always get a connection. Can just say it like that because, it's not easy to get. W- WhatsApp is not Telegram. Yeah. but here, yeah, you get it from the source, like absolutely native. the thing that I love about this, Yam- I just wanna… Mm … is that you can talk about, WhatsApp, and then you have this like channel with WhatsApp that's read only, so view only on the desktop. So you can see your chat with WhatsApp. Like, you can, you can actually see it on the desktop so that you know what you talked about. That's incredible.

Yam PelegDoes it have access to my other chats? Like, can I ask questions about-

Alex VolkovYeah, yeah … I

Yam Pelegdon't know, see what I asked

Alex Volkovfolks are asking, Tony's asking in the comments, "How many agents does it have access to on Muse?" so this is a very interesting thing. Unlike Grok Bot, where Grok Bot, like we told you, is a series of AI agents with their own definition to talk to each other, there is one main chat here in Muse. That's a very good decision for many folks 'cause they're not ready for multi- multi-agents, main chat. And then there is, side chats. You can open as many side chats as you want, and any one of them can be like a specific agents. All of them have, very interesting, heartbeat thing. So there is a heartbeat. th- there is a heartbeat, and there is a Chiron watch every minute.

Alex VolkovSo this is the, the thing that makes sure that it listens to our show and puts stuff up on the, Gyron, Chiron, I don't know how to say this. you can define this, here. You can say, "Hey, I want heartbeat every 15 seconds," 15 minutes. But I do think that at some point, though, those agents will differ in the UI because the underlying layer will be pretty much the same. They do the same stuff. They read your email, they read your documents, et cetera. The UI is gonna matter. The stuff that I noticed is the productivity. Let me scroll up a little bit. where is this? Where

Yam Pelegis this? I, I, I just wanna say, because it's- This okay, continue, continue.

Alex VolkovI was just browsing, and I saw, "I can plan Emma's birthday party for this weekend." My daughter, Emma, is turning eight this weekend. Out of nowhere, and obviously, I inputted my memory, so it knows the dates, et cetera. it also knows all your contacts if you want to, which is also a novel thing. You can upload your contacts connector via your iPhone, so like it has native iPhone connectors. You can connect Apple Health. You can connect Reminders. This is all novel, by the way. I don't know of any other agent right now that you can natively connect via your iPhone connectors to the Apple ecosystem. That's great. So it knows my dates.

Alex VolkovIt proactively said, "Hey, I can help you plan Emma's party for this weekend." I told you that 2026 is the year of proactive agents, and this proactivity just blew my mind. That's what I want. I want my agent to know things about me and then suggest things to me. That's what I want. Grok Bot, I need to set this up. Instinct is pretty good at this, but mostly they do-- they go off ba- reading the email. This seems like it goes further Ream, comments about this before we go to the Meta- Yeah, yeah … Facebook data, thing?

Yam PelegLook, I think the main question is Which one supports which connections? Because you don't have full control. Now, it's not open source that you can just do whatever you want, connect it to whatever you want. That's, like, the power of OpenClaw was because it's a completely, open and free thing that you can just customize and connect to whatever you want. Now, I think the c- the, the most important question is, okay, what connection do I get from Groq, and what connections do I get here? And I don't know, I'll, I'll decide. I mean, what you just explained is, is really, sounds really good, but-

Alex VolkovYeah

Yam Pelegdo I have this elsewhere or no? I, I mean, what, what's the actual difference?

Alex Volkovdifference on, on features, we cannot go into all differences right now, but, like you said, native connectors, I think is a differentiator. WhatsApp, Facebook Marketplace is a big one. Tons of people buy shit on Facebook Marketplace. if you use other AI agents, Facebook can block you because that they don't want other agents. Facebook has a competitive advantage in, you cannot put them on WhatsApp, for example. It's not as easy. so th- there's that. The iPhone connectors is unique. I haven't seen this in any other places. the way they do this is, very interesting. So here you can connect to Apple Health, your contacts, for example, so it would know all your contacts.

Alex VolkovBut I think that the, the way they're building this is very important, and Zack, went on Alex Heath's, e- podcast and talked about specifically the safety stuff, okay? least privilege, thoughtfulness, et cetera. They have a built-in sentinel, I believe that it's called. this is the architecture, by the way. They have another LLM running there, reading all the network requests coming out and in, making sure that they're not, LLM, injections. They have a injection security. If you can prove that you got, prompt injected via an email, you'll get, like, $130,000. They have very mi- … a lot, a lot of money there for incentives of whether or not this, you know, this can be fucked with, or somebody can

Alex Volkovsend you an email and, and screw you.

Wolfram RavenwolfMeta had also the safeguard models and so on, so it's probably using those. And the funny thing or ironic thing is that Peter Steinberger was offered to join OpenAI or join Meta- Yeah … and he talked to, Zuckerberg. So Zuckerberg seems to be a great fan of OpenClaw because this looks exactly like with the soul MD, the memory MD, with the heartbeat. It feels exactly like if you took OpenClaw, put it in a, in a hosted environment and improved the UI. I think,

Alex VolkovI think the level to which Peter Steinberger changed the world is un- Described. It's really strongly resembles multiple of the best things about OpenClaw. It also removes all of the horrible things about OpenClaw. Speaker 4: You don't need a Mac Mini at all. They have their own computer. Here, I wa-- the last thing that I wanna talk about, not the last thing, but the, the most important thing I wanna talk about here is the secure stuff. So we can talk about browsing, et cetera, how they save your, keys. They don't save them in plain text. the LLM does not see your keys. It's in secure storage. Muse Confidential VM, I think, is the most important thing

Alex Volkovthat we need to talk about here. they announced Muse Confidential VM. They say, "We believe that people should be able to choose a fully private mode for personal agents, where even Meta or any other service provider cannot see or grant access to your information." This is why we're investing significantly in Muse Confidential VM and plan to deliver this capability later this year. Muse Confidential VM is intended to be cryptographically and verifiably prevent Meta from accessing data in your VM. already using the system with a small group of trusted testers, and we've begun making our design and the source code for the system available to external auditors.

Alex VolkovWe're in the process of taking auditor feedback. Once launched, we have a continuous audit of the system that will be visible and inspected by anyone. Experts will be able to confirm that Meta does not have the ability to access data within the VM environment. We also welcome security privacy expert to reach out to us about early access." Zuck, talked about on the podcast is Moxie Marlinspike, the founder of Signal, which is notoriously known as one of the most, secure messaging platforms out there, was recruited by them. Nat Friedman recruited Moxie Sp- Marlinspike to work on the secure VM. This, to me, kills all of the stuff about, "Hey, Meta," blah, blah, blah,

Alex Volkov"take your data," just completely… Th- this is more secure than my Mac Mini. Because Mac Mini is not verifiably secure, and it's not being continuously tested. I think that this, to me, puts to rest about data. They know the meme. They understand the people know that, Facebook has access to everything, and, they take this very, very, very seriously, more seriously than anybody that I've seen before. Zuck is not by mistake after 18 years still at the helm of one of the biggest companies in the world. And the super intelligence efforts that, Meta is going towards, Zuck is very, like, open about them. Personal super intelligence should be personal and beholden to you, and, Muse

Alex Volkovseems to be, like, the first step there.

Wolfram RavenwolfI'm super excited. You know how deeply invested I'm in the Hermes eco- ecosystem with all the patches and so on, but I will share this, architecture diagram with my agent and see if it can improve our own setup that way. Yeah. And I think the, the confidential computing aspect, that is super important and a good precedent that hopefully others will follow because we are sharing so much personal… The most personal data we are sharing with our own agents. Everything is in there. Everything is in there. Health data, personal problems, everything. And this is, like a, like an attorney, like a psychologist, that kind of information, it should get more protection.

Wolfram RavenwolfThat is something rarely discussed, but having ways in that direction, moves in that direction, I fully support that.

Alex VolkovYeah. So folks, give, M-Muse a try at muse.ai and tell us, what you think. as a reminder before, we're almost at the end of the show. Token Juice, if you want DeepSeek for free, but you're okay with sharing your data, Token Juice, from Aaron Batie, a friend of the pod, will get you that as well. if you wanna join us, September and see Pitbull, Mr. 305, Mr. Worldwide, please, please join. This is gonna be super cool. I wish it was in Miami also, not in San Francisco, because this would fit. And, I think we covered pretty much everything. with that, I think we've covered, like, a very decent chunk for this week. Wolfram Ravenwolf, Nisten Tahiri, LDJ confirmed.

Alex VolkovLDJ's last name is confirmed. and, and Yam Peleg, thank you so much. We also had Chris Alexius here from NVIDIA talk about DeepSeek. if you missed any part of the show, the full replay is going to be available on ThursdAI.live, and, please check it out, vibe coded with Fable myself. And, if, if you want the links to sources like the Nisten simulator of DeepSeek layers, or Nisten, please add the Astra Mars, simulation as well in the link. we will send them to you- Oh … on ThursdAI.news newsletter and podcast, which you can also subscribe to on ThursdAI.news or ThursdAI.live. We will be here next week.

Alex VolkovThank you so much for joining, everybody. Have a good one. Bye-bye.

Nisten TahirajBye, everybody

Every Thursday since GPT-4

The highest-signal weekly AI news show.

One conversation a week with the people who actually run the models. Follow the podcast, or get the newsletter for free.