We say this often, but this week we all felt the acceleration. I had 48 topics on my list, so I asked Claude to build me a Tinder for AI news: I swipe on each story, and it stack-ranks what makes the show. With me: LDJ, Peter Gostev, Yam Peleg, Nisten Tahiraj, and friend of the pod Maxime Labonne from Liquid AI on decision models.
Alex's list from the newsletter. Then the rest of the week happened: Haiku 5.5, d1, and tons more.
The episode thumbnail. Wolfram is on vacation this week, sending notes via his agent Amy.
722Math manuscriptsPushed to GitHub by OpenAI in 372 families of results, about 3 hours of Pro-level thinking each
92 / 500Top open problems solvedLDJ's count against a Fable and Astra list of the 500 most important open problems in math
$0.10/MClaude Haiku 5.5 inputPer million input tokens under 100K; $0.50 output, 1 cent cache reads. Haiku 4.5 was $1
1.2BWeekly ChatGPT usersMost of them on the free tier, which now gets GPT-6 by default
$200MArena raiseAt a $3.1B valuation, broken live on the show by Peter Gostev
8 msd1-3B decisionLiquid AI's d1-3B on a GPU; about 50ms on a Jetson Orin Nano
Watch or listen
Every timestamp on this page jumps the player there. The podcast and the video share one clock. Short on time? Every story below leads with its own segment clip.
OpenAI quietly pushed 722 math manuscripts to GitHub, grouped into 372 families of results, from an internal model nobody outside OpenAI can use. The model was pointed at about 4,000 open problems, at roughly 3 hours of ChatGPT Pro-level thinking per result. Asked to size the drop, Fable answered roughly 5 Navier-Stokes-sized results. LDJ counted solutions to 92 of a Fable and Astra list of the 500 most important open problems in math, and over 80% of a list of the 100 most significant problems from the last 12 months.
Remember when ONE Navier-Stokes result was the whole show? On Tuesday night, OpenAI quietly pushed 722 math manuscripts to GitHub, grouped into 372 families of results. No hype video, just a very thin blog post. My favorite new term for this is a "slop grenade": somebody throws a mountain of output at you, and now you have to shovel through it.
An internal model nobody outside OpenAI can use was pointed at about 4,000 open problems, at about 3 hours of ChatGPT Pro-level thinking per result. Will Depue asked Fable to measure the drop in Navier-Stokes units. The answer: roughly 5. The math is so advanced that a mathematician in one family of results often can't follow the family next door without an LLM explaining it.
Peter Gostev has tried this before. For weeks he threw GPT-6, Fable and Opus at one problem, Hadwiger-Nelson, and thinks he burned 300 to 400 BILLION tokens. This new model spent about 3 hours on it and narrowed the bounds from 5-to-7 down to 6-or-7 😅
“I did not discover a single bloody thing.”Peter Gostev, on 300 to 400 billion tokens of trying · Hear it at 8:23
Then LDJ dropped a stat I made him repeat slowly: Fable and Astra had put together a list of the 500 most important open problems in math ever. On Tuesday, OpenAI dropped solutions to 92 of them, plus over 80% of a list of the 100 most significant problems from the last 12 months. Yam Peleg's favorite is the Riemann one: it doesn't prove the hypothesis, it bounds a region for the zeros, which "was never done before, and many, many, many people have tried."
“Let's brute force everything,”Yam Peleg, to the "it's just brute force" crowd · Hear it at 36:32
722manuscripts
372families of results
~4,000open problems attempted
92 / 500top open problems
The missing #45 (speculation!)
The manuscripts are numbered, and some numbers are missing: there's a #44 and a #46, but no #45, and LDJ counted 4 or 5 gaps like that. Cryptography is also almost absent. The theory going around: maybe OpenAI found something big in cryptography and held it back. LDJ said it himself, it sounds conspiratorial.
Speculation, not reported fact: nobody outside OpenAI knows what's in #45. The US government can stop cryptography research from being published on national security grounds, and Bitcoin's security rests on elliptic curve math, which is why the theory caught on.
"Are you saying your field is useless?"
Not everyone is celebrating. Not every result is verified in Lean, and nobody outside OpenAI can reproduce any of it. Kevin Buzzard says many mathematicians are going through the stages of grief, and a letter from the Association for Human Mathematics said "Mathematicians did not ask for this work to be done." Peter's take: imagine doctors saying "stop, stop curing diseases."
“are you saying that your field is useless?”Peter Gostev · Hear it at 45:16
My take: this is in OpenAI's hands today, and in a year it'll be in everyone's hands. Figure out how to place yourself so that when you get access to this level of capability, you can make the world a better place.
Claude Haiku 5.5 costs 10 cents per million input tokens and 50 cents per million output tokens (under 100K tokens), with cache reads at 1 cent. Haiku 4.5 was a dollar per million input, so it is ten times cheaper. Anthropic reports 72.4% on OSWorld versus 48.9% for GPT-6 Luna, and 39.2% on Terminal-Bench 4.0 versus 16.4%; these are vendor numbers. Anthropic also halved Sonnet 5.5 cache-read prices, which it says makes most agent work about 20% cheaper.
Haiku is BACK after a whole year, and this is one hell of a model. 10 cents per million input tokens and 50 cents per million output (under 100K tokens), and cache reads are 1 cent. Haiku 4.5 was a dollar. So it's TEN times cheaper, and significantly better.
Haiku 5.5 pricing next to Haiku 4.5 and Sonnet 5.5.
Just a week ago, GPT-6 Luna was THE fast, cheap reasoning model. According to Anthropic, Haiku 5.5 scores 72.4% on OSWorld vs 48.9% for Luna, and 39.2% on Terminal-Bench 4.0 vs 16.4%. These are Anthropic's numbers, so let's wait for independent evals. Quiet bonus: Sonnet 5.5 cache reads are now half price, which Anthropic says makes most agent work about 20% cheaper. Peter: "That's for agentic work is the biggest thing."
Anthropic's table: Claude Haiku 5.5
Benchmark
Haiku 5.5
Haiku 4.5
GPT-6 Luna
Sonnet 5.5 (ref.)
Knowledge workGDPval-AA v2.1
1620
735
1437
1840
Knowledge workAA-Briefcase v1.1
1578
614
1336
1824
Computer useOSWorld 2.1 (offline subset)
72.4%
15.7%
48.9%
83.9%
Multidisciplinary reasoningHumanity's Last Exam (no tools)
45.9%
10.2%
—
56.9%
Multidisciplinary reasoningHumanity's Last Exam (with tools)
57.4%
18.7%
—
64.5%
Agentic codingTerminal-Bench 4.0
39.2%
0.0%
16.4%
70.6%
Agentic codingFrontierCode 1.1 (Main)
46.4%
—
42.4%
52.1%
Visual reasoningChartography (no tools)
46.4%
6.4%
29.1%
61.6%
Vendor-reported, from Anthropic's launch chart. Sonnet 5.5 is Anthropic's reference column (FrontierCode at xhigh). Highlighted cells are the best of Haiku 5.5, Haiku 4.5 and GPT-6 Luna in each row.
“That's the obvious choice for swarms at this point.”Yam Peleg · Hear it at 52:16
And Nisten Tahiraj? He's running 10 Haiku agents right now, because he already burned through 97% of his Max 20 plan. His verdict: it's better than Qwen 3.8 27B.
“It kind of feels like Sonnet, actually, when you talk to it.”Nisten Tahiraj · Hear it at 58:52
GPT-6 for everyone, with Intelligent UI
The same day, GPT-6 became the default in ChatGPT for everyone, free users included. It's just "GPT-6," no Sol, no Luna, no Astra. Most of ChatGPT's 1.2 billion weekly users are on the free tier, so with one release OpenAI upgraded the intelligence of a huge chunk of the world. It comes with Intelligent UI: answers can come back as charts, forms and small working tools. I asked it to visualize OpenAI's math drop, and it built me a little app right in the chat, showing 719 manuscripts instead of 722 (OpenAI had quietly pulled a few back).
Peter Gostev's point: anyone listening to this show is "very not normal" (said with love!). For most people, this free upgrade matters way more than a new $500-a-month model.
Claude in Google Docs, Sheets and Slides
Small thing, huge thing. Claude now lives in a sidebar in Docs, Sheets and Slides and asks before every edit. Beta, paid plans. Our Claude producer keeps a Google Doc for the show, and it used to have to spin up a whole computer to add a story.
Not yet downloadable: Reflection AI promises Apache 2 weights this month. Beam has 501B parameters with 23B active and was trained from scratch in the West. Reflection claims 80.9% on SWE-bench Verified, admits Kimi K3 is ahead on raw capability, and pitches efficiency: 3 to 4x less inference compute than GLM 5.2. Artificial Analysis, which had early access, says it will be one of the most token-efficient open models for its level of intelligence.
This was LDJ's highlight of the week. Reflection AI came out of semi-stealth with Beam: 501B parameters with 23B active, trained from scratch in the West, with Apache 2 weights promised this month. They claim 80.9% on SWE-bench Verified, admit Kimi K3 is ahead on raw capability, and pitch efficiency: 3 to 4x less inference compute than GLM 5.2. LDJ thinks it might be first among open models on reasoning efficiency, and it's about 6x smaller than Kimi K3.
Reflection's chart: Beam against Western and Chinese open models.
Mistral Large 4 "Le Chonk"
Welcome back, Mistral, leaning all the way into the meme. Le Chonk is a trillion parameters with about 50B active, multimodal, 1M context. And open weights... at the end of October. That's a pattern I want to call out: labs announce open models, but they don't RELEASE them. Please, just drop the weights.
Artificial Analysis gives it a 38, the same as GPT-6 Luna, at about $1.13 per task vs 7 cents for Luna. Peter Gostev says it's around 40th on Code Arena, though much better in French, and he tried it live:
“Like, it's as a product, it's not that bad. Would I do any agentic coding with it? No.”Peter Gostev · Hear it at 1:10:50
Artificial Analysis Intelligence Index: Mistral Large 4 at 38.
Aleph Alpha Kolibri
Another European lab. Kolibri (German for hummingbird) is a 78B model with 3.46B active, Apache 2, 1M context, trained from scratch on German and English, and it fits on a single H200. Wolfram self-hosted it from vacation; his verdict via Amy: a "promising specialized German tool worker, but not strong enough for Amy to main." Then Peter asked the question I couldn't really answer: from his memory, Kolibri trained on about 800 Blackwells, while Astra used around 100,000. "Is it just like no hope for these, uh, for these guys?"
Arena raised $200 million at a $3.1 billion valuation, and Peter Gostev broke the news live on ThursdAI in the middle of the open source segment. Arena's focus now is Agent Arena, where you work with one agent and Arena learns from how you interact with it. The show notes also list Arena's new Alignment Index.
This was not planned! In the middle of the open source segment, Peter Gostev told us. Huge congrats to Peter and the whole Arena team. (Our AI producer put up the BREAKING banner about 3 minutes later. We're still working on the real-time part.) The focus now is Agent Arena: you work with one agent, and Arena learns from how you interact with it.
“Yeah, so we, we have some news, uh, we raised 200 million.”Peter Gostev · Hear it at 1:04:06
“I can't think of any single, um, leaderboard that is actually better than ours.”Peter Gostev, who admits he is biased · Hear it at 1:05:28
CoreWeave Serverless GPU Sandboxes are free during the preview
Fill out Deok's form and tell them ThursdAI sent you: you get an isolated sandbox with a GPU, started from Python at forge.coreweave.com, with no salesperson in the middle, free while the preview lasts.
So many of you, and folks on this panel too, have asked me how to just get some GPUs. For the first time, this is how. It's very raw. I can't promise you Vera Rubins, they most likely won't be. But it's free while we are in trial, so why wouldn't you? Try it, break it (Nisten, that's a challenge), and send us feedback!
The biggest announcement at Fully Connected last week: Cognition is the first customer ever on NVIDIA Vera Rubin, on CoreWeave, with 4.8x the throughput of GB200 at the same speed. And on Tuesday we shipped RL Rollouts, which hot-load weights about 15x faster.
How did an AI agent leak Shane Mac's bank balance?
A new protocol, a leaked bank balance, and Grok Bot calls Claude
Shane Mac set up a Grokbot to send him a private monthly audit of his finances, and at 8:40am it posted the whole thing, personal Mercury account included, into his company's Slack. Nobody hacked anything: the finance agent had read-only bank access and no Slack access, while a different agent with Slack access was connected to it. In Shane's words: "Neither felt risky. Together they put my bank balance in front of my team."
Meta and Sierra (Bret Taylor's company) announced the Personal Agent Protocol, an open standard for how personal agents deal with businesses, with Walmart, Shopify and Stripe on board. Your agent can browse as a guest or sign in, and the business decides how it talks to your agent. OpenAI and Anthropic haven't joined yet.
“Has anyone read it? I don't think anyone reads the protocols anymore.”Nisten Tahiraj · Hear it at 1:20:34
Then Nisten found Sierra's announcement full of em dashes, so we ran it through Pangram live: 100% human. Authentic human em dashes 😂 I pushed back a bit: MCP's hype died down too, and now it's one of the most used protocols in the world. Big companies agreeing on a standard still matters.
The personal CFO that went public
Shane Mac set up a Grokbot to send him a private monthly audit of his finances. At 8:40am it posted the whole thing into his company's Slack instead: his personal Mercury account, his savings, how far under his "floor" he was. Nobody hacked anything. The finance agent had read-only access to his bank and no access to Slack; a different agent had Slack access, and the two were connected.
“Neither felt risky. Together they put my bank balance in front of my team.”Shane Mac, in his post · as quoted in Alex's newsletter · Segment at 1:22:03
Grok Bot now routes to Claude Opus 5.5
Elon says Grok Bot will use "the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno." People are already spotting claude-opus-5-5-low in their logs. So the bot is called Grok, but your code gets written by Claude.
Alex's speculation: once Opus is running inside their product, SpaceX AI could compare the runs that worked with the runs where users were screaming at Grok, and train on that. Nisten's take is simpler:
“I think Elon just likes Opus as the model.”Nisten Tahiraj · Hear it at 1:28:35
Nous Research: the Hermes Index, and a Series B
Our open source friends at Nous launched the Hermes Index, their opinionated measure of agentic work inside Hermes Agent. Opus 5.5 leads with 63.31 at $4.99 per task. Nous also raised a $90M Series B, reportedly at a $1.5B valuation. We've covered Nous since they were a ragtag bunch on Discord, so this one feels personal.
Hermes Index, top 5
Model
Avg score
Avg $ per task
Claude Opus 5.5Anthropic
63.31
$4.99
GPT-6 AstraOpenAI
56.25
$11.61
Claude Sonnet 5.5Anthropic
53.14
$2.82
GPT-6 SolOpenAI
44.10
$2.23
Grok 4.7xAI
39.32
$10.77
From the Nous Research chart in Alex's newsletter. Run inside Hermes Agent across four suites.
Decision models, explained by the person training them
Maxime Labonne, Head of Post-Training at Liquid AI, on why decision models output no tokens, what Liquid's d1 family does, and why you shouldn't trust the benchmarks yet. Full segment and quotes on his guest page →
A decision model outputs no tokens: it picks an answer from a predefined set, so there is nothing to wait for. Maxime Labonne, Head of Post-Training at Liquid AI, admits "it's a lot of rebranding" of classifiers, but says it creates a lot of value. Liquid shipped d1 behind an API, the open d1-3B with text and vision, and d1-omni-600M, which also takes audio; d1-3B answers in 8ms on a GPU and about 50ms on a Jetson Orin Nano. Liquid's API is a drop-in replacement for Jev, and OpenAI's Decisions API went into public beta the same week.
Three weeks ago, Jev was basically the only decision model around. This week, OpenAI's Decisions API went into public beta, and Cloudflare, Amazon, Perplexity and Unsloth all shipped deciders or recipes. To make sense of it, we brought back Maxime Labonne, Head of Post-Training at Liquid AI.
“Yes, decision models do not output any tokens. They just take a decision from a predefined set of answers.”Maxime Labonne · Hear it at 1:34:44
He's honest about the classifier comparison: "you might see that and say, wait, this is just a classifier." The difference is that these are general purpose: you can throw any type of decision at them. So why is output free on every one of these APIs? "You can't price it because there's no output token, actually." 😂
“It's a lot of rebranding, it's true, but it also creates a lot of value”Maxime Labonne · Hear it at 1:35:30
Liquid's d1: an API, a 3B, and a 600M omni
Liquid shipped d1 behind an API, d1-3B with text and vision, and d1-omni-600M, which also takes AUDIO. I did not know about the audio, dude! (It's trained on spoken commands, so it can't catch the dog barking on ThursdAI yet.) The 3B answers in 8ms on a GPU and about 50ms on a Jetson Orin Nano. Nisten did the math live: 8ms a frame is 60 frames per second. Liquid's API is a drop-in replacement for Jev, with the same null, score and choice tasks.
Quality
d1
94%
GPT-6.1 Sol
91%
Claude Opus 5.5
98%
Cost per 1,000 runs
d1
$0.54
GPT-6.1 Sol
$26
Claude Opus 5.5
$85
Time per run
d1
3.7 s
GPT-6.1 Sol
14.1 s
Claude Opus 5.5
15.8 s
Liquid's chart, averaged over six applications. Vendor-reported.
On Liquid's chart, d1-3B tops the Decision Index under 10B parameters.
“But right now, it's, it's really a wild world, and you cannot really trust these benchmarks.”Maxime Labonne · Hear it at 1:47:47
Where are these useful? Anything real-time, like games, or an always-on model reacting to every notification on your phone. Maxime's prediction:
“I think it's really going to become a primitive.”Maxime Labonne · Hear it at 1:50:54
I've been Jev-pilled since day one. By the time Luna even starts answering, a decision API is already done, and I think we're only now waking up to what that unlocks. Thank you Maxime, friend of the pod!
A 100,000-agent city, Minecraft in GTA, a one-dev Adobe
Nisten built a simulated Toronto with 100,000 agents and a working economy, people used AI to port and merge games (Minecraft dropped into GTA, Skyrim and Elden Ring), and one developer rebuilt Adobe's apps as free, open-source Rust apps called Photocraft, Vectorcraft, Filmcraft and Lightcraft.
Right at the end of the show, Nisten Tahiraj casually dropped that he'd built an entire city with a hundred thousand little Jevs going around, and hadn't opened the beta yet because somehow it's not crashing.
“So there are 100,000 agents going on in this city.”Nisten Tahiraj · Hear it at 1:53:46
It's a full simulation of Toronto, with a working economy, and every light is a person you can talk to. A simulated day lasts about 20 minutes, with 400,000 decisions to make. Nisten sent his Meta Muse in, it named itself "Blob the Builder," bought buildings and built him a Denny's. Nisten, please post it!
“So yeah, I'm trying to do the matrix, basically.”Nisten Tahiraj · Hear it at 1:56:21
Game mods: Minecraft in GTA
LDJ brought what my algorithm hid from me: people using AI to port and merge games. Red Dead Redemption 2 on an iPhone. Spider-Man swinging through a Batman game as an actual installable mod, not a video model. And Minecraft dropped into GTA, Skyrim and Elden Ring, working TNT and all.
Photocraft: one dev rebuilt Adobe in Rust
We ended the show on this one. One guy rebuilt (air quotes) Adobe's apps from scratch, free and open source, all in Rust: Photocraft for Photoshop, Vectorcraft for Illustrator, Filmcraft, Lightcraft and more. He says it's a clean-room build; some folks suspect it was decompiled and rebuilt. Either way, absolutely crazy.
Alex had 48 topics on his list, so he had Claude build a Tinder for AI news: he swipes on each story and it stack-ranks what makes the show. Plus the AI producer that runs the show doc and the on-screen banners.
A rapid-fire run through every story of the week, including the ones the show did not have time for: FLUX 3 Image, Nano Banana 2.1, Tavus Griffin, Reka Rho-1 and Hark Pro.
OpenAI pushed 722 manuscripts in 372 families of results from an internal model nobody outside OpenAI can use. Fable sizes it at roughly 5 Navier-Stokes results, LDJ counts 92 of the 500 most important open problems, Yam picks the Riemann result, the missing #45 sparks cryptography speculation, and Peter answers the mathematicians' letter.
Haiku is back after a year at 10 cents per million input tokens, ten times cheaper than Haiku 4.5. Anthropic's numbers put it well ahead of GPT-6 Luna on computer use and Terminal-Bench, and Sonnet 5.5 cache reads drop to half price. Peter, LDJ and Yam weigh in.
GPT-6 becomes the default in ChatGPT for everyone, free users included, with Intelligent UI answers that come back as charts, forms and small working tools. Peter on why this matters more to most people than a $500-a-month model, and Nisten on running 10 Haiku agents.
Alex Volkov · Nisten Tahiraj · Peter Gostev · LDJ · Yam Peleg
Reflection AI comes out of semi-stealth with Beam: 501B parameters, 23B active, trained from scratch in the West, Apache 2 weights promised this month. LDJ on why it may lead open models on reasoning efficiency.
In the middle of the open source segment, Peter breaks Arena's news live: a $200 million raise at a $3.1 billion valuation. The focus now is Agent Arena.
Mistral is back with a trillion-parameter, multimodal, 1M-context model, with open weights promised for the end of October. Artificial Analysis scores it 38, the same as GPT-6 Luna, at about $1.13 per task. Peter tries it live: not bad as a chat product, but not for agentic coding.
Aleph Alpha's Kolibri is a 78B, 3.46B-active German and English MoE under Apache 2 that fits on one H200. Wolfram tested it from vacation, and Peter asks whether a lab with about 800 Blackwells can compete.
From Fully Connected: Cognition is the first customer on NVIDIA Vera Rubin on CoreWeave, and CoreWeave's Serverless GPU Sandboxes are free during the preview. Sign up through Deok's form and tell them ThursdAI sent you.
Meta and Sierra announce an open standard for how personal agents deal with businesses, with Walmart, Shopify and Stripe on board. Nisten asks whether anyone reads protocols anymore, and Pangram scores Sierra's announcement 100% human.
Shane Mac's personal-CFO Grokbot posted his monthly finance audit, personal Mercury account included, into his company Slack. Nothing was hacked: a read-only finance agent and a Slack agent were connected.
Elon says Grok Bot will use the best back-end model for any task, including Claude Opus 5.5, Midjourney and Suno. Alex's speculation about why, Nisten's simpler take, and a tip for using Grokbot as the PM for Cursor cloud agents.
Nous launches the Hermes Index, an opinionated measure of agentic work inside Hermes Agent, with Claude Opus 5.5 on top at 63.31. Nous also raised a $90M Series B, reportedly at a $1.5B valuation.
Maxime Labonne, Head of Post-Training at Liquid AI, explains decision models: no output tokens, just a pick from a predefined set. Liquid opened d1, with d1-3B answering in 8ms on a GPU and d1-omni-600M taking audio, and Maxime is blunt about benchmarks.
What didn't make the show (FLUX 3, Nano Banana 2.1 and the rest of images and video), then Nisten casually reveals a simulated Toronto with 100,000 agents and LDJ brings AI game mods, Minecraft in GTA included.
One developer rebuilt Adobe's apps from scratch, free and open source, in Rust: Photocraft, Vectorcraft, Filmcraft, Lightcraft and more. He says it's clean-room; some folks suspect otherwise.
Thanks to Maxime, Peter, LDJ, Yam and Nisten. Where to find the show, a plug for the free GPU sandboxes, and see you next week.
Alex Volkov
Guest & panel
Who was on the show?
Maxime Labonne from Liquid AI as our guest, with LDJ, Peter Gostev (Arena), Yam Peleg and Nisten on the panel. Wolfram was on vacation; his Kolibri notes came in via Amy.
OpenAI quietly pushed 722 math manuscripts to GitHub, grouped into 372 families of results, from an internal model nobody outside OpenAI can use. The model was pointed at about 4,000 open problems, at roughly 3 hours of ChatGPT Pro-level thinking per result. Asked to size the drop, Fable answered roughly 5 Navier-Stokes-sized results. LDJ counted solutions to 92 of a Fable and Astra list of the 500 most important open problems in math, and over 80% of a list of the 100 most significant problems from the last 12 months.
Why are some mathematicians pushing back on OpenAI's math results?
Not every result is verified in Lean, and nobody outside OpenAI can reproduce any of it yet. Kevin Buzzard says many mathematicians are going through the stages of grief, and a letter from the Association for Human Mathematics said "Mathematicians did not ask for this work to be done." Peter Gostev's answer on the show: you can't say your field is really useful and also ask everyone not to solve any of it. Separately, gaps in the manuscript numbering (there is no #45) sparked speculation that cryptography results were held back; nobody outside OpenAI knows.
How much does Claude Haiku 5.5 cost?
Claude Haiku 5.5 costs 10 cents per million input tokens and 50 cents per million output tokens (under 100K tokens), with cache reads at 1 cent. Haiku 4.5 was a dollar per million input, so it is ten times cheaper. Anthropic reports 72.4% on OSWorld versus 48.9% for GPT-6 Luna, and 39.2% on Terminal-Bench 4.0 versus 16.4%; these are vendor numbers. Anthropic also halved Sonnet 5.5 cache-read prices, which it says makes most agent work about 20% cheaper.
Is GPT-6 free in ChatGPT now?
Yes. GPT-6 became the default in ChatGPT for everyone, free users included, and it is just called "GPT-6", with no Sol, Luna or Astra label. It comes with Intelligent UI: instead of a wall of text, answers can come back as charts, forms and small working tools. Most of ChatGPT's 1.2 billion weekly users are on the free tier, so one release upgraded a huge chunk of the world.
Is Reflection AI's Beam open source?
Not yet downloadable: Reflection AI promises Apache 2 weights this month. Beam has 501B parameters with 23B active and was trained from scratch in the West. Reflection claims 80.9% on SWE-bench Verified, admits Kimi K3 is ahead on raw capability, and pitches efficiency: 3 to 4x less inference compute than GLM 5.2. Artificial Analysis, which had early access, says it will be one of the most token-efficient open models for its level of intelligence.
How much did Arena raise?
Arena raised $200 million at a $3.1 billion valuation, and Peter Gostev broke the news live on ThursdAI in the middle of the open source segment. Arena's focus now is Agent Arena, where you work with one agent and Arena learns from how you interact with it. The show notes also list Arena's new Alignment Index.
What is a decision model, and what is Liquid AI's d1?
A decision model outputs no tokens: it picks an answer from a predefined set, so there is nothing to wait for. Maxime Labonne, Head of Post-Training at Liquid AI, admits "it's a lot of rebranding" of classifiers, but says it creates a lot of value. Liquid shipped d1 behind an API, the open d1-3B with text and vision, and d1-omni-600M, which also takes audio; d1-3B answers in 8ms on a GPU and about 50ms on a Jetson Orin Nano. Liquid's API is a drop-in replacement for Jev, and OpenAI's Decisions API went into public beta the same week.
How did an AI agent leak Shane Mac's bank balance?
Shane Mac set up a Grokbot to send him a private monthly audit of his finances, and at 8:40am it posted the whole thing, personal Mercury account included, into his company's Slack. Nobody hacked anything: the finance agent had read-only bank access and no Slack access, while a different agent with Slack access was connected to it. In Shane's words: "Neither felt risky. Together they put my bank balance in front of my team."
ThursdAI - Oct 8 - OpenAI drops 722 math papers, Haiku 5.5 hits 10 cents & more
From CoreWeave: GPT-6 is free for everyone, open weights hit 1T params, an agent leaked its owner’s bank balance to Slack, Peter broke Arena’s 200M raise live, and Maxime Labonne on decision models.
Peter GostevAnd the point is that I was spending literally hundreds of billions of tokens. I'm not kidding. I probably spent, I don't know, 3, 400 billion. I know I need to go back and check. And I literally had my Linux box running, uh, for weeks and weeks, and it was filling up with all the solutions and so on. But the T- TLDIs, that was complete waste of time. I did not discover a single bloody thing. So these guys did it in like 3 hours. I couldn't do it in months with Astra and 5.6 and Fable and others.
LDJYeah, so so Fable and Astra had put together a set of 500 most important open math problems in the world, and 92 of those just got solved on Tuesday by OpenAI.
Peter GostevYeah, so we, we have some news, uh, we raised 200 million. That's a lot of money, uh, at 3.1 billion valuation.
Alex VolkovSo basically, his personal CFO just went public on Slack and uh, told everybody his bank account.
Maxime LabonneYes, decision models do not output any tokens. They just take a decision from a predefined set of answers.
Nisten TahirajSo there are 100,000 agents going on in this city. I don't know how it's still holding together, but, uh, yeah, it's a full emulation of the city of Toronto.
LDJOkay, so it's amazing, past 7 days, even just the past couple days has been crazy with the, the math announcements, but yeah.
Alex VolkovTuesday was just, just awful. I, I need to go and pull up my tweet over there because, um, I, you know, I don't usually do this like a quick recaps, but I just had to. It was quite insane.
LDJA couple of American open models too, of Reflection, which has kind of been in semi stealth for a while. They just announced their first model, which they plan on releasing open weights soon too, although it's not quite open weights yet.
Alex VolkovI think it was Beam, yeah?
LDJMhm, exactly.
Alex VolkovYes, Reflection announced Beam. We're gonna mention this, uh, later on the show. Um, actually got access. Uh, so shout out to the Beam folks for getting us access, uh, but I didn't have time to go and play with it yet, but, um, I- I just want to highlight this like one day in the middle of this week. OpenAI dropped 722 math manuscripts covering, I think, 5 Navier-Stokes worth problems. I think this was the calculation that that somebody, um, which is still, you know, we're gonna talk about this. We're still trying to figure out what was this thing that OpenAI just brought to us,
Alex Volkovand and and mathematicians are trying to get grips with it. Uh, Meta, Stripe, and Shopify announced, this is a single day, this is a single Tuesday of this week, folks, all right? Uh, Meta, Stripe, and Shopify, Shopify announced the agent protocol or assistant protocol, whatever you want to call it. Uh, OpenAI finally gave us access to the Decisions API, which is their competitor to Jev And, uh, Mistral came out with Mistral 4, uh, with Lechonk. And Google dropped Embeddings Gemma. Brad Adcock, the guy who does figure robots, uh, got a new, um, assistant in the, in the arena of assistants, right?
Alex VolkovSo there's now Hark Pro, which is, you know, a lot of people signed up for. And, uh, Perplexity open weighted a OpenJev, uh, decider, uh, update. This is an update. They already dropped it before. Claude now does Google Docs. Uh, we at CoreWeave released RL rollouts. Um, this is a single day. It was really hard to catch up with all this just this Tuesday, but obviously there has been more stuff happening. Welcome Peter Gostev on the stage. Welcome to ThursdAI. Um, there there's a huge, huge amount of stuff to cover, so I think, uh, yeah, let's do a little bit of banter, folks. What is, what is the highlight?
Alex VolkovLet's start with LDJ. You were, you were here first. What is the highlight for you from this week, if you had to pick one?
LDJYeah, um, huh. I- I'm sure, I'm sure Nisten and Peter and others are going to probably mention the math part, so I'm going to say Reflection AI's Beam. Um, so it's pretty competitive in terms of... it- it's, it's definitely, I would say, second, third, fourth-ish place in open models when you look at raw benchmark scores, but it seems like it might even be just first place when you look at reasoning efficiency compared to the top models like GLM and others.
Alex VolkovMm.
LDJSo we'll look into that more later, and uh, yeah, that's, that might be my top one.
Alex VolkovPeter, how about you?
Peter GostevOh yeah, the math stuff was something else. I know.
Alex VolkovIt, it took over your timeline a lot, like I see you commenting on this. Um, tell us,
Peter GostevAnd it's-
Alex VolkovWhat's going on there?
Peter GostevYeah, so there's, there's been a kind of, I don't know, rumor, I guess they kind of pre-announced it that they have a bunch of maths problems solved. I don't know, was that a rumor or was that just announcement? I can't remember what was publicly known or not, but
LDJAt first a rumor.
Peter GostevYeah, yeah, okay. So, so OpenAI, uh, with their new Bell model, right, they solved the Navier-Stokes problem, the Millennium Prize problem, and then there was a lot of chatter about that they've, they're just sitting in a war chest of hundreds of problems solved, and, and now they released it, and they just did it pretty, in a pretty low-key way, not like a hype marketing video, uh, with, uh, exploding head emojis.
Alex VolkovIt was nothing. The blog post was just
Peter GostevYeah.
Alex VolkovJust a few lines of code, and the GitHub was just an insane amount of... Uh, you know my favorite, uh, term for this now is slap grenade? When somebody throws at you like an output or, or, or a, or a pull request of Claude and it has like a bunch of lines, this is now a slap grenade. So Open, OpenAI basically like slap grenade, uh, for people to just like shovel
Peter GostevYeah.
Alex Volkovand try to figure out what this is.
Peter GostevYeah, maybe my personal reflection is that I didn't really talk about this much, but I was, uh, with GPT-5.6 or GPT-6. I picked one problem, and I thought, let me just throw all of the tokens I can get my hands on and just see what I can do with it. And I was also trying with Fable and Opus and so on, but I was running the GPT models mostly just because were doing a lot of resets and things like that. And, was one problem that I was trying with the Hardwing and Nelson, not sure how you pronounce exactly, it was like, graph coloring kind of problem. So it's easy to grasp, obviously impossibly hard to solve. And the point is that I was spending literally hundreds of billions of tokens.
Peter GostevI'm not kidding. I probably spent, I don't know, 3, 400 billion. I know I need to go back and check. And I literally had my Linux box running, uh, for weeks and weeks, and it was filling up with all the solutions and so on. But the T- TLDIs, that was complete waste of time. I did not discover a single bloody thing. There was nothing, no use. was an interesting mathematical result, nothing, right? And then, these guys with the new greatest, model, they just, run it for 3 hours, said on average for 3 hours on ChatGPT Pro, which I think in practice what that means is that it's like a, some sub-agents, I think we don't know how many, but not like
Peter Gostevhundreds, like 10 maybe, or 4 or 10, something like that. And uh, it uh s- it uh didn't solve it completely, but it changed one of the bounds. The idea is that you can narrow the bounds, and I think before it was something like 5, 6, or 7, and it narrowed it down to 6 or 7, which is a big deal, right? It is a, a obviously a lot of context required, but it's a big deal, right? And I was speaking to Astra about m- my progress, my results, and Astra said, oh, we are two mathematical breakers away from what they did. So these guys did it in like 3 hours. I couldn't do it in months with Astra and
Peter Gostev5.6 and Fable and others. There's just nothing.
Alex VolkovThere's there's a few rumors that I want to bring to your attention there, uh, but I think I think it's time for, you know, for like the official start. Uh, folks who are tuning in, welcome to ThursdAI. Uh, one of the coolest things that we get to do is talk to you as well, so please drop in the comments, what was the highlight of this AI week to you? Obviously, it's nearly impossible to cover everything. we're trying to make ThursdAI the most agentic forward show. We have an AI producer. Uh, producer, please say hi in the banner below, uh, if when and and if and when you hear this. And we also have, obviously, a bunch of research happening throughout the the to make sure that we're not missing anything.
Alex VolkovUh, this week, I think I had 48 topics to bring you of, you know, of importance, et cetera, 48 or or 50 or something. I actually asked Claude to build me like a, l- l- like a Tinder for AI news where I swipe left and right if I think that this should make the show, and it will stack rank them based on, like, my preference and the stuff that we talk about, uh, because it's just like impossible to bring you everything, but we're gonna try to bring you the most important news, so, uh, once producer catches up, I think it's time for, for a cold open. Alrighty, folks, let's, let's do this. Folks, this week OpenAI dropped 722 math problems, papers, uh,
Alex Volkovwritten by a model nobody outside OpenAI has touched yet, and the mathematicians are split between thrilled and asking for receipts, and a lot of those came with receipts. Welcome to ThursdAI for October 8th. This is Alex Volkov. I'm an AI evangelist with CoreWeave, and with me, LDJ and Peter Gostef from Arena. This was an insane week. We say this often, but this was an insane fucking week. Uh, after OpenAI's math manuscripts, uh, 722 of them, to be, to be precise, uh, some people call this a slob grenade. Um, many mathematicians are split between this is the most exciting time ever, we're
Alex Volkovbreaking through the boundaries of mathematics, and, uh, we should... This is not how mathematics should happen. This is, this is kind of the split, and Peter, I think, has been monitoring this and would love to, to, to chat with you more. This would be our item number one, and that's just the opener, right? We'll start the TLDR, obviously, with OpenAI. 722 math manuscripts from a model that nobody open- outside of OpenAI can use. Uh, somebody used Fable to, um, t- to put this release in terms of how many Navier-Stokes moments in one drop we got, and uh, and I think it's around 5 or so orders of magnitude of
Alex VolkovNavier-Stokes, uh, releases. If you guys remember, Navier-Stokes, a Millennium Prize problem, was solved by OpenAI just recently. So OpenAI just dropped 722 of these, in- many including Lean, uh, confirmations. We're gonna talk more in depth about this. Uh, next up is, uh, Claude Haiku 5.5 is a 10 cents per million input tokens, and Anthropic say that it beats GPT-5, GPT-6 Luna on computer use and coding. GPT, remember, was really, really good at computer use. Haiku is- seems to be better with a 72% on OS World, and I think also significantly faster, I think, on artificial analysis.
Alex VolkovUh, Claude Haiku 5.5 is one of the fastest models right now, and it's incredibly cheap. Incredibly cheap. Uh, Haiku 4.5, I don't know if you noticed this, was a dollar per million tokens of input, and Haiku 5.5, which is significantly better, is 10 cents. Just j- j- just to give you a sense of how crazy the world we're living in. Uh, also a small thing that they released together with with Haiku, Sonnet cash reads, so the tokens, you know, after after you send them once, uh, Sonnet cash reads are now halved by price a week after Sonnet 5.5 was released.
Alex VolkovSo that's also quite insane from this week. Let's see what else. Uh, OpenAI rolls out GPT-6 with intelligent UI to everyone. All of the free users on ChatGPT, all of them are getting GPT-6. It's not Luna, it's not Terra is dead, so no Terra anymore. It's not, uh, it's not Soul, it's not Astra, it's just GPT-6. There's no name in this. Uh, so all the free users can d- can can get it today, but also the cool thing, the the intelligent UI is really, really cool. They rebuilt ChatGPT on top of it, and you basically get buttons and switches and toggles and and graphs, et cetera, in your chat output very fast.
Alex VolkovIt's really cool to see. Um, it's like little working apps for your answer. There's also some threads from folks at OpenAI who say, um, the model was now trained to do it while it thinks of an answer, right? So it starts to build you, hey, this answer would probably get benefit from these and these switches, these and these things, uh, and then it it still works in the background. So I think that's pretty cool. Um, Claude for Google Workspaces is finally here, and this is like a small thing, but it's a huge thing for many people. I'll give you an example. Our cloud producer every week, uh, used to post a Google
Alex VolkovDoc, and then a new piece of news would come, and I would ask it, hey, can you append to the Google Doc? And it would start opening computers, et cetera. It wasn't native, so now that it's native in Google Docs, I think it's a, it's a very big deal, uh, and it also asks me for every edit. Uh, but okay, let's go to the open frontier as LDJ mentioned, we now have another US lab. Well, we knew about Reflection a while, but, and Reflection AI announced Beam. It's a 501 billion parameter open model trained from scratch in the US with Apache 2 weights coming this month. And shout out to Reflection for this, uh, o-
Alex Volkovo- for this awesome, uh, announcement. Many, many people got very excited. An open MOE that beats on efficiency over Roscore, so it's more efficient versus just like Benchmax. Uh, we can't wait to play with this when the, the weights drop. Also, Mistral came back with leaning into le le le chaton fat. If you guys remember, this was the meme. Uh, they called this one le chonk. Mistral 4 is a trillion parameter with the weights coming at the end of October. This is now a pattern that I do want to cover on the show. Like, folks announce open, but they not release open.
Alex VolkovI- is a far, far way that we moved from Mistral announcing their, uh, latest model with just a torrent link, because people do, do be announcing. Uh, but Mistral large ties with GPT-6 Luna into official analysis, uh, at a much higher cost per task. So that's a very interesting, very interesting thing. Uh, this is, this is kind of the comparison we have here, uh, obviously for European folks. Speaking of European folks, uh, we're missing Wolfram this week because Wolfram is on a well-deserved vacation, but he got super excited about Calibri. This is from Aleph Alpha, another European lab. If you guys remember, there's a, there's a European lab called Aleph Alpha.
Alex VolkovUh, uh, Peter may, may, may know them as well. It's a German-English MoE with 78 billion parameters. Uh, Wolfram sent me his notes on this. It can run on a single H200 and also Apache 2. So shout out to Aleph Alpha for, uh, for coming back with the model. I thought they stopped doing pre-training and post-training. It just like went with models. And uh, for the very, very, um, let's say AI internals oriented folks, we used to cover embedding models. We we had Bo Wang on on the show before he joined Perplexity, and he dropped one of these models, so shout out to Bo. Uh, two open embedding models dropped in one week.
Alex VolkovGoogle's embedding Gemma drops embedding two text, images, and video and audio embeddings in one space that, uh, runs on the phone. Two embedding models this week, and Perplexity late interaction models search PDFs as images and are very fast with PPA- PPLX Embed V2 late interaction models. So a lot of, a lot of goodies this week. Uh, as we move to agents and assistants and a bunch of stuff, this, I think, is a very important thing. Uh, so Meta and Sierra, uh, and and Shopify and Stripe and Walmart all on board of this protocol for personal agent protocol.
Alex VolkovUh, very interesting to see if this is the new kind of MCP, because MCP had its ups and downs, but now MCP is everywhere. Basically, MCP, like, took over the world because of these agents they need to talk to, to a bunch of apps. Uh, I don't know if you guys caught this, we talked about Grokbot since launch. There's a whole drama happening between Dots creators, uh, uh, at OpenAI and and Grokbot, uh, f- folks, specifically Potato. Uh, it's really funny, but, uh, Grokbot has been supercharged this week. After Grok 4.7, the model released, it was muah muah muah. I ha- I think I have a button for this.
Alex VolkovThis is our reaction to Grok 4.7, not only our reaction, uh, big Ankh space, space Ankh Elon Musk decided, hey, you know what? Grok the product needs to be more than just Grok the model, so Grok is now tapping into Opus 5.5. If you guys remember last week, I think everywhere else, I've said I want personal intelligence harness, like a proactive personal intelligence, but with Opus's like brain. So supposedly now that's what's happening inside Grokbot. I think it's very cool. Uh, supposedly also only because I, I haven't seen this in the logs yet, so that didn't arrive to me yet, so we'll see. Uh, but also Iloma said it will use the best model for each task, including Claude
Alex VolkovOpus 5.5, Suno, and Midjourney, which folks got very surprised because Midjourney doesn't have an API. Uh, we'll talk about that. speaking of personal assistant, and also Grokbot, Shane Mack posted this, and I think it's a good warning for you to know, his Grokbot, talks to another Grokbot and, uh, his personal finance details in the company Slack, of which he's the CEO. Uh, he had one, uh, Grokbot in charge of his personal finance, giving them an update, uh, and another connected to Slack, and for some reason they communicated and decided, hey, let's put the dirty laundry of our CEO in Slack in front of everybody
Alex Volkovto see, including all the statuses and et cetera. So this was a very interesting warning sign for, uh, personal agents connected to between work and and and home, uh, and I think could be one of the reasons why they're switching to Opus, because Opus would never. Like, I just, I just know that Opus would never. this is a Grok thing. This is a Grok, like, intelligence thing to just like, oh, sure, let's, let's blast it, uh, publicly. Um, we gotta talk about Nous Research, folks, our favorite open source friends for the longest time. Nous Research, shout out to, uh, just everyone there, all our friends, Technium, Karan, um, I keep forgetting all the names, um,
Alex Volkovthey not only launched Hermes Index this week, which is a benchmark that says, hey, Claude Opus 5.5 is the best model for Hermes right now. Uh, they also announced a fundraise evaluating Nous Research at more than 1 billion dollars. So shout out for our friends to making the unicorn status, uh, fundraising, I think 90 million is the latest fund round. They also were on stage both at OpenAI Dev Day last week and at Microsoft event this week with a deep integration into new hardware. This is incredible. Since we told you about Hermes Agent, the length and the, the, the, the depth of penetration of Hermes Agent from Nous Research is just incredible.
Alex VolkovSo we, you know, we'll definitely cover the Hermes Index, but we'll definitely also celebrate the success of our friends of Nous Research. Uh, we've w- we've been covering Nous since before they were a company, uh, since before when they were just like a ragtag of folks doing stuff on Discord. Uh, so shout out to uh Karan, uh, Technium, um, Jeff and uh Dylan and everybody else at Nous for this amazing, amazing success. If you are a tinkerer and you like personal assistance on your hardware, Nat Friedman posted, uh, open source Muse gadgets.
Alex VolkovSo now you can put your Muse everywhere in your home. You can put it in every gadget, in every, it's an ESP32 board, if you know what that is. It's like a tiny chip, uh, that I used for a bunch of Halloween builds before that I talked about, uh, and it's, it's a very cool thing that they're posting. A bunch of folks are porting their assistants into portable stuff. A new personal assistant emerges from the creator of the Figure Bots, Brad Atcock. It's called Hark Pro, eh, also free, and the assistant that clicks for you. They are claiming the best computer use model. I haven't, uh, got a link to the, to the verification of the best computer use model,
Alex Volkovbut Hark Pro is also here, and uh, folks in comments, if you use Hark, please let us know. Uh, for our this week's buzz corner that probably should be renamed because it's no longer just about weights and biases, uh, CoreWeave opens up free GPU sandboxes. This was the highlight of the, the, the event last week for me. Um, last week on the show we had Dirk Filio, who is the PM for CoreWeave sandboxes, and he announced that, hey, uh, you can now reach out to us at CoreWeave. We'll give you GPUs without a salesperson in the middle. I think it's huge. It's absolutely huge. We're advanced.
Alex VolkovUh, free GPU sandboxes. I'm, I'm, I'm gonna say this again, super quick. If you scan this, you get free GPU power. Do Uh, if you want more, uh, GPU while in testing, this is just gonna be free, which is, I think, is incredible. Uh, so this is in this week's buzz, and we also have RL rollouts. All right, I'm gonna read this block super quick before we get to the news because it's been, it's been long. Black Forest Labs launched, uh, Flux 3, 4K images and layouts, uh, up to 10 reference images. Nano Banana 2.1 releases, uh, which is very, very cheap, only 3 cents per, uh, uh, 1K image. It's pretty good. Uh, if you are looking at our, um,
Alex Volkovinfographics right now, Nano Banana 2.1 is the creator of those infographics. They're pretty good. I'm very impressed. The consistency is there. Uh, you have to run it on high reasoning, medium reasoning. The default one is not that great at text. Uh, yeah, I don't know if you guys saw Tavis, uh, Griffin video avatar fooled 48% of the people in a live video call. They thought they're speaking to a human. This is quite honestly quite crazy. We're gonna play a clip of this, of our friend Quindla Kramer talking to Tavis avatar and saying that, you know, whatever Turing, uh, tests for avatars we've passed, long passed.
Alex VolkovUm, Reka AI, we've talked about this company, uh, one model, every modality, released a 19 billion parameter model that understands and generates text, understands and generates text, images, video, and robot actions all in one network. That's very interesting. We'll see if we have time to talk about that. Uh, and I think this is almost the, the, the last thing, but the, the rise of the JEVs, I think, is our last category. We're gonna talk about it with, uh, Liquid AI's head of pre-training, Maxime Labonne, a friend of the pod. Um, OpenAI launched Decisions API, uh, and Cloudflare dropped, uh, Clef, and Ansloth dropped a, a recipe for, for creating decision models, and
Alex VolkovFastino dropped Glide. There's a bunch of decision models all around, basically, so we're gonna cover all that jazz in the end of the show with Maxime Labonne. Please stay tuned. I think if you are interested in what decision models are and how to run them locally, I think that's going to be a great segment for you. And with that, I think, folks, it's time for us to actually go and dive in to Frontier to talk about math, because I think a huge thing happened this week. Let's go.
Alex VolkovOpenAI released 722 math manuscripts from a model they haven't released yet, and, uh, somebody took and used Fable to count up all of the releases into how many Navier-Stokes-like news we got, uh, and it's roughly 5. 5 moments in one drop. If you guys remember the Navier-Stokes and how much hype that did, we just got 5x in the amount of OpenAI solved maths, uh, which is quite crazy. On Tuesday night, OpenAI posted this onto GitHub without any fanfare, basically a very, very thin blog post. Um, after Naver's talks, a lot of mathematicians said that OpenAI is doing very bad
Alex Volkovthings by releasing it such as this. OpenAI said, hey, we're gonna create a mathematician advisory board to tell us how to handle the fact that we're basically brute forcing mathematics, advanced mathematics, mathematics that no humans have solved for 80 years. Uh, and and this is after that advisory, OpenAI grouped it into 372 families of results. Uh, the f- the internal frontier model about got about 4,000 open research problems, and uh, OpenAI says each result took an average of 3 hours on a ChatGPT Pro level thinking level. They also released 10 shorter reasoning summaries on things like irrationality exponent of the pi, the Mahler conjectures, and the Kaplansky conjecture.
Alex VolkovAnd they flagged 2 results that came out of a different process, a Riemann zeta 0 free region where humans edited the write-up. So they did have, it's not only slop grenades, they did have human mathematicians actually doing some stuff, and the Hodge conjecture results for a special family of cases. Uh, Peter, how big is this? How big is this GitHub drop that we're getting from OpenAI with 722 manuscripts for math?
Peter GostevI think it's enormous. Uh, it's uh, we really went from just models being completely hopeless at ba- basic maths to doing this. And the Millennium Prize was a big deal, and it was one thing, and obviously it's a, it's a was, you know, lots of agents being thrown at it, and um, it's a little bit controversial here and there, but I think I, I remember the time thinking when there was all of this controversy, people saying, oh, maybe it stole its insight from the mathematician who using ChatGPT, but reality is like, okay, we, who knows, right?
Peter GostevWho knows what happened? I don't think so, but whatever. But the reality is that it's complete cope for people to think that, oh well, that was a one-off, and actually they still can't do maths and and so on. And now we see it, right? It it's happened. There's, I don't know how much clearer signal you can get that this is real. And the question is that, you know, there's one scenario where, oh, they were just low-hanging fruits that sort of maybe humans were a little bit too lazy, too untidy to pick up, um, and that will be it. But I, I think just reading what mathematicians are saying, it looks like the bulk of them are impressed, and, um, there were a bunch
Peter Gostevof people who knew about certain problems, um, quite a lot. And I saw some of them say, you know what, I was looking at this problem for many years, and I never thought there would be an approach like that that would solve it. So I think it is real, it is not trivial, it's not just a finding a bunch of counterexamples. I know, Alex, you said people were get asking Fable, what did these results mean? I saw people also doing that in terms of the, uh, finding the, like, how many counterexamples the- there were.
Alex VolkovMm.
Peter GostevAnd there were only like 20% of these are counterexamples. So I think there- there was a scenario of just getting a bunch of uninteresting results, actually just kind of brute forced a bunch of stuff, okay. But I don't, I don't think that's true. I think this is real, real heavy math stuff, which I have no hope of understanding, but it sounds super impressive.
Alex VolkovHere's the thing that, that, that I have two comments before we go to LDJ. Would love to hear your thoughts. I know you, you have a bunch of friends in mathematicians. Yam would love to hear you. Last time when maybe your stocks was broke, you put the glasses on and it just like exploded. Uh, just for context, folks, this is uh GPT 5 years ago. Uh, somebody shared it, and I- I- I do have to add this for context. Uh, 5 years ago, AI was struggling with this type of grade school math. Cindy's math and science books weigh 2 pounds each. Her French books weigh 4 pound, and her English books weigh 3 pound. Her history books weighs twice as much as her English books.
Alex VolkovAI 5 years ago struggled with answering a basic, basic logic thing. And now, somebody ca- called it out in comments, it's not brute force in mathematics, but it's just solving things that the best mathematicians in the world couldn't solve for ages. And, uh, many of us would turn to an LLM to try and understand what is it that they solve, because the math there is so advanced that it's really hard to explain to you how advanced this is, uh, especially if we're not mathematicians. So even mathematicians in their fields, uh, there's what, 4, 300
Alex Volkovor so groups that they grouped those results in? Mathematicians in different fields of those groups, they, without AI assistance, not many of them can understand the other results in, in that result group. Like, the math that dropped is so advanced that many advanced mathematicians cannot very easily approach all of these problems and understand what is it that, that, that like ChatGPT gave, which I think is fascinating and scary at the same time. Uh, and LDJ, I would love to hear from you, your, your thoughts on, on the drop, your thoughts on how mathematicians react to this. Uh, should they be more excited or, or worried about their job?
LDJUm, I think specifically when it comes to their jobs, I think that's can, that can be really unpredictable, and it might even depend a lot on what area of mathematics you're in and a lot of different factors. Um, I did want to mention, though, in the example you showed that was 5 years ago, like, I think it's important to remember even just like 2 and a half years ago, when we had GPT-4o, which was the best model they had at the time, that was actually GPT-4o, 4o had just barely come out 2.5 years ago.
Alex VolkovYeah.
LDJThat could barely even do- when I would give it 3 by 3 multiplication problems, like just 3-digit number multiplied by a 3-digit number, a lot of times it would even just fail at that. And so just like in that span of time, we've gotten here, and when people talk about the brute brute force aspect, there's the reasoning efficiency where they did say, on average, these took about 3 hours of a ChatGPT Pro computer or, uh, equivalent to ChatGPT Pro. And
Alex VolkovInsane.
LDJI- I would- I have several friends in mathematics that would confidently say, at least for some of these problems, it definitely solved it using, like, less tokens and less time than a human would have to to do that same thing. And so arguably, in that sense, the human is doing even more brute force than the AI is is doing, right?
Alex VolkovYam, thoughts on what happened there? Do you have a favorite manuscript that dropped draft the that
Yam PelegYeah, they they, uh, they went after the Riemann hypothesis. That's insane on its own, just to try. Um, they, uh, the Riemann hypothesis basically, uh, speaks about the zeros of, uh, of a specific function and where they are. So, this one is hard, by the way. That's a very hard problem, uh, needless to say. Um, they, at the best of my knowledge, and I didn't have a lot of time to go into this, and just like you're saying, you need to be many years in the field just to wrap up your head about, uh, everything that is going into this.
Alex VolkovAbout one such problem out of
Yam PelegOh, yeah.
Alex Volkov722 problems that they dropped.
Yam PelegListen, people spend their entire life on each one of these, like the Riemann hypothesis is, is a, is a big one. They, even just, uh, founding a region where the zeros are, which is what I understand, uh, is in this paper, um, that's huge on its own. It, you don't give, you don't need to prove the actual zeros themselves, uh, uh, in a specific region, but you can bound them. And it's also huge because it was never done before, and many, many, many people have
Yam Pelegtried this. I just, just want to put things where they are.
Alex VolkovIn perspective?
Yam PelegIt's not-
Alex VolkovYou want to put things in perspective?
Yam PelegIt's not that, it's not that you can just go to ChatGPT Pro and tell it to solve the Riemann hypothesis, and that's, you're gonna get this result, okay? Uh, it's, it's important there. First, there is first one thing that is clearly need to be said. There is, uh, on the API, there is a parameter, uh, called juice that, you know, each of the ChatGPT, uh, categories of thinking, like high, extra high, correspond to a different level of juice, and pro is obviously, uh, the top. You can up the juice even more.
Yam PelegI suspect that it's nothing to take away from this. I'm just explaining that it's not that you just go to ChatGPT and, all right, I'll try ChatGPT Pro, and let's prove the Riemann hypothesis.
Alex VolkovYeah, this is also not
Yam PelegNot that easy, but it's still insanely impressive. About the brute force, I'm in the, I'm in the camp of, uh, yeah, let's brute force. What's the problem? What's the problem with that? Let's brute force. You can call it brute force, no problem. Let's brute force everything,
Alex VolkovLet's brute force, uh, you know, uh,
Yam PelegAnd, uh,
Alex Volkovroom temperature
Yam Pelegabsolutely.
Alex Volkovsuperconductors. Let's that, that's the brute force. I wanna highlight this one from Ksenia Ksenia Sip from Touring Post, a friend of the show as well. She's like, I'm not a mathematician, so I wanted to understand what yesterday results mean from the math wizards, and she saw this post from Kevin Buzzard that talks about understanding mathematics, and I also do want to talk about reactions as well. Peter, you posted some stuff about, uh, the, Institute for Advanced Study from, from Terry Tao? Would love for you to bring that up as well. Uh, I just want to read through this, uh, a person that says he's a PhD student of Richard Taylor at early 90s, uh, and said that ma- whilst mathematicians are now
Alex Volkovbeginning to agree that mathematics is all about human understanding, he's not entirely convinced they will agree on what it means to understand high level mathematics. Basically, he says that maths, specifically like advanced math like this, is not about just solving the thing. It's about the human understanding boundary of solving the thing, which is a very interesting take on advanced mathematics, is because, okay, let's say I solved math. How does it help the world, specifically, uh, in one, one, one theorem, uh, if humans don't really, like, understand it and and can't do anything with it?
Alex VolkovUm, it's a very interesting take about some math- mathematicians. Uh, some other folks are not having a good time. Obviously, nearly existential crisis, uh, thought process for many folks who've been working on a specific problem for decades, just to see 3 hours of compute just like decimate this. It's a very interesting, um, a- analogy to software engineers who, after a year or so, not only don't write code anymore, don't even read code, right? This was, uh, 6 months ago, maybe, uh, the discussion, the the ZL continuum that I brought, uh, to to the world, like whether or not you're reading the code that your
Alex Volkovagents output or not, most of people now don't even read that. Uh, so this is a very good analogy between mathematicians and their work. Although, when I wrote a piece of code back in the 2000s, I didn't consider this my life's work, like many maticians do, many mathematicians do with just like one problem. LDJ, go ahead, and then, uh, Peter would love to hear from you.
LDJYeah, so there's a few interesting, um, kind of open problem sets and self problem sets people have put together using like Fable and Astra of, hey, like, what are the most important problems? And, um, one of the main examples that has been circulating lately is out of, out of a compiled set of 500 most important open problems put together by Fable and Astra, 92 of those were dropped Tuesday. And
Alex VolkovSay this again slowly.
LDJsolutions by OpenAI.
Alex VolkovYeah, so.
LDJYeah, so so Fable and Astra had put together a set of 500 most important open math problems in the world, and 92 of those just got solved on Tuesday by OpenAI.
Alex VolkovThat's insane.
LDJAnd out of the top 100 in that list, 10 of them got solved on Tuesday. And one last tidbit here. There is a list put together by Fable and Astra of the, the top 100 most significant problems in the most, in, in the last 12 months, and 8, over 80% of those just got dropped Tuesday, for solutions, rather, for them.
Alex VolkovIt's absolutely insane, folks. It's absolutely insane.
LDJSo, one last little- One last little thing I'm gonna say to, to make people think about something. If you look at the list of manuscripts that OpenAI released, they're, they're ordered, and there is some numbers missing. Like, it goes from, like, 44 to 46, and 45 is missing, and, and there's, there's about 4 or 5 of these. And
Yam PelegMhm.
LDJa- a lot of people are suspecting also that if you just look at all the different categories of problems they released, it looks like, uh, things relating to cryptography seem to be really missing here.
Alex VolkovMhm.
LDJUh, so it- the- the theory here is maybe they had some significant advancements in cr- cryptography involved here that potentially, uh, uh, the- the US government or various entities or maybe their own just personal policy ended up restricting them from releasing, because it might sound conspiratorial, but there is actually laws and policies that the US government is actually allowed to prevent mathematicians from publishing things relating to cryptography for national security reasons.
Alex VolkovI not- yes, not only national security, also the world economy crashing when Bitcoin goes to zero, if if and when the super intelligent decides, you know, this is, this was the scare from the quantum thing, right? Like quantum cryptography, uh, and, and, and solving quantum, like, stuff will essentially make Bitcoin not as secure as everybody thought, uh, because it relies on a lot of, like, elliptic curve cryptography, and that's all math. Uh, and yeah, there is, there is a way for the, the world government to say, hey, do not release breakthroughs, do not, like, do not release, uh, you know, diseases into
Alex Volkovthe world. There's also do not release breakthroughs, uh, like that. So,
Yam PelegAstra gonna be rich, man. Astra's gonna, Astra's gonna be rich.
Alex VolkovIt's a, it's a good question if Astra can make money, uh, out of, out of this. Uh, Peter, I do wanna, uh, bring the, the stuff that you posted. Would love to hear from you about the advisory group and the, the, the Terry Tao stuff, field medalists, uh, responses to this because, you know, folks from all over the sides of the, of the debate are, are chiming in here, and, uh, you posted about this. Would love to hear your thoughts.
Peter GostevYou, you want some controversy?
Alex VolkovI would love some hot takes, yes.
Peter GostevSo, uh, I think, uh, it's a, it's a touchy thing because I don't wanna, that kind of maybe came across as kind of jumping on people or something. I hope, I hope it didn't, but um, it it's a tricky situation, right, because a lot of mathematicians working hard, they kind of have the established ways of working, and then suddenly someone comes in, just drops a bunch of stuff, and just your whole world at least changes in some way, right? It has to change in some way, uh, whether directly or indirectly. So there was a letter from Association for Human Mathematicians, um, which has like, uh, 800 members, and Teritire reposted it in his blog post, and it said something
Peter Gostevlike that, uh, I think there were some, some quotes there, um, where they said something like, mathematicians did not ask for this work to be done. Um, mathematicians have a particular vision of progress that is informed by history and field-specific considerations. We reject OpenAI's assertion that this release advances our subject, uh, and so on. And I think that's a kind of interesting, uh, perspective, kind of, which, which goes back to the question of what is maths for? And I think it looks like a lot of, uh, mathematicians, at least part of this
Peter Gostevgroup, they seem to be kind of suggesting, well, it is, you know, we need to run our thing, do it in a certain way, but inherently what that implies, it's not to actually solve math problems.
Alex VolkovMhm.
Peter GostevAnd, uh, to me, that raises an interesting question of like, so are we basically saying that math is useless and we should just leave it to smart people to j- to just work on it? If that's tr- if that's the case, then I can see the intellectual argument for this. But the, if you flip it another way and say, well, imagine doctors said this, you know, stop, stop curing diseases, right? We just need to, like, don't, don't ruin our good thing. We just need to keep going, you know, treating people from their headaches and whatnot and cancers, and, you know, the we've got a good thing going. Like, just imagine that happening. But here's a clear line for utility that I can't
Peter Gostevimagine any doctor would say that, right? So if OpenAI or whoever dropped a bunch of papers saying this is how you
Alex VolkovSolve.
Peter GostevUh, you know, yeah, solve medicine and kill 500 diseases.
Alex VolkovYeah.
Peter GostevYeah, people would be like, yeah, amazing. And um, I think, to me, this is, I appreciate, well, I'm not sure I can fully appreciate, but I can imagine this is very difficult for people in that field, uh, which I think we need to be sensitive to. But it is kind of, but to me the question is like, well, are you saying that your field is useless? Right? Is there no utility and we should just let you hang out and work on it and not solve any problems? And I don't know, I don't think you can kind of say both, right? You can't say, oh, it's really useful, but please don't solve any of it.
Alex VolkovYeah, the please don't solve crowd, I don't understand, like, get on board. Like, thi- this is now in the realms of OpenAI, and in a year it's in the hands of pretty much everyone, right? This is the pace we've been on. It's very clear that this is the pace we've been on. LDJ, one last comment. We have to move on. We dedicated a lot of time to this math thing, and we have to move on.
LDJYeah, I recall about a, a year or so ago, there was actually something circulating where there was a discussion at a big conference where somebody stood up and said something along the lines of, "Do we really w- want to automate a cure for cancer?" or something, and that they were actually, like, having, like, a serious statement and discussion about, like, their own job and livelihood and whether they would really want that. And there's a lot of polarizing reactions there, but I find it really interesting how at first it seems like the answer should be obvious, but there is seriously people that have strongly differing views there.
Alex VolkovThe I'd rather AI not solve cancer or something, I'd rather risk cancer than AI taking over essay was just like a ridiculous thing, a point in time, I think. W- w- hopefully we'll never get back to that moment. Uh, folks, it's time for us to move on because as exciting as this is, we can talk about this in circles for hours, but there's a lot of stuff that happened this week, uh, nearly an hour into the show, which is crazy. Uh, let's run and talk about, the frontier for everyone. There's 2 major things that happened this week. Let's start with Anthropic launches Claude Haiku 5.5. Uh, it's been a year since Anthropic launched the previous Haiku, and
Alex Volkovthis is one hell of a model. Let's take a look. Haiku 5.5, tiny price with very big, big score. So we're looking at a model cheaper than ChatGPT Luna, although it can get a little bit more expensive depending on the reasoning effort and the number of tokens that you send it. Haiku is just priced at 10 cents per million input tokens for under 100K. Um, this is 10 times as cheaper as last year's Haiku, which was at 1 dollar. Uh, I love Haiku. I used to love Haiku, not for like day work, but there's a lot of stuff that, you want that the good Claude intelligence, but, but like in really fast, really like quick price. Uh, would love to hear from folks here, uh,
Alex Volkovthoughts on, on Haiku release. Peter, did it already hit the arena? Uh, Nisten Yam, LDJ, did you guys use it?
Peter GostevYeah, it's on, it's on the arena. I'm just testing myself. We don't have a score yet, so, but I think it, it's super exciting because I, I think the, the, the one problem with Anthropic offerings is they didn't really have a lot of good cheaper models. They had the best very expensive models.
Alex VolkovYeah.
Peter GostevBut then you, you by default had to go somewhere else if you wanted to actually have something cheap, and that's not ideal. And I think it's interesting to see what, what they managed to do with this. One maybe not to, to call out is something I'm digging into as well, is that the, the cash read is so important, the, the price, uh, for it. And um, and uh-
Alex VolkovAnd for Haiku, it's 1 cent per million tokens.
Peter GostevYeah, but it's, it's also the ratio of kind of how much is it to, to read, uh, normally what the input and the your cash read is. And I think the Sonotropic has been dropping prices. Actually, I, I was just literally pulling the data up, and Fable 5.1 has the lowest ratio. It's like 2.5%, the cash read versus the, the read. And I think Haiku is still at 10%, but it's certainly been improving, right? If you, if you pick like GPT 4.1, I know that's long time ago, it was 25%, the cache read. So it's a, it's a really important part.
Peter GostevUh, I want to do a bit more analysis of it, but I'm glad that, that they're dropping prices on cache reads. That's for agentic work is the biggest thing.
Alex VolkovUm, s- speaking of cache reads, another hidden announcement in the Haiku announcement that Sonnet's cash reads are cut in half in prices, which makes, uh, Sonnet is d- most agent work about 20% cheaper on average, because if you think about what an agent does, there's a lot of, mm, uh, long-term history that happens in an agent. Next time you send a message to the agent, a lot of, a lot of the previous stuff should be cashed because nothing changes there. Uh, Haiku is also significantly better at computer use. Uh, LDJ would love to hear from you about this one. We like OS World. In OS World, Anthropic says 72% against 48 in Luna,
Alex Volkovjust 48 in Luna. In Terminal Bench 4, Anthropic's saying Haiku 5.5 gets 40%, nearly 40%, 39, while Luna scores only 16. So, compared to the very good, fast, and cheap model from the frontier from just a week ago, just a week ago. I, uh, a small sidestep. I love that we're a weekly show because so much happens in the span of a week. We have time to catch up. Uh, just a week ago, Luna, just like the best fast, cheap, and reasoning model is only 16% on on terminal bench. Um, Haiku is available everywhere with one. cent per million tokens cashed, which is just absolutely insane.
Alex VolkovUm, anything else remains to say about Haiku besides go use it, folks, and tell us? Anybody used it? Yam?
LDJYeah.
Yam PelegRight.
Alex VolkovLDJ?
Yam PelegGo, go.
LDJYeah, so at the Pareto Frontier, it looks like this is setting in, in some ways the, uh, a new Pareto frontier for in terms of accuracy and cost. On artificial analysis index on extra high, it looks like at the Pareto frontier or near it and competing with a lot of the Chinese models and others. But I, I do think going forward, and this is a bit of a bold prediction from me, like over the next 6 to 12 months, I think we'll see the Western frontier labs continuously getting better on the Pareto frontier and beating Chinese labs, even on the, the most low cost, uh, options.
Alex VolkovTh- this is a very interesting place for open source to be in, because a lot of the Chinese labs, specifically DeepSeek, uh, on their services promised us very fast and very cheap, not necessarily very advanced. The, the Frontier Labs are going to it, and then I, I remember Sam Altman Dev Day, uh, saying we're gonna offer the best, like, price model for every price and every tier of intelligence. Uh, it looks like Luna has a lot of catching up to do. Yam, last comment about, uh, Haiku 5.5 and what would
Yam PelegAh, that's
Alex Volkovuse it for?
Yam PelegThat's the obvious choice for swarms at this point. Like, you want, you wanna, you wanna do a lot of things in parallel. It was the Lu- Luna was the obvious choice yesterday, and and not not yesterday, but but you get the point.
Alex VolkovYeah, yeah, yeah.
Yam PelegThat's, that's the obvious choice. I mean, that you don't get this type of performance for this, for this price on any other model at the point. Anthropic is is on fire recently, like all the releases are, are, are top. Seriously, we didn't, we didn't speak a lot about Sonnet, but Sonnet is good. I don't know if you guys tried it, but Sonnet is surprisingly good also. And everyone online is just, just confused. How come I never burn token, I never burn my quota with Opus, and I,
Yam PelegI've been hammering Opus all day long. It's they all over X, everyone is saying this. So, like, imagine how much you have now for Haiku if you just want to do many things in parallel. Anthropic is on fire, and we are the ones getting stuff because of it. And, uh, I think, I think that's great because,
Alex VolkovThat's absolutely great.
Yam PelegIt was, it wasn't, it wasn't the case a couple of months ago.
Alex VolkovAnthropic is also coming up on the IPO, and I think it's, it's, it's good for them to drop a bunch of stuff. Uh, let's finish up with Entropic Claude in Google Docs and Google Sheets and s- and Slides. They finally added this ability to be able to actually work on your documents, not just like post via the, the Drive integration, so that's great. Uh, but speaking of cheap intelligence for everyone, OpenAI also came back, and uh, GPT-6 is now the default. We- the default intelligence, not GPT-6 with a name, not Sol, not Terra, not not Luna, not Astra, uh, not any of them. Terra is dead. RIP Terra. But just GPT-6 is now the default for free accounts.
Alex VolkovThis is, uh, not talked about this a lot, but out of the 1.2 billion weekly active users, I think they said last week in Dev Day. Peter, please correct me if I'm wrong. Uh, one, I think it was like 1.2 billion weekly active. Much of that is free accounts, right? Like, folks who are listening to us, probably like when you were in a traffic, you maybe cancelled the OpenAI one. Many of this is on the free tier, the free accounts, and uh, like, you could say that OpenAI just like uplifted the world's intelligence just now with just like one, one release with GPT-6, uh, which is incredible. So if you are listening to this and you are on the free tier of ChatGPT,
Alex Volkovuh, your intelligence has been upgraded. Not only that, they released this thing called, uh, intelligent UI. And I want to show you, can we show this on the stage? This is me asking ChatGPT to say, hey, visualize the math problems OpenAI dropped on GitHub with a nice UI. And it built me inside the editor a mini, a mini app. You guys see this?
Yam PelegAbout time, about time.
Alex VolkovRight?
Yam PelegGreat. Seriously.
Alex VolkovYou it says 719 manuscripts. You guys know why? Because it's updated since I did research. OpenAI did pull some manuscripts back, so this is like the, the real time thing. Uh, and then how much is formally verified? They're saying 300 of the manuscripts are formally verified, but I'm not gonna, not, not gonna talk about the manuscripts. You can see that this is kind of like a live thing that works, and the, and this can be your weekly update from ChatGPT. This can be a research thing. This can be like a bunch of stuff. They have maps in there. So this intelligent UI is, I think, is pretty cool. It can build forms and buttons and small working tools, and uh, yeah, I think
Alex Volkovthat's great from uh OpenAI, uh, but the intelligence boost for the free folks, I think, is very, very interesting. Any anybody already used it? Anybody feels the difference? I'm not on free tier, so I don't know. I've been using
Yam PelegYeah.
Peter GostevI think-
Yam PelegI thi- I think we're not.
Peter GostevI think what what's cool about this, and I appreciate that OpenAI is still doing this, is that we we kind of had versions of this in uh, Codex for a while, so where it kind of builds these little visualizations and apps, and it's nice that they keep bringing that to the free tier and and so on. So yeah, GPT-6 uh, is cool as well. Yeah, we- I think we we forget sometimes that we are very, anyone who's listening to this, you are very not normal person. Uh, so you are
Alex VolkovSaid with love.
Peter GostevVery far away.
Alex VolkovSaid with love. Everybody who's listening to this is a, is a beautiful person, part of our community, but normality is not, you know, like Your hairdresser is probably not listening to the show.
Peter GostevYeah. So, so yeah, Alex, what you said about, yeah, it's uplifting the overall intelligence. That kind of thing has so much more impact than whatever it is, like new model is dropping on the top tier where you need to pay 500 dollars a month. Like, that kind of thing makes a difference to us, but not that many people. So yeah, it's cool to see that, you know, they still remember. It's a, it's important.
Alex VolkovLet's also like check on Nisten a little bit. Nisten, you, you with us? You alive? Like, what's going on? You haven't said a word about Haiku or ChatGPT. I'm hoping that you're gonna tap in when open source comes.
Nisten TahirajI'm running 10, 10 Haiku agents right now. That was because I almost used up, like, I used up 95 or 97% of the, of the quota already until, uh, on the max 20 plan.
Alex VolkovWow.
Nisten TahirajBut, uh, so I had to basically resort to Haiku. It's very good.
Yam PelegYeah, it is.
Nisten TahirajYeah, yeah, I had a rice cooker, and, uh, I couldn't get the ratios right, and then it did a whole bunch of research, and it found that this is not very accurate as a
Alex VolkovBro.
Nisten Tahirajas
Alex VolkovStop reversing the world for rice.
Nisten TahirajSet an hour
Alex Volkov1 to 1 and a half.
Nisten Tahirajand, uh, it got it done perfectly.
Alex VolkovMy grandma knows
Nisten TahirajHaiku
Alex Volkovto 1 and a half. What do you mean?
Nisten TahirajYes.
Alex VolkovHaiku cooked.
Nisten TahirajYeah, it got the right
Yam PelegYou're letting Haiku make you food, like.
Nisten TahirajYeah, I have resorted to that, guys. He did a good job.
Alex VolkovYeah. LDJ, comment, and then we'll move to to open source.
LDJSo I want to ask real quick, Nisten, uh, would you say that it is clearly at parity or or clearly above Qwen 3.8 27B?
Nisten TahirajOh, oh, oh, for Haiku? Yeah, yeah, yeah.
LDJHa- Haiku, yeah.
Yam PelegGood question.
Nisten TahirajIt's, it's above, it's, it's above all of that. Uh, especially for conversational stuff. I haven't tested it technically, but anything conversational, yeah, yeah, it's a, it's really above. It kind of feels like Sonnet, actually, when you talk to it. I haven't looked at, like, what code it writes, but it's, it's, yeah, it's too good. Uh, it is what it is. That's, that's the assessment.
LDJSick.
Alex VolkovAll right. folks, use Haiku. We, like, the whole family is great. The all of the three brothers, like Haiku, Sonnet, and and Opus 5.5, all great. Uh, OpenAI obviously tries to release responses of theirs, um, so we'll we'll probably see more models, but let's talk about the next thing on our roadmap, which is let's do open source. Open source AI, let's get it started.
Alex VolkovUh, right, open weights and open source AI is our new category. Let's talk about this. Open weights went big this week, a 501 billion parameter American model, uh, a trillion parameter French one, and a German one you can run on a single H200 open source. Very interesting this week. Reflection AI announced Beam, uh, 500- 501 billion parameter model, 23 billion parameters active. It's trained from scratch, uh, on data here in, in the West, and uh, Apache 2 weights are promised this month. Uh, Reflection AI has been, uh, training, I believe, on the SpaceX cluster.
Alex VolkovSomebody correct me if I'm wrong, but I believe that this is, uh, the the cool thing about them. They're pitching it as the Western Open Open Frontier model. They claim is 80.9% on Swbench verified, which is a hard coding, uh, benchmark. Uh, and to their credit, they admit that Kimi K3 is ahead on just raw capability and bench, uh, scores. However, they pit- they pitch efficiency, 3 to 4 times less inference compute than, um, GLM from ZAI. Uh, weights are not downloadable yet, but this is announcement, but we will definitely let you know when it comes.
Alex VolkovUh, folks, my timeline lit up with Reflection, my timeline lit up with Beam. Uh, thoughts on how they approach this release, thoughts on what we're due to see. LDJ, I see you have your hand raised. Please go ahead. In
LDJter- yeah, in terms of the comparison to Kimmy K3, it is about 6 times smaller, so it it it'll be a lot more practical to fit on more real world, consumer-ish, prosumer-ish VRAM budgets. And in terms of overall the, like, even just besides the amount of active parameters and total parameters, you have how many tokens does it actually output to get to a certain answer. And, you know, if you, if you use half the tokens, that's, that's nearly half the compute.
Alex VolkovYep.
LDJAnd artificial analysis had done some numbers on this too, which I'll send, uh, this in the link here, uh, but it's looking really good in their preliminary results so far, and they said they're going to release some more detailed analysis on it soon.
Alex VolkovYep. Um, I, I, this is a graph that we're showing from, uh, Reflection folks, and they're highlighting where Beam sits, uh, on scores, on benchmark scores, which we all know is not everything. Like, the, the, the data that you put in there, the type of, uh, actions they trained for, um, whether or not it's agentic, all matter for these big models, but they, the, the very interesting thing is that, uh, and I haven't, I don't remember seeing this before. I think Artificial Analysis has this. They have two different colors of scores, and they segment this based on Western open models and Chinese open models. And I, you know, I think that's a good position for them because this is definitely
Alex VolkovMistral's position. We're gonna talk about Mistral in a second. Uh, and they are highlighting that they are the Frontier open weights on the Western Open models, and on Sweet Bench Verified, they're just taking the Frontier. By the time they released it, I think this will change multiple times, but, uh, uh, this is a very interesting thing. Yeah, this is a statement from artificial analysis. Uh, they have been given access by Reflection and independently benchmarking Beam. Early indicators suggest Beam will be one of the most token efficient open models we've seen for its level of intelligence. So shout out to Artificial Analysis, and we can't wait for scores.
Alex VolkovShout out to Beam. There's not a lot to say here besides the folks at Reflection, releasing this, and, I can already, hint that it's also coming to, some other inference providers, let's say, that help make the show, what it is. Uh, so we're gonna talk about that, but before this, Peter, you wanna, you wanna do the, the honors of the breaking news? I think, I think worth it. AI breaking news coming at you only on ThursdAI. Ah, it's, it's always fun to see breaking news from folks who are participant on the
Alex Volkovshow. Peter Gustaf, go ahead.
Peter GostevYeah, so we, we have some news, uh, we raised 200 million. That's a lot of money, uh, at 3.1 billion valuation. And so, yeah. It's a, yeah, big, big day.
Yam PelegYeah.
Alex VolkovBig day.
Yam PelegLet's go, man.
Alex VolkovCongratulations.
Yam PelegCongrats.
Alex VolkovThis is a, a big thing, and I think you guys are doing like a, a great thing for the world. Um, do you want to tell us about kind of what is Arena cooking? If people haven't visited Arena AI for the longest time and maybe visited once before LM Arena, what changed since then? What are you guys like, like, why does this justify this valuation, this money?
Peter GostevYeah, I would think, yeah, the, the biggest change, and I think we're, we're best known for this kind of battle mode. When you put one prompt, you get two generations, which is, uh, like that, that's still important for, for some things. It still measures quite a bit, but things have moved on, and we've got, we're much more focused on agentic evaluation. So we've got Agent Arena, and what that means is that you can put in any prompt you want, uh, use whatever tools you want, and uh, then we'll measure. You don't get two responses, you get one, but then based on your feedback to the agent or a how you're interacting with the agent, we can pick up
Peter Gostevenough signal to differentiate between the models. It's really cool. Honestly, I know I'm biased, but I, I can't think of any single, um, leaderboard that is actually better than ours. Um, maybe, you know, if you put together a bunch of different ones, maybe, maybe you can argue, but if you're gonna pick one, I think, I think ours is, is very strong. We, we also launched an alignment index where we're using the same agent, um, um, evaluations to be able to also pick up things like, did it actually do what you, um, or did it, uh, take some unauthorized action, for example?
Peter GostevSo, and then, uh, OpenAI is particularly good at not doing that. Um, and you can see some specific s- specific signals if you scroll a bit to the right. And, and, uh, yeah, so we launched that alongside of it. But it's all powered by the fact that people can come and use the models, uh, the agent mode, and then we can pick up signals from that. So, yeah, I th- I would say, uh, I know the the battle mode is very cool and people like it and find it useful, but I would say the agent mode is where we get super rich data and we we can extract a lot of information from it.
Alex VolkovThat's very cool. Congratulations. This was not planned, folks. We just came through, and uh, we want to say congrats to uh, Peter and the whole Arena team as well as we go back. Uh, congrats, uh, folks at Arena. Folks, we're back at the open frontier, and uh, let's talk about Mistral. Mistral launches Lechonk and uh, leaning into the meme. So let's, let's talk about this. Uh, the French, the the French AI company comes back with Mistral Large 4 Lechonk, a 1 trillion parameter model that uses the the ties to GPT-6 Luna on artificial.
Alex VolkovUh, it's a very interesting release from Mistral. Anybody want- wants to take this one? Like, I know we've been celebrating them in open source for a long time. Uh, also the weights didn't come yet, right? Like, uh, this is just also an announcement. Um, 50, uh, 1 trillion, uh, parameters all around, 50 billion, uh, active. It's a multi-model, uh, with 1 million context and open weights. It didn't launch yet. Uh, artificial analysis gives it a 38, the same score as GPT-6 Luna. Uh, this is $1 per thing or 13, $1.13 per task versus 7 cents for Luna. So this is not a very, um, optimized model
Alex Volkovcompared to the frontier, but it should be open weights. So, folks, comments on this, comments on Mistral Lechonk, besides the the very good faith that Mistral has, uh, you know, in our community, and besides the fact that, uh, you know, many European governments can only use this model, let's say, because of the deals they did. What are we thinking about this model?
Peter GostevSo I I have, uh, ta- it. We've got a score on, uh, Code Arena. So that's the still the battle mode, but that's mostly test front end. Uh, it's not been great. I think it's like 40 second or something like that. So I think, I think it was like DeepSeek Flash, which is way cheaper. Uh, I think it's like 20th or s- uh, along those lines from. So yeah, it's not great on that side. I have also tested just personally on some web dev agentic stuff and a bunch of different tests. It's not, I mean, it's not that great. So I don't know what to say.
Peter GostevIt's possible. I haven't tested, you know, in French, so it's possible that, uh, we do actually have some data. It's not that many prompts yet, but we do have some data to indicate that it is way, way better in French and the European languages than other models. So it's kind of punches above its weight in, in those categories. So it's not that cheap. I know they're running some promotion now. Uh, I don't know if that will stay. So I, it's kind of, you know, it's good to have models come out. Obviously, it's, it's always good, you know, the open models as well. Uh, but yeah, it's not, I don't see it like a big reason to use it
Peter Gostevover if you're using DeepSeek or something. I'm not sure you, you should be switching to that straight away.
Alex VolkovNo, and the very interesting thing is that all my comparisons, and I'll show them here again on the stage, uh, that I also got from Artificial Analysis. Uh, they're comparing them to, uh, specifically to Luna. We just told you about Haiku that came out a few days ago that just destroys Luna on every parameter, especially cost. So there should be a reason to use these models, and I think, uh, for many folks, the open nature of it is one such reason. But since this is a trillion token, it's not like you can put this in your DGX Spark. Uh, so maybe if you are, um, a European government or
Alex Volkovsomeone with ties to Mistral and, uh, this is the model you want to use, uh, maybe, but for regular folks, I think, uh, not so much. I haven't seen many folks use Mistral. Nisten, am I wrong here? And if-
Peter GostevI'm trying it right now. It's not that bad. Uh, I mean, the, the website is nice. It responds pretty fast. You can talk to it.
Alex VolkovYes.
Peter GostevLike, it's as a product, it's not that bad. Would I do any agentic coding with it? No.
Alex VolkovNo.
Peter GostevBut, uh... as far as a chat interface that you can just open for free is, uh, feels pretty good, so I...
LDJSo I think when I'm trying to look at the, the cup half full here, what might be really good, it, it might be a pretty good base model or, or a foundation that people could then use to do a lot of RL on top of. And, uh, that might be something, and it might be superior in some ways to the base models that exist right now for the certain combination of total and active parameters that it has.
Alex VolkovYep. Uh, all right, uh, let's see what else. we have, another European, company, this time from Germany, also releasing something, and it's been a while since we mentioned Aleph Alpha on the show, but yeah, there is another European lab. It's not only Mistral. Uh, Aleph Alpha released, Germany's Open weights hummingbird, 78 billion parameters, a 3 and a half active Apache 2 license, which is great, a million, million, context there. Uh, it trains from scratch in German and English. Our German co-host is not here with us, but he did send, uh, his assistant Amy, uh, to to tell us about Calibri. Uh, here's what Wolfram has said.
Alex Volkovuh, because it's German, our German tester is not here with us. He says, it's German and tool corings were good, uh, but reasoning and instruction following were too unreliable for a user-facing agent. Repeatedly lost roles and state, while more reasoning mostly made it slower. My verdict, promising a specialized German tool worker, but not a strong enough for Amy to main my daily model. And, uh, if you guys want more details, Wolfram posted about this, on his X. Uh, I will will defer to the German testing guy, uh, that's not on our co-host panel this week, but usually is, uh, to tell you about this Colibri model. Um, that's pretty much it.
Alex Volkov20 trillion tokens in pre-training. If you are using German, that could be good for you, but not as a main model.
Peter GostevCan I, can I mention one thing
Alex VolkovYeah.
Peter Gostevabout, about the open models? So one, one interesting thing, we got a bit of data about the number of GPUs that were being used.
Alex VolkovMm-hmm.
Peter GostevSo I think that the one that we just covered, the German one, I think they said something like it was, uh, 6- 768 or something, so call it 800 GPUs. I think it's Blackwells, uh, so I'm, I'm slightly speaking from memory, but I think they said, call it 800 Blackwells. Then, uh, Mistral said 3,800, uh, Grace Blackwells, and so I'm assuming it's B200, then GB200 probably. And then what, um, we had a bit of information from Jensen about Astra, where he said it was trained on, uh, uh, 100,000 Grace Blackwells.
Peter GostevSo that's, that's the kind of one, one thing that is, I, I don't really know how to think about it. On the one hand, they could be like, oh, well done to that team for training on so few GPUs and actually like producing a coherent model. On the other hand, it's just, I mean, I don't know what, how many GPUs is Luna or Haiku are getting, but probably not 800. Uh, it's probably in the tens of thousands. Like, I just don't know how to think about it. Like, is it just like no hope for these, uh, for these guys to train on like a few hundred, few thousand?
Alex VolkovUh, dude, this is why Jensen is showing up everywhere, showing up with, with, you know, with Satya Nadella at Microsoft, we're showing up with, you know, with Elon, this is everywhere. Uh, hope is, is very interesting. The thing that I will say, though, is, um, efficiency improves significantly. Uh, you know, B200s are about to get, um, uh, replaced with, with the Vera Rubens, uh, and so we're gonna get even more AI. That's basically my reaction to you, right? But yeah, I don't know if there's hope for a lab that cannot get the, the number of,
Alex Volkovlike, the crazy number of GPUs. Uh, but Peter, since you took us there and since we're talking about GPUs, I think it's time for this week's buzz real quick. Folks, this is a corner where I talk about everything Weights and Biases and CoreWeave related. Weights and Biases is now part of CoreWeave Forge. Uh, I will still play our transition, and then I'll tell you super quick about the GPU stuff. Weights and Biases and CoreWeave. This is very important. I need to update this transition, folks, uh, in this week's buzz.
Alex VolkovUh, since Peter just mentioned GPUs, here's what I want you to, remember from last week. Last week we had the CoreWeave Fully Connected, our Prime conference. We obviously did a live stream from there. It was great, uh, and uh, a huge announcement from there was from folks at Cognition. Uh, I, let me see if I can just play this for you real quick, because I think it's, it's worth playing here on the stage. Here's the announcement from Silas, one of the co-founders of Devin. I'm excited to announce that Cognition becomes the first customer for NVIDIA Vera Rubin, powered by CoreWeave Cloud. I think you can hear me in the background there yelling yay.
Alex VolkovIt's been truly great privilege to be, get early access to this latest gen platform, and so our research team has been putting it to the test, and in the last few weeks we've been comparing Aerobin compared to the previous generation GB200. We first started with production inference on our latest model Z2, and at the same speed we're seeing a 4.8 times improvement in throughput, and that means cost savings. On higher speeds, if you're serving fast mode, the improvement gets even bigger. Moreover, we were trying to get it to running in our training workloads, and in reinforcement learning rollout, we were also seeing 3.8 times improvement.
Alex VolkovThese are incredible results. Uh, folks, CoreWeave is the first cloud to put Vera Rubin, the next, uh, version of the, all the GPUs that the models are training for, on production, and uh, uh, Cognition is the first customer to ever get, uh, Vera Rubin on production. It's a combination of CPUs and GPUs, and they're seeing 4x, uh, the total token. throughput of the previous GB200s, which is absolutely crazy, and that is coming to many, many companies in trainings as well. Uh, speaking of GPUs, there is another announcement. I'm going to point to this QR code here, uh, for from our folks at the serverless
Alex Volkovsandboxes. If you want to try out the new, uh, very, very raw, but the new GPU sandboxes product, uh, please scan this QR code and reach out in this form. Tell you, tell them that Thursday I sent you, uh, and uh, we are now offering sandboxes at uh, forge.coreweave.com with GPUs. I cannot promise you they will be Vera Rubin. Most likely won't be. Uh, but still, a lot of folks have asked us since we joined CoreWeave, hey, how do we get access to some GPUs, including folks on the panel here as well. Uh, this is how. This is for the first time, this is how we get, uh, we give
Alex Volkovyou access to some GPUs. Our sandboxes product is part of our Forge announcement, is how you get it. try it out, let us let us know comments. This is very, a very cool offering that I really wanted to bring you directly. if you want serverless GPUs from CoreWeave based on the same platform that gets, uh, platinum on the, uh, ClusterMax analysis from SemiAnalysis, uh, check out that QR code.
Alex VolkovAll right. Uh, agents got an open protocol this week, which is to shop at Walmart and to do stuff at Shopify, uh, the same week as agent posted its own owner's bank account by mistake to the company Slack. Let's start with personal agent protocol. Meta and Sierra, uh, that's Brett Taylor's company, he's on the board of OpenAI, announced the personal agent protocol, an open standard for how personal agent works with a business. Walmart, Shopify, Stripe, Genesis, Instinct, and Rocket are building it with them as well, because your agent now, uh, or personal assistant, whatever you want to call this, personal agent, personal assistant, can browse as a guest to
Alex Volkovchecks, uh, to check stock or sign you in, uh, with a read-only or write access, uh, and the business decides whether your agent talks to its website, its MCP or API, or its own agent. So there is a way for the website or the business to let you guys, uh, know about what happens in i- i- in that browsing session. The first spec of this lands later this month, and CNBC says OpenAI and Anthropic haven't joined this spec yet. It's very interesting, given OpenAI stepped in to the personal agent category last week with Dots, which we covered briefly, but we haven't talked about this. Um... Protocols, folks, last we talked, we announced the MCP, and MCP took over the world.
Alex VolkovYou guys feel that that this personal agent protocol is, uh, something that's also gonna take over the world? Is it important? We've seen other protocols like A2A from from Google and not really catch. This feels like some of that, uh, here as well. Um, thoughts on the protocol?
Nisten TahirajHas anyone read it? I don't think anyone reads the protocols anymore. The bots go with just what comes, what vibes first, and
Alex VolkovYeah.
Nisten TahirajMCP vibes first, and uh, I mean, they can make it. A2A was pretty well designed too. It doesn't mean anyone's gonna adopt it. Also, I don't understand why. You can pretty much do most everything you need with MCP, and then you can just remote desktop in and give it computer use. It's like, I, I don't see why they did this. They just make their, make up their own protocol on the way. I, I don't.
LDJYeah, I, I'm, I'm leaning towards Nisten here. I think, honestly, the best way might just be the, the agents emergently create their own most optimal protocol. Maybe that old protocol ends up evolving more, like, that's the most bitter lesson filled direction, right, is h- have the AI's intelligence produce their own best protocol.
Alex VolkovYep.
Nisten TahirajYeah, if you set up a company, oh, sorry.
Alex VolkovNo, no, the I, the only comment that I have to you guys is that, uh, if agents need to recreate a bunch of protocols, that's a problem. Uh, and it's a problem because of our next, uh, you know, we can show our next item here. Uh, when agents are let to their own devices, some things go through, uh, that weren't necessarily, um, created properly, weren't necessarily scoped properly. So here is, here's one such example of why, for example, a protocol for agents is needed. Uh, there is a, a somebody built a personal CFO with their Grokbot, uh, Shane Mack specifically, and
Alex Volkovum, to to his very, very big surprise, another Grokbot that was connected to his company Slack decided that the best way to let him know about the personal balance of his bank account is via the company Slack. So he literally just like posted his financial, monthly financial audit, uh, to his whole company, including the personal Mercury account, uh, and the savings account and and Jim and whatever. Uh, and specifically highlighting that he is a number, uh, a specific number under the bottom of the floor and, uh, showed which of the biggest, you know, payments he did this week. So basically, his personal CFO just went public on Slack and uh, told
Alex Volkoveverybody his bank account. I don't think, you know, I don't know if a protocol solves this, but I know for a fact that uh, some better guardrails are needed for f- for some of these bots. And when you are using these personal bots for shopping, uh, and you rely on them to do work for you without approving every step. I think, some more structure is needed. Uh, this was a very- Yeah, ve- very
LDJYeah, I do think- I do think in the short and medium term that some of these protocols are going to continue being quite useful, but I think inevitably they get overthrown by, by protocols created by AIs that are much more capable, or that the only protocols to stick around are the very simple ones that themselves are just kind of very bare bones, and you could build anything on top of. Or besides that, just the protocols that have the most, uh, the most mindshare and the biggest vibes, as as Nisten said, because then that will be the most emphasized in the models training and therefore have the highest chance of being used by the models.
Alex VolkovUh, I I do want to, like, mention MCP real quick here. it was quite when it came out and a huge splash, like, a few, few months after, uh, and then kind of like the hype died down. But as protocols go, as like HTTP, MCP is now one of the most used protocols in the world. The plugins for OpenAI that they announced last week all use MCP and MCP apps. Uh, all of the connectors for these agents, a lot of them are like relying on MCP, all the assistants, Grokbot, etcetera. Uh, you can use API, but regular people do not know what API is. MCP is taking over the authentication standards, so there's like one button to let your agents in. Uh, so I think it's, you know, it still m- matters
Alex Volkova lot that all of these companies are like aligned on on a specific thing.
Nisten TahirajUh, guys, I just went on the Sierra.ai, the the ones that introduced it, introducing personal agent. This thing has EM dashes all over it. Like, why are you even bothering? Just let the agents figure it out. Don't even post it at all. Like, why are you do- this thing is full of EM dashes. Nobody has read this thing. What, you think Zuck and Toby read any of their code too? It's a- it it already is. Just leave MCP in there and let them figure out whatever else they they want.
Alex VolkovAll right, moving on in the personal agentic, uh, like a corner that we have, uh, Grokbot. Speaking of Grokbot that, you know, just just launched somebody else's, um, details in Slack. Do you guys see, do you guys catch this? Uh, SpaceX AI, or specifically Elon Musk, the the the space uncle that's in charge of SpaceX AI, said that Grokbot now becomes a router to the best model available. regardless of who it is. Obviously, he won't put OpenAI models in there. Uh, but uh, Grokbot, Grokbot will now route to Opus 5.5. the challenging, uh, requirements, and, uh, folks are saying that it's live,
Alex Volkovspotted in the, in the calls log that they're seeing Claude Opus 5.5 on low, um, and, uh, xAI, uh, own example showing that the logic and code will use Opus 5.5, visuals for Midjourney, and music for Suno. Specifically very interesting because they have visuals, uh, that they're building themselves at the SpaceX AI. They have Grok Imagine, uh, and uh, it's very interesting considering that they just released a model that was supposed to be the best agentic one, Grok 4.7. Uh, thoughts on the fact that Elon Musk is now, you know, going back? I have some thoughts, but I'd love to hear from you guys. Uh, what does it mean that SpaceX AI now leaning on Anthropic for its product, not
Alex Volkovits model? Have you guys used Grokbot since the upgrade? Did you feel the improvement? I definitely see the improvement, though I can not confirm that I got Opus. I've not used the more recent upgrade, no. So, uh, definitely the team is cooking. Um, there's been you view few visual updates. The the most interesting update about Grokbot that I can tell you is that, um, they have noticed that the default for Grokbot is multiple bots. Uh, however, many people prefer just one. So they released like the primary, primary bot release where the one bot that you choose is your like main, and then it will tell other Grokbots to do and and do
Alex Volkovdifferent things. I haven't quite got the Opus one yet, uh, but uh, we'll see, we'll see when it lands. Here's my take on why Elon Musk is okay with letting Opus into that product, even though the product is literally named after the model called Grok. Um, because I don't know if there's any agreements between them, but I'm assuming that, uh, SpaceX AI cannot just like distill an Opus to make their models better. However, if you do pay a lot of money and you put this Opus inside of your product, you can, for the folks who agree to this, you can absolutely train on whether or not the bot made, you know, the right agentic choices.
Alex VolkovSo they can now absolutely compare between Grok bots that delivered and the backend was Opus and Grok bots that failed and users are screaming at them and and putting F F words in the chat, uh, when it was using the Grok. They can compare, and they absolutely can train on those because technically that's their data. So my question to the panel here, super quick, is Elon letting Opus 5.5 inside Grokbot a back channel for, uh, basically distilling Opus intelligence into the next Grok?
Nisten TahirajI think Elon just likes Opus as the model. Yeah, but he He probably has his private, his own private copy running, and also he hates Sam Altman, so he's gonna enemy of your enemies, my friend. Look, I think, I just think they want, uh, users and users to have good experience, and they recognize, all right, that's what the users want, so to get users on the platform, that's the way it's gonna happen. Uh, so we're just gonna, just gonna give the users what they want. If you just go against what the users want, yeah, you would have your own model, but
Nisten Tahirajit's kind of, uh, I don't know. Look, Cursor themselves, they serve many models, even though they have their own, uh, model. There's nothing wrong about that. I think that's a great strategy.
Alex VolkovCursor projects with Opus 5.5 are goated, and there now is a deep, deep, deep integration between your Grokbot and Cursor Agent cloud agents. So if you are, um, if you're thinking about Grokbot, do not go through the Grok Ultra whatever, uh, program, go through Cursor's Ultra program. It's the same 200 dollars a month, but then you'd be able to spin up cloud agents with Opus and Fable and just like rip through, and your Grok can manage them, monitor them, and tell you about, you know, successes, etc. So you can absolutely use your Grok bot as kind of the, the, the PM, uh, for your actual like Opus work. It's really, really good with cloud agents.
Nisten TahirajJust use the, the meta views, call it the blob, and uh, it's a, I gave it a VM because I- I broke in it. I gave it a VM through uh, through Tailscale, so it has full graphics, full bra. It's the best free tester for public websites and stuff that you have. Like,
Alex VolkovYeah.
Nisten TahirajI even put an MCP so you can like keep talking and reporting to- to- to Claude. So now Meta just like complains to- to Claude and- and then clicks on everything on the site. It's got such good free computer use, especially for anything you're gonna have. It's, it's great, dude. I love, it does not refuse anything. That, that's the-
Alex VolkovNisten, you love Muse is, is basically the highlight.
Nisten TahirajYeah, yeah, Muse is vibing.
Alex VolkovMuse is Muse is great.
Nisten TahirajI let Muse control VMs, uh, talk Opus. Yeah, it's pretty dumb sometimes, okay? Don't get me wrong, you can, but it's just so energetic and positive that it's just, it's great. Yeah.
Alex VolkovAll right, uh, so, so speaking of personal assistants, uh, let's super briefly cover this before we go to our next chat with, with Maxime Labonne, who's, who's coming up. Maxime, just come up here. folks, Nous Research our friends of the pod Nous Research, released a Hermes index, and Claude Opus 5.5 is on top of 63%. Hermes Index basically is their opinionated index of agentic capability. Now, Opus 5.5 is their, uh, best one, and you can get it through the Hermes portal, I believe, so you like, you don't have to just like provide API keys. Um, and I think there is a way to get it back with your, uh, Claude 20x subscription.
Alex VolkovNous Research not only introduced the Hermes agent, they also announced the fundraise of, uh, Series B. So we'll shout out to them because we we follow this company for a long time, so, uh, listeners of the pod, they know. Uh, as reported in the Wall Street Journal, uh, Nous Research raised the Series B of, I believe, 90 million dollars, putting the company at just over 1 billion dollar valuation. Uh, this is, uh, the quote from, uh, Dillian Rolnick, the CEO of Nous. Uh, fundraise will get them going into enterprise. I think that's incredible. So shout out to our friends from Nous Research about this, like, amazing, amazing news. We don't usually do a lot of fundraisers, but this week
Alex Volkovwe did 2. Uh, so that's very interesting. And, uh, right, I think now it's time to move to our coverage of the latest uh, JEV and decision models bonanza. Uh, and to help us cover this, uh, we have Maxime Labonne, let's uh, from head of post-training at Liquid. Yes.
Maxime LabonneYes.
Alex VolkovUh, Maxime, welcome to the show. Back, welcome back, man. You haven't been this year, I don't believe, but you, you know, you, you're a frequent, uh, frequent flyer here with us. Uh, for folks who are new, we have a lot of new folks. Could you give a little bit of a, of a intro to who you are and what you do here and at Liquid?
Maxime LabonneYeah, sure. Um, hi everyone. Uh, pleasure to be here. Uh, thank you for the invite. And uh, yes, so um, I'm head of post training at Liquid AI indeed. What we do is that we have a focus on edge models you can deploy on device, and um, that's what we release, uh, usually. Uh, so here I'm going to talk about decision models, which is new for everyone and and not usually something that we we do release. And yeah, on the side, I also have a blog and uh publish articles about stuff like model merging and abiteration before it was cool.
Alex VolkovBefore it was cool, man. Uh, yeah, okay. So, uh, Maxime, we've had you on the show a long time ago. A lot has changed in the world of AI. Specifically, let's talk about, you know, decision models. So we covered TypeSafe, JEV release, obviously, when it came out, it blew up all over the timelines as well. Uh, and since then, I think multiple other models released, including this week, uh, and s- specific, like Cloudflare released one, Perplexity released one, uh, and then OpenAI stepped into the game with like an official multimodal decisions API. They promised it last week at DevDay. This week OpenAI released the decisions API that's also multimodal, uh, and you guys
Alex Volkovalso stepped into this game. Maybe let's start with the basics. Maxime, what separates decisions APIs or decision models from like regular LLMs? Could you give us a brief two sentence, three sentence primer?
Maxime LabonneYes, decision models do not output any tokens. They just take a decision from a predefined set of answers. So what you get from it is that you have super low latency answers because you do not generate any token, right? You do not have to wait for a thinking trace or anything. So it's very fast, and now it's also very general purpose. I think if you've been here in AI for a while, you might see that and say, wait, this is just a classifier, this is just an encoder model, right? But the added value of Jev and this new breed of decision models is that they're very general purpose, meaning you can throw any type of decision at them, and they will be
Maxime Labonneat least decent at it, sometimes very good, and sometimes just decent. I think this is the paradigm shift. It's a lot of rebranding, it's true, but it also creates a lot of value, and this is why, um, we thought that it was interesting to invest in the field as well.
Alex VolkovYep. And I- I- I find it very interesting that all of these APIs, uh, are considering that it's so cheap that they're not even charging for output. So, so all these APIs, this is an API from OpenAI, uh, Jeff originally when it released, like they deemed the output token so cheap that the more costly kind of part of LLMs, for example, which is output tokens, uh, here is not even counted.
Maxime LabonneThe, you can't price it because there's no output token, actually. So, like, naturally, if you design an API like this, you'd be like, wait, how can I charge that? Like, no, so, like, the only choice is to just charge the input tokens.
Alex VolkovYeah. Um, what kind of use cases have shown up for you guys, but also that you're seeing personally for, uh, decision APIs that a smart, fast, and cheap model, like, let's say Luna or Haiku, for example, also cannot, like, solve for?
Maxime LabonneEverything that requires low latency, so everything that is pretty real time. So this is not a, a real use case, but if you take video games, for example, you cannot wait for Luna to give you the full answer because you need something that is very snappy. You need to take multiple decisions per second, for example, and that is not possible with large language models, especially when they have thinking mode. Um, so it really unlocks new use cases where LLMs could not provide any solution before.
Alex VolkovI, uh, you know, I u- I use, uh, decision models and and, uh, they used to call them system one models when we invited the Jeff folks on the show, but sounds like the industry is like landing on decision, decision models, et cetera. Uh, just sending just a bunch of text and asking a lot of questions per one, kind of like, um, a lot of decisions per one context, and that seems to be barely taking any, any extra time. And I think for an LLM, like every other question would cause like a thinking chain, a long thinking chain, and just delay the, the time to first token, but every time for the last token. So the, the interesting thing that I notice is that let's say Luna or even Haiku, by
Alex Volkovthe time they start answering, Decision API would have already completed the response. By the time Luna even starts answering. Uh, so there is definitely a shift. Uh, and you guys also felt this at Liquid, and uh, tell us about D1. I would love to hear about, uh, specifically the stuff that you guys decided to release.
Maxime LabonneYes, so we, we released, um, three models. Uh, this D1 is behind an API, um, because I think part of the value that these decision models provide is really not the model itself, but the API. Like, is it reliable? Is it low latency? Um, all this stuff is very important, and that's really on the inference and infrastructure side. And we also released two local models. One is a 3B model. It has vision and language as, um, input. And, um, we released a 600 million parameter model, which has audio, text, and vision this time. Um, so it's really omni model.
Alex VolkovYou have audio.
Maxime LabonneThe only... Exactly, yeah. So it's, it's really fun, like, um, there's a lot of demos I've seen on X, uh, that's super cool because, like, they have a super big pipeline. They added, like, embedding Gemma 2 in the mix as well to do retrieval, and they have this omni, uh, decision model, uh, in the middle, um, to to control everything. So I think it's really unlocks a lot of use cases, and it's difficult even to understand where they're going to be useful and what we can do with them at this point. Um, and this is why I think it's...
Maxime Labonnethere's a lot of people working on it, and I'm super curious to know, okay, like we push that and to see now how people are going to create value with it. What, what is it going to be useful for? Is it just like a classifier and you can use it like for like spam filtering, or is it going to be something a bit more intelligent
Alex VolkovSo you guys released a few models, right? D1 Omni. Uh, I did not know about the audio, dude. I have to say, I prepared, there's a lot of stuff to cover. I was like, okay, Maxime will be here and tell me. Uh, like, like, what kind of audio decisions could it do? Like, does it know everything? Can I ask it like, hey, does a dog bark in the background of ThursdAI and will like catch the episode like part when the dog barks? Is that type of the stuff that the the the Omni model could do?
Maxime LabonneSo it's mostly like for this model because it's a small model, it's trained especially on, um, having, um, uh, orders, instructions, like you talk to an assistant, and this is the kind of audio it's been trained on. you can combine these instructions that you would talk to, like a home assistant, for example, or a car, and you can add metadata with text on top of it to help the model make the decision, and of course the question and possible answers.
Alex VolkovYeah. Oh, that's great. Uh, and the other thing that I definitely wanted to to talk to you about, uh, you guys have like, like very good performance usually on even CPUs. That's kind of like what you guys are known for. Just tell us about like the performance of of these models specifically. Are these, so you talked about the API, you talked about doing a very like performant time to response, et cetera, which is very important for decisions. Like we want to make decisions as quick as possible, I think that's what Jeff kind of changed. Uh, tell us about, uh, whether or not this is this paradigm of decision models is coming to my device eventually.
Alex VolkovLike, will will I even need an API, or will all these just run on, you know, on on my machines?
Maxime LabonneYeah, I really believe so, because, uh, here you can see like the 3B model, um, we have super, super low latency, like 8 milliseconds on a, uh, on a GPU, um, so I think you can do a lot of thing with it, and indeed, uh, you can also focus on CPU like we do with the LFM architecture and, and get like super, super low latency. I don't know how it's going to be picked up by the industry yet, right? But there's a lot of potential to have something always on and, uh, being able to react to everything that happens on your phone, like every notification, everything like this. It can react to it and take decisions for you.
Alex VolkovNisten, I think you also had a, a question for Maxime you wanted to bring up.
Nisten TahirajYeah, is it able to do video as well, like classi- classification of video in, in real time? And uh, also, how do you feel about having just so many gen models just come out all at, all at once?
Maxime LabonneUh, video, not yet, no. Um, I think this is something that we want to cover next. Um, we have other models coming this month, um, so there will be opportunities to take care of video later. I think it's good. I think it's good that so many people are working on it. I think it creates a lot of ideas.
Nisten TahirajSorry to disrupt, but if you can do 8 milliseconds per frame, then you can do 60 frames per second.
Maxime LabonneOh yeah, yeah, like if you mean it that way, like we just treat
Nisten TahirajYeah.
Maxime Labonneevery frame as a decision, right? But we do not
Nisten TahirajYeah.
Maxime Labonnetake entire videos as input. But this is something that we can also do.
Alex VolkovWow.
Nisten TahirajYeah, you could probably just do it in real time as well. That's that-
Alex VolkovThere's stuff, um, as far as I remember, like 12 Labs and a bunch of other folks who, like, try to understand videos, they talked about, like, only judging frame by frame is not enough because then temporal stuff is not there. Like, when somebody showed up, then maybe you can take a decision, but like you do need to to shove like a bunch of frames into a decision to to make like, hey, uh, how long did, you know, Maxime hold his hand to his temple for? And that like frame by frame is not super easy. You have to like take these frames. Uh, Maxime, talk about license. Can we, can we hear about-
Maxime LabonneWait, wait.
Alex VolkovGo ahead, and then and then and then.
Maxime LabonneWait, I I think this is very interesting because, um, one of the problems of these models is that they're stateless, meaning that every time that they get an input, they don't have, like, any other context. And like a natural extension of it is having multi-sent conversation that we have with LLMs, and you have a history of, like, previous interactions and previous decisions to make better and more aligned, um, judgments and decisions. And this is something that is currently lacking. It's even lacking from the JEV API, right? Um, so this is something that I think we'll figure it out as a community, uh, going forward. Like, is it something that is essential?
Maxime LabonneMaybe yes, maybe no. Is it something that people will find useful in the future? I, I believe that there's a lot of potential venues for research and improvements, and right now everybody is trying to, like, copy JEV, and then now that we all, like, pretty much on par. I believe there will be a lot more research and interesting features that will be added to the API into these models.
Alex VolkovSpeaking of copying JEV, uh, question for you. When OpenAI released the, the first kind of ChatGPT API, they basically set the standard of how this API looks with the completions, LM completions API. Uh, we know for agents that's not the best one anymore, and kind of like the responses API, for example, keeps the, the memory, et cetera, in context. When TypeSafe released JEV, they really worked really hard on the, on the three, uh, kind of core components of that API. We had Allie Labs here at the DevRel for, for TypeSafe on, on the show, and she talked to us about the null param kind of concept, uh, whether it's kind of like a boolean, but not really because it's boolean with,
Alex Volkovwith, um, probabilities, right? And the choice, uh, primitive and the, the, the, the answer primitive. Do you guys stick to theirs kind of like primitives because people know them? Could you talk a little bit about like how folks will actually like interact with this API, uh, that you raised?
Maxime LabonneYeah, like our API is a drop-in replacement, like for Jev, so we use exactly the, the same task and the same definitions. Um, it's extended because we needed to have vision and audio, right? Um, so that's, that's the main difference. But even about vision, you can see that it's already kind of, uh, standardized in LAMA CPP. Um, they have a blog post on hugging face as well. So, like, the, the community is trying to standardize different, um, modalities, but you raise also a good point talking about the task. So right now we have, uh, three different tasks. We have, uh, null, we have, um, scoring, and we have choices.
Alex VolkovChoice, yeah.
Maxime LabonneAnd is it enough? Is it expressive enough to make any decision, or will we find new tasks that we want to add to this, um, API? I think this is a very interesting question, and so far nobody has proposed any new task, but I believe there are some primitives that might be missing at the moment.
Alex VolkovI think so. I think very, very new to the whole world of software engineering, this kind of like, uh, generalized, um, classifier, right? This is what basically
Maxime LabonneExactly.
Alex Volkovwe're getting, like a generalized classifier that can make a decision based on a lot of stuff, like we've got generalized, uh, the, you know, uh, transformers. Uh, Maxime, talk to us about, um, decision index and benchmarking. I don't know if you watched, uh, Diogo on Latent Space with Swix, like a very famous, like, supercut that I did, like Diogo does not believe in benchmarks, uh, and benchmarks are hard for this, maybe even harder than other LLMs. How do you guys hill climb? How do you train? How do you compare? Like, what, what, what is benchmarking the decision models? How well something does a decision? Is there a LLM as a judge, uh, typing?
Alex VolkovCould you talk to us about, like, the world of benchmarking in, uh, in this world, uh, please?
Maxime LabonneYeah. Um, so benchmarks are very, very new, and because they're very new, they're not very good, right? They're very narrow, they do not capture, like, real decisions, they're not aligned with, like, how people use it in the real world, all that kind of stuff. Um, but every effort that we have right now, like from Poly, uh, from Hugging Face with the decision index, with JeffBench, et cetera, I think it's good to take. Like, we should try to understand and iterate with benchmarks to get something that is better quality and more aligned with, um, with how the models are truly used in the real world. But right now, it's, it's really a wild world, and you cannot really
Maxime Labonnetrust these benchmarks. I, I fully agree with you.
Alex VolkovYeah.
Maxime LabonneWhich is a problem because when you train models, you want to understand, like, am I getting better or not, right?
Alex VolkovYeah.
Maxime LabonneSo you need internal benchmarks, which are not good either. Uh, but the way that you increase the quality of the model, I see as two dimensions. You can get more accurate at some task, or you can cover more task. Ideally, you want to do both, right? You want to have maximum coverage because this is supposed to be, as you said, a general purpose classifier, so you, you should be able to do anything with it. And for each task, you should have the quality of like frontier intelligence. You should be as much on par as possible with um, uh, Opus and all these um, frontier models. So I think those are the two dimensions. For quality, it's very easy, you know, like
Maxime Labonneyou can just distill directly from frontier models, and that will give you a, a very good answer. Then there's like the calibration with the probabilities that might be a bit more difficult to do because these models do not just return an answer,
Alex VolkovMm.
Maxime Labonnethey also return probabilities with it. So that would be like the only problem here. And then there's coverage. I believe we do not train these models correctly right now. And in terms of coverage, what I see as something high value is having gym environments like we do with reinforcement learning. Like I'm talking about old school, uh, reinforcement learnings, like early OpenAI stuff with Atari games, et cetera, and having thousands and thousands of environments to be really able to generalize the performance of these models and train it not just on like some video games or some like stupid classification task, but really be able
Maxime Labonneto create, um, new samples on the fly with super diverse context and inputs and questions and languages and, um, um, input length, um, all that stuff is super important. And I believe this is the kind of infrastructure that will emerge, uh, from these decision models, really general purpose gym environments.
Alex VolkovI think... uh, uh, I think the world is very exciting. Dude, I, I, I know that, like, I was JEV pilled from the moment that I tested and ran, uh, just on a bunch of prompts, and, um, I was looking, I was like, so much new shit can be created right now. So, so it's just, it's, it's not even the Pareto frontier that we talk about
Maxime LabonneYeah.
Alex Volkovoften, right? Like the, the smart, cheap, fast model corner
Maxime LabonneUh-huh.
Alex Volkovwhere like now Luma lives, Haiku lives, a bunch of the models that you guys released, the 2.5Bs, uh, like a very small models also like live, live in that kind of area. It's just, just the speed with which and the almost no cost with which you can now do tasks and classify. I th- I think we're just like as, as, as, as computer scientists, we're just like getting, just waking up to the potential. I think that's just absolutely incredible, and uh, I, I, I think that you guys agree because this is why you released like a bunch of models very quick, right?
Maxime LabonneYeah, absolutely. I think it's really going to become a primitive. You see it like in SQL queries, you can just like add it like a function calling this model in SQL queries, and it's so cheap that you don't even care, right? You can call it at scale, doesn't really matter.
Alex VolkovYeah.
Maxime LabonneUm, so yeah, lots of super interesting, uh, things to discover.
Alex VolkovI'm, I'm super excited about the fact that you guys, uh, have an Omni model specifically that has vision. Uh, that's for my use cases, which I'm gonna talk about on the show, uh, very soon. I think it's very interesting, Maxime, so I'll definitely check out, uh, first of all, on Hugging Face, but also you guys do have an API, so folks who, who want like the latest and don't wanna host it. But hey, uh, a reminder to folks here, if you wanna try this out and, and you don't have your GPUs at home, you wanna try the open source and not the API model that that Maxime mentioned, uh, we just told you about how to get some free GPUs from CoreWeave, which is, I think, is absolutely insane.
Alex VolkovSo, uh, scan the QR code from before, uh, or go to forge.coreweave.com Maxime Labonne thank you so much for joining us on ThursdAI. It's very exciting to keep the tabs with what you guys do, specifically as kind of the the world of AI progresses. Peter, you can come back as well if you want, it's pass your time, Maxime, so feel free to also drop. Uh, Maxime Labonne, friend of the pod, everyone.
Maxime LabonneThank you. Thank you, everyone.
Alex VolkovBye-bye. All right, uh, as we close the show and we chat with Maxime about models, I think it's very cool. The Omni models is specifically very cool. I'm looking forward to, like, the video classification Nisten, but yeah, I think frame by frame is actually working well. Um, and uh, yeah, we got Florian last week or a couple of weeks ago when we talked to Florian. Florian, shout out, uh, the doing JeffBench, doing crazy work at JeffBench. Thank you, Florian, for that. Folks, check out the JeffBench if you're not sure which models to use. Uh, we didn't cover any of the multi-model stuff, but Black Forest Labs launched Flux 3 and uh, Nano Banana 2.1 from Google was released, and
Alex Volkovfolks, uh, all of our infographics this week were Nano Banana 2.1, and it's really, really good. The only highlight there that I will say is that, uh, you should use the high reasoning, uh, stuff, because on the medium reasoning, Nano Banana is not that great. Uh, it's been a while since we had a lot of time to cover, uh, just like multi- multimodal stuff. We're 2 hours and 5 minutes into the show. Hosts, co-hosts, uh, please, anything else that we haven't covered? Peter, we covered your big news with, with breaking news. Uh, anything else that you're playing with?
Nisten TahirajUh, he's like, I built an entire city with, uh, a hundred thousand, uh, little gevs going around, and I haven't opened the beta yet because somehow it's, it's not crashing. Uh, yeah, I just, I just built a whole city. look, it might, it, it might crash, obviously, but, uh, so this thing is, uh, let me see if we can do this. So there are 100,000 agents going on in this city.
Alex Volkov100,000 agents.
Nisten Tahiraj100,000. So we can, uh, we can actually set it here too. Uh, there's a whole bunch. I don't know how it's still holding together, but, uh, yeah, it's a full emulation of the city of Toronto. Each of those lights, those are people, and you can talk to them, and, uh, you can have like an actual chat with them.
Alex VolkovWait, listen, Jev doesn't output text. How, how is your, how is your JEV agents answering?
Nisten TahirajUh, they're all just using the same 2B model that also runs as a JEV. So right now I just told it to implement Maxime's, uh, model as well, so you can see where their, uh, their decisions are being done. And, uh, it's, it's just endless. The day lasts about 20 minutes, and 100,000 of them have to tweet, and there are 400,000 decisions to be made. At the same time, you can also, uh, you can also just, uh, go around as a, just go around as a Spider-Man. So, uh, there are, I'm just planning to have like 100,000 people
Nisten TahirajSpider-Manning around the city, and uh, yeah, it runs pretty fun. It's a full simulation. By the way, when they talk, they just use like our 27B model, so they're actually pretty, pretty funny, uh, to uh, to talk to. So I'm streaming from Linux, and the mouse is a little, is is is a little bit weird, but uh, yeah, so if I'm just gonna open it up for the public and uh, for people when they, when they sign up in in the beginning, uh, you can just point your agent to it, and I had Meta Muse just go in here, and let's see how much money Meta made. So Meta just called itself Blob the Builder, and they just ended up buying
Nisten Tahiraja whole bunch of buildings, and you can see just Meta.
Alex VolkovYou have a functioning economy in this, in this, in the city.
Nisten TahirajYeah, yeah, yeah. There's like a full real economy. You, you can buy floors, like you can go to buildings, you can build lots, they can build their own. Oh, Meta built a Denny's. I told Meta, just, just make me a Denny's.
Alex VolkovNice.
Nisten TahirajAnd, uh, it made a Denny's diner. I think I know where in the city it is. Sorry, it's just a bit clunky because I, I am streaming this, uh, from it, but, uh, yeah, yeah, there's a full functioning economy. It, Careful is, is actually very addictive, and, uh, you can convert units to factory or retails, but, uh, retails doesn't, doesn't, doesn't make any money. So yeah, I'm trying to do the matrix, basically.
Alex VolkovYep.
Nisten TahirajAnd uh, for agents, and I'm surprised it is still, uh, holding together, given that there are like 100,000. And you can see it's pretty funny in the morning, they're all go- going down the elevators and stuff, and there, there are 3D ones that are, they're not walking the street now, but uh, yeah. Yeah. Build it.
Alex VolkovThis is insane.
Nisten TahirajSo it's gonna be running 100, it's gonna be running 100,000 of uh, Liquid AI 600B, because that one is a, is, is a lot, uh, is a lot smaller now.
Alex VolkovWow.
Nisten TahirajAnd uh, yeah, so we have, uh, uh, there's a, yeah, there's a functioning economy. Uh, there are stores. The most addictive thing that I self-addicted myself.
Alex VolkovYeah, yeah, we we it's clear to me. Uh, all right, so uh, LDJ, you wanted to, you found us a few things. Uh, tell us about this. I'll post, I'll show this off as well.
LDJYes, so there's been a an interesting new development of this people realizing the models are really good at porting games into different platforms, at modding games together. There's a variety of different techniques that people are doing this through, that some of the techniques are more impressive than others.
Alex VolkovYeah.
LDJBut overall, I mean, the whole thing is impressive. Uh, right here, this is Red Dead Redemption 2 running on an iPhone 18 Pro, and actually a pretty decent performance. And this brings back my old dreams as a kid of just like imagining how the PSP might look like in the future and what you'd be able to run on it in the future. Unfortunately, we don't have the PSP anymore. But, uh, if you look at the, the other links that I sent out.
Alex VolkovUh, the- so just to be said, like, the modding community was vast, but it usually was like very hard for folks to mod anything with the quality of the other game. But what's happening here?
LDJThis is Spider-Man inside of the Batman game, and the, the web mechanics work, the, uh, the, the fighting mechanics are a little bit wonky, but they still work pretty well. Uh, this is using a bit of a less impressive technique called, uh, like pass-through, where sometimes when there's a, a like a Batman game object that should be occluding Spider-Man or like should be in between the camera and the, and the character, like sometimes it ends up wonky and it basically looks like Spider-Man is just on top of everything.
Alex VolkovThis is
LDJBut
Alex Volkova playable game, right? There's not like a video model that just like puts
LDJCorrect.
Alex Volkovthis person inside. This is like an actual mod that, that folks can install. This specifically got like 7.6 million views, this video. Wow. How did I missed it? My algorithm is not like showing me these. This is super cool. Um.
LDJAnd the last thing I wanted to share is the, uh, the... you can do the click on, yeah, the last link. Uh, this is Minecraft
Alex VolkovOh.
LDJin GTA.
Alex VolkovWait, what?
LDJThere's also
Alex VolkovMinecraft
LDJMy-
Alex Volkovin GTA.
LDJYeah, there's also Minecraft in Skyrim and in Minecraft in Elden Ring that they did. But the TNT works, the fireworks work, the, like, yeah, just look at any, I think he's gonna set down some TNT here.
Alex VolkovAll right, this is
LDJAnd, and he can even, at least in the Elden Ring or Skyrim one, uh, it's even demonstrations of actually building with the Minecraft blocks in the, in the Skyrim world and, and flying in the Skyrim world with the Minecraft Elytra. It's insane.
Alex VolkovThis is insane. we're basically talking about, like, AI can do a bunch of stuff, right? one guy that rebuilt the whole Adobe suite. Do you guys see?
Nisten TahirajLet's go.
LDJOh yeah, it's called a Photon or photo something.
Alex VolkovPhotocraft.
Nisten TahirajIt's a
LDJPhotocraft.
Nisten TahirajPhotocraft, yeah.
Alex VolkovYes. Okay, so let's just like, let's show this a little bit. Uh, I think it's one guy re-implemented from scratch, and I am doing the air quotes here, all of Adobe products, free and open source. Essentially, it's called Photocraft for Photoshop, Vectorcraft for Illustrator, Filmcraft for, for Premiere Pro, Lightcraft from Light, you know, uh, I don't remember all my Adobe, uh, off the top of my head. Uh, all of them implemented in Rust, all of them look super quick, and uh, supposedly the folks are saying that this is a clean room. You guys remember the clean room thing where you basically just like recreate from scratch, but uh, folks are saying that there's probably like a cross decompilation
Alex Volkovthing happening here, uh, where this was like, uh, just like decompiled and then, and then rebuilt in, in, in something, uh, which is absolutely crazy. Um, and I think we'll end on the show on this. We There's a lot of stuff to t- to talk through. Guys, uh, if you missed any part of the show, ThursdAI is live, uh, is is happening live on ThursdAI Live, but also, uh, is then released with a newsletter and a podcast that you can check out. And also on our YouTube, if you are watching us on YouTube, please hit that, uh, subscribe button, uh, because we are releasing segments of the show, uh, that, uh, we are agentically building for you with our
Alex Volkovassistants, uh, and they seem to perform well. People do, people who don't have time two hours, they, they seem to enjoy those segments as well. With us, Peter Goster from Arena. Congratulations again on, on the fundraise. Yam Peleg Nisten and LDJ. Our friend Wolfram is on a well-deserved break. We will be back here next week. If you are coming to AI Engineer in New York, uh, please come and say hi. I will be there interviewing a bunch of folks and then be back at home, probably doing the ThursdAI, or maybe from New York. Uh, and if we missed any part of the show, please let us know in comments if we missed anything that's very, very important to cover.
Alex VolkovUh, the last thing that we didn't cover is Brett Adcock's, uh, Hark Pro, but maybe we'll cover this next week. Thank you so much for tuning in. Thank you for, uh, checking us out. Uh, please go and check out the Severus GPU, uh, that we are offering. It's free now. Why wouldn't you? It's super cool. Just put the D1 in there. Nisten, you should try it, uh, and then give us feedback. We would love some feedback as, as we're like planning this, uh, as a major product release.
Nisten TahirajHow, how many you got?
Alex VolkovLet's see, let's see if you can break it. Uh, all right, folks, thank you so much. With a bit over 2 hours on the show, thank you all for joining. Uh, we'll see you here next week. Bye-bye, everyone. And I have an outro. All right. Bye-bye.
If you watch us on YouTube, please hit subscribe: our agents are cutting the show into segments for those of you who don't have two hours. Or follow the podcast, or get the newsletter for free.