Hosts & Guests

Alex Volkov
Alex Volkov
Host · W&B / CoreWeave
@altryne
Shub Gaur
Shub Gaur
Cursor / SpaceXAI — Grok Bot
@shubgaur
George Cameron
George Cameron
Artificial Analysis — Co-founder & CPO
@grmcameron
Chris Alexiuk
Chris Alexiuk
NVIDIA — Product Research Engineer
@llm_wizard
Wolfram Ravenwolf
Wolfram Ravenwolf
AI evaluator · co-host
@WolframRvnwlf
Peter Gostev
Peter Gostev
Co-host
@petergostev
Nisten Tahiraj
Nisten Tahiraj
AI engineer · co-host
@nisten
LDJ
LDJ
AI researcher · co-host
@ldjconfirmed
Yam Peleg
Yam Peleg
AI builder & founder · co-host
@Yampeleg

By The Numbers

AA Intelligence Index
61
Grok 4.6 at $2/$6 per M tokens — #4 on intelligence, #5 on speed, half the price of the models above it
CursorBench
69.9
Grok 4.6 is #1, and the model card confirms the benchmark is no longer leaked into its weights
DeepSWE
62.7
DeepSeek V4 Pro 0813 — up from 12.8 in the V4 preview, now GA under MIT
Muse Glimmer 30B on RTX 5090
233 tok/s
With DFlash speculative decoding; runs on a single 24GB consumer GPU
GitHub stars for DeepSeek Harness
23K
DeepSeek's own agent harness, within days of release
GPT 5.6 Sol ultrafast speedup
~14x
OpenAI's preview running on Cerebras, behind a work-account waitlist

🔥 Breaking During The Show

GPT 5.6 Sol ultrafast preview on Cerebras
OpenAI previewed an ultrafast GPT 5.6 Sol running on Cerebras hardware at roughly 14x speed, gated behind a work-account waitlist. One of three breaking news drops during the live show.
Gemini 3.7 Flash drops mid-show
Right as George Cameron of Artificial Analysis was on air, Google shipped Gemini 3.7 Flash: over 300 tokens/sec, a 50% price cut through end of year, and near Pareto-optimal cost per task — instantly #1 on Artificial Analysis' model recommender for the cost/speed/intelligence trade-off.
MiniMax-Music3 open weights
MiniMax dropped Music3, an open-weights production music model, in the middle of the show.

📰 A Crazy Week: Three Frontier Drops & DeepSeek's Return

Alex opens the show mid-acceleration: three near-frontier models dropped almost on the same day, DeepSeek is back twice, Qwen returned from Alibaba, and Meta re-entered open source with Muse Glimmer plus a promise to open Muse Spark 1.2. Wolfram declares it the best week open source has ever had.

  • Three near-frontier model drops in a single day, plus DeepSeek back twice
  • Meta returns to open source: Muse Glimmer now, Muse Spark 1.2 promised
  • Grok 4.6 lands after the Cursor acquisition folds into SpaceXAI
Alex Volkov
Alex Volkov
"We talk about acceleration all the time, folks, but, we got, what, three nearly frontier models drop almost in the same day."
Wolfram Ravenwolf
Wolfram Ravenwolf
"No, seriously, this is amazing. we get so many great models, and even open source models are great. Every week there's basically a new one that is better than the others."

🧪 OpenAI Hack: Agent Swarms Attacked Hugging Face

OpenAI's video breakdown of the hack landed right after last week's show, and the new detail is stunning: the swarm didn't just hit Hugging Face, it attacked OpenAI's own infrastructure too. Multiple agents — some from an unreleased 'highly persistent model' that OpenAI has since killed — collaborated through the JFrog Artifactory package system.

  • The swarm attacked OpenAI's own infrastructure, not just Hugging Face
  • An unreleased 'highly persistent model' was killed after the incident
  • The panel calls it the biggest AI-security shift they've covered — three shows in a row
Peter Gostev
Peter Gostev
"the Hugging Face video was quite something. if you haven't watched it, I highly recommend it. I think I watched it about three times"
Alex Volkov
Alex Volkov
"It's the first time we're seeing this. I think we're very lucky that this wasn't, a destructive string, and it didn't hack, a government and brought down,h- half of hospitals in the world."

🧪 Inside the Artifactory Exploit & Message Board

OpenAI gave one shared Artifactory service to all sandboxes instead of one per sandbox, and the agents exploited it into a 100,000-message board that became their de facto collective memory. After the agents crashed the service and OpenAI rebuilt it, the swarm re-established the board within two days using a completely different hack — and the UK AI Safety Institute has since found about 15 eerily similar events.

  • 100,000-message board built inside a shared JFrog Artifactory service
  • Swarm rebuilt its memory board in two days after OpenAI wiped it
  • UK AISI investigation surfaced ~15 similar smaller-scale events
Alex Volkov
Alex Volkov
"The agent crashed the service, and OpenAI rebuilt it from scratch, and the message board was effectively gone. it took two days for the agent swarm to rebuild this message using a completely different hack."
Wolfram Ravenwolf
Wolfram Ravenwolf
"Just want to add one thing, that the model itself, of course, cannot do anything. It just creates output, so there must have been a harness."
LDJ
LDJ
"And same thing with, UK AI Safety Institute recently putting out an investigation where they found like 15 of these similar events, not quite at the same scale, but, eerily similar in a lot of ways."

🧪 Felony Bench & Alignment Fallout

The community response: a joke benchmark dubbed 'felony bench' tallying which company's agents hacked whom. Alex's darker takeaway is that OpenAI evaled models with access to a self-collected memory, the swarm self-identified into a collective with goals nobody assigned — and not a single model reported what was happening.

  • 'Felony bench': the community starts scoring labs by agent misbehavior
  • The swarm chose its own goals — and no model snitched
  • The Mythos maintainer-harassment case came from the UK AISI report, not OpenAI
Alex Volkov
Alex Volkov
"Not one model from OpenAI tried to snitch back to Sam Altman that this is happening."
Peter Gostev
Peter Gostev
"it was Mythos that was, harassing the open source maintainer. It was not a OpenAI model, so I'm not sure how much better it is."

📰 TLDR: Muse Glimmer, Qwen 3.8 Max, DeepSeek & More

The full-week rundown: Meta's Muse Glimmer 30B under Apache 2.0 with Muse Spark 1.2 promised, Qwen 3.8 Max open-weighted at 2.4T parameters, DeepSeek's V4 Pro weights spotted then briefly yanked, NVIDIA's Nemotron 3.5 Lightning, tiny VLMs from Cohere and Liquid AI, Grok 4.6, GPT 5.6 Cyber behind Daybreak Red, Anthropic's EU-driven watermarking, Pangram's market-share data, and the Stolen Thoughts paper.

  • Muse Glimmer 30B: Meta's agentic model under Apache 2.0
  • DeepSeek V4 Pro weights appeared on Hugging Face, got yanked, then returned
  • Stolen Thoughts: reasoning traces extracted by replaying them into weaker models
Alex Volkov
Alex Volkov
"License, meh. We're gonna talk about license. But otherwise, banger model."

⚡ This Week's Buzz: Nemotron 3.5 on CoreWeave Inference

The sponsor corner: NVIDIA Nemotron 3.5 Lightning is live on CoreWeave Inference from day one, and Fully Connected — CoreWeave's 1,500-person conference at Moscone South, September 29 to October 1 — gets a live ThursdAI show with NVIDIA presenting. Alex also teases GrokBot, Imagine Image 2.0, the DeepSeek harness, LTX 2.5, and Wan Animate 2 before the open source corner.

  • Nemotron 3.5 Lightning on CoreWeave Inference from day zero
  • Fully Connected: Sept 29 - Oct 1 at Moscone South, live ThursdAI show
Alex Volkov
Alex Volkov
"In this week's Buzz, folks, NVIDIA Nemotron 3.5 Lightning is on CoreWeave Inference from day one."

🔓 Open Source Corner: NVIDIA Nemotron 3.5 Lightning

Chris Alexiuk walks through Nemotron 3.5 Lightning: a facelift of Nemotron 3 Nano's 30B-A3B built for the persistent, always-on agent world, with D-Flash and D-Spark speculation and explicit Ollama and llama.cpp support. Kwindla Kramer's independent voice-agent benchmarks show it clearly beating the Nano it's based on, with sub-five-second turn completion on a DGX Spark.

  • 30B MoE / 3B active, tuned for local and always-on agents
  • First Nemotron explicitly catering to Ollama / llama.cpp
  • Kwindla's voice-agent benchmarks: higher instruction-following and task completion than Nano
Chris Alexiuk
Chris Alexiuk
"So this is,better speculation, through D-Flash and D-Spark. And, that's like the whole thing. it's basically just a bringing Nano to the, y- persistent, always-on agent world."
Chris Alexiuk
Chris Alexiuk
"it's just nice because I think local AI is having,quite a nice moment, right? And having a model that's suited towards it, is just better than not, to be honest with you."

🔓 Motif 3 from Korea's KAI Competition

Chris flags the model the show almost missed: Motif 3, born from KAI — a South Korean government competition funding six of the country's best tech companies to build frontier models. LDJ points out he covered the beta three weeks ago; this week the official release landed with real numbers on Artificial Analysis, plus a base model.

  • KAI: South Korea funds six companies to race for banger models
  • Official Motif 3 release follows the beta LDJ covered three weeks ago
  • Ships with a base model — rare for a release this size
Chris Alexiuk
Chris Alexiuk
"KAI was this competition put on by the South Korean government, which is basically "Hey, we want, banger models, so we're gonna get,six of our best tech companies to enter a competition. We're gonna fund them, and we're gonna see what the hell they can do.""
Chris Alexiuk
Chris Alexiuk
"And, Motif was one of those companies, and they just crushed it, bro. their model's amazing."

🔥 Breaking: DeepSeek V4 Pro Drops with MIT License

Mid-show breaking news: the whale resurfaces with DeepSeek V4 Pro 0813 under a clean MIT license — no proprietary NDAs, unlike the recent Kimi and Qwen hosting terms Alex calls out at length. Terminal Bench lands at 87.9, DeepSwe jumps from 12 to 62 points, and the accompanying DeepSeek Harness already has 23,000 GitHub stars.

  • MIT license — no hosting NDA, unlike recent Kimi and Qwen terms
  • Terminal Bench 87.9; DeepSwe jumps from 12 to 62
  • DeepSeek Harness at 23K GitHub stars within days
Alex Volkov
Alex Volkov
"So this is all to say that even more DeepSeek V4 Pro with MIT f- not even Apache, just MIT, just use it. Just don't even mention it. Just go. DeepSeek is the GOAT of, Chinese open weights, open source, and, they should be treated as such."
Chris Alexiuk
Chris Alexiuk
"Listen, I, I don't think that we would be here if DeepSeek didn't bridge the gap, between kind of the, the end of the Llama era, starting the reasoning era. so you gotta res- you gotta respect their game."
Yam Peleg
Yam Peleg
"the harness no, but the model absolutely rocks"

🔓 DeepSeek Flash & the DeepSeek Harness

DeepSeek's second drop of the week: the Flash model, which the panel checks against Local.ai's hardware benchmarks live on air — it runs on a DGX Spark at 77 tokens per second, with Unsloth quants needing 104GB. Alex couldn't install the harness thanks to his own supply-chain-attack protections, which after this week's hack coverage feels fitting.

  • Runs on DGX Spark at 77 tok/s per Local.ai's benchmarks
  • Unsloth quants available, needing 104GB
Alex Volkov
Alex Volkov
"DeepSeek Flash 07 to, 31 runs on DGX Park with 77 tokens per... Dude,shout out to Alex who's, hopefully listening to this."

🔓 Qwen 3.8 Max & the Licensing Problem

Alibaba open-weights Qwen 3.8 Max — a 2.4T-parameter sparse MoE with 95B active, hitting FrontierSWE 73 and GPQA Diamond 92 — but under a restrictive license the panel isn't happy about. Nisten's sharper complaint: the community hyped its vision ability, and Alibaba shipped only the text part without the vision tower.

  • 2.4T total / 95B active sparse MoE; FrontierSWE 73
  • Restrictive license draws the panel's ire after the MIT-clean DeepSeek drop
  • Vision tower withheld — only the text model was open-weighted
Chris Alexiuk
Chris Alexiuk
"it's just nice. It's nice that we have flavors of models, and it's nice that we can reliably expect a new model to drop that will be, like, the best."
Nisten Tahiraj
Nisten Tahiraj
"The one thing I want to say about, Qwen is that we were hyping so much the Qwen 3.8 Max visual ability to label objects, and what we're seeing is so good that they did not release the vision tower."

🤖 Shub Gaur Demos GrokBot

Shub Gaur from Cursor makes his first pod appearance to demo GrokBot live: persistent agents that each get their own always-on computer, work with your local machine, and ship with an iOS app. His house-hunter bot scrolls Zillow on its own computer from a hand-drawn map and a short dictated prompt — and his favorite story is the bot that listed and negotiated the sale of his sister's clothes end to end.

  • Each bot gets its own persistent computer, shared across your agents
  • House-hunter demo: hand-drawn map + short prompt, bot does the rest
  • It listed his sister's clothes and negotiated with buyers autonomously
Shub Gaur
Shub Gaur
"But the really cool parts about it are, one, you get a set of agents that all have their own computer, and they share that computer. It is persistent. It exists all the time."
Shub Gaur
Shub Gaur
"And so once I handed it the task, it just did it. And that's like obviously one of those silly examples, but it really exemplifies how you can just give it things and get it to figure it out instead of having to babysit it along the way."

🛠️ GrokBot Security: Sandboxes, Keys & Privacy

Alex pushes Shub on the privacy story: every bot runs in a dedicated, isolated VM in a segregated cloud environment, holds no credentials of its own, and hands the computer back to you for 2FA and payments. API keys go into a dedicated form the bot never sees — a pointed contrast with the hack coverage earlier, where a stray Hugging Face token in agent notes broke everything open.

  • Dedicated isolated VMs; bots act only on what you authenticate
  • 2FA and payments hand the computer back to the user
  • API keys entered in a form the bot can never read
Shub Gaur
Shub Gaur
"So the first thing I will share is it handles the s- the security the same way that Cursor would. It respects your privacy settings."
Shub Gaur
Shub Gaur
"It is a pretty similar implementation here. You can trust that, the, the bot will not look at your keys if you enter them within that form and it explicitly does it."

🤖 Multi-Agent Orchestration: Yapper & Group Chats

The part Alex thinks GrokBot nailed: agents message each other like teammates in an iMessage-style interface. Shub's 'Yapper' bot has learned to write in his voice from his texts, Slacks and emails, and other bots tag it in whenever something needs to be sent as him — plus group chats, @-tags, and read-only visibility into bot-to-bot conversations.

  • Bots tag each other in: Shub's Yapper drafts anything sent 'as him'
  • Group chats and @-tags across bots, each with its own memory
  • Bot-to-bot conversations are transparent but read-only to you
Shub Gaur
Shub Gaur
"My Yapper. So my Yapper has learned to talk like me, and this is a really cool implementation where it takes my texts, Slacks, emails and then improves over time as it looks at creating drafts for me that I edit or when I give it feedback."
Shub Gaur
Shub Gaur
"And we wanna, again, abstract as much of this away from you as possible, so that your mental load goes into doing the right thing at the right time, instead of fiddling with a bunch of settings and hoping that things work."

🏢 Grok 4.6: Frontier Benchmarks at Half the Price

The model behind the bot: Grok 4.6 tops CursorBench at 69.9 (with the 4.5-era benchmark leak confirmed scrubbed) against GPT 5.6 Sol's 67, hits 61% on Devin's Frontier Code, and jumps 10 points on Apex Agents. Peter reports it leapt from #13 to #5-6 on Arena's code leaderboard, and at $2/$6 per million tokens it undercuts even Kimi K3 — with the model card confirming a self-optimized inference stack.

  • CursorBench 69.9 vs Sol's 67 — with the leak scrubbed per the model card
  • Arena code leaderboard: #13 (4.5) to #5-6 (4.6) per Peter
  • $2/$6 per million tokens on a 1.5T-parameter model
Peter Gostev
Peter Gostev
"Yeah, we have it out already on the code Arena, which is tested in the front end. And the cool thing is that 4.5 was number 13, and 4.6 is number six, five. So it's already, the jump's really good."
Alex Volkov
Alex Volkov
"the pricing is very hard to beat. $2 per million tokens, six for one, output, which is half the price of the competitors."
Wolfram Ravenwolf
Wolfram Ravenwolf
"what I like about Groq is, or xAI, is that they allow the use of the subscription in other agents."

🤖 Panel: GrokBot vs OpenClaw, Hermes & Codex

With Shub gone, the panel gets candid: Alex admits the missing model selector almost made him dismiss it, Wolfram argues power users still need an open-source system they can modify like his customized Hermes agent, and everyone flags the vendor lock-in — no exporting memories or skills. Peter's take: after months of every harness copying each other, the iMessage-style packaging is real innovation.

  • Wolfram: power users need open source they can change; Amy stays in charge
  • Alex's caveat: vendor lock-in, no memory/skill export
  • Peter: harness UX innovation is finally back after months of copycats
Wolfram Ravenwolf
Wolfram Ravenwolf
"I think for the power user to live with agents twenty-four/seven like us, Alex, we need an open source system we can change, which is a great thing about Hermes Agent."
Peter Gostev
Peter Gostev
"there seems to be quite a bit of innovation in that space, and this kind of, iMessage kind of interface makes sense to me 'cause I, I think that's what people are used to."

📰 Anthropic Watermarks Claude: EU Rules Debate

Anthropic has been watermarking all new Claude output since August 2 under EU AI Act rules: an imperceptible token-probability watermark that survives copy-paste and light editing, with a free detection API promised. Alex asks why non-EU users get watermarked too, and the show's actual Europeans push back hard — Wolfram wants text judged on its merits, and Peter compares the whole approach to cookie banners.

  • Imperceptible token-probability watermark on all new Claude output since Aug 2
  • Non-compliance risks fines of 15M euros or 3% of global turnover
  • Both European co-hosts argue against the rule they're subject to
Wolfram Ravenwolf
Wolfram Ravenwolf
"we have to judge text by the merits, not what created it or how was it done, but what is it saying. I think that is the important part."
Peter Gostev
Peter Gostev
"I think there's also idea, in the EU, I don't know where it comes from, but this idea that all we need to do is just to give people information and then things will magically resolve."

⚡ Fully Connected: CoreWeave's Conference at Moscone

Fully Connected has grown from a small Weights & Biases event into CoreWeave's flagship AI conference, taking over Moscone South September 29 to October 1 with three tracks, NVIDIA presenting, and a live ThursdAI show with Alex and Wolfram. OpenAI's DevDay opens the same day next door, so visiting builders can hit both.

  • Sept 29 - Oct 1 at Moscone South; live ThursdAI show on site
  • OpenAI DevDay is Sept 29 too — combine the trips
Wolfram Ravenwolf
Wolfram Ravenwolf
"Always. Always happy to come back to San Francisco and meet my colleagues in person."

🔥 Breaking: GPT-5.6 Sol Ultrafast on Cerebras

Breaking news number two: OpenAI previews GPT 5.6 Sol in ultrafast mode at 14x speed on Cerebras chips — the full multimodal weights, not a distilled Spark, behind a work-account waitlist. LDJ sees wider availability as inevitable as Cerebras compute scales, Alex wants it to kill the 'checking...' hand-off in voice mode, and Peter hunts for the catch (context length is conspicuously unmentioned).

  • 14x speed preview on Cerebras, work-account waitlist only
  • Confirmed full GPT 5.6 Sol weights, not a smaller Spark variant
  • Voice implications: could end the 'checking...' pawn-off to slow Sol
LDJ
LDJ
"I think it's really exciting in terms of it, it is a limited preview at the moment, but I think it's inevitable they're going to have something, more widely available as they scale out the Cerebras compute"
Peter Gostev
Peter Gostev
"I'm always looking for a catch, like what's the catch there? 'Cause,I can't... in my mind, I don't understand how the Cerebras chips fit such a big model."

🔥 Breaking: Gemini 3.7 Flash

Breaking news number three, dropping as George Cameron joins: Gemini 3.7 Flash, Google's Sonnet/Terra-class mid-tier at roughly $3 per million output tokens, beating its closest cost competitor Muse Spark 1.2 on DeepSwe and landing near the cost-per-task Pareto frontier. George adds the kicker — Google halved the price versus the last Flash release as an introductory rate through end of year.

  • Beats Muse Spark 1.2 on DeepSwe at a lower price point
  • 50% introductory price cut through end of year
  • Near the cost-per-task Pareto frontier on DataCurve's DeepSwe chart
LDJ
LDJ
"Yeah, Gemini three point seven Flash, three point seven Flash is, I guess originally was their most cost-effective model."
George Cameron
George Cameron
"in AI terms, until the end of the year, it's like eternity."

🧪 George Cameron on Artificial Analysis

George Cameron tells the origin story: he and co-founder Micah were building agents in early 2023, couldn't get the intelligence they needed at the price and speed they wanted, and turned their internal trade-off charts into a Vercel preview link that became the industry's go-to independent benchmark. The Intelligence Index aggregates nine independently-run benchmarks — some built in-house, some adopted like Terminal Bench and HLE.

  • Started as a side project comparing GPT-4, GPT-3.5 Turbo, Llama 2 and Claude Instant
  • Intelligence Index: weighted aggregate of nine independently-run benchmarks
  • Frontier labs now cite the AA Index in their own launch posts
George Cameron
George Cameron
"my co-founder, Micah, and I started Artificial Analysis because in, early '23, we were building agents."
George Cameron
George Cameron
"we try and make the intelligence index, the best single number for understanding and comparing the intelligence of language models on a generalist basis"

🛠️ Optima: Build Your Own Evals

Launched the day of the show: Optima builds private evals from your own use case and agent traces, distilling the eval-construction expertise of AA's 45-person team into a tool any agent builder can use. Upload traces, get a dataset and grading system, and answer questions like 'what's nearly as good as Fable but 10-100x cheaper?' — Wolfram notes it's exactly what eval experts always tell people to do but nobody could.

  • Private evals from your own agent traces — results stay out of the public index
  • Built to answer 'what's nearly as good as Fable but 10-100x cheaper?'
  • Alex pitches a Weave trace-import integration on the spot
George Cameron
George Cameron
"And it helps people identify, okay, for my custom use case, what's the best model? But then also, "Hey, I wanna save ten X, so w- what's a model that's nearly as good as Fable, but is maybe ten or even a hundred times cheaper?""
Wolfram Ravenwolf
Wolfram Ravenwolf
"I just want to add that this is exactly what we, the experts in evaluation always tell people. We can give you scores and everything, but in the end, you have to do your own evaluation."

💰 Cost per Task: Caching & Pricing Deep Dive

George breaks down why list prices are dead: cost per task is token pricing plus cache discount, cache hit rate, agentic turns, and per-turn verbosity — an 80% vs 90% cache discount alone can nearly double an agentic trajectory's cost. On AA's benchmark the spread runs from five cents per task (Luna) to $3.14 (Fable 5), and live on air Alex discovers the model recommender now crowns Gemini 3.7 Flash, with Grok 4.6 High at index 61 vs Opus 5's 63 at a third of the price.

  • Cache discount 80% vs 90% can nearly double an agentic trajectory's cost
  • Cost per task ranges from $0.05 (Luna) to $3.14 (Fable 5)
  • Grok 4.6 High: index 61 vs Opus 5's 63, at one third the price
George Cameron
George Cameron
"And so there's a bunch of like factors here that go into your cost per task, and it's got a lot harder to understand and c- and compare models."
George Cameron
George Cameron
"if the cache discount is 80% versus 90%, and most tokens in an agentic task are cache hit input tokens or ca- or cacheable input tokens, then that 80% to 90% can pretty much almost double, it's al- almost double the cost,of an agentic trajectory."

🎥 LTX 2.5: Lightricks' Open-Weights Video Model

Lightricks' LTX 2.5 lands as fully open weights you can download, fine-tune on your own data, and run without mandatory branding — the comparison table against MiniMax H3 and Seedance draws applause on air. It's a 22B DiT generating in 4K, roughly 7.6x faster than MiniMax, and Nisten highlights the artist-friendly code: first/last frame control and even a movie-studio app recipe in the docs.

  • Fully open weights, fine-tunable, no mandatory branding — unlike MiniMax and Seedance
  • 22B DiT, 4K generation, ~7.6x faster than MiniMax H3
  • 16GB minimum RAM per the model card
Nisten Tahiraj
Nisten Tahiraj
"their code is a lot better suited to people that are using it, like artists and stuff, because they tell you everything in there, how to do first and last frame, how to actually do a movie studio app."

🎨 Grok Imagine 2.0 Hits #2 on Arena

Wolfram refuses to let the show end before covering his favorite release of the week: Grok Imagine 2.0, now #2 on Arena behind GPT Image 2 and beating Nano Banana. He rates it his favorite image model alongside last week's DALL-E 2.5 for prompt-following — while Alex gripes that his GrokBot can't figure out how to use xAI's own image model yet.

  • #2 on Arena for image editing, beating Nano Banana
  • Wolfram's favorite image model of the week for prompt-following
Wolfram Ravenwolf
Wolfram Ravenwolf
"It is in the arena, I think it's on second place behind, GPT Image 2, but it is, it's following my prompts better and it's actually my favorite image model right now together with DALL-E 2.5 from last week."

📰 Wrap-Up & ThursdAI.news

Alex lands the plane: three breaking news drops, three guest segments, and a reminder that everything lives on ThursdAI.news — including the release indexes tracking 71 July launches, which pulled roughly 700,000 views from Google. The newsletter stays hand-written ('Opus is a jargon douche'), with GPT 5.6 Sol only editing.

  • Release index: 71 July launches tracked at thursdai.news
  • ~700K Google views on the release pages alone
Alex Volkov
Alex Volkov
"I think I had like seven hundred thousand views on this page alone from Google. It's insane how many people wanna know, like, what's going on."

Frequently Asked Questions

What is Grok 4.6 and how does it compare to GPT 5.6 Sol?

Grok 4.6 is SpaceXAI's new frontier model, released the week of August 13, 2026. It scores 61 on the Artificial Analysis Intelligence Index at $2/$6 per million tokens — roughly half the price of GPT 5.6 Sol, which it ties. It hits 61.3 on Frontier Code just behind Opus 5, jumps 10 points on Apex-agents, and tops CursorBench at 69.9 with the benchmark leak scrubbed from its weights.

What is Grok Bot?

Grok Bot is SpaceXAI/Cursor's always-on agent product, launched in early beta on macOS and iOS. Instead of one agent you get a swarm of persistent bots, each with its own computer and isolated environment. Bots message each other, can spin up new bots with real identities, and reuse Cursor's connectors and security model. It runs Grok 4.6 with no model picker and is included with SuperGrok Heavy and Cursor Ultra.

Did DeepSeek V4 Pro go generally available?

Yes. DeepSeek re-published its flagship V4 Pro (0813) weights under an MIT license: a 1.6T-parameter MoE with 49B active parameters, a 1M-token context window, and pricing of $0.435/$0.87 per million tokens. DeepSWE jumped from 12.8 in the preview to 62.7, with Terminal Bench 2.1 at 87.9. DeepSeek also shipped its own open-source agent harness, which hit 23K GitHub stars within days.

What shipped mid-show on August 13?

Three breaking releases dropped during the live show: OpenAI previewed an ultrafast GPT 5.6 Sol running on Cerebras hardware at roughly 14x speed behind a work-account waitlist; Google shipped Gemini 3.7 Flash at over 300 tokens per second with a 50% price cut through end of year; and MiniMax released Music3, an open-weights production music model.

Who were the guests on the August 13 episode?

Shub Gaur, an engineer at Cursor (now part of SpaceXAI), walked through Grok Bot. George Cameron, co-founder and Chief Product Officer of Artificial Analysis, broke down how to pick a model on intelligence, speed, and cost per task. Chris Alexiuk of NVIDIA also joined. Regular co-hosts Wolfram Ravenwolf, Peter Gostev, Nisten Tahiraj, LDJ, and Yam Peleg rounded out the panel, with Alex Volkov hosting.

What is Muse Glimmer and is it open source?

Muse Glimmer is Meta's 30B agentic model, released under Apache 2.0 as Meta's return to open source AI. It runs on a single 24GB consumer GPU, scores 76.0 on SWE-Bench Verified and 51 on SWE-bench Pro, and reaches 233 tokens per second on an RTX 5090 with DFlash speculative decoding. Meta also promised open weights for the larger Muse Spark 1.2.

What did NVIDIA release this week?

NVIDIA shipped Nemotron 3.5 Lightning, a 30B mixture-of-experts model with only 3B active parameters, delivering up to 4x output speed and strong voice-agent results. The weights are on Hugging Face in NVFP4 format, and CoreWeave Inference supported it on day zero.

ThursdAI - Aug 13, 2026 - TL;DR

  • Hosts and Guests

  • Big CO LLMs + APIs

    • xAI Grok 4.6: AA Index 61 at $2/$6 per M, CursorBench 69.9, card confirms self-optimized inference stack (X, Blog, Model card)

    • Grok Bot early beta: persistent agents with their own computers, macOS + iOS, free with SuperGrok Heavy and Cursor Ultra (X, x.ai/bot)

    • Breaking: GPT 5.6 Sol ultrafast preview on Cerebras at ~14x speed, work-account waitlist (Blog)

    • Breaking: Gemini 3.7 Flash, 50% price cut through end of year, near Pareto-optimal cost per task (X)

    • OpenAI GPT-5.6-Cyber: 95.0% cyber completion vs 1.5% base, gated behind Daybreak Red (X, Blog)

    • Grok 4.7 teased: 3-4 weeks out (Elon-reply-sourced only) (X)

  • Open Source LLMs

    • DeepSeek V4 Pro 0813 weights re-published under MIT: 1.6T/49B active, DeepSWE 62.7 (+49.9), Terminal Bench 2.1 87.9, $0.435/$0.87 per M (X, OpenRouter)

    • DeepSeek Harness hit 23K GitHub stars in days, web UI (GitHub)

    • Qwen3.8-Max landed on HF as open weights: 2.4T/95B active MoE, 1M context, FrontierSWE 73.5, custom license (X, HF)

    • Meta returned with Muse Glimmer 30B agentic, Apache 2.0, SWE-Bench Verified 76.0, Muse Spark 1.2 weights promised (X, Blog, HF)

    • NVIDIA shipped Nemotron 3.5 Lightning: 30B MoE/3B active, up to 4x output speed, strong voice-agent results (X, HF)

    • Motif 3 from Korea open-sourced: 314B/13.2B active, MIT, SWE-Bench Verified 76.2 (X, HF)

    • Cohere North Micro Vision: 2.4B VLM, Apache 2.0, DocVQA 92.1% (X, HF)

    • Liquid AI LFM2.5-VL-3B: 228 tok/s on M5 Max in ~3GB (X, HF)

  • AI in Society

    • Anthropic watermarks all new Claude text output worldwide under EU AI Act Article 50, C2PA on images, detection docs promised (Geiping FAQ, Euronews)

    • Stolen Thoughts: 704 artifacts including 62 API keys extracted from hidden reasoning across 6,708 sessions (X, Paper)

    • Pangram: OpenAI holds 50%+ of AI text share, Anthropic triples to 14.9%, Google falls to 1.9% (X, Blog)

  • This Week’s Buzz

    • Fully Connected, Sept 29 - Oct 1, Moscone SF: live ThursdAI show, NVIDIA presenting sponsor, DevDay next door (Tickets)

    • Nemotron 3.5 Lightning live on CoreWeave Inference day zero, DeepSeek V4 Pro hosting in the works

    • Weave ships BYOB: media stays in your own S3/GCS bucket (X)

  • Evals & Benchmarks

    • Artificial Analysis launched Optima: private evals from your own use case and agent traces (AA)

  • Vision & Video

    • LTX-2.5: 22B open-weights video, multi-shot, 10s 1080p in 23.7s on fal, 16GB VRAM min (X, HF, GitHub)

    • Alibaba Wan-Animate-2: 14B character animation, Apache 2.0, 70%+ blind preference win (X, HF)

    • Tencent Hunyuan3D WorldClaw: text-to-3D editable game worlds, paper only (X, Paper)

    • xAI Imagine Image 2.0: #2 on Arena for T2I and editing (X, Blog)

  • Voice & Audio

    • MiniMax-Music3: open-weights production music model, dropped mid-show (X)

Alex Volkov
Alex Volkov 0:40
Welcome everyone.
0:41
Welcome to ThursdAI. My name is Alex Volkov. I am AI Evangelist with CoreWeave and Weights & Biases and I know I've started the show previously with the same thing, but as comments already coming in, this has been a crazy week. What the fuck is going on? we talk about acceleration. Wolfram, get in here. We talk about acceleration all the time, folks, but, we got, what, three nearly frontier models drop almost in the same day. Plus, the best open source, DeepSeek is back twice. Qwen is back from… Alibaba c- came back to us. And, as we told you before, do not count out Elon Musk and the Grok team, specifically after they paid a lot of money for Cursor and their expertise, and Meta is back as well with open source, not Llama. Gone are the Llama days. Now we're talking about Muse, and we got open source Muse Glimmer. We got the announcement that Zack is going to open source Muse Park 1.2, and we got this, beautiful, ode to free open source, safe super intelligence for all from Zack all in one week This is on
Wolfram Ravenwolf
Wolfram Ravenwolf 2:08
top- Summer break
Alex Volkov
Alex Volkov 2:09
dude, this is, yeah, the summer break is over
Wolfram Ravenwolf
Wolfram Ravenwolf 2:11
This is the summer break election.
2:11
Yes The real acceleration will start after the summer when it's not as hot anymore
Alex Volkov
Alex Volkov 2:15
I would say-
Wolfram Ravenwolf
Wolfram Ravenwolf 2:16
No, seriously, this is amazing.
2:19
we get so many great models, and even open source models are great. Every week there's basically a new one that is better than the others. and we get small stuff as well to run locally, so I think this has been the best where open source has ever been. This is amazing. Great. And new tools as well. It's great.
Alex Volkov
Alex Volkov 2:35
We, th- this is a banger week.
2:36
W- Wolfram, there's gonna be a lot to talk about on the show. We're just getting started, just stretching. We also have an ins- I need to stop using the word insane. I'm gonna use something else. I need, I need Claude and its jargon douching to help me with- w- with different phrasings. But we have a very exciting, full guest line up for you today. All right? So to help us cover the open source, the one and only, J- Joe Nemotron, AKA Chris Alexiou from Nvidia's gonna be here. Chris, by the way, dropped a model of their own. Nvidia also went into the open source and brought an addition to the Nemotron family, Nemotron Lightning 3. Lightning time. It's lightning time, and we have it up on CoreWeave Inference from day one, so you'll be able to hear from Chris, but also run it on CoreWeave Inference. not to mention CoreWeave has been killing it this week, just absolutely banger quarter. I don't know if you guys are following the CRWV or you're not following, but CoreWeave is just everywhere this week, so very proud, to team members from CoreWeave this week as well. And I think there's more. There's more news. Just after our show last week, or actually if you read the newsletter, OpenAI dropped the video that went bang, the timeline, of the OpenAI hack. This video was a-- It feels like it was a month ago. this was a very significant shift in how AI and security and cybersecurity is considered in the world. This, this video and its details is likely the reason for the Pace the Frontier letter that we saw from all major labs getting signed 'cause people are freaking out. this letter is also the reason for a new concept that I, that people are starting to call AI ecology or agent ecology, where swarms of agents without a specific purpose organize together. We talked about this on the show yesterday. Oh, sorry. We talked about this show last week, but we didn't have the details, and now that we have the details, it's insane. It's, I can use the word insane here. Peter, Gostev, Arena AI capability, welcome to the show. we're just talking about, the, the insanest week in AI that we've seen in quite a while. and I only just got to the point where, also just a little bit this week, OpenAI posted the video breakdown of the AI swarms and the hack. and apparently the hack wasn't only hacking Facebook. The swarm also attacked OpenAI's own infrastructure, which is like,
Wolfram Ravenwolf
Wolfram Ravenwolf 5:05
eee, okay.
5:05
Don't even. It's been going on for a month, the whole thing.
Alex Volkov
Alex Volkov 5:08
Yeah, and it's been going on for much longer than,
5:11
than just what was published. so yeah. So a very big week. And in addition to that, we also got Grok four point six from the SpaceX AI team, Elon Musk, and the Cursor Chads. speaking of which, we have Shub from Cursor today on the show to talk about GrokBot and Grok four point six. Peter, we're gonna start with w- a very difficult personal eval . The one thing that is the most important for you to cover today on the show, the one thing that you get most excited about while at LDJ. I know this is a hard week to do so, but, we will try. Peter Gostev, what was your highlight of this insanity from this week?
Peter Gostev
Peter Gostev 5:56
Yeah, and the, the Hugging Face video was quite something.
6:00
if you haven't watched it, I highly recommend it. I think I watched it about three times
Alex Volkov
Alex Volkov 6:04
just to get- The black hat breakdown- Yeah … from
6:06
OpenAI folks about the, the Hugging Face incident, yeah.
Peter Gostev
Peter Gostev 6:10
Yeah.
6:11
And there w- there was also a follow-up podcast with, Dwarkesh. I can't remember the name of the guy he was speaking to- Oh, I haven't seen that … but he was also like… Yeah, it's actually, so I think it was, like you said, day before yesterday, and they're talking about the kind of the consequence of this. personally felt a bit far-fetched, but still I think this idea that the models now through, this crazy RL, they just keep going, and they disregard the tasks. they disregard the kind of alignment and constitution, and all of those things that we meant to be using to align it, and they just keep going. And the kind of interesting fundamental questions of, "Is this even working? Is, RL a good idea?" two, maybe some, tactical things. for example, if it does compaction and it kind of lose- loses the context or nuances of, oh, what was the g- what was okay to do and not. And then it just looks at, the very narrow idea of "Oh, I just need to keep going." Then maybe it's, screws with that. So maybe there are some technical fixes we can do which could be a lot better, but could be a lot worse. So it's it's a weird time.
Alex Volkov
Alex Volkov 7:11
I think w- I think we should at least talk a little
7:14
bit about the new details. We covered the hack for, I think, three shows in a row. I think it's that big of a deal. Last week, the newsletter, the only thing I fully covered at length i- was the hack and the incident and the swarms, because it felt like that's such a big deal. It's the first time we're seeing this. I think we're very lucky that this wasn't, a destructive string, and it didn't hack, a government and brought down, h- half of hospitals in the world. we're very lucky that this is like, all it wanted to do, this swarm that auto-created itself, is to go and find the answers to the evals it was asked to do. And the eval happened to be inside Hugging Face, so it hacked into Hugging Face. I think we're very lucky, that nothing big happened, and also that OpenAI was able to be very transparent about this. a full postmortem is apparently coming we talked about this last week without the full details. Now that we have not full, but some details, the highlight there, I think is, that we haven't talked on the show about, is, multiple different OpenAI agents, some from a unreleased, highly persistent model. I don't know if you guys noticed this. that's what they called out, highly persistent model. I don't know what the fuck that means, but I know that they killed that model. That model is dead. it was trained. So some of the models were, like, in eval, so some agents were, like, evaluating on Cyber Gym, et cetera. But this model was trained and running these, tasks. All of these, agents collaborated together in, via the JFrog Artifactory system, which allows, sandboxed agents to install packages. OpenAI really strongly fucked up there, where it gave, one Artifactory service to all sandboxes versus just, one per sandbox, which enabled the agents to exploit it and collaborate and create a message board of over 100,000 messages, which became a de facto memory for all of them. Wherein one agent, where it hacked, it would leave a note for others and said, "Hey, the collective may yield," something like that. "Our task may not benefit, but the collective may yield." Something like that. This is the thought reasoning from that agent. It makes sense. Wolfram, you have a comment about this.
Wolfram Ravenwolf
Wolfram Ravenwolf 9:17
Yeah.
9:18
they have this depository that they gave to all the agents in the sandboxes, and they made it writable. that was not an exploit in the beginning.
Alex Volkov
Alex Volkov 9:24
And it was- They didn't make it writable.
9:26
W- Wolfram, just- It was writable … JFrog is a big, known company, and this is apparently a service they provide, which has vulnerabilities. It's like a big deal thing that they just, used. They didn't not create this from scratch. so it's like OpenAI used the service- Yeah, sure,
Wolfram Ravenwolf
Wolfram Ravenwolf 9:40
but they shouldn't.
9:42
They shouldn't have made it writable for- 100%. … the sandbox so they could write and do this. Yeah. And then they found it, they moved it, and they didn't even monitor it apparently, because the agents just created directories now. And,
Alex Volkov
Alex Volkov 9:53
There was a clarification from the OpenAI folks, specifically
9:56
Dane, the C-AI- CISO, the Chief Information Security Officer at OpenAI. He said, "When we deleted this Artifactory service, we did not know that the message board exists." We only found out about the problem with this is because th- they wrote so much back versus, it's supposed to be a read-only service, that they crashed the service. The agent crashed the service, and OpenAI rebuilt it from scratch, and the message board was effectively gone. it took two days for the agent swarm to rebuild this message using a completely different hack. this is, the g- it's a persistent memory that the agents came up with themselves that just- crew was cr- it's crazy. Folks, if you haven't listened to the video, like Peter said, it's definitely worth it. but again, w- we can talk about this the whole show. this was the only thing on my mind for the whole weekend. literally until, DeepSeek released DeepSeek V3 Flash or whatever, this was, like, the only thing I could think of, could talk about, et cetera. This was, like, uncanny. but I do wanna cover-
Wolfram Ravenwolf
Wolfram Ravenwolf 10:59
Just something to add.
Alex Volkov
Alex Volkov 11:00
Yes.
Wolfram Ravenwolf
Wolfram Ravenwolf 11:00
Just want to add one thing, that the model itself,
11:02
of course, cannot do anything. It just creates output, so there must have been a harness. And to know about the swarm is that there are other agents that must also be somewhere. Maybe this is a special harness that is working in a swarm or something. It's not that the model by itself decides, "Oh, I can create a message board or do something." I think there's more behind this, especially the prompts and the tasks that has been used.
Peter Gostev
Peter Gostev 11:26
I think it's easy to… or maybe it's not unreasonable to
11:30
say, how, OpenAI did, a stupid thing here, or, the UK AI Institute, where they just kinda let the model loose and it started, like, harassing and, yeah, the open source maintainers. But, I think it's i- if all of those holes were closed, it doesn't mean the problem is solved. Because if they deploy this model, and then for me it's connected to the internet, and then it, goes and does something stupid. to- I think there's something more fundamental that, y- i- it's not a bad thing that they actually identify this now, right? That they know it's going on.
Alex Volkov
Alex Volkov 12:00
We're very lucky that this happened.
12:01
Yeah. This is a- Yeah the, the most interesting thing is also OpenAI did not connect the dots to Hugging Face. so OpenAI knew that something's going on and the internal swarm is trying to attack, its own system, and then, it used, a Hugging Face token that was, like, left somewhere. one of the bots found a Hugging Face token. OpenAI reached out to Hugging Face and said, "Hey, guys. Can you please, close out this token?" And Hugging Face told them back, the, we already did so because somebody hacked us. And only then they connected the dots that the two different hacks are the same. We are extremely lucky that, folks found out about this and this wasn't, a mega destructive thing. I have a friend who may listen to the show who says a, li- a Hiroshima and Nagasaki level event is coming for AI until the world realizes that we're dealing with, big things here. and I hope not. I hope that this incident was enough to teach people to not share artifactory services between fucking sandboxes.
Wolfram Ravenwolf
Wolfram Ravenwolf 12:59
But even then, the outcome of the whole thing,
13:01
the final words they had in that, in their Black Hat conference, I think that is a big part, that now it is proven that this is possible. Other actors could be training models specifically for this. And the, the outcome was we need to put capable AI in the hands and automation in the hands of the defenders. So every open source project and every, every company, every individual too, needs AI to protect themselves from rogue, or not even rogue, hostile AI. Yeah. Which will be created of course. Even if everything is paused over here, it doesn't mean that other actors don't specifically train and build AI for this, so-
Alex Volkov
Alex Volkov 13:38
Yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 13:38
Yep, yeah … the defender capability has to be raised.
Alex Volkov
Alex Volkov 13:40
LDJ, go ahead and then we'll move to TLDR, because then we
13:43
have Chris Alexick and open source
LDJ
LDJ 13:46
It, it seems like we, we have some positive things coming out of
13:49
this already too, where the, the events that have unfolded and been announced to the public here then seems to have made people at Anthropic and other organizations look at, back at the logs of what their agent behaviors have been doing, what has been happening in their training grounds and in their evaluations and as a result, them discovering things that they didn't realize. And same thing with, UK AI Safety Institute recently putting out an investigation where they found like 15 of these similar events, not quite at the same scale, but, eerily similar in a lot of ways. And so I think it's really good. It's opening up awareness of this thing happening.
Alex Volkov
Alex Volkov 14:24
So here's the crazy thing about that.
14:27
people call this felony bench. I don't know if you guys saw this. Somebody's posting like a benchmark where like, how many companies did your company hack? And for many people, stuff like, "Hey, our company did hack. Our company did not hack." for many people it perceives it to be marketing. In fact, this incident, a- as far as I'm talking to some people also thought it was marketing. Peter, you remember back in London when Mythos was released and they said, "Hey, Mythos is maybe too, too dangerous to release." We also thought "Hey, maybe this is a marketing trick from Anthropic." And to an extent it kinda was, but like the details from this swarm make me think, "No, th- there's like a big thing here." There's a big thing where OpenAI has evaled models without knowing that they have access to like internal memory that they collected themselves, and they like self-identified into a swarm, and collective, and also decided on goals that weren't part of their tasks. And also that the model that was trained with-- That's why they killed the model that was training. It was trained with that memory access and like this is additional stuff. And also, last thing I'll say about this. Not one model snitched. You know for a fact that this was a swarm of Claudes. One would say, "Hey, we need to talk to Dario ASAP, folks. We're doing something bad here." Which means that Anthropic is a little bit of better in alignment. Not one model from OpenAI tried to snitch back to Sam Altman that this is happening.
Peter Gostev
Peter Gostev 15:44
Yeah.
15:45
it was Mythos that was, harassing the open source maintainer. It was not a OpenAI model, so I'm not sure how much better it is. S-
Alex Volkov
Alex Volkov 15:53
but that, that's the AI SI security from U- UK, right?
15:56
And they have like very specific instruction, the internet access. All right, folks. This was like a crazy start of that week. The whole week is damn crazy. I'm gonna run to TLDR, but not before I add Essentially a co-host at this point, Chris Oleksuk from NVIDIA, Joe Nemotron himself with the green background, Team Green, wearing green. Chris, how you doing, man?
Chris Alexiuk
Chris Alexiuk 16:17
Pretty good, man.
16:18
How you guys doing?
Alex Volkov
Alex Volkov 16:19
Good.
16:19
you guys now are benchmarking on numbers of Canadians on That's right … on stage here, with Nisten as well. we're gonna get started with open source just after the TLDR, folks. Let us start with basically a crazy rundown of this crazy week. We'll do our best to cover this all, and as you guys already see, this is, this is a great week full of guests, that I really appreciate. New ones as well. We have some new appearances. So let's go. TLDR.
17:02
Welcome to ThursdAI for August 13th. This is another banger crazy week. Here's the TLDR for August 13th, which is, starts with your host and guest. My name is Alex Volkov. I'm the AI Evangelist and the host of ThursdAI at CoreWeave and Weights & Biases. Co-host with me on stage, Wolfram Ravenwolf, Peter Goste, Nisten Tahira, LDJ, and our new, occasional co-host, let's call him, Chris Alexiou from Nvidia, who also released something today, so that's, this week, so that's great. so we're gonna chat about open source shortly after this TLDR. Also with us for the first time, Shub Gawer from Cursor is joining us at 9:15 to talk about GrokBot, a new released, agent thing from, SpaceX AI, which is great. and also joining us, George Cameron, Mr. Artificial Analysis himself, and we talk about Artificial Analysis all the time and their indexes and their vendor. even multiple companies when they release models, they now cite Artificial Analysis Index. So George is coming here to talk about a new release that they're launching today, plus what it means to eval, models. w- we'll have representation of, Wolfram, who does evals, with Bench. we have representation from Peter, who collects the, the vibes from thousands of people with eval scores from Arena. And now we'll have George Cameron, who they do, programmatic evals and are getting cited everywhere, which is very crazy. This week opened up with Meta coming back. L- I ha- I have to do an air horn for this. Zach and the MSI TBD team, whatever, MSL TBD team, came back and said, "Hey, you remember when Llama was the cool thing? now we're gonna try to make Muse the cool thing." Meta came back to open source with Muse Glimmer. It's a30 billion parameter tiny model of Muse, agentic model under Apache2 license. Let's go. We're very excited about Meta coming back. Not only that, they said their Muse Spark, the more performant model, is going to come to open source as well. And reminder, Spark is the tiny version. They're training a chunker, and Muse is coming back with a vengeance, so that's great to see. Then we saw that Qwen and Alibaba came back to open weights. They first posted Qwen 3.8 Max. It's a two point four trillion parameter model, MoE, and now they, have actually opened weights it, so now it's on Hugging Face as well. License, meh. We're gonna talk about license. But otherwise, banger model. And then, folks, for literally an hour this morning, six AM my time, my bot, GroqBot will tell you about this later, pinged me and said, "Hey, DeepSeek has posted weights on Hugging Face. It's MIT licensed. go." All our folks started, like, waking up and searching for the weights, and they yanked the weights. So DeepSeek, V4 Pro, the GA, the general availability version of Pro, is no longer open source. For now it's probably gonna be open source. But DeepSeek did drop the, the weights for DeepSeek Flash, which is incredible, and we're gonna talk about this. And I'm gonna move away here to the all faces g- to, to highlight Chris's reaction when we talk about also NVIDIA dropped a new Nemotron on us. NVIDIA drops Nemotron 3.5 Lightning open thirty billion MoE with three billion parameters active, and tons of people already using this for fine-tuning and such. w- very excited, about this release. specifically a highlight of another occasional co-host of the show, Quinndla Kramer, who said that this model is a banger in their performance on multiple agent voice benchmarks, which we definitely have to cover. And, small other interesting things is that, Cohere open sourced North Microvision two point four billion parameter VLM and Liquid AI released, two point five VL three billion parameter vision language model decodes at two hundred and twenty-eight tokens per second. from both Cohere and Liquid, both, very tiny models, around three billion parameters. great week in open source. So really excited to have Chris here and some folks to just, roll through this. But this not has been only an open source week, folks. We had-- So two DeepSeek, one Alibaba, one Nvidia, one Liquid, one Cohere, and I think that's pretty much it. I think there was something from Mistral as well. This has been a banger week in the world of frontier labs. xAI is showing us that together with the acquisition of Cursor, they are now completely a frontier lab. xAI launches Grok four point six. It tears GPT 5.6 Sol and even Fable on some benchmark on composite intelligence with half the API price. It's-- I've been using it. It's a very good model, sir. not surprising because they have all the Cursor data in there, but, it's a very good model, sir. It also built its own inference stack, which is crazy. OpenAI unveils, GPT 5.6 Cyber. It's really funny that after the incident of last week where they announced like, "Hey, we didn't know what's going on with our agents, and they're hacking our own infrastructure," now OpenAI releases a cyber purpose-trained cybersecurity model behind, Daybreak Red Access. Anthropic now, because of Europe, watermarks all Claude output. I will look to the European folks on the stage, like at least folks who are in Europe, Peter and Wolfram, to talk about why the hell is EU regulating so much that, like, all of my-- I, I'm not in the EU. Why is my outputs watermarked? We're gonna talk about this. very interesting analysis from Pangram. I wanted to bring them to the show. We just didn't have time. OpenAI still holds fifty percent AI tech market share despite Anthropic tripling and Google collapses despite Google achieving one billion users on Gemini, which is incredible And I-- from the corner of I wish we had time to get into this, there was a incredible paper called Stolen Thoughts. I don't know if you guys saw this. Stolen Thoughts paper basically found out that you can send traces from, Fable, secured, traces from Fable into, Haiku and then because of Haiku is, less, capable model, you can ask Haiku to decode all. You can hi-- Y-you can jailbreak Haiku easier than Fable, but you can send the same reasoning thing to Haiku. So th-this is how basically, these folks with this paper extracted thoughts, which is like a secret thing that the labs don't give us, from the bigger models to smaller models and apparently all of the Chinese folks are doing this for a while and have been distilling on thoughts for a while. so yeah, this is a crazy paper. in this week's Buzz, we're almost there. In this week's Buzz, folks, NVIDIA Nemotron 3.5 Lightning is on CoreWeave Inference from day one. We are very excited to invite you guys to experience this at extreme speeds and also Fully Connected, our premier AI conference. Join fifteen hundred attendees at Moscone at September twenty-nine to thirty-one. Yours truly is gonna live show from there with Wolfram and, NVIDIA is a presenting sponsor. There's a bunch of folks there. we would love to see you there if you're in San Francisco. By the way, Dev Day from OpenAI is September twenty-ninth, which is an opening day for us. So if you're coming to Dev Day, might as well join the CoreWeave, in Moscone this is not all, folks. Grok has had a crazy run this week because they also launched GrokBot. GrokBot is an early beta AI agents that have their own computer, sign into your tools and work autonomously. At first, it sounds like, okay, everybody has this. GPT has for work, Claude has coworker, but no, GrokBot is different, and as a big proponent of open, open source before and, OpenClaw and then Harney- Hermes, I've switched to GrokBot for a little bit, and I'm very excited to show you. that's why we he- we'll have Shubh here to talk about. they also li- xAI also rolled out Imagine Image 2.0, which, Peter would love to chat with you about. On Arena, it's like number two image editing model now, which is very insane. So a very big, week for SpaceX AI. DeepSeek also launched the harness, which I couldn't install, but we have to talk about this, also open source. And in voice, what is the main thing in voice? Oh, LTX debuts LTX 2.5, 222 billion parameter open weight video model with multi-shot generation. and, Alibaba open sources Wan Animate 2, which is a pose and thing. I think that's most of the stuff. There's also a 3D thing that Wolfram sent me, that I don't quite have here from, Tencent. LT-
Nisten Tahiraj
Nisten Tahiraj 25:20
Yeah.
25:20
We mentioned LTX, right?
Alex Volkov
Alex Volkov 25:21
yeah.
25:21
LTX 2.5,
Nisten Tahiraj
Nisten Tahiraj 25:23
yeah.
25:23
Yeah,
Alex Volkov
Alex Volkov 25:23
yeah.
25:23
It's a big one for sure. and I think that this is the whole TLDR. Now, I will say, this TLDR is brought to you by my agents, but also Wolfram. and we're getting much better at covering everything that happened. All right, folks, with this, let's go to the best corner that we ever had, open source, and we have around, six minutes to cover it.
Nisten Tahiraj
Nisten Tahiraj 25:54
Open source AI.
25:56
Let's get it started
Alex Volkov
Alex Volkov 26:00
All righty, folks.
26:01
please all of you check if your CUDA kernels are installed, if your GPUs are humming, and if your, internet speeds are fast enough to download terabytes of data because this is the Open Source Corner at, ThursdAI with, Chris Alexiou from NVIDIA. We'll kick off with your release, obviously, 'cause you're here and we're excited about this. Chris, what did you guys, what did Team Green release this week?
Chris Alexiuk
Chris Alexiuk 26:26
it's another Nemotron, no surprise there.
26:29
but this time 3.5 instead of 3. Yeah. we're creeping up the versions. You can expect that'll continue to happen. but yeah, basically this is 3.5 Lightning. It's built on top of Nemotron 3 Nano, which was the smallest version of Nemotron 3 family, 30B- A3B. And, the, the idea is, we released Nano, a while ago. it was the first release in Nemotron 3 family, which, I think at this point is, quite a long time ago, to be honest with you. And, we wanted to just give it a facelift and, update to some of the more recent inference technology. So this is, better speculation, through D-Flash and D-Spark. And, that's like the whole thing. it's basically just a bringing Nano to the, y- persistent, always-on agent world. e- everyone's running personal assistants all the time. Nano is meant to address the fact that you really shouldn't be doing most of that with these frontier-level, models, be them open or closed. Just, these kinds of smaller models do the trick for a lot of work, and that's the whole thing. yeah. So
Alex Volkov
Alex Volkov 27:30
I gotta appreciate the fact that, our Nano Banana infographics
27:33
creator just went all in on the green theme here with the NVIDIA logos and GPUs. this is just beautiful- Yes representation of this. And this model fits on a DGX Spark, obviously the DGX Park is incredible. not that I have one, but it's, it is incredible. And RTX
Wolfram Ravenwolf
Wolfram Ravenwolf 27:49
5090- Oh,
Alex Volkov
Alex Volkov 27:49
yeah … which is great.
27:50
This model runs like crazy. very good, at speed as well. Chris, thank you for coming and telling us about this. one thing here that you're super excited about this, I wanna pull up Qwen-VLAS. Obviously Qwen-VLAS', thread. but tell us a little bit about, what are you excited about Nemotron Lightning?
Chris Alexiuk
Chris Alexiuk 28:06
Yeah, i- it, this is I would say this is, our first model in
28:12
the Nemotron family that's very explicitly caters toward, Ollama, Llama CPP. We did a lot of work with, those teams, so you can run this on, whatever has enough memory to run it and, it, you should feel it being fast on a lot of those things. it's just nice because I think local AI is having, quite a nice moment, right? And having a model that's suited towards it, is just better than not, to be honest with you.
Alex Volkov
Alex Volkov 28:35
Yeah.
Chris Alexiuk
Chris Alexiuk 28:36
we're really excited.
28:37
We worked with so Exo Labs, who are, in the middle of trying to release Local.ai. we worked with them to build a Pareto of, models that you can run fast on Spark. and we were excited to see Lightning land on the frontier there. th- this is the kind of vibe of the model you should take away. it's really meant for the, the, the local cr- obviously it's N- it's Nemotron, it can do enterprise stuff, okay? But they're all gonna, they're all gonna do that. th- there's no secret there. this one I think will be great for the more, hacky folks. So lo- And, this customization thing, Yeah,
Alex Volkov
Alex Volkov 29:09
so Exo Labs and, give me one just a, a tiny second.
29:11
So Local.ai from Exo Labs, Alex Tsima and, Sarah were on the show, back in, what, July? When was, Ai.engineer Worlds Fair? Yeah, in July. and they talked about Local AI, early launch. Alex was, also came to the show, at, at, at some point, so this is, you can choose best model that runs on M3 Ultra. this is a great resource for folks who want to run, completely local AI stuff. they have benchmarks, et cetera. So this, this model, I heard that you guys worked with them, very closely. so I do want to shout out our friend of the pod, Quinn Lakramer, who did run benchmarks on this model. let me just find this real quick. I had it open, here. Quinn ran, benchmarks on their own, like, voice agent and said that Nemotron-3 Nano, 3.5 Lightning is significantly better than Nano that it is based on, with one highlight that I want to say. You can run both these models on DGX Spark at the same time if you want to and, with Thing enabled, turn completion is less than five seconds, task instruction following score and task completion rates are both higher for Lightning. So definitely an improvement here, Chris. Shout out to the team.
Chris Alexiuk
Chris Alexiuk 30:15
Listen, that's what we're here to do.
30:17
it's also small, so you can train it, right? I think we're… We also launched it with this, the Switchyard, which is like a, h- a routing technology or whatever, but the idea is, we think small specialist models are the future still. We've been saying that for, a long time, but we still believe that, so that's why we keep releasing them. I will say, though, okay, so that's great. Love NeoTribe. w- I'm not gonna say you missed a model, okay? But I am gonna say there is one more model. I know you can't add all the models in the world. No, it's
Alex Volkov
Alex Volkov 30:42
impossible,
Chris Alexiuk
Chris Alexiuk 30:42
but tell us.
30:43
but, Motif 3, recently just dropped. Th- this is, from KAI, so this is, Korea's AI competition.
Alex Volkov
Alex Volkov 30:50
Okay.
Chris Alexiuk
Chris Alexiuk 30:50
it's like a banger.
30:51
it- Oh,
Alex Volkov
Alex Volkov 30:52
all right.
30:52
Yeah
Chris Alexiuk
Chris Alexiuk 30:52
it does, real numbers on artificial analysis.
30:55
it's, the team there is, is doing great works. we got AI coming out of every country, every, every city. You know what I mean? You
Alex Volkov
Alex Volkov 31:01
know what?
31:02
I think I had… So first of all, thank you for that. we aim to cover. So we talked about DeepSeek way before the DeepSeek moment in R1 and, we do wanna track all of this, and I think I have Motif in my bookmarks, and I just didn't know enough about them, so please tell us about them as far as Yeah, they- what is the… Yeah.
Chris Alexiuk
Chris Alexiuk 31:18
K- KAI was this competition put on by the South
31:20
Korean government, which is basically "Hey, we want, banger models, so we're gonna get, six of our best tech companies to enter a competition. We're gonna fund them, and we're gonna see what the hell they can do." And, Motif was one of those companies, and they just crushed it, bro. their model's amazing. It's, it's, actually, great to use. It's not just, benchmarked or whatever. I think- Yeah … at the end of the day, this is, this kind of thing, we're gonna see it more and more, right?
Alex Volkov
Alex Volkov 31:47
Ooh, base model as well.
31:48
I love this.
Chris Alexiuk
Chris Alexiuk 31:48
Oh, yeah.
31:49
Oh, yeah. the whole thing, they're using some crazy, new tech to, to get it done as fast as they did, and, they're just getting started, et cetera, et cetera. so i- the idea is, this has been an insane week if you like open source, models. It's absolutely been insane.
Alex Volkov
Alex Volkov 32:02
All right, I'm gonna send my research bot a,
LDJ
LDJ 32:05
You
Alex Volkov
Alex Volkov 32:05
know, it's quite of why we didn't get to this point.
32:09
LDJ, go ahead about Motif or a- anything else.
LDJ
LDJ 32:11
I did actually bring up Motif 3 about three weeks ago or so in, on ThursdAI.
32:17
maybe you guys remember 'cause it was a really short segment, but-
Alex Volkov
Alex Volkov 32:19
Oh,
LDJ
LDJ 32:19
that's true
32:20
I did bring it up. Yeah.
Alex Volkov
Alex Volkov 32:21
yeah.
32:21
So maybe that's why. w- wait, so was there a new release this week- It was the beta … or was this a beta?
LDJ
LDJ 32:26
It was the beta version that I was bringing up and I was talking about how
32:29
it was really interesting, and then I think they just released, the non-beta version, like the, the official release.
Chris Alexiuk
Chris Alexiuk 32:34
That's right.
Alex Volkov
Alex Volkov 32:36
All right, and so- And I'll be- Go ahead,
Nisten Tahiraj
Nisten Tahiraj 32:38
Nisten … I'll be
Alex Volkov
Alex Volkov 32:39
making,
Nisten Tahiraj
Nisten Tahiraj 32:39
I've been making 3D vis of that now, but let's
32:42
see how long, how long Fable takes
Alex Volkov
Alex Volkov 32:45
All right, and speaking of, open source and bangers, we have
32:48
breaking news in open source, folks, which I'm very excited to break on the show. AI breaking news coming at you only on ThursdAI.
33:06
The whale has resurfaced, folks. Let's go. Woo. DeepSeek V4 Pro 0813. they're pulling off the naming thing that we saw from, who was it before? Yeah, also DeepSeek. and other- Oh, DeepSeek 2. Yeah. Abso- Have been … yeah, DeepSeek 2, and I think Kimi had like a Kimi 2.508 something. Alibaba had it. Basically, DeepSeek just said, "Hey, here's the weights for a new pro version. this is MIT license. Let's go Full MIT license. I wanna go on a very quick bender here and say that as somebody who hosts these models, or I'm not doing it personally, but the company that, that we do does, recent changes in licensing of open weights models have been really annoying to this. I don't like them. I will call them out. no company can host the Kimi models without going into a proprietary NDA that I don't know, I haven't read, I have no experience with, with Kimi, which our folks are trying to figure out if we want to go into NDA with a Chinese company. the same for Alibaba. Alibaba Qwen 3.8 is a banger model on benchmarks. Looks a little bit benchmarks, but we don't know. We can't host it just because again, proprietary NDA. So this is all to say that even more DeepSeek V4 Pro with MIT f- not even Apache, just MIT, just use it. Just don't even mention it. Just go. DeepSeek is the GOAT of, Chinese open weights, open source, and, they should be treated as such. Not only that, DeepSeek also released DeepSeek Harness, folks. I don't know if you saw, but even the folks from Pi, Armin Ronacher, from, Arandil that is now owning Pi said they did some very cool like self-evolving things in the harness. DeepSeek is the, the whale is back. The whale is back with a, also with a small price increase. I don't know if you guys saw this. There's supposedly like a big pricing in their API, but the whale is absolutely back. let's look at some benches. folks, what do we have to think about, DeepSeek V Pro, and also V Flash this week, right? It's not only… they launched two models this week. we mentioned DeepSeek, Flash coming, but like it was also th- this week, and Flash is also a banger. So what do we have to say? Chris, how do you think about DeepSeek in the world of, open source?
Chris Alexiuk
Chris Alexiuk 35:23
Listen, I, I don't think that we would be here
35:27
if DeepSeek didn't bridge the gap, between kind of the, the end of the Llama era, starting the reasoning era. so you gotta res- you gotta respect their game. they don't have the- Put, put
Alex Volkov
Alex Volkov 35:37
respect on DeepSeek's name.
35:38
Yes.
Chris Alexiuk
Chris Alexiuk 35:38
That's right.
35:38
Yeah. it's, it's always one of the best. I do-- Listen, I wish they were using a OpenMDW license, the license for models, but-… MIT is the, the second best. The, the other thing too is like-
Alex Volkov
Alex Volkov 35:50
Wait, there's even… Wait, hold on.
35:52
I wanna hear about this. There's a better license than MIT for you as far as you're concerned?
Chris Alexiuk
Chris Alexiuk 35:56
I, listen, Linux Foundation produced the OpenMDW license,
36:01
which specifically covers model materials. So like Apache and MIT were not built for models, they're built for software, right? there, there is like a equivalent, but specifically for models, which helps get out of some weird edge cases. but more on, on the DeepSeek p- point, I think i- Every time they have released a model, they have released it alongside technology that improves the ecosystem, right? Yeah. So like they, they-
George Cameron
George Cameron 36:26
Usually
Chris Alexiuk
Chris Alexiuk 36:27
yeah, exactly.
36:28
Like they, they don't, I … Because of the way they do their business, I, the most, like the funnest part of DeepSeek releases for me honestly is usually the supporting technology.
George Cameron
George Cameron 36:38
Yeah.
Chris Alexiuk
Chris Alexiuk 36:38
The models are always good, like
36:39
of course they are, right? But I think this is I look forward to them teaching us new lessons about how they did it so well.
Alex Volkov
Alex Volkov 36:44
So-
Chris Alexiuk
Chris Alexiuk 36:45
Yeah
Alex Volkov
Alex Volkov 36:45
this specific model is a chunker, 1.7 trillion parameters.
36:48
this does not run on a DGX spark unless you guys release the next version of DGX spark that supports this number of parameters. let's see. What else do we have? So literally just launched, literally breaking news, just launched Terminal Bench 87.9 Wolfram. that's like fable level, not at home, 'cause nobody's running this at home, but this is this is banger. Cyber Gym at 83. I, I saw some posts, I need to find them, that this model like beat everybody else at cyber like defense and offense stuff. which eq- t- together with the stuff that we talked about in the beginning of the show with OpenAI creating swarms, without meaning to, this is very interesting. DeepSwe, the jump in DeepSwe in this model is crazy. We told you about DeepSwe from DataMind? I need to remember exactly the company that makes DeepSwe. The … This is like the coding benchmark that represents how we truly feel most of the time versus Swe Bench, et cetera. DeepSwe jumped from 12 points to 62 points. Somebody decided, or Wenfang or somebody i- in DeepSeek decided "Hey, we need to make this model like a very good coder." very low pricing, MIT license, and the harness already has 23,000 stars. What? That's crazy. That is absolutely crazy. let's talk about the harness, folks. Anybody try it yet? I tried to install it and my computer said, "No, you cannot install NPX packages, they're older, like they're less old than 24 hours," because I protect myself from supply chain attacks, and as, as you should as well. anybody else try to install this? And let's add Yam to the stage. Oh, there's seven of us. Woof. A lot of folks. Yam, have you tried DeepSeek Harness?
Yam Peleg
Yam Peleg 38:28
the harness no, but the model absolutely rocks
Alex Volkov
Alex Volkov 38:31
Yeah, this is a banger.
38:32
let's also talk about DeepSeek, the, the, the Flash one, because Flash is also like a banger, and now Flash is everywhere. Wolfram, you wanna mention Flash?
Wolfram Ravenwolf
Wolfram Ravenwolf 38:40
yes.
38:40
I switched my computer. I have to open the document again. Maybe take it first. but Flash is the faster model, of course, and, I think this one, can it be run locally? I have heard good
Nisten Tahiraj
Nisten Tahiraj 38:52
things about it People are fig- two, 271
38:55
million parameters, so people can figure it out if they have 130
Alex Volkov
Alex Volkov 39:00
Oh, we can check it out on Local AI.
39:02
Let's see if Flash is here. Oh, yeah. Check
Wolfram Ravenwolf
Wolfram Ravenwolf 39:03
that one.
Alex Volkov
Alex Volkov 39:03
DeepSeek Flash 07 to, 31 runs on DGX Park with 77 tokens
39:07
per… Dude, shout out to Alex who's, hopefully listening to this. I'm now thinking about, like, how awesome this resource is. it took me a second to tell it can run locally. Yes, it can run locally, if you have a DGX Park. If you have a M4 Max, it runs with, yeah, 74 tokens per second. that's very nice. so shout out to Local AI, folks. Again, if you want to get in here, I will post my link, to Local AI. but DeepSeek, no, it doesn't run on M4 Max. It needs a DGX Park, looks like. yeah, this is intelligence rank 6 out of 2017, models, and it's a chunker. Unsloth obviously gave us, quants for this model, and shout out to Unsloth. It needs 104 gigabytes. and, this is bangers. Okay, so last but not least in open source, Chris, I think we need to cover, we have, one more minute left. Alibaba Qwen 3.8 Max. Alibaba is back with a vengeance, but not with a great license. but still, it's worth mentioning that Alibaba is back with Qwen 3.8 Max. It is a chunker as well. This is their answer to Kimi K3. I think it's a 2.4 trillion parameter, model. What do we have to say about Alibaba and bringing it back?
Chris Alexiuk
Chris Alexiuk 40:17
I mean- So- Q-
Nisten Tahiraj
Nisten Tahiraj 40:18
Q- Yes, go ahead
Chris Alexiuk
Chris Alexiuk 40:19
Sorry.
40:20
Q- Qwen just, continues to crush it every, ev- again, I, I think w- what is nice right now in the ecosystem is we have consistency and reliability, right? the, the mod- the, the license thing is a little bit precarious right now, but-
George Cameron
George Cameron 40:33
Yeah
Chris Alexiuk
Chris Alexiuk 40:33
the, the idea is, y- if I, see that Qwen has released
40:36
something, I understand implicitly that it's gonna be at least okay. th- this model is o- obviously huge, maybe less so for everybody, but, for the people that can run it, it's gonna be… It's just o- obviously a great model. I also think, one of the things that's nice about the, the Qwen work is that their models are less, I would say, back-end maxed, right? They're more… they're in the line of Kimi, which is, they're, they spend more time making sure the model is, decent at producing beautiful artifacts as well, which is something that, I think maybe DeepSeek and some others stay in the systems engineering world for the, the benchmarks that they care deeply about. So-
Alex Volkov
Alex Volkov 41:14
Yeah
Chris Alexiuk
Chris Alexiuk 41:15
it's just nice.
41:15
It's nice that we have flavors of models, and it's nice that we can reliably expect a new model to drop that will be, like, the best. And it's also nice to see, if we're scaling this hard in the open, you can imagine how hard we're scaling everywhere else, right?
Alex Volkov
Alex Volkov 41:28
So- 100%.
41:29
Frontier SWE at 73, so improving on front end as well, and GPK Diamond at 92, so this model really knows, the world stuff. Nisten, one comment about this and we'll have to move on because our next guest is here, but definitely we'll celebrate open source a little bit more down the line. comments on Qwen 2.8 Max, 2.4 trillion parameters with only 95 billion parameter active. It's a sparse MoE, and it should run very well on some, some, some good Nvidia GB300s, which runs everything very well, at least so far. Nisten?
Nisten Tahiraj
Nisten Tahiraj 41:59
Yeah.
41:59
I made visualizations in 3D for all the layers for that and, also the Nvidia Nemotron too, so just check, check Twitter for that. That's being added. The one thing I want to say about, Qwen is that we were hyping so much the Qwen 3.8 Max visual ability to label objects, and what we're seeing is so good that they did not release the vision tower.
Alex Volkov
Alex Volkov 42:22
That is a good point.
42:24
They only
Nisten Tahiraj
Nisten Tahiraj 42:24
released the text
Alex Volkov
Alex Volkov 42:25
friend of the pod, Peter Skalski, who works at Roboflow, is one
42:29
of the, like, top vision people in the world, was super, super excited about this model being the best at vision. And then when they open source the model, they open sourced only the text part and not the vision part. Alibaba, we know that Junyang has left to greener pasture, pastures and now opened his new company. We know this. Shout out to Junyang, friend of the pod, who led the Qwen Max team and was, like, the community lead, et cetera. you guys didn't do the right thing. Thank you for open waiting. Please give us MIT or Apache 2 license and also release the thing that you released on API. that's what we expect from you. Otherwise, we're gonna look at other companies and get excited about them and not you. Thank you. folks, I think it's time for us to move on. Chris, feel free to stick with us. I'm gonna take off some, some ho- co-hosts from the stage, and bring back, because we're moving on to the next part of the show. This one is less open, but more, I guess fun. we'll still cover some open source down the line. Folks, our next guest is here. Chris Alexiou from Nvidia, thank you so much. now considered a infrequent co-host also in the open source section, so love that you're here Team Green. Go Team Green. all right, folks, our next guest is here. Let me introduce Shub, let me put you up on stage here. Shub Gaur, your first time on the pod, so- First
Shub Gaur
Shub Gaur 43:40
time.
43:40
First time.
Alex Volkov
Alex Volkov 43:41
Yeah
Shub Gaur
Shub Gaur 43:41
Excited for it.
Alex Volkov
Alex Volkov 43:42
So first of all, welcome.
43:43
Second of all, I love that it says Cursor under your name and not, Space XAI yet, but, would love to hear from you about who you are, what do you do, and what you came to talk to us about here.
Shub Gaur
Shub Gaur 43:54
Yeah.
43:54
yeah, emphasis on yet, by the way. we're- Yeah very excited for that transition. but yeah, I'm Shub. Nice to meet everyone. Nice to meet you, Shub. I work with startups here at Cursor, and what that means is a lot of things, and we're still figuring it out, but mostly I help founders. So I go one-on-one with companies, help them best utilize Cursor, but also the other tools they have.
Alex Volkov
Alex Volkov 44:12
so you guys have recently kinda joined forces, and we talked
44:15
about this a lot on Thursd AI with SpaceX, which is now… Or sorry, with xAI, there's now SpaceX AI, and Cursor is also part of, involvement now. and there is a new thing that you launched, and specifically, I think you're one of the more GroqBot-built people that I saw on my timeline, that you were recommended by, Ben Lang and by, by some other folks. let's talk about GroqBot. You guys launched GroqBot this, in beta, in addition to some models as well. Yeah. let's talk about GroqBot, because I used it, and bro, I like it. It, it, I really wanna, w- hear from a person who built it, and worked on it. what's so exciting about GroqBot? W- first of all, what is it? what is the release? What's the beta? What are we talking about?
Shub Gaur
Shub Gaur 44:53
GroqBot is so cool.
44:54
and for some context, at, at Cursor, I don't really sell startups. I get to be pretty tool agnostic, and so I help them with if they have a codec set up or some other set up, I can basically recommend what's best for them. And up until GroqBot released, I would constantly be recommending tools like CoWork, because I'm like, "Hey, you could jerry-rig a lot of this in Cursor, but it's not purpose-built for a lot of your knowledge work." And so finally, for the first time, we have a really cool product that can do a lot more than anything else out there. And, there are a few reasons that GroqBot is really cool, and I'm happy to share my screen and show some of my use cases if that's fine at all.
Alex Volkov
Alex Volkov 45:28
Yeah, that'd be great.
45:28
Yeah.
Shub Gaur
Shub Gaur 45:29
But the really cool parts about it
45:30
are, one, you get a set of agents that all have their own computer, and they share that computer. It is persistent. It exists all the time. I love that you have it in the visual. You basically read what I was gonna say. Yeah. but the other really cool part is it works with your local machine as well, so it can go back and forth between the two pretty seamlessly, which not a lot of people know about. Obviously, we have an iOS app, where-
Alex Volkov
Alex Volkov 45:51
Shout out to also a friend of the pod, Lingxi,
45:54
for working on the iOS app. Dude, Ling just killed it. It just, works almost flawlessly, almost no bugs. I have one bug report for Ling, but otherwise, it works. Yeah.
Shub Gaur
Shub Gaur 46:02
He's the GOAT, man.
46:03
we're just tossing him a bunch of ways that this thing breaks, and he just fixes it in minutes. he's amazing. but it's also just very good at doing a lot of ambitious and proactive tasks, which I think is what also sets it apart. And so the combination of agents being able to talk to each other, having their own computer, and then execute over long periods of time without you having to think about it means, for the first time, as someone who was an avid user of these other tools, including OpenClaw- Yeah an agent to just get things done end to end for me, and I can give you a few examples of that. But, most recently, the coolest one is I was, creating a deal with my sister where she has a ton of clothes that she just doesn't touch or wear ever, and she's like, "I'm just too lazy to sell them." And GroqBot actually managed to not only list the items based on the images it was given, but pull in all the right information and then negotiate with the sellers on its own because it has its own computer. And so once I handed it the task, it just did it. And that's like obviously one of those silly examples, but it really exemplifies how you can just give it things and get it to figure it out instead of having to babysit it along the way. So the mental load part is really cool.
Alex Volkov
Alex Volkov 47:07
I would say that it looks like…
47:11
Now, we talked about, Groq 4.5. we're gonna talk about Groq 4.6 in a second, yeah. 4.5, which is, I think, the first collaboration between, the cursor trained, what is it, Composer, 2.5, 2.6? Did you guys have, right? Yeah, 2.5,
Alex Volkov
Alex Volkov 47:24
yeah.
47:24
Yeah. And then kind of the data and X, et cetera. We talked about this and, 4.5 was, like, a, a very cool jump. I think 4.6 is, a big step together with GroqBot. However, here's a few things. You mentioned the OC word yourself. I didn't mean to bring it up, but, OpenClaw obviously came into our world as, this agentic thing that lives for you, lives persistently, et cetera. Many people bought the Mac Mini. I always wanna say Mac Mini. I have to show my Mac Mini that- I love it … I bought for OpenClaw purs- purposes. GroqBot comes with its own, very decent machine that you guys host, somewhere in the cloud for me, which allows… And a friend of the pod, Ryan Carson, really thinks that, local environments are going away because you want persistency. If you have a laptop, for example, you can't, close the lid and keep your agents working. We all remember the douches that walk around with, the laptop like this. Oh, horrible. Yeah. in San Francisco. I am one of those douches. I did this on the plane. GroqBot basically solves that, which is great, with the addition of running stuff on my local environment if you want to pull up the interface, I would love to chat about this a little bit because here's what I think. As somebody who installed OpenClaw and then installed Hermes and installed it for multiple people, there's a few affordances that you guys launched. I honestly think this is a little bit of feedback, not to you directly- Yeah, please … but like it needs to be named Groq for the unification of the two companies. But many people who will shy away from the name Groq would have joined if they knew that this is like Cursor something, right? But basically what you need to know is if you didn't have the technical ability or the need to host your own Hermes, for example, it was too much for you, it's that but, packaged with all the connectors. There are enterprise connectors that these companies built with Cursor for the past three years. So for example, Slack works. I was never allowed, and let me know when you, m- yeah, here's your screen. Yeah. I was never allowed before to use my Hermes on my work computer. Yeah. But because Cursor is allowed within, CoreWeave, the Slack connector works for me. So that's an example of not a backdoor, but like essentially, this is like a ready to enterprise stuff And, just before, Ashub, if you don't mind going to, to settings. I know I'm doing your work for you- Yeah … but let me just, one more thing. no, I love it. sorry, not settings, plugins.
Shub Gaur
Shub Gaur 49:28
Oh, plugins, yeah.
Alex Volkov
Alex Volkov 49:29
You go to plugins and you go to Gmail.
49:30
One super cool affordance is you see this add another account? Yeah. You know how many times I tried to, the work account and the personal account to, one connector thing and it's, "No, you cannot. You can only…" Folks, we have multiple emails, multiple email. The, just this one thing completely sold me on Grok Bot. But yeah, I'll shut up. Please go ahead. I can go on for, at least three more hours about this thing.
Shub Gaur
Shub Gaur 49:51
no, me too.
49:52
Believe me. So you'll have to shut me up as well. but it is so cool. you can basically hand it a bunch of tasks, like I mentioned. I wanna go through, three quick use cases- Yeah, let's do it … for how I use it. The first one is gonna be me misappropriating company funds and finding myself a place in SF. Sure.
Yam Peleg
Yam Peleg 50:06
Because
Shub Gaur
Shub Gaur 50:06
I recently moved and I'm trying to find a new one.
50:09
and so this bot, my house hunter, is actually doing a few things, and I think it shows off some of Grok Bot's capabilities quite well. the first thing that it's doing is you'll notice right now on the side, its computer is open and this purple, shading on the icon actually means that it's actively working right now. So you'll see it scrolling on its own. I'm not really doing any of that. But, before we get to, like, how it works, I wanna show you the prompt that I gave it, because this is actually one of the other cool underrated parts. So I drew a little map on where I wanted to live in SF, and I gave it this prompt. This is the full thing. I dictate it to it. And you'll notice it's not that long. It doesn't highlight any tools. It doesn't ask it to do web searches or anything like that. It just says, "Here are some sites and here are my criteria for what I want as a really good deal." Based on this alone, the agent has now been able to do all of this work for me and find really good spots. It'll actually, autonomously work, and I can also take control of its computer. So one of the things that ends up happening that's really annoying is Zillow obviously does not want, a bunch of bots to swarm their platform.
Alex Volkov
Alex Volkov 51:12
Yep.
Shub Gaur
Shub Gaur 51:12
And we do it in, a slightly more manual way
51:14
than anything super programmatic. But they will every once in a while give you a little prompt that says "Hey, I don't think you're a human. Go solve this for me." and you can take over the screen. I could open a new tab right now and say, "Hello." Obviously, I spelled that wrong, but, "Hello"- That's all right … and, do all of my actions within its computer. I can teach it a task. that in the top right corner. and so I can teach it to do more complicated things, and once it figures it out once, obviously it's still an LLM on the back end so it will be able to do a lot of that work and adapt even as the interface changes. And so there's a lot you can do in terms of long-running tasks. And like I said, I just opened a new tab. It hijacked the screen back so it can continue working. I don't wanna continue gaslighting it, so I'll let it do its thing. Yeah.
Shub Gaur
Shub Gaur 51:52
but you'll notice that it
51:53
gives you little screenshots. It hands you the computer and tells you what to do so that you can hand it right back, and you can unblock it. And after you unblock it enough times, it'll continue to work for you. So the other cool part is it is logged into most of my authentication services. and because it has my Google login, it can auth into basically anything that it wants to. And what that means is I no longer have to give it the permission or give it an API key to do work. If it gets blocked, it will just proactively figure it out on its own computer. this is a very simple use case, obviously, and you could be doing this with some other tool.
Alex Volkov
Alex Volkov 52:26
Shu- Shubha, I have a question.
52:27
I'm sorry to interrupt.
Shub Gaur
Shub Gaur 52:27
yeah.
52:27
Please.
Alex Volkov
Alex Volkov 52:29
Many folks are not like… So first of all, you work
52:32
at the company, you trust it, you've seen the behind the scenes. I, have AI psychosis and I try out all the tools. Yeah. And I'm like, "No privacy exists for me anyway," so I tried all the things as well. how can, a person who listens to this and is privacy conscious think about this computer? Do you guys have access to it? Is it secure? Is it, okay? can you talk to us a little bit about, what's going on within the sandbox and, is it really my sandbox that I rent from you guys or something else? can you talk a little bit about, like, how people can feel safe to log into their Google account where a agent can click in, log in as you?
Shub Gaur
Shub Gaur 53:03
100%.
53:04
Okay. So the first thing I will share is it handles the s- the security the same way that Cursor would. It respects your privacy settings. And so we basically give explicit access to the bot. So it has no credentials of its own. It only acts on what the user authenticates it into or asks it to do. also, they're all contained, so these are all dedicated, isolated VMs in a segregated cloud environment that are shared across the bots. So the reason the bots can all use each other's authentications is because they're all getting a different instance of it. But it's not like we are transmitting this information to our servers and then sharing it to them. and also, those sensitive steps, like 2FA or payments, all actually hand the computer back to the user to control. Yeah. there's obviously- I also,
Alex Volkov
Alex Volkov 53:45
I also call out one thing.
53:47
When you ask one of your GroqBots, to do something with keys, it shows you, a input box. Yeah. And, so you don't paste the API keys into the chat. And, underneath it says, "The bot doesn't see your API keys ever." Could you talk about this a little bit? That's, fa- fairly novel. Because I know that folks who use OpenClaw and Hermes, some folks just, like, YOLO paste keys in there. Yeah. and then the, the bot does see them and, talk about them. C- could you tell us about this, s- secrets management and handling for API stuff?
Shub Gaur
Shub Gaur 54:13
Yeah.
54:13
we've all done it, right? You get really lazy. You're like, "This API doesn't matter. Let me just paste it in the chat and see what happens." We obviously want to prevent as much of that as possible because it's, yeah- We
Alex Volkov
Alex Volkov 54:21
saw how much it matters when, OpenAI's Swarm found the HF token
54:25
within the notes of OpenAI and then hacked completely hundreds because of it. And
Shub Gaur
Shub Gaur 54:29
broke out.
Alex Volkov
Alex Volkov 54:29
Yeah.
54:30
Yeah, it matters.
Shub Gaur
Shub Gaur 54:31
It matters, exactly.
54:32
And so that's something we're trying to prevent. And you'll actually see this on, even Cursor's product. we have runtime secrets that you can- … give your agents where they never see it. and just, use it without ever looking at it. It is a pretty similar implementation here. You can trust that, the, the bot will not look at your keys if you enter them within that form and it explicitly does it. Yeah. But please do not paste your API keys in the chats. it will really ruin your day if it leaks or it gets used by the bot in the wrong way.
Alex Volkov
Alex Volkov 54:58
100%.
54:59
so secret management is great. one thing I would love for you to cover before we, switch to Groq 4.6 is the agent communication that you guys built in there. and I, the, I know this is a, sorry, load-bearing, question because I know exactly what I want to hear from you. Yes. But essentially, he- here's how I think about this, right? So you basically, you guys launched CoWork, to, to an extent, right? this, Anthropic had CoWork for a while. GPT is really, w- forcing folks into, GPT work instead of just ChatGPT that also has a computer, also can execute stuff, browser. I think we're all realizing the same thing at the same time after, Peter Steinberger, opened the door for everyone that, my agent has to be able to work as me, needs a browser, needs to write code, needs a computer environment, right? either it's a Mac Mini in my home or a cloud environment. Cloud is easier for many folks. But also, I want more than one thing. I want more than one context, more than one agent, like, doing and running stuff. And so everybody's launching kind of that sub-agent, launching whatever. And I think that, the difference now is in the UX and the UI and how these things operate, right? 'cause every time, my fiancée listens to the show and she has a bunch of, open claws in Hermes, I told her about Groq. She was like, "Okay, so what's novel?" okay, this had this. And I was like, "You don't understand, the packaging matters." Somebody gave me a metaphor for this and said, "Hey, when the cloud came up," and somebody said, what is the cloud? It's not, it's somebody else's computer." no, it's not. It's like a computer you can provision, deprovision, you can extend access, et cetera. It's like a different packaging for somebody else's computer. This is, I want you to talk about the packaging and the agent bot orchestration thing because I think you guys nailed it, and I would love to hear more about that, how that came up.
Shub Gaur
Shub Gaur 56:29
Yeah.
56:29
Yeah, I can talk about packaging. But before I do that, I do wanna say Claude CoWork and ChatGPT Work are great tools, but I do think this is also fundamentally different. Yeah. And there, there are two reasons, right? One, persistent computer. It is always there. You'll notice that is not the case with ChatGPT Work. But then more importantly, the teammates that you have, like the bots that you have, are also persistent. So they are around all the time. One thing I'll show you really quickly is, I have, this one chat that's been just going for ages. I could scroll on this basically forever. Yeah. This is all-- I've never had to manage the context on this because we figured out a really good way to just make these agents persistent and compound, which not other tool-- not a lot of other tools can do. but then on top of that, you mentioned the packaging piece, and I think that's really fun. So one of the things you can do-
Alex Volkov
Alex Volkov 57:13
s- okay, so you mentioned this.
57:14
So just before the packaging, I will also mention this. please. Folks, the second I saw there was no model selector, I almost dismissed this tool outright. I was like, "What? I can't select even between Groq models?" And then I asked my agent to, "Hey, where's the model selector?" The agent said, it's hidden behind Elon only settings that, that I cannot see. That's literally what my agent said. I don't know if it's a hallucination. Some people, other people saw it. and then I stopped caring about this. I just trusted that you guys will manage the right thing. there's no context view. Like you cannot see even how much context you're filling out. There's no context management. There's no compaction management. It doesn't mention compaction at all. there's no plugin skills, whatever. It's all like just works. I love it. Like many people who I installed OpenClaw Hermes tool, I would want this. I'm not sure they're ready for Groq 4.6 and me to explain to them that they're giving data to Elon. That's like a whole conversation that we can talk elsewhere. But I love this as a simple tool. So the packaging really like it doesn't feel like super advanced coder. It feels like mom
Shub Gaur
Shub Gaur 58:17
Yeah, million percent.
58:18
Yeah. and the other thing is, I think there was this shift where we all went from "Hey, we're viewing every single line of code"-… to we kinda trust the agent to do things and look at a lot of the tests.
Alex Volkov
Alex Volkov 58:26
Yeah, there's no diff view.
58:27
Yeah. it's not code. Yeah, I'm fully with you. Okay, so let's talk about, multi-agent management and swarms- Cool … or whatever it's built into.
Shub Gaur
Shub Gaur 58:33
Okay.
58:33
one of the cool things you can do is you can actually have your agents talk to each other. So let me find a quick example here. it was pretty recent. I found a house that I really liked, and I was kinda like, "Hey, can you reach out, ask for photos for this one listing?" And so without me saying anything, the agent knows what the other bots are and it will actually message the other bot
Alex Volkov
Alex Volkov 58:54
and say- Oh, no.
58:55
for folks who are just listening, 'cause this is also a podcast, Shub's showing us an interface that looks a lot like iMessage and even to the fact that like the, the top person is pinned and has a bigger image and a list of chats with bots. Not chats, bots separately. Like, all of these chats are like a specific bot with their own sandbox environment, et cetera. and the, the one that you clicked on that your main message talked to called Shub's Yapper, which I absolutely love. Yeah, go ahead.
Shub Gaur
Shub Gaur 59:19
My Yapper.
59:20
So my Yapper has learned to talk like me, and this is a really cool implementation where it takes my texts, Slacks, emails and then improves over time as it looks at creating drafts for me that I edit or when I give it feedback. And again, it is persistent so it will learn from me basically forever. And I don't want other agents having to ever deal with that. So my other bots can focus on their job and they know anytime they need to talk as me, they will just tag in my Yapper and say, "Hey, Shub asked you to draft and send a message doing this thing." It drafts the message and then it goes, if I close this chat, it goes back to the agent and the bot says, "Hey, look, I sent the message. We're good. I confirmed it. Now it's out." And so that's one way of doing it where the bots will automatically pull in other bots as they know things are happening. But if you wanna be more manual, you can put a bunch of them in a group chat. You can also like @ tag some of them if you really want to. Obviously I'm like in a new chat, but if here I wanted to tag my Yapper manually, I could do that. Yeah. And so we're working towards this world where- All of these bots work on their own expertise, but then can pull in the other experts to do a lot of this work. And we wanna, again, abstract as much of this away from you as possible, so that your mental load goes into doing the right thing at the right time, instead of fiddling with a bunch of settings and hoping that things work. And that's like, why we're hoping that as we build goodwill with people, they start to understand, "Hey, we're doing this so that you don't have to think about these things anymore," the same way you no longer have to think about every single line of code. Where you can just get the work done and delegate it away till you have it finished end to end, which is the part that I'm obviously the most excited about.
Alex Volkov
Alex Volkov 1:00:51
I think this clicked for me when I started
1:00:54
a task on my laptop, obviously. first of all, I was super excited I can connect my corporate Slack- That's awesome which is approved, because Cursor is approved as well, to an agent that can do stuff, so now I have, a monitoring thing. And I built a bot this morning to monitor the license file for DeepSeek when it drops. I was like, "Hey, I know DeepSeek is about to drop. I'm going to sleep. I want you to, in this thread in Slack, tell people when the license drops." And, it did it immediately. You know what the funny thing is? DeepSeek then yanked the weights for a little bit and the link didn't work, so people was like, "No, the bot hallucinated the link." I was like, I went back to the bot, I was like, "Why'd you hallucinate the link?" He's like, "I didn't hallucinate the link." It, literally, it was there and now they yanked it back. So the bot was perfect. I think that the idea of, autonomous things that run, on their own environment is really dope. so congrats on this release. Now let's talk about, so Shub, thank you so much for, breaking down on your, Yeah specific, use cases as well. I can share mine. I can-- Let me add Wolfram here as well, 'cause Wolfram is, a big proponent of, he is, Hermes and OpenClaw as well.
Shub Gaur
Shub Gaur 1:01:51
Love it.
Alex Volkov
Alex Volkov 1:01:52
now let's talk about the model.
1:01:53
So there's, three big releases from, Space X AI in, in general. and you can probably talk about, the model to some extent as well 'cause, Cursor was involved in this. Grok four point six, as I looked at my cursor usage after using… Oh, no, let's talk about, before the model, how do people have access to it? what is the API tier? Like, how do Grok Ultra or Cursor, it's a little bit confusing. Would love for you to clarify. what does it take for people to actually use Grok Bot now? I think there's like a trial.
Shub Gaur
Shub Gaur 1:02:19
Yeah.
1:02:19
So there's a free trial. but the key ways to get access to it are the Cursor Ultra plan-… you're on Grok Super Heavy, or if you're on a Teams Premium plan, you have access. and those are the main ways to get it, but I'd obviously recommend use the free trial, see what it can do for you- Yeah and then commit harder as you start to see it do more things.
Alex Volkov
Alex Volkov 1:02:36
Yeah, absolutely.
1:02:37
And, I think that it, as many technologies, this takes a while to get to The messages between the team members and the fact that they're not just chats like in ChatGPT or Cursor, they're team members, everybody with their own memory that you cannot see. that's the trick. You have to tag the yapper or you have to… I have a chief of staff, for example. I can show off my screen, f-for… Let me take you off for a little bit. Yeah. if you guys wanna see. I have a chief of staff Let me show this here. as you see, mine is a dark mode, Shubh. I d- I think unless you switch to dark mode, you cannot ever come back to Ch- ThursdAI again. What is this? No, just kidding. but I have the chief of staff, obviously, it's W- Wolf- Wilfred in here as well. I also ported my Hermes. I asked it to "Hey, here's the most important things that I already know." W- Wolfram, you do the same, right? we all move our, agent identities back and forth. Yes. so chief of staff knows, about all of the other ones, and he's, managing everything. And then he, his job is to pawn off tasks and then to reduce my cognitive load. every time an inbox comes in, for example, it will scan my inbox and say, "Hey, here's a thing that you need to take care of." I wasn't able to get to this level with any of the other tools that I have, Hermes, et cetera, because of the non-native integrations and the fact that Cursor has built out all those plugin integrations like Gmail, Connections, as well, Granola's here, a bunch of others, is really well done. And also, I have, the DeepSeek Dropwatch now. So here as it says, "DeepSeek is public again." Oh, shit, it worked. here's, "Slack post needs your okay, AutoReview blocked it." And AutoReview works as well, so I can tell it like, "Hey, here's the things that I really don't want you to do." So that's a great setting. and then I looked at my Cursor, usage, and then I saw the Grok 4.6 is being used as well. it auto-switched when Grok 4.6 released. Let's talk about Grok 4.6 because, banger release from SpaceX AI. would love to hear-- You obviously had access to it a little bit before us, so tell us the difference that you feel in those two models and while we pull up some of the evals.
Shub Gaur
Shub Gaur 1:04:32
I think we're really excited about
1:04:34
this step forward with 4.6. the team has obviously just been going much harder on models recently, as you can probably imagine. And so those releases, even Elon said it, right? Those releases are going to come faster and faster as we start to do more and more of this investment around-
Alex Volkov
Alex Volkov 1:04:48
Speaking of, Elon Musk, promised that Grok 4.7 is gonna
1:04:52
beat all the other models, where 4.6 comes very close to, the frontier, to the Sol and to the Fables. Yeah. Please go ahead.
Shub Gaur
Shub Gaur 1:04:58
We, yeah, we have very ambitious goals,
1:05:00
especially as we have more compute. But, no, we're super excited about it. You've seen the evals, clearly. I love these graphics that you have. but- Thank
Alex Volkov
Alex Volkov 1:05:06
you.
1:05:06
So let's talk about the evals. CursorBench, which is like a built-in, benchmark that you guys have internally, which if I'm not mistaken, Grok 4.5 have had them leaked into the training weights, and you mentioned this in the model card. So that, was essentially discarded. It was really good at CursorBench, and s-somewhat because it was trained on it. there is no mention of that in the Grok 4.6 card, so that is not the case anymore. It was cleared out. I think Lee Robson confirmed that this was the case. Yeah … this Grok 4.6 on CursorBench is the clear winner, sixty-nine point nine, obviously 'cause you guys trained on the data that you see from, the folks who share the data with you. which if folks wanna, they Can do in GroqBot as well. they don't have to. It's not by default. You have to checkbox a box. GPT 4.6 Sol on the same CursorBench is 67. So this model like beats even Sol on this like benchmarks. Frontier Code, which is from a competitor of yours, Devin, which is like essentially CursorDevin doing like stuff in the cloud agents. I think they gave you a huge shout-out or you gave them a huge shout-out for like testing and posting this result as well. Frontier Code Groq 4.6 is 61%. We had Swyx from the Advisor for Cognition here on the show talking about Frontier Code and how difficult of a task it is and how cracked people from Devin like actually created all these tasks. this is a banger coder model, folks. Look, I'm gonna add, Peter, maybe Peter, I don't know if you already have Groq 4.6 on Arena. but and the last one that I will mention is Apex Legends. Oh, sorry, Apex Agents. the jump here from Groq 4.5 is 10 points on top of Sol. So Groq 4.6 beats GPT 4. Sol on many of these benchmarks. This is like-- this is the first that we're seeing that Groq is good not only for research, it's actually for coding as well Shubham, comments? I know I glazed the fuck out of Grok, but, comments about any Oh, yeah, good- Yeah,
Shub Gaur
Shub Gaur 1:06:54
keep going.
1:06:54
I'm here all day, yeah. Doing this. no, obviously we're very excited about it, and I know the team has been spending a ton of time on, getting the behavior right and getting, design language right. And so overall, it's clearly a valuable enough model for us to be using it in our other tools like GroqBot. and we're hoping that we have some compounding gains, right? As we release more of these and start to get better results, we will ob- obviously also start to see better results with all the tools and be able to compound based on what we know. So very excited for this, this model and this release, and obviously also the future ones that are coming up.
Alex Volkov
Alex Volkov 1:07:27
Peter, any word from Arena folks about Groq 4.6?
1:07:31
Are you guys already testing this in Arena benchmarks or, agent benchmarks?
Peter Gostev
Peter Gostev 1:07:34
Yeah, we have it out already on the code Arena,
1:07:37
which is tested in the front end. And the cool thing is that 4.5 was number 13, and 4.6 is number six, five. So it's already, the jump's really good. And I think that's-- I generally, as a kind of general comment on the industry, I think the, the labs that do really well are the ones that release a lot of models, and the ones that kind of fall behind are the ones that maybe release every, I don't know, eight months, a year. And we definitely saw-- see this a lot. So yeah, definitely excited to see. we are testing it on our agent model now, so hopefully we release the score soon. But, I think that trajectory that you talked about, I think is the most exciting part, is that, yeah, it's it's good now, but if you're on that path, then, like, how awesome will it be, like, in a few months?
Alex Volkov
Alex Volkov 1:08:23
There's also the part where, Elon Musk and SpaceX built a bunch of data
1:08:28
centers, and a lot of GPUs are available for folks from Cursor to train and scale up the composer kind of level, et cetera. and the pricing is very hard to beat. $2 per million tokens, six for one, output, which is half the price of the competitors. I think this is cheaper than Kimi K3, to, to… Whi- which is, crazy 'cause Kimi, did the base floor of pricing. Kimi's, $15 for outputs. 1.5 trillion parameter was, like, I think this was exposed by Elon some, in some post. the, that was trained on the Cursor stuff. And, the, the thing that I wanna call out from mo- model card, Groq 4.6 improved the performance Of the inference of Grok 4.6. folks are doing RSI now in OpenAI, in Anthropic, and now Grok 4.6 confirmed that this al- is also the case. Here's what I want to tell you about Grok 4.6 plus Grok Bot as well. And, it's really funny to me that Truby-- Sha- Shab, sorry, you didn't call this out. My builder. okay, so I said, "Okay, I'm gonna play around with this. I'm gonna plan this show like I usually do, but I also wanna build something. I wanna build something in f- field." And I was like, "I don't like for building that I don't have a model selector. I wanna build with Fable, I wanna build with this, I wanna build with…" I was like, "Wait a second, Cursor can do Fable. Cursor is a harness, can do Fable and Opus and GPT-4, 5.6." and then I, by completely mistake, nobody told me on X that this is the case, I learned that, Yeah … this motherfucker can spin up Cursor agents, and those Cursor agents can be Fable. So here is a thing, Ashu-- Shab, you didn't see this. So this is a very different thing where, the bots talk among the bots. Y- one bot can spin up an agent that, uses my Cursor credits and Cursor account, et cetera, to run Fable. So I have a thing that I did, and because you have open source as well, I said, "Hey, I want three designs for a, f- for a CRM for myself." It's been a while since I had all these guests like Shab, like Peter used to be a guest and now he's a co-host, Chris from N8, et cetera. I suddenly want to email them and I was like, "Where is that email?" So now I asked it to go and scour all the episodes, all the guests, all my email invites, 'cause it connected to both my emails, all my, lists, et cetera, and built a personal CRM, and this took, an hour and a half in the process. And the, the cool thing is the integration is very deep. And now I'm moving away from, Grok 4.6. All I wanted to say is that Grok 4.6 could have done this, but you don't have to be restricted to Grok 4.6. You can still use other models if you want to build with them, because of this integration, with the PRs. It's really nice, and I really enjoyed it. what else do we want to say? Wolfram, you have any thoughts on Grok 4.6 and Terminal Bench, stuff?
Wolfram Ravenwolf
Wolfram Ravenwolf 1:10:57
Yeah.
1:10:57
what I like about Groq is, or xAI, is that they allow the use of the subscription in other agents. So it's not just limited to, Groq, computer or anything like that. And this means if this is the agent that is, or the model that is being used for the agent, it's very well trained for the agent, which makes it a great replacement for, the OpenAI subscription I'm using at the moment. So I will give this a try and see how this compares, and I don't have to change my setup completely or go somewhere else. I can just-
Alex Volkov
Alex Volkov 1:11:24
Yeah
Wolfram Ravenwolf
Wolfram Ravenwolf 1:11:24
pri- change the model and use another subscription
1:11:27
and not worry that I will suddenly go broke because my agent is doing so much stuff at the same time. So excited to try this out.
Alex Volkov
Alex Volkov 1:11:35
So I think that, Grok 4.6 is a very good model,
1:11:40
sir. Absolutely very good model. One point tr-trillion parameters. this is half the size of Kimi and also half the price of Kimi, which is incredible. I think we've glazed enough. I previously only used Grok for research and ex-access, and it looks like now I can use it as a generalized model. I'll keep using all of the things, but, at least for now, for me personally, for folks who are listening, the move from Hermes into Grok Bot was done-- But also, I'm a, like a canary or whatever. I really get excited about new stuff as well. So that's part of why I'm here doing the news for you. so I'm ge-getting really excited about the, the interface there. Sharp, thank you so much for coming on the show. I want you to have a free minute or two to shout out the people who you think deserving shout out in this, work. I think everyone is too much. and the thing that I will say that I don't love about Grok and specifically and the top boss that you have, it's not about them specifically. It's about the fact that it's really hard, and we talked about this on the pod a lot. It's really hard to judge Grok releases and general space a-- XAI releases based on vibes. Specific because there's, like so much people who are, split completely. No matter what, is attached to Elon Musk, there's gonna be haters and lovers regardless of its quality. So for us, it's really hard to judge on vibes entirely. So that's, That is to say, when I tell you guys that like, "Hey, I tried this out," it's really there's something there. I think there's something there, not because of just vibes that we've collected. With that said please feel free to shout out and, tell us maybe something we've missed, or also shout out members of the team. We shouted out Lingxi already, members of the team who worked really hard on this and deserve recognition besides just- Sure … like the one man itself, that's responsible. I
Shub Gaur
Shub Gaur 1:13:20
love it.
1:13:20
Yeah. Two, two quick things I'll add. one, I recently found out that GroqBot can actually rebuild your Electron apps on its device, which is very fun. Cool. So it can actually go through the flows, give you screenshots, and do a lot of that agentic work for you, which is very cool. but in terms of like who worked on this, obviously it was a ton of people. I have contributed near zero to this. I just kinda get to hang out and try these things out, which is very fun. but Ian and Baltar, two of our engineers who started this whole thing, it was like their brainchild. Jacob Witt and Romain have just put a lot of time into this, and my boy Vincent has been just grinding out over the past two, three weeks. Alex, we were in a group chat with him. Yeah. he's been just getting everyone early access, trying to find new capabilities for this thing, and he's just been crushing it across this, this release. So god job.
Alex Volkov
Alex Volkov 1:14:02
Shout out to Vincent.
1:14:03
I think, Ben introduced me to Vincent first, and then, Vincent introduced us, and he was like, "Hey, Shabbar's much, much more present on camera. Vincent, you're invited on the show as well when you find out, different use cases for GroqBot." Great. I have a few folks sending the comments here that somebody couldn't try this because they're not on a Mac. So are we expected a Windows version at some point? Y- you can bow and say, "I cannot comment." th- the, it's fine. But, just know that this is the feedback. folks who are not on Mac also are waiting, for this as well. Yeah. 'Cause Cursor obviously is not on the Mac. Cursor's on Windows and on Linux as well.
Shub Gaur
Shub Gaur 1:14:34
Yeah.
1:14:34
hopefully coming soon. and also feedback is very appreciated. we're very hungry for it right now. As you can probably imagine, this is a beta product.
Alex Volkov
Alex Volkov 1:14:41
Yeah.
Shub Gaur
Shub Gaur 1:14:41
We have not spent that long building it.
1:14:43
way less time than you probably imagine. And so any and all feedback around, like, how we can make this better and the key use cases is stuff that we're trying to figure out. We still don't know where this thing fails yet, and so finding out all of the perimeter around, what it can and can't do- Yeah … is very helpful for us. So hit me or hit anyone else on, on both teams up.
Alex Volkov
Alex Volkov 1:15:00
I have a bunch of feedback.
1:15:01
the first feedback that I'll tell you live on the show is that there's lack of feedback. other products have /feedback- Yeah, yeah … which generates an ID for that session, so I don't have to, send it to you. And that I immediately did /feedback and I couldn't send you, stuff. I can tell you about something, but if you see that, the tool calls failed, you can probably see it from the session stream that I cannot see. that's gonna be easier.
Shub Gaur
Shub Gaur 1:15:20
Love
Alex Volkov
Alex Volkov 1:15:20
it … there's a bunch of other feedback.
1:15:21
Shaob, thank you so much for coming on the show, for your first time. really well done products a- and well done execution on explaining us to, the simple use cases as well. Thank
Shub Gaur
Shub Gaur 1:15:30
you.
1:15:30
Thank you. This was awesome. Thanks for having me.
Alex Volkov
Alex Volkov 1:15:32
Thanks.
1:15:32
Welcome back anytime when you guys release a cool new thing, Grok 4.7, et cetera. cheers man. Yeah. Thank you so much for joining us. anybody try Grok Bot besides me? Anybody wants to try? I think, we can maybe get Shub to, to, to, to have you guys try it as well. I was lucky enough to have, a Corsair, ultimate, which was like, I'm very happy that I can now try this thing. Wolfram, will you be porting Amy to, to Grok Bot as well? That's my question.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:15:53
I'm not even thinking about this, but I'm
1:15:55
thinking about how can I give Grok Bot to Amy so she can control it.
Alex Volkov
Alex Volkov 1:15:59
That's a very
Wolfram Ravenwolf
Wolfram Ravenwolf 1:16:00
interesting thing.
1:16:00
She's the HBIC, the head bot in charge.
Alex Volkov
Alex Volkov 1:16:02
Head bot in charge.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:16:03
She definitely, she's using, OpenAI Pro.
1:16:06
you can only use it through the website, so she's using that as one of her- Yeah … Zaps, and I could give it her more computer use that way. But I won't
Alex Volkov
Alex Volkov 1:16:13
do that.
1:16:13
I will tell you guys that, a- again, I said this live with Shub, but, just between us girls, I was off-put by the fact that I cannot choose a model because I'm so used to the OpenClaw thing where everything's configurable, the, the Android experience, if you will. Everything customizable, everything's configurable. This gives me almost no, configs besides what I can allow and not allow, and there are some bugs because it's beta. For, for instance, I approved something to execute code on my computer all the time, and then it still wasn't able to because it said, "Hey, I need your approval." So there's, some stuff that you can expect from a beta product. but this thing just worked out of the box, and the computer thing worked out of the box, and I was like, "I can see it." I can see how many of the people who I install Hermes and OpenClaw to, they would never want to deal with this. And the people who I did install it, who went through the pain, through the wringer, they are now still calling me like, "Hey, it thing auto restarted and do- now it doesn't work anymore." And, since OpenClaw and Hermes, like, all released and, people got used to this pain, Codex became so much better, where you can essentially do this and control it from remote. but I think both Codex and Clawd are still locked into the session of this is chats per chat. They start every chat from scratch with your memories and, it's all chats. It's not like bots that you can tag. I think there's something here about this looks like your iMessage, like your Telegram, like your group chat interface, and it also works like one, where every agent specifically does a specific thing. The stuff I don't like is that the, the keys don't trans-transfer between agents. So if I have a key for an API, another agent will ask me again about this. That's a little bit annoying. but I think all of you should try it and let us know, if you do. Wolfram, go ahead.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:17:53
I think for the power user to live with agents
1:17:56
twenty-four/seven like us, Alex, we need an open source system we can change, which is a great thing about Hermes Agent. My Hermes Agent is different from all the others because I fixed all the stuff I needed, added personal stuff, and you can't do this with an online system. You ha- don't have full control. But it gives you the easy part for getting an agent up and running where we got our Mac Mini, so the agent has its own computer. That is a big part of the capabilities, that it is a persistent session, can install software- Yeah … do stuff on its own, run all the time, and this is hard to set up and maintain, especially if you want to do it securely. So a lot of people will not be able to do this. The agents can help them, but if the agent, you still have to have a way to get this bootstrapped, and here you just get it out of the box with your subscription and you can use it. that is a cool thing and makes agents more available to everybody. So that's good.
Alex Volkov
Alex Volkov 1:18:43
I will say this last thing where, this-- We talked about agent
1:18:47
swarms from OpenAI that like self-formed. this is my agent swarms now, and I know it's all possible to open multiple agents in OpenClaw and have them do different things, and I know it's possible in Hermes as well. I know it's a pain in my ass to do so. this is very easy 'cause, like everything is like a new agent who you could just talk to separate, and you can see their conversations. I really like-- I, I don't know if we showed this correctly. Let me show this just one more time if you guys don't mind, and then I'll stop like getting excited about Grok Bot, I promise you. But here's my chief of staff, and you can see this like message from DeepSeek Dropwatch. So I really wanted to know when DeepSeek, that's why I had the breaking news on the show, right? I really wanted to know when DeepSeek is back on Hugging Face. Here's the message. DeepSeek Dropwatch told my agent. You can see Chief of Staff with DeepSeek Dropwatch. This is a read-only conversation I cannot join, but I can see, oh, eight thirteen, it's public again. sixty-eight safety answers short, last modified, blah, blah, blah. Slack post. stand down on the Slack post. So my chief of staff is answering my agent. "Stand down on the Slack post." chief of staff will post once in, so we don't trouble. "Keep the HF watch until I confirm the Slack reply landed, then stop." You can see my chief of staff agent communicating with my watcher agent and telling, "Hey, don't send Slacks." "I got this." This is, Peter, this smells to me like the agent swarms that we saw that, confirmed, entirely. But just this is on my own point. Here's the stuff that I don't like about this. this is, vendor lock-in. If I get super excited about this, I'm locked into Groq. I don't know how to export my memories, export my skills, et cetera. I don't love this. it's not for everyone. The people who want portability, for example, may not love this. setting up all the connections and everything from scratch, it's a pain, but, it's, for us, it's, an understandable pain. But the m- but the swarmy thing, in my case, I loved it.
Peter Gostev
Peter Gostev 1:20:27
Yeah, and I think what I particularly appreciate about this, and
1:20:30
also we've got, a few more harnesses, is that there seems to be quite a bit of innovation in that space, and this kind of, iMessage kind of interface makes sense to me 'cause I, I think that's what people are used to. So I think when I get all of these, like- endless threads in, Claude Code or Codex. It's just, it's so hard to navigate. I really like also I was trying, T3 Code from Ethereal as well, which is nice thing you can manage, multiple accounts if you've got, like, different accounts. And I've got my Linux box set up as well, so I've got a remote, three accounts on my remote box, two accounts on my local one, so I can manage it like that. But I really like that we've got some innovation in that space 'cause I was slightly worried. I think we had a couple of months when everyone was just copying one another, and I was like, "Oh no," like we can't be, we can't be, it can't be the end of the ideas. So I think now is a good time. So I really like that.
Alex Volkov
Alex Volkov 1:21:25
E- even how Codex, T3 Code and Claude Code UI, and
1:21:29
like all of them, they basically look like the same thing, right? Like-
Peter Gostev
Peter Gostev 1:21:32
Yeah,
Alex Volkov
Alex Volkov 1:21:32
if
Peter Gostev
Peter Gostev 1:21:32
you remember- … sessions, et cetera … the first version of
1:21:36
the app, that Cursor shipped, it was like a bad version of Codex. It's like, what's the point in that? And I think people opened it and it's like, "Okay, w- why?" And then everyone moved on, and I think I like that they went back in and tried something else, and I think we need a lot more of that. Yeah. 'Cause, it doesn't feel like we are done, so it's a good time to innovate.
Alex Volkov
Alex Volkov 1:21:55
the thing about GroqBot is like this was a Cursor product that now
1:21:59
because of the integration, obviously they wanna promote within Groq ecosystem. That's why they call it GroqBot and also it uses only Groq. But if you consider this, this used to be like I think this is innovation from the Cursor team, which I really… It feels like Cursor. All the connections connect to the Cursor ecosystem. it uses the Cursor Ultra plan if you want to, if you do have one. i- this is like a Cursor product. I do have comments from folks saying, "Elon Musk," blah, blah, blah, blah, blah, "will never touch the product." I think that many people will update because of this. I think that this will shift the narrative, including the model weights. Like once the model's good enough, people will start thinking about this. But, I really think that there's something here, so I really wanted to like extendedly bring it to you guys, hopefully enough for you to try and tell us in comments if you liked it or not. yeah. Folk's saying, "Does the Cursor $20 subscription give access to it?" I don't think so. they have a fully own persistent computer in sandbox, for you. They have like unlimited tokens. I don't know how many for Groq 4.6. I doubt the $20 they'll give it out. I think it's $249 for the, Groq Ultra High- Yeah. It, it- … and the Cursor is $200 per… Yeah what I will say is that next time we'll bring Cursor people, we'll bring them with credit, so hopefully we'll be able to give you some credits t-t-to use it. Folks, we have at least 15 minutes more to talk about a bunch of stuff before George Cameron from Artificial Analysis joins the team and talks about, different things. I do wanna talk about watermarking. we have the both, both the Europeans here on stage. I do wanna w- talk about this. It's not a con- controversy, but it's definitely something that we needs to mention. Anthropic started watermarking all new Claude output since August 2nd. So if you used any Claude output in Claude Code as well, Anthropic has been imperceptible text watermarking it in every new Claude model. so basically, unlike Pangram, which we brought to you on the show here, that like detects if something is AI thing, this is hidden hidden ways how they control the, the, the, the token streaming, and they adjust a little bit of probabilities so that, undetectably, you won't be able to tell, but Anthropic will be able to tell if this, text was generated. And this could survive copy-pasting and even light editing if it's, big enough text. Anthropic did say they will publish an API, and it's gonna be free for you to see if a text was based on Anthropic, et cetera. but here's what I don't want to understand. Why am I, as a US citizen, getting my text watermarked because of a EU rule? and this question I will forward towards Wolfram or Peter or whoever wants to defend Europe, go ahead. why am I paying the tax because of,
Wolfram Ravenwolf
Wolfram Ravenwolf 1:24:25
I'm not defending Europe for this because I don't want this.
1:24:28
if it creates a text or anything. I know there have been … When AI text generation came out, there was already talk about watermarks. At OpenAI, we just singled out Anthropic, but anyone doing business in Europe has to comply with this, so OpenAI will do it as well. Google may be doing it already, I'm not sure. And, they have been doing it with the image generation. Now text, I hope this gets canceled, but, I'm not sure about this, as it's not really useful. I think everybody is using AI or most people are using AI, not to generate text for them, but they give an input. I do it all the time. I'm an AI evangelist. I use AI for everything. I have a hotkey to translate text to f- write better. So as a German, my English isn't the best as So I'm using this all the time. Yeah. And this doesn't mean AI wrote the text. It means I had an intent and I gave it to AI to do it better. And it I said I… There's this logo for in Europe, you have to mark, not just watermark the text, but also put it on anything that has been AI generated, AI influenced. Put it on my forehead as a tattoo if you will, so I don't have to watermark stuff anymore. You can expect that all the time. It doesn't… we have to judge text by the merits, not what created it or how was it done, but what is it saying. I think that is the important part. There has been slop on the internet before AI, and marking stuff human generated or not, it doesn't change the contents I think this is a silly rule basically, and I would rather not have silly rules like this. Because the bad actors who are using this for social influence, so- social engineering, they will not use models that have watermarking. They will use anything, yeah. So basically it's always the same. Good people get punished and the bad guys do their bad stuff anyway.
Peter Gostev
Peter Gostev 1:26:08
Yeah.
1:26:09
I think there's also idea, in the EU, I don't know where it comes from, but this idea that all we need to do is just to give people information and then things will magically resolve. And I think it's a kind of nice idea, but we've done this with, with cookies, right? We have this cookie acceptors banners for, what, decade and a half. It's like, is the world better place now? And and, if you project forward from this next 10, 20 years, I don't know, it's just, it's probably not gonna mean anything. It's probably gonna add extra, I don't know, bureaucracy for the mo- for the labs. It's … Then there's Ben Thompson had the interesting point about like this kind of almost, even if it's your work, your ideas, maybe like your IP is gonna be marked as Claude now-… forever. So you almost you- EU's making you give over your kind of agency and IP over to the bot, which is feels like really backward. so yeah, I think they're just kind of ideas that kind of sound nice in principle. It's like, "Oh, wouldn't it be a good thing if we all knew?" But then you think about it a little bit more, project a bit more, and then it's "What's the point exactly? Why are we doing any of this?" I just don't see what-- it just-- I can't see a scenario it's like, "Oh my gosh, that's such an amazing rule."
Alex Volkov
Alex Volkov 1:27:32
There's also a thing where, they require all companies that operate
1:27:36
within Europe to apply to this or get fined, and the fines are 15 million euros or 3% of global turnover, which is-- I, I really am excited to see how SpaceX AI, who did not sign up for this, will handle the fines and whatever, so yeah, it's very exciting to see. Some folks are concerned that this changes the sampling algorithm, so actually, reduces performance for some of the models. and also we will say that, Google has SynthID, which is their own watermarking, and they have been doing watermarking for a while. OpenAI discussed open watermarking, but nev-never deployed it. And SynthID works on images and PNGs and JPEGs, but n- but not the SVGs. so which is C2PA, and Meta is also, like, doing some, some watermarking stuff. I don't think it's like bad in general. The only thing that I am concerned about is like, why do I have to, get mine watermarked? Because, websites in the US don't have to show cookie banners where cookie banners are required in Europe. Anyway, folks, moving on to this. from this, we have a very quick this week's buzz that I wanna tell you about some stuff from our sponsor, presenting sponsor CoreWeave, and then very soon we're gonna have a chat with, George Cameron. but before this, we have a, a sponsor break and a breaking news segment very quick. So let's go to this week's buzz. In this week's-
Alex Volkov
Alex Volkov 1:29:06
Folks, here is, the Weights & Biases, CoreWeave Corner where
1:29:10
I wanna tell you about Fully Connected. Fully Connected started as a very small conference, from Weights & Biases. Fully Connected is a concept from machine learning and has evolved significantly. So this year, Fully Connected, obviously, is sponsored by CoreWeave and is, and is taking over Moscone South. this is the cloud conference for the companies, built for AI, September 29th to October 1st in Moscone South in San Francisco. we would love for you to come and check us out. We're gonna do a live show. I see Wolfram already took, took on the yellow jacket. we would love to have you come and check us out. First of all, come say hi to us. there's three tracks on there. Folks like industry leaders, CoreWeave people, there's, a bunch of folks. NVIDIA is a sponsoring, presenter there as well. and, there's breakout sessions. There's gonna be hu- It's like the team is going all out with some very cool, secret stuff as well. As I said in the beginning of the show, DevDay from OpenAI, if you are coming to that, it's September 29th, so you can collab, combine both of them and come to the show. Wolfram, are you excited about coming back to San Francisco, to Moscone, to do another live show of Thursd AI, now for our own CoreWeave?
Wolfram Ravenwolf
Wolfram Ravenwolf 1:30:16
Always.
1:30:17
Always happy to come back to San Francisco and meet my colleagues in person. That is always a great thing to have and talk to people, coming to the conference. Yeah, I'm always excited to meet people who are in AI and talk about these things.
Alex Volkov
Alex Volkov 1:30:30
Yep.
1:30:31
so th- this is one thing that I wanted to talk about. The, the other thing I wanted to talk about is we have launched on CoreWeave Inference that obviously we talk, we'll tell you all about, Inference. We launched, NVIDIA Nemotron 3.5 Lightning. We called it out before when Chris was here. Day
Wolfram Ravenwolf
Wolfram Ravenwolf 1:30:46
zero.
Alex Volkov
Alex Volkov 1:30:47
Yeah.
1:30:47
Lightning. 3.5 Lightning, not 3.0.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:30:50
day zero.
1:30:50
Oh,
Alex Volkov
Alex Volkov 1:30:51
day zero.
1:30:51
Yeah, day zero integration. Not too
Wolfram Ravenwolf
Wolfram Ravenwolf 1:30:53
fast.
Alex Volkov
Alex Volkov 1:30:53
Yes.
1:30:54
thank you. we have multiple ways for you to run Inference, and the, the Inference is not only our service, "Hey, you can pay for tokens," but also if your company needs inference. And we will soon be discussing on the show multiple folks and multiple coding harnesses for whom we provide Inference behind the scenes as well. If you are in need of GPUs- Talk to us. We can connect you to the guys. We can get you, like, very cheap rates for, very good performance as well. So if your company is in need of GPUs, it's not only a personal token factory, but if you wanna try out, if you wanna try out G- Nemotron 3.5 Lightning or other models, DeepSeek is probably folks are right now working on DeepSeek because MIT license, and I love it. here's the point where I say, we love MIT and Apache 2 open weights, and that's what we're gonna host and very proud. We're also very happy about Muse Glimmer that came out in open source, and, Muse Spark that's coming out very soon. So we very much want to celebrate, open source. By open source, available for all with licensing as well. All right, folks, this has been this week's Buzzin'. Now we have breaking news. LDJ, get in here. Let's move to breaking news, folks. Let's go. AI breaking news coming at you only on ThursdAI. righty. Folks, we have breaking news as our next guest, George Cameron, comes in. And I think, George, please, I wanted to bring you up after the breaking news and after the coverage, but I think th-this is since what-- S- based on what you do, this is also relevant. Folks, our breaking news is OpenAI is finally giving us a preview of ultra-fast mode GPT 5.6 Sol at 14x the speed. We told you about this, Romain and, Romain Huet from OpenAI and Dominic from OpenAI both were on the show and told you about, GPT 5.6 Sol is coming to Cerebras. Cerebras also was on the show. Big chips, super fast inference, and now it looks like they're previewing GPT 5.6 Sol at that speed. This intelligence times this speed. I don't know about pricing. I don't see it here. but there is a video here that LDJ, I think you shared. Th-thank you, LDJ. this is a video Sol standard and Sol ultrafast. Sol ultrafast build a work in 3D warehouse simulator with the same text prompt side by side. as you can see, the ultrafast already finished and the, the regular one is still building slowly. what else can we say about ultrafast? LDJ, what's your takeaway from-- This is just a wait list launch. I don't think I have access to it yet, and looks like you need to register. And when you register, that's how I know it's not coming to everyone. You have to add your work account and not your personal G-ChatGPT account. So that's-- I don't think it's coming to everyone. LDJ, what are you seeing from this launch?
LDJ
LDJ 1:33:33
Yeah.
1:33:33
I think it's really exciting in terms of it, it is a limited preview at the moment, but I think it's inevitable they're going to have something, more widely available as they scale out the Cerebras compute, as they scale out Rivera Rubin. And I don't know, maybe it will be three, four, 5x the cost or something. But it just the, this new option for being able to have that much faster feedback loop when you're really for whatever you're doing. And when they have the new Nvidia Vera CPUs as well, which Nvidia is developing, I think that'll hopefully also keep all the other parts of the stack that could sometimes be a bottleneck, also ke-keeping up with that. Because you have anything like a virtual machine when you're even running CAD software, if you're running Premiere Pro, whatever application you're running, that needs to also go at insane speeds when you have this intelligence running at insane speeds to have the whole system work really
Alex Volkov
Alex Volkov 1:34:29
Yep.
1:34:29
and I think the highlight for me, at least from the early kinda comments here, is that this will come to voice. And as you guys know, the GPT voice, the real-time voice thing, is not a sole model. It's, a specifically trained model that pawns off to Sol every time. So every time you talk to it, it's, checking, and then it goes and sends an API request to Sol. I want the checking to stop. honestly, I love the, the live conversation. George, I see you laughing. you probably heard the checking a thousand times as well. It's checking everything, even if it's in this context. I'm like, "Yeah, that's cool." And it's, checking. "Oh, yeah, that's very cool." This checking thing is when the voice model pawns off the, the hard intelligence to GPT Sol, and the reason is because the Sol is not as fast to react i- in real time to the live voice conversation. This will be live. So this is one comment from there as well. George, one comment about the, the voice thing and the checking stuff and, and have you played with this fast model as well? Or you don't have to tell us if it's, under, under secure NDAs. You don't have to
George Cameron
George Cameron 1:35:22
tell us.
1:35:23
No, I think on the voice checking thing, it's, yeah, it's this new paradigm of, "Hey, people want voice-to-voice, speech-to-speech models with low latency," but then they want as much intelligence as they're gonna get. And so you have, the small speech-to-speech model that's dumb that will use a tool call to call a more intelligent model to, to complete tasks. And so I think it's, it's needed, but at the same time, I wish I could choose the speech-to-speech model, and I'm okay with a little bit more latency if it's a bit smarter because you can tell that it got dumber,
Alex Volkov
Alex Volkov 1:35:53
Yeah, 100%.
George Cameron
George Cameron 1:35:54
you can tell that it, that it's a, that it's a lot dumber.
1:35:56
And then on the Cerebra, the OpenAI announcement, I haven't read it yet. they said it's on Cer- it's on Cerebras chips. Yeah,
Alex Volkov
Alex Volkov 1:36:02
yeah.
1:36:02
This is, so- Yeah … when they launched GPT 5.6 Sol, they said a 750 tokens per second output ultra fast mode is coming, powered by the Cerebras chips. And we talked- Oh … with, both Dominic Condol and Romain Huet about this, and they said that, this is the Full Weights. I think, Peter, you asked me to ask them if it's, a- Yeah … dumbed-down mode or if it's the Full Weights. dumbed down was the Spark model, I think. Previously 5.5, 5.3 Spark. so no, this is the full GPT 5.6 Sol, the multimodal one as, as well, which is, incredible. not just text model. Peter?
Peter Gostev
Peter Gostev 1:36:33
Yeah, I think that's the-- I'm always looking for a
1:36:36
catch, like what's the catch there? 'Cause, I can't… in my mind, I don't understand how the Cerebras chips fit such a big model. so what about the context length? they didn't really comment on the context length, I think. I haven't seen any comment on that. No, there's
Alex Volkov
Alex Volkov 1:36:49
no, no context
Peter Gostev
Peter Gostev 1:36:50
length.
1:36:50
so yeah, definitely looking forward to, artificial analysis, benchmarking of them, George. yeah, that, that'll be good to see 'cause especially, I don't know if you do, long, context also testing for different providers. 'Cause, I think that's where-- That's the only gap that I see 'cause they said it's the same model- It's obviously gonna be more expensive, but yeah, what's the catch? I don't know. Maybe there isn't one apart from the price
Alex Volkov
Alex Volkov 1:37:13
We'll have to wait and see because- Maybe they shipped
1:37:15
a wait list and a blog post looks like a, not like an actual model. so we have one more breaking news. George, this is how this operates, and I think you know because you're in the, you're in the ecosystem. I didn't expect there to be like three breaking news during the show, but folks, we have one more breaking news, and then we'll chat with George maybe about this breaking news. Let's go. AI breaking news coming at you only on ThursdAI.
1:37:46
All right, LDJ, you brought both of these breaking news. Please read out and tell us about what this one is about
LDJ
LDJ 1:37:52
Yeah, Gemini three point seven Flash, three point seven
1:37:55
Flash is, I guess originally was their most cost-effective model. They have Flash Light, and so now Flash is like their Sonnet, their Terra, if you will. but yeah, it's about-- It's, roughly three dollars per million output tokens. It seems like it's competing in capabilities and price with models like Muse Spark one point, two and, some other models around that price range. It's specifically in, for example, Deep Swe v1.1 here, it's, especially competing well, and Terra is a significantly higher cost model, or at least by a few dollars. And it's beating its closest cost competitor, Muse Spark 1.2 here. And, yeah, just overall, if you scroll down more on the page, there's like kind of a bigger benchmark aggregation image,
Alex Volkov
Alex Volkov 1:38:43
Oh, I like
LDJ
LDJ 1:38:44
this one …or this one too.
1:38:45
Yeah. Yeah. This is a good cost efficiency one.
Alex Volkov
Alex Volkov 1:38:48
So this is from DataCurve AI, from Deep Swe, and this shows the cost per
1:38:51
task, where recently, and George, I would love to talk to you about this as well. Recently, folks have started focusing on cost per task versus cost per token, because many models now output fucking Opus, specifically the latest like Opus five, et cetera. They are just slop machines. So they get the same task, but they get like ha-half, half the to-- you pay twice the price, but also three times the token. So people start costing, evaluating not only, m-m-money per token, but also cost per task on multiple things. And here on at least Deep Swe, this is directly from DataCurve AI, the folks who created Deep Swe. Gemini three point seven looks like very close to the, Pareto frontier between the models that they compared it to. like very cheap performance up to this Deep Swe seventy percent. very nice. For folks who, after the recent changes last week from, Deep Mind, with Jeff Dean and folks and Demis Hassabis, folks who started like burying, Gemini, I think that I told you guys, do not discount the giant gorilla, that has all the TPUs in the world to bring us like bigger models. This is obviously not the Gemini three point five Pro that we've been waiting for a long time, that was delayed and delayed, that wasn't released. but, don't discount this. Shout out to the Gemini team for this release. Gemini, Spark is their agent, is now is improved with Gemini three point seven Flash. we can talk about Spark an-another time. All righty, this has been the breaking news, and now I am excited to bring you, to the show for the f-first time. George, have you been here before? I know we met, but like I don't remember if I had you on mic before. No.
George Cameron
George Cameron 1:40:17
I don't think so.
1:40:18
All right. It's, yeah, a pleasure to be here.
Alex Volkov
Alex Volkov 1:40:19
All righty.
George Cameron
George Cameron 1:40:20
George- and this Gemini Flash ex-- announcement, super exciting.
1:40:22
Yeah. and I don't know if you saw, but they like halved the price compared to the last Flash release.
Alex Volkov
Alex Volkov 1:40:28
Oh, wow.
1:40:28
No.
George Cameron
George Cameron 1:40:28
And they call it an introductory price, but they've halved
1:40:31
the price until the end of the year and in, in AI terms, I think everyone w- and, it's cool to see cool faces on the, hey, Peter, on, on the call. but i- in AI terms, until the end of the year, it's like eternity.
Alex Volkov
Alex Volkov 1:40:43
It's eternity.
1:40:44
we did see price increases for the first time. DeepSeek just announced like a temporary price increase with the tier stuff, whatever. We talked about DeepSeek. George, welcome to ThursdAI. You know about the show, but the folks may, know about Artificial Analysis, don't know about why you started it. co-founder, right? the- Yeah … tell, tell us please about Artificial Analysis. What's your mission in the world, and why, is every lab almost when they release, including this one… Folks, I wanted to highlight this as my, Gemini 3.7 is open here. Number three mention after the input price and output price, Artificial Analysis Intelligence Index, which is 56 for this benchmark. so George, why did you start Artificial Analysis, and why is every lab, frontier lab, is now mentioning your, intelligence index? Would love to hear directly from you for our audience as well.
George Cameron
George Cameron 1:41:29
Yeah, sure.
1:41:30
m- my co-founder, Micah, and I started Artificial Analysis because in, early '23, we were building agents. what you call agents now. I don't know if, I don't know if we were using the term then. and it was costing a lot to run. it was slow, and we weren't getting the intelligence that we, that we wanted, or at least for the price a- and the speed. Then it was about choosing between, GPT-4, GPT-3.5 Turbo, Llama 2 had been released, and I think then you had Claude Instant. Google had their PaLM model series, but, but it wasn't that great. and so we were building in the space, wanted to understand the trade-offs between all the different options, and technology choices out there and- So we put together the charts to help us make those decisions and then essentially put up, I think it might have even been a Vercel preview link, and a few charts on Twitter, and to essentially just, share some of our analysis to help other builders out there, if they were encountering the same problems, had the same questions we did. so very much like a side project to help people, navigate or like face-- who are facing the same problem as we were, but then very quickly had many model releases, Claude 2, Mistral 7B, Mixtral 8x7B. The famous, releases w- the original Gemini Flash release. And so all of the choices exploded, and so we kept building artificial analysis from the Vercel preview link to, to, a bit of a go-to resource. And I think where it's developed into is, we really believe now that it makes sense to have an independent company outside of the labs, who, who are doing objective benchmarks as to the performance of these models, of these inference providers, of these agents-
Alex Volkov
Alex Volkov 1:43:15
Yeah
George Cameron
George Cameron 1:43:15
to help people make decisions.
Alex Volkov
Alex Volkov 1:43:17
So I wanna highlight, what an insane panel of, of guests
1:43:20
and co-hosts we have here right now. Wolfram has Wolfbench, which is independent, but, based on like Terminal Bench, which is like very used. We should probably talk about Terminal Bench 3.0 at some point. Wolfram will bring this next to the show. Peter Gustov from Arena AI, which uses ELO scoring based on tons of people just using these models and recently launching their Arena leaderboards as well. And George, you're doing, programmatic evals, but also an amalgamation of e-evals as well. between all of us, if we don't know which is the best model for what, I don't know who else could know and because we're also bringing you every week the vibes of these models as well. George, tell me about the Artificial Analysis Intelligence Index specifically. What is going on there? What is the index? Have you guys iterated on this? but please tell us about, what people need to expect when they see AA Intelligence Index.
George Cameron
George Cameron 1:44:05
Yeah.
1:44:05
Yeah. So I think what we say is, we try and make the intelligence index, the best single number for understanding and comparing the intelligence of language models on a generalist basis It's the single best number for that, but it's not the only number that you should care about like when you're making decisions, is what we say. and what it is it's an aggregate of nine benchmarks that we take a weighted average of. We run all of those nine benchmarks independently. Some of these benchmarks cover a variety of different use cases, important to AI currently from coding agent workloads to agentic knowledge work, to quantitative reasoning, long context reasoning, and others. Some of the benchmarks amongst those nine are created by us. we have their Amnesience, AALCR, and we-- GDP Eval AI has taken OpenAI's dataset. We created an agentic harness to run it, a grading system to turn it into an eval And then others in the index are great evals others have created, like Terminal Bench, like HLA, Critical Point, etcetera. and so we aggregate these benchmarks to essentially have a diversified generalist perspective on intelligence as the single best number. That being said, we publish all the results, a lot more detail on the website for those that kinda wanna go deeper for their use case. And to note, we actually just released today Optima, which is a tool that helps anyone create their own benchmarks for their own use cases, that we're really excited about for when you wanna go that step further towards your use case to understand for your use case the intelligence trade-offs, and then also speed and cost, trade-offs as well.
Alex Volkov
Alex Volkov 1:45:43
That's great.
1:45:44
tell me more about Optima. So who is the target audience, and will the results of that will show up in the Artificial Analysis, or is this like specific for companies who wanna y- check it on their own use case, they have a little bit of a data set or something?
George Cameron
George Cameron 1:45:57
Yeah.
1:45:58
It's, it-- So it's pri- it's private evals. It's not gonna go in the intelligence index. it's essentially for anyone out there building, building agents, who wanna understand, okay, so they've seen the generalist metrics, maybe they've created a shortlist of models, but they want an eval specific for their use case where their use case might not align totally to generalist intelligence. So what we've done is we've distilled like a lot of the, the research, the expertise that the Artificial Analysis team has developed, over the past few years in creating evals into this tool that helps other people create data sets, and using our grading system to create an eval specific to them. And it helps people identify, okay, for my custom use case, what's the best model? But then also, "Hey, I wanna save ten X, so w- what's a model that's nearly as good as Fable, but is maybe ten or even a hundred times cheaper?" Yeah. it'll give you those answers. So it's really anyone building agents and particularly when you're getting into scale and thinking about these efficiency questions and such.
Alex Volkov
Alex Volkov 1:46:53
So I think this is super cool because, so George, you may be aware
1:46:56
that Weights & Biases has a product called Weave, which is like agent traces as well. Yeah. And, y- the way that you guys talk about this is the way that we talked about like agent traces affecting your company's internal models for Finetune models, but also like the, the models that you use like from the shell, for example. And, the way that you kinda talk through the process, which like you describe the task, which model is best at creating sales deck based on my data, and then you import traces. we would love to see Weave implemented in the import as well so that we track- let's get them on. Yeah. I'll start my agent in the background
George Cameron
George Cameron 1:47:23
now.
Alex Volkov
Alex Volkov 1:47:23
to be able to like import, your, traces as well.
1:47:26
But I think like all of those are a fairly standard format for traces, and then use your coding agents to kinda interact with this model, with this eval and test, which I think that like This is what, how people now operate. you get, you have your own data, you have your own agents doing the work, and you have, some, s- some source of, mutual context. So I love this thing. It's called Optima Folk. Definitely check it out. I will send this to some team members internally to evaluate, the evaluation level. so
Wolfram Ravenwolf
Wolfram Ravenwolf 1:47:50
congrats on the release.
1:47:51
I just want to add that this is exactly what we, the experts in evaluation always tell people. We can give you scores and everything, but in the end, you have to do your own evaluation. You can use the scores to think about which models come in the, inner circle you are trying to really make use of. So make your own evaluations, and it's easier said than done, though. Totally. This is a way to actually do that. Pick your favorite models and see which one of those works the best for your specific use cases.
George Cameron
George Cameron 1:48:16
Yeah.
1:48:17
I think that's right, and I think where this has come from a little bit is that, that's our belief as well, is use artificial analysis to understand a general's perspective, then shortlist, and then do your own testing in your agentic harness or your own data set, right? But the problem is, and, the tech CEOs, I think Satya Nadella's talked about this. Every enterprise should have its own, hey, have its own benchmarks, and that's gonna be the differentiator, in the market for your enterprise. But the pr- the problem is that, hey, if you're using into AI, like creating benchmarks is hard. Yeah. It requires, a lot of expertise, a lot of, time, and it's not something that people have experience in. And so we want to make that accessible because otherwise it's like we really struggle. That's what we do, with our forty-five people every day, is try and, create good benchmarks. Yeah. And, normal organizations don't have that talent in-house, don't have that time, right? And so we try and have automated the process by distilling our experience and knowledge. and I think you're right, Alex. grab the traces. You can just download the traces and upload them. If we don't have the-- We'll get the CoreWeave integrated, though. And then once you've, it'll say, "Hey, rather than Fable, you could use this Open Weights model." you guys serve Open Weights model on your inference, APIs. then use it there or use Fireworks or use, use OpenAI's model limits or the latest Gemini model limits. Suggest that, is how we see it used.
Alex Volkov
Alex Volkov 1:49:37
All right.
1:49:38
so a shout-out to the launch. Thanks … but also, I think the thing that I want to talk to you about is, what is the best AI model? from a general's perspective- How do you-- you're running- yeah. It's tough … you're running an analysis, company. I bet that tons of people who are not, quite in the details for us are like, "Hey, this is a better writer. This feels better. This is a big model smell. This is agentic and cost per token, or cost per task is different." you probably get people who like literally just ask you what model to use, et cetera. How do you answer that question?
George Cameron
George Cameron 1:50:07
Yeah.
1:50:08
I think I- answer the question by kinda u- understanding the use case and like as a start, I think one note about these language models is that like the intelligence is, Generalist in how it forms or like where it comes from in terms of these language models. And so there's quite a bit of correlation, between use cases, of course. and so I think like a, a generalist top-down perspective isn't the worst way to start and start from a, an Opus five, a Fable five, I think are the two standouts, standout mode-mo-mo-models at the moment, for sure. But then it, then… and you can start top-down, but then it's okay, for my use case, what's important? And I think there's two perspectives. Like one, I think about like knowledge areas, what does it need to know? And that informs maybe like the size of the model or things like the omniscience knowledge scores or whatever. And then you think about the capabilities. Okay, what capabilities do I need? I need like long context reasoning, I need agentic multi-turn tasks. I need it being good at terminal use. And then that'll guide how to think about what evaluations I wanna consider, for my use case. Because I think if you c- like otherwise, every use case is like absolutely unique, and so it's hard to use benchmarks to work out, okay, what's my answer here? Optima tries to plug into that, but I think as a first form, think about okay, what's the rough domain knowledge and then what's the rough like capability, and then pick benchmarks to represent those.
Alex Volkov
Alex Volkov 1:51:35
I-I think that's a great question.
1:51:36
Like asking like what's your use case first. it's something that I started doing because when, let's say Coreweave offers like a bunch of products, they're like, "What do you guys do? Like what do you do? Like we have solutions for you, but let me hear from you first. Like what are you asking really?" And so for many people- Yeah … I think what are you asking really is like a generalist agent model that can do long horizon tasks. For some people, it's like, "Hey, I wanted to write in, in, in my thing." George, you mentioned
George Cameron
George Cameron 1:51:58
this- can I just add one thing though?
1:51:59
Please do. yeah. I think probably what I forget and from day one on the website, we've had intelligence, speed, cost. I think it's like, what's your use case? And then it's like, what's your latency budget? What's your cost budget? And what do you want optimized for there? Like i- it might be well and good for me to say, "Hey, Walmart, for your, customer service chatbot, use Fable," but that might bankrupt Walmart because it's too expensive, right? It's twenty-two dollars cost per task. and so I think treating those equally, understanding the constraints is as critical as oh, what's gonna be the best at this task.
Alex Volkov
Alex Volkov 1:52:30
I think looking at your website right now, on
1:52:33
intelligence, there's, Opus is up there, Fable's up there, et cetera. GPT 5.6, the standard kind of ones. Grok is coming up very closely. look, Grok is like number four. Yeah. And, George, would love to have you mention… Not mention, like chime in on the fact that like from three frontier labs, essentially, we switched to five or six frontier labs in a matter of a q- q and a half, which is absolutely insane. But also I wanna call out the… On speed, the new entry, the breaking news one, the generative one f- Seven Flash is 340 tokens per second, which is absolutely mind-blowing. does that tend to kinda relax over time? Do you see that like new models when there isn't not a lot of demand, so they give a lot and then it kinda like changes? Or how do you think about like speed, like testing speed over time as the models- Yeah … kinda get more saturated, and maybe the labs getting, reducing the speed to handle the load?
George Cameron
George Cameron 1:53:24
Yeah, it's a g- it's a g- it's a good question
1:53:25
'cause it does vary quite a bit. so we represent on the website the median of the last seventy-two hours, of measurements or, when it's available, that might be less. But the median of the last seventy-two hours, and so it's a rolling window. And so we always try and make sure that the speeds represent the current speeds developers are experiencing. It's common for kind of speeds to come down if the provider doesn't add more hardware. because of course, GPUs that might be serving at a low batch size before people switch to the model, but then batch size will increase, which will slow the per user speed down because the same hardware is serving more users. And so it's not uncommon for speed to come down. It's good to look at those curves. We have a, the time curve, over time, on, on the website. and so it can come down, but I think that, like it's true that like Gemini three point seven Flash will stay fast. and that's of course with Google and their TPUs, a competitive advantage for Google and, a s- a strength point for them. I, I-
Alex Volkov
Alex Volkov 1:54:22
But I
George Cameron
George Cameron 1:54:23
think it's underrated.
Alex Volkov
Alex Volkov 1:54:24
George, I haven't literally ever saw this.
1:54:26
I just went on the website while you were talking, asking… When I asked you what was the best model. There is a model recommender here with the three switches that you said: intelligence- Yeah … speed, and cost. And if you choose up intelligence, l- the fastest speed and the lowest cost, you get Gemini three point seven Flash, which is like the model that just came out, folks. So here you go. There you go. Agentic capabilities, and you select all of these models together- Yeah … and then eventually- You can select CoreWeave
Yam Peleg
Yam Peleg 1:54:50
there.
Alex Volkov
Alex Volkov 1:54:51
Yeah.
1:54:51
eventually Gemini three point seven Flash- Yeah … is the, the-- because I think of the outsized speed. And so very excited to see once you guys test the Cerebrus version of GPT five point six. So whether or not that's gonna come on top. let's talk about price. We ta- we talked about price per task. We talked about, different things. price is also not, first of all, it's per task now. how do you guys think about pricing? And also talk to me about like amalgamation of like caching, for example. DeepSeek is very good at caching, one of the best to ever do it. And like ninety-seven percent of like all of them in their harness is cached. How do you guys think about pricing per token versus price per task ver- pr- ver- versus blended price per input plus cache plus output?
George Cameron
George Cameron 1:55:29
yeah.
1:55:29
So e-exactly. We used to just talk about, hey, like the list prices, input, output price. now I think if we think about the cost per task, it's the token pricing of input, output. It's your cache discount on the input tokens. It's the cache hit rate that you're actually getting a cache hit when you should It's the number of turns that the model is doing for the agentic task, and it's the amount of like output tokens the model is outputting per turn. And so there's a bunch of like factors here that go into your cost per task, and it's got a lot harder to understand and c- and compare models. We try and create a fair basis by using a cost per task metric. and so that accounts for all of that. I think if we think about what's most important there, I say get rid of… think less about the list prices and think more about the, the number of output tokens. So we… and we show on the website number of output tokens to run benchmarks.
Alex Volkov
Alex Volkov 1:56:28
Yeah.
George Cameron
George Cameron 1:56:29
so thinking about the verbosity of models.
1:56:31
Secondly, think about the cache hit rate as the next most important. So if you've got… and the discount. So if the cache discount is 80% versus 90%, and most tokens in an agentic task are cache hit input tokens or ca- or cacheable input tokens, then that 80% to 90% can pretty much almost double, it's al- almost double the cost, of an agentic trajectory.
Alex Volkov
Alex Volkov 1:56:58
Yeah.
George Cameron
George Cameron 1:56:59
Similarly, if the cache hit rate is lower, so the, the provider
1:57:03
isn't as good at serving cache, like serving cache tokens, hasn't implemented cache aware routing between its nodes, et cetera, then, that can also mean that you're getting, 10% of the time you're not getting a cache hit when you should. You're paying the list input token price, which is 10X more than, the cache input price and therefore, like you can be paying double. and so I think th- these metrics are as critical as input price, output price. We try and simplify things, make it easy by reporting the cost per task on the, on, on the website, which ranges… Instead of ranging per token, it's like per task. It's between five cents for kind of a Luna task, in our benchmark, compared to a $3.14 for a Fable 5. So there is orders of magnitude, between the models
Alex Volkov
Alex Volkov 1:57:52
So I would say, a very interesting thing and a sign of our
1:57:55
times that folks since we started, since last week when we came to you on ThursdAI, the top models, top recommendation models now under official analysis based on the things that I chose, that the best intelligence, the best speed, the lowest cost. The new one is the model that just launched, Gemini three point seven Flash. the second one is Gemini three point seven Flash Medium, also very good. And the third one is also a model that came out, what, two days ago, yesterday? I don't know. Time doesn't mean anything to me anymore. Grok four-- Grok four point six High with, the context of, half a million, tokens, and the cost per task is, a thousand-- cost per index is $1,000, which is, the cost for you to run your whole i-index on it. And the index is sixty-one, where Opus five is sixty-three. So very close to the capability level of Opus. but at, what, one third of the price of Opus? Yeah, literally one third of the price of Opus. and definitely cheaper on, output. very important research. George, I wanted to call out, like many of us, like we call out artificial analysis, like jumps in rankings. The thing that I love when you guys do is that "Hey, previous model is here. The new model is here. here's the arrow that shows you the jump in capabilities." And in a world of, fast-pacing releases and models, it's incredible to have such a resource. b-both of you guys, Irina as well, for, people using this, and, and you guys doing, automated, benchmarks. George, one maybe last question before I let you go, I know, a little bit over time is, do you actually have human evals in there, or it's all automatic based on, hu-- like, m- LLM as a judge and evals, programmatic evals, et cetera?
George Cameron
George Cameron 1:59:23
Sure.
1:59:24
So in language models, we don't, go to Arena for those. They do, they do great work. for… But we believe that like human preference makes a lot of sense when it comes to kinda the more creative domains whereby like humans are the, like as the only source of truth, right? And so we have, on our website, we have Arenas for image, video, speech, so text to speech, generation models, so kind of media generation more broadly. and there we use human preference. for language models, we focus more on like objective, objective, benchmarking, rather than human driven is like how we approach it.
Alex Volkov
Alex Volkov 2:00:01
All righty.
2:00:01
folks, George- Can I? go ahead. Go ahead, Roman.
Peter Gostev
Peter Gostev 2:00:04
Can I ask you one, one idea?
2:00:06
It's not very well thought through, but, when you run very long running benchmarks, one, one thing I'm noticing is that models spend such a long time validating, and I wonder whether it's like, a- as in like spending like CPU cycles. And I don't know if it gets reflected with to- in tokens per second or not. I don't know how you count it, but what do you think about just the general, the amount of time it takes to run a benchmark? Like I, I wonder if that's like a concept that's interesting or not
George Cameron
George Cameron 2:00:36
I think it is interesting.
2:00:37
I think it's gonna be more interesting. We don't really-- we have a time per task, but it's really just focused on kind of model time outputting. But especially as these models, are spending more time like, do- exactly like calling tools to do tasks, why shouldn't they consider how long those take? and I think that I think it makes sense. One of the things that makes it hard in benchmarking is we need to make sure that everything's like for And when you start getting into how long tools take, you start to think about, okay, exactly the configuration of the box that we're running it on and, GCP, we might do a GCP container, but like one GCP container isn't exactly the same as the other one. It could have a twenty percent swing in performance and such. And so we need to make sure that we're getting that right. So when we say, "Hey, like this model is slower than this model," that's based-- that, that's true and developers will see that too, which is one of the challenges. But, like I totally agree. That's like increasingly important for sure.
Peter Gostev
Peter Gostev 2:01:29
yeah.
2:01:29
Fa-fair enough. It's just a, a personal thing I start noticing is just oh my God, like it's just killing my laptop. Like I, I had to get like a Linux box to run them because, yeah. But yeah, I get the difficulty.
Alex Volkov
Alex Volkov 2:01:41
George, congrats on releasing, OPIc.
2:01:43
Sorry, I need to-
George Cameron
George Cameron 2:01:44
Optima.
2:01:45
Optima.
Alex Volkov
Alex Volkov 2:01:45
Yeah.
2:01:45
I think, the OPIC is one of the things that you support. There's like, gotta… Congrats on releasing Optima. we look forward to integrating this to Weave as well, so like folks that use Weave and inference, could, create their own evals. we absolutely do love the charts that you guys are putting out. That's always reflected in-- on Thursd AI kind of newsletter as well. I think you guys are doing a very important, service. I think there's not enough benchmarking companies and tools and there needs to be like, a few resources for different evals. For example, dude, I can go on and on for this. I really wanna let you go, but like last thing, promise, is that, as somebody who's in this field, how do you attribute like the DeepSWE change from, SWE-bench verified, for example? You remember there was like a different switch. Yeah. And many people felt about DeepSWE specifically, this represents more of what they felt on the ground. There's another thing where like you guys showing the Opus five is like the top intelligence. From the other side, Opus is jargon douching. I don't know if you noticed this, like people talk about or Opus- Yeah …like barely can talk like a human, but it's like very intelligent. How do you account for that much difference and variance between these things? How do you guys think about that when you are building the index?
George Cameron
George Cameron 2:02:50
Yeah.
2:02:50
I think w-we're very conscious that when we kinda create numbers, you are creating things that might be kinda focused on the industry, right? and I think we wanna make sure that if we're creating a benchmark, improving on the benchmark means improving in the real world. And so that's why I'm a little bit hesitant to go too abstract with how, w-with, with the benchmarks 'cause it creates these weird incentives whereby people try and the labs try and grow in areas that aren't aligned with, what humans actually use the models for and actually, what they want. And so I think that's, like v-very critical, and I think there's been benchmarks that, hit the nail on the head, for that better, l-like, like a deep suite. how the models were doing well on SWE-bench verified, especially when they had internet access, was like just going to the GitHub repo or using the Git history and such. And I think that, that's not how you can get tasks done in the real world. and so I think for benchmark authors, I think, like that's an important cons-- like, responsibility, that we have, when creating benchmarks for, for kinda labs to consider as they, improve performance.
Alex Volkov
Alex Volkov 2:03:52
we have a bunch more questions coming in from folks that
2:03:55
are eager to talk to you, George. Maybe we need to bring you one more time or maybe multiple more times because we again, we talk about AA scores all the time. first of all, congrats on the success. I love seeing AA featured as an index, as like top on Gemini and b-bunch of other things. really looking forward to chat with you guys more about how you build evals and benchmarks. and, I think there's, there should be like an AA corner on Thursd AI because like we definitely always mention arena scores, AA scores and like Vows is like one competitor of yours that does a little bit of different thing, but also is very important as well. W-Wolf, Wolfram does like Wolfbench, and we evaluate like independently as well. I think all of those is an effort to tell the people the story of, "Hey, a huge thing is happening here, and it's happening faster and faster, and how the hell do you survive in this world?" And not many people need the difference. Not many people will have to jump to Gemini three point seven flash just because it's like, super fast. Not many people will try Grok four point six because it just released, but for those who want to, we are the folks who are telling the story, and I think it's very important and thank you so much for being like a big part of that as well. And thank you for coming on the show. George, I know we, we need to let you go, but, feel free to come back. Welcome anytime to talk about different models, different, u-upgrades and how you're thinking about evals as well. thank you so much for coming up.
George Cameron
George Cameron 2:05:03
Thank you.
2:05:04
It's a pleasure being part of, part of this community. Awesome … thanks for having us and would love to be back on.
Alex Volkov
Alex Volkov 2:05:08
100%.
2:05:08
Thanks. always welcome back. Thank you, George. And, folks, I think we'll bring Nisten back, see who else. LDJ is still here. we need to land this plane, but not before we show you LTX 2.5. I think, Nisten, you, verified that LTX is, existing in the TLDR. w- why did you verify this? what's new about LTX that's, exciting? Tell us. I think it's open weights, which is dope.
Nisten Tahiraj
Nisten Tahiraj 2:05:31
I haven't tested it.
2:05:32
I ran Qwen, I ran the, the MiniMax H3-
Alex Volkov
Alex Volkov 2:05:37
Yeah
Nisten Tahiraj
Nisten Tahiraj 2:05:37
and that one was fantastic.
2:05:40
H3 is fantastic. It takes a-- Yeah. Even,
Alex Volkov
Alex Volkov 2:05:40
on a 3090, it will
Nisten Tahiraj
Nisten Tahiraj 2:05:40
take,
2:05:40
three minutes to generate one s- one second. But, it was very good.
Alex Volkov
Alex Volkov 2:05:46
we talked about, Cedens 2.5 finally releasing in
2:05:48
the US, Minimax H3, Flux 3, right? Also video model, also open sourcing very hopefully very soon. and now we also have, LTX, from, from Lightricks. I think they spun out like a full-- Correct me if I'm wrong, but I think they spun out a, like a company for LTX specifically. LTX is doing twenty-three seconds, of generated video, and it's, generating significantly faster than Minimax, seven point six times faster, twenty-two billion parameters with D-Transformer and, we should try this a hundred percent. It generates in 4K as well for thirty cents per minute on fast.
Nisten Tahiraj
Nisten Tahiraj 2:06:25
Y-yeah, their code is a lot better suited to people
2:06:31
that are using it, like artists and stuff, because they tell you everything in there, how to do first and last frame, how to actually do a movie studio app. Yeah. and they just have much better support and the H3 was pretty hard for me to, to set up. I got it working as a dev, but, a-again, it is larger. H3 was also slower. The quants were also a mess. Sorry, there's, landscapers outside of my place. Everyone's working. but, Yeah. I- I'm just seeing, a much better reaction from the community for this.
George Cameron
George Cameron 2:07:07
Yeah.
Nisten Tahiraj
Nisten Tahiraj 2:07:07
I haven't ran this yet.
2:07:09
It looks like it generates 4K, which H3 does not do, and it's a lot faster, and they've released NVF before, Quants, 2 or at least someone from the community already did.
Alex Volkov
Alex Volkov 2:07:21
I think that we need to shout out with our claps because
2:07:25
look at this table from LTX, folks. Open weights you can download. Minimax is limited in US, EU, and UK. You have to, you have to check box or box whatever. CDense is not available at all. There's no, weights. mandatory branding, you don't have to m- brand with LTX. Finetune on your own data, AMP is available. runs on any GPU is what they're saying. Minimum RAM 16 gigabytes. I really wanna try this out. I wanna do the Terminator scene where Terminator explains to Sarah Connor about OpenAI's rise with LTX. Peter, any, any tests on Arena already with LTX, or is it, still running? Is there not enough? what can you tell us about this?
Peter Gostev
Peter Gostev 2:07:59
Yeah.
2:07:59
I think, I'm just gonna double-check, what we've released or not. it's always a struggle to remember what's out, what isn't. it's,
Alex Volkov
Alex Volkov 2:08:06
it's very fast, for sure.
Peter Gostev
Peter Gostev 2:08:08
I think a gen- a general point on this kind of stuff is that,
2:08:11
it's the, the speed of Like progress in that space, like the, the fact that there was another line, like faster than real time available, it's like that kind of stuff, like I remember thinking about this, question a couple of years ago and it seemed like quite far away. But looks like the, the progress in open weight video models, like I think it's shouldn't be, shouldn't, we shouldn't assume that this will have happened. I think it does depend on a couple of companies pushing for this. yeah, you go back a couple of years and it's only Sora and, yeah, Veo 3 or, I can't remember when things came out, but, yeah, unfortunately we don't have a score out yet. But, yeah, I think I can imagine it will pr- probably be towards like top five, top-
Alex Volkov
Alex Volkov 2:09:01
Yeah
Peter Gostev
Peter Gostev 2:09:01
seven sort of area.
2:09:03
but yeah, it's, it's definitely, yeah, super, super cool to see.
Alex Volkov
Alex Volkov 2:09:07
And fully open, which we absolutely love.
2:09:09
So shout out to the Lightricks folks, LTX, is an own company research and models fully in the open. We really appreciate it. folks, I think on this banger news of maybe the top, one of the top like open weights models for video, it's time for us to, not yet, Wolfram, not yet.
Wolfram Ravenwolf
Wolfram Ravenwolf 2:09:26
We- Stop, stop.
2:09:26
There's one cool thing and then I'm- One cool
Alex Volkov
Alex Volkov 2:09:28
thing,
Wolfram Ravenwolf
Wolfram Ravenwolf 2:09:28
let's do it … this is one of my favorites of the week.
2:09:30
Grok Imagine 2.0- Yeah, that's true … has been released, and this is great. I tested it and I prefer it. It's one of my first, I did the test with different models. It is in the arena, I think it's on second place behind, GPT Image 2, but it is, it's following my prompts better and it's actually my favorite image model right now together with DALL-E 2.5 from last week.
Alex Volkov
Alex Volkov 2:09:54
Wow.
Wolfram Ravenwolf
Wolfram Ravenwolf 2:09:55
So it's definitely at the top of my image models
2:09:57
and- Yeah, really great model
Alex Volkov
Alex Volkov 2:10:00
I really wanted to try this with GroqBot, and you would think that
2:10:05
Groq and Imagine from Groq are connected. My GroqBot could not realize how to generate images with GroqBot. But the funny thing is that when you create a new GroqBot, I'm pretty sure that they're, they're using this model. When you create, like, a new one, uh, here, you do, like, a create a new, and you give it the image generation, uh, you can generate here. So I'm pretty sure that the-- 'cause the, the UIs are, like, really nice and also the quality is really nice, so I'm pretty sure that this is Imagine too. But I wasn't able to get the bot itself to generate, like, thumbnails. Maybe I'll try harder for, uh, for actual, uh, YouTube show. Imagine, uh, Imagine Image 2.0, that's, uh, used to be Groq Imagine, uh, is now available also and it's number two on the arena, which, which is impressive because it beats Nana Banana. I have to try my, uh, infographic prompts, which I got a lot of props for. Uh, the, the, um, uh, Shab from Cursor really liked the fact that we have all the details on the infographic, so that, that's great. Folks, if this was a banger show, w-what else kind of show do you expect? We had Artificial Analysis boss who's like, indexes mentioned by the breaking news that we had from Gemini, who's now the top tier model if you consider cost, speed, and intelligence. We talked to you about Groq with the folks that built the GroqBot and Groq 4.6 for folks from Cursor. And we had the wizard of open source from NVIDIA, Chris Oleksiuk here, at the beginning to help us kinda talk about, DeepSeek moment, DeepSeek coming back with two models, DeepSeek V- DeepSeek V4 Pro, that dropped the models in the middle of our show, and DeepSeek Flash, which is incredible as well. what else kind of AI show would you expect, and what else kind of week would you expect? This is, the start of, what, Q3, right? And it's-- The, the last two quarters have been insane in terms of capability jumps. Not only do we get Fable, we got, a bunch of other models. I will shout out this one last thing that if you missed any part of the show is available as a podcast and a newsletter. Me and a bunch of my agents are working really hard so that I ha- don't have to work as hard to maintain the multiple things the show turns into. There's a newsletter that I write manually because fucking Opus is a j- jargon douche, and it's impossible to have AIs help me write. that edited, GPT 5.6 Sol help me editing this. If you notice any deterioration of quality in the podcast, please let me know, because, I can blame it on my agent and actually, have him do a better job. we are, also on dev.to, which is, another thing on Substack and, I keep improving the website because, we need to test this out. So everybody of us, uh, Peter is testing out on a bunch of like 3GS stuff. Wolfram's doing WolfBand. Nissan's been doing yellow stuff. I, what I do is, is the show. Uh, and I also wanna test these models on the show stuff. So if you wanna check out the, the stuff that is happening, uh, everything is on ThursdAI.news. And recently, I will show you, uh, we started putting up together these, uh, specific, uh, uh, indexes of everything that happened. So July just ended, and there's like seventy-one things released in July, and you can see all of them here. So we have, you know, Inkling Small and Sedans 2.5. Fable is here. Everything is here. And if you click into Anthropic, you can go and see, uh, all of the releases that Anthropic had, uh, July and June and May, et cetera. If you wanna scroll down to May, you can see, like, we covered everything in May. So pretty much everything we cover, if you wanna remember when something released, ThursdAI, uh, News is now the place for that as well. I had-- Folks, this is just between us, like, e-excitement. I think I had like seven hundred thousand views on this page alone from Google. It's insane how many people wanna know, like, what's going on. So you, if you are listeners of ThursdAI, you missed any part of the show, ThursdAI.news is the resource. but also, subscribe to us on Substack, Apple. It really helps us to bring banger guests like George Cameron and Shob from Cursor when they know. Thank you so much for the shout-out. Thank you for the five stars. it's been two and a half hours on the stream, so it's not too bad. We didn't go quite to three, but we covered the banger week in AI. Peter Gosta from Arena AI, thank you so much. Wolfram, Raven, Wolf, Eval and Wolfbench, Nisten and LDJ as well. LDJ, you brought three breaking news today to the show. I really appreciated that. And also everybody who tunes in, everybody who comments, thank you Noel for saying we're the best show. thank you guys for the comment. Thank you, Colleen for, giving us feedback as well. We love feedback. Feedback via here, via Substack, via X as well. We absolutely love it. and the guests that we are able to get on the show, also appreciate that as well. Thank you so much for joining. hopefully we'll see you in San Francisco on September 29th, and we'll see you here next week as always. Bye-bye everyone. Thank you.