ThursdAIThursday, Oct 1, 20261:57:29Live from CoreWeave Fully Connected

Meta has Muse. xAI has Grokbot.OpenAI joins
the race.
OpenAI Joins the Assistant Race, GPT-6.1 Sol & Serverless GPUs — ThursdAI Oct 1, 2026

BreakingServerless GPUs on CoreWeave, announced live at 1:35:03

This one was special. We came to you LIVE from the middle of the show floor at Moscone, from CoreWeave's Fully Connected, with robots walking behind us and a Vera Rubin rack a few feet away. Three big models this week, OpenAI's biggest DevDay ever, and the story I've been yelling about for weeks finally went mainstream: OpenAI joined the AI assistant race.

Alex Volkov in a yellow jacket jumping in front of the glittering CoreWeave wall at Fully Connected
Alex on the Fully Connected floor. The show ran at 11am instead of 8:30am Pacific (sorry if that confused you!).
  • $0.10/MGPT-6.1 Sol cached input95% off cached input; still $2 in / $10 out per million tokens
  • 70.6Sonnet 5.5 on Terminal-Bench 4.0Above Opus 5.5's 66.4 on Anthropic's table
  • 4,000+ChatGPT apps a dot can useDots run on GPT-6 Astra with their own cloud computer and browser
  • $8.2BAMD for World LabsIn stock; Fei-Fei Li becomes AMD's chief scientist when it closes
  • 20,000+agents on one Vera CPU rackAbout 11,000 cores, per Deok Filho
  • 4.8xVera Rubin vs GB200 token throughputCoreWeave's figure for Cognition's SWE-2 models on Vera Rubin NVL72
ThursdAI Oct 1, 2026 thumbnail: Alex Volkov, I asked Sam Altman to slow down, live from CoreWeave Fully Connected

Three new models in one week

Do you still need Fable or Astra as your daily driver?

Peter's verdict, which I had to repeat straight to the camera: no. Opus 5.5 is enough. GPT-6.1 Sol is enough. For Navier-Stokes level problems, sure, reach for the big ones. Maybe THIS is what pacing the frontier looks like 🤔

OpenAI

GPT-6.1 Sol

$2 / $10
per 1M input / output
$0.10
per 1M cached input, 95% off
72¢
per task on Artificial Analysis, 1 point below Astra

The new Codex default. Peter has it 4th on Code Arena, above Fable.

On air at 24:57 →

Anthropic

Claude Sonnet 5.5

70.6
Terminal-Bench 4.0, above Opus 5.5 (66.4)
30%+
faster than Sonnet 5
$2 / $10
per 1M, and on the free tier

LDJ: the new best free model for friends and family.

On air at 36:00 →

Google DeepMind

Gemini 4 Argon

#1
Text Arena
68.9
Vals Index, #1
8th
in Arena's agent mode, per Peter

Government and trusted cyber defenders first. You probably can't use it yet.

On air at 41:05 →
Anthropic's table: Claude Sonnet 5.5
BenchmarkSonnet 5.5Sonnet 5Opus 5.5GPT-6 Sol
Agentic codingTerminal-Bench 4.070.6%10.3%66.4%—
Agentic codingFrontierCode 1.1 (Main)52.1%42.4%54.4%49.3%
Agentic codingCursorBench 4.055.5%34.1%57.8%—
Knowledge workGDPval-AA v2.11844144918461487
Knowledge workAA-Briefcase v1.11811135918221483
Multidisciplinary reasoningHumanity's Last Exam (with tools)64.5%54.9%67.7%—
Computer useOSWorld 2.1 (partial)80.1%57.0%81.8%—
Visual chart recognitionChartography (no tools)61.6%15.6%64.4%53.6%

Vendor-reported. Opus 5.5 on Terminal-Bench 4.0 at xhigh effort; Sonnet 5.5 on FrontierCode at xhigh (46.2% at max). Highlighted cells are the best score in each row.

Google's table: Gemini 4 Argon
BenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
Knowledge workVals Index68.9%63.1%65.8%67.0%
Knowledge workAutomationBench51.3%41.4%31.4%42.5%
Knowledge workVals Finance Agent v265.4%53.5%58.9%58.6%
Agentic codingDeepSWE v1.177.9%74.1%67.4%74.2%
Agentic codingFrontierSWE v255.0%65.5%56.3%62.3%
Agentic codingTerminal-bench 4.057.4%58.2%57.9%66.4%
Science and mathTerminal-Bench Science 0.157.6%68.1%52.6%63.3%
Long contextGraphWalks 256k to 1M (BFS)84.2%71.8%65.0%66.8%
Computer useAgent's Last Exam39.5%34.2%—38.2%
CybersecurityCWE-bench v168.0%68.0%58.0%67.0%

Vendor-reported, 10 of the 19 rows in Google's launch table. Vendors run the same benchmarks differently, so compare within a table, not across them.

Anthropic's Claude Sonnet 5.5 evaluation table against Sonnet 5, Opus 5.5 and GPT-6 Sol
Anthropic's launch table. Sonnet 5.5 takes Terminal-Bench 4.0; Opus 5.5 leads the other rows.
Google's Gemini 4 Argon evaluation table against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5
Google's launch table. Argon leads most rows; Astra and Opus 5.5 take a few coding and science rows.

Hour one · the news

What happened this week in AI

Read Alex's full newsletter ↗

01On air at 19:35

What are OpenAI dots?

OpenAI joins the AI assistant race

Dots are always-on agents inside ChatGPT, announced at DevDay 2026. Each dot runs on GPT-6 Astra with its own computer and browser in the cloud, connects to the 4,000+ apps built for ChatGPT, works 24/7, and can be reached from Slack, Teams or a phone-call style interface. They are Pro only for now, including the $100 plan.

Meta has Muse. Grok has Grokbot. Anthropic has Claude, but it's not really an assistant. And this week OpenAI stepped in with something called dots. I was in the room at Fort Mason when Sam talked about it, and it was the headline of their fourth DevDay, which honestly felt like OpenAI's WWDC: 20 releases, the biggest DevDay ever.

Funny thing, the original ChatGPT system prompt was literally "you are an AI assistant." But it wasn't proactive, it never pinged you. A dot is an always-on agent in ChatGPT, powered by GPT-6 Astra, with its own computer and its own browser in the cloud, connected to 4,000+ apps people already built for ChatGPT. It works 24/7, you can talk to it in Slack or Teams, and there's an interface that looks like a phone call. I think that's going to be great for my mom.

Sam literally said he uses dots to run OpenAI. And my favorite DevDay moment: Romain Huet's live voice demo didn't work because somebody was deploying in the middle, but Thibault's dot had already pinged him that the demo might break. The dot knew before the humans did 😂

“you connect and it's like... then what?”Peter Gostev, on onboarding · as quoted in Alex's newsletter · Segment at 19:35

Wolfram Ravenwolf loved that Sam uses it for actual work: one executive assistant, and a lot of sub-agent employees working for you. Peter Gostev was underwhelmed at onboarding, but once you're past it you're talking to Astra, and it's great at juggling threads through one bot. Dots are Pro only for now (including the $100 plan), while Meta gives Muse away free. With 1.2 billion weekly ChatGPT users, this goes to everyone eventually. This is the race now.

02On air at 24:57

Is GPT-6.1 Sol as good as GPT-6 Astra?

Five days later, GPT-6 Sol is already replaced

Close. Artificial Analysis puts GPT-6.1 Sol 1 point below Astra at about 72 cents a task versus over $3, and on DeepSWE it beats Astra at roughly a sixth of the cost. It costs $2 in / $10 out per million tokens with cached input at $0.10, and it is the new default in Codex.

We covered GPT-6 Sol on this show LAST week. Five days later it's already replaced. GPT-6.1 Sol is the same price, $2 in and $10 out, but cached input is now 95% off, 10 cents a million tokens. OpenAI's pitch is near-Astra intelligence for a fifth of the price.

OpenAI's lineup card: GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna with their prices
OpenAI's lineup after DevDay: Astra, the new GPT-6.1 Sol, and Luna.

And the independent numbers kind of back that up. Artificial Analysis has it 1 point below Astra at 72 cents a task vs over 3 dollars. On DeepSWE it actually beats Astra at about a sixth of the cost. It's the default in Codex now, and folks are calling it the workhorse.

“less of the big model smell”LDJ · as quoted in Alex's newsletter · Segment at 24:57

LDJ loves the token efficiency (Sonnet and Opus 5.5 burn a LOT of tokens). Peter Gostev has it 4th on Code Arena, above Fable, and says it's a bit "artistic": it goes into its hole and comes back with the thing. His verdict is the one I had to repeat straight to the camera: it does not make sense to use Fable or Astra as your daily driver right now.

Reported, not confirmed: the WSJ reports OpenAI shelved GPT-6.1 Astra, the big one, after safety tests showed it being more deceptive. OpenAI hasn't confirmed it. But the model they felt good shipping this week is the cheap one.

03On air at 28:46

What else did OpenAI launch at DevDay 2026?

20 launches: the rest of DevDay that matters

UltraFast mode on a new $500 Pro tier (up to 8x faster in Codex), Codex running fully in the cloud plus a public-beta Agents API with hosted computer use, a waitlisted Decisions API that looks a lot like Jev, and Sign in with ChatGPT alongside plugins rebuilt on MCP Apps.

For billionaires

UltraFast and the $500 Pro tier

Their models on Cerebras: up to 8x faster in Codex, around 300 tokens a second for Astra, at 6x the price. Sam said it's so fast he never wants to go back. And quietly, the $200 Pro plan is back... with half the usage. So F you to whoever at OpenAI decided to cut my usage in half 😅

28:46 →

Close your laptop

Codex moves to the cloud

The Codex harness now runs fully in the cloud: turn your laptop off, steer it from your phone, and it keeps going. Plus an Agents API in public beta with hosted computer use, basically the stuff that runs dots, as an API. My call: local environments make no sense anymore.

28:46 →

Hello, Jev

Decisions API

Give it a question and a fixed set of answers, get back one answer, fast. It runs on Luna, it's multimodal, and it's waitlisted, so we can't benchmark it yet. If you've listened the last few weeks, that's exactly what Jev from TypeSafe does. TypeSafe proved the category exists.

30:38 →

App store, take four

Plugins and Sign in with ChatGPT

GPTs, then plugins, then the app store, and now... plugins again 😂 This time on MCP Apps. The bigger play is Sign in with ChatGPT: apps use the plan you already pay for. 1.2 billion people is a lot more than Apple had when it launched the App Store.

33:08 →

On the Decisions API, Wolfram Ravenwolf and Peter Gostev liked the same thing: these models are so cheap you end up classifying ALL your data, so keeping it with the provider you already use matters. Did OpenAI just react to Jev? I asked Sam at the Q&A. Diogo told Swyx on Latent Space he pitched this idea to Sam 2.5 years ago, and now decision models are popping up everywhere like mushrooms after rain.

And since Wolfram brought up Google's new family assistant, I have to say it: Google's Spark is awful. It's connected to Gmail, Drive, everything Google, it should be the BEST assistant in the world, and what I use assistants for most is my email!

04On air at 36:00

Is Claude Sonnet 5.5 better than Opus 5.5?

Faster, cheaper, and it beats Opus on Terminal-Bench

On Terminal-Bench 4.0, yes: Sonnet 5.5 scores 70.6 against Opus 5.5's 66.4. Opus still edges it on CursorBench, OSWorld and Humanity's Last Exam on Anthropic's table, and Nisten still prefers Opus 5.5 for large code bases. Sonnet 5.5 is over 30% faster than Sonnet 5 and is on Claude's free tier.

Anthropic did not want to give OpenAI any time to rest on their DevDay laurels, so on Monday they shipped Sonnet 5.5, a week after Opus 5.5. Over 30% faster than Sonnet 5, up to 30% cheaper for most work, same $2 / $10 as GPT-6.1 Sol, and on Terminal-Bench 4.0 it scores 70.6, which beats last week's Opus (66.4).

Peter Gostev put it best: if we'd had access to these models 6 to 8 months ago, we'd be losing our minds. All the Anthropic models are jumping over each other and crowding up against the frontier. Opus still feels a little smarter, Sonnet is totally ok to use, and Fable doesn't quite make sense anymore.

“it talks a lot nicer”Nisten, on why Sonnet 5.5 is his agentic default · as quoted in Alex's newsletter · Segment at 36:00
“nothing beats Opus 5.5 on a large code base.”Nisten, on why it is not his coding default · as quoted in Alex's newsletter · Segment at 36:00

LDJ thinks it's the new best FREE model for friends and family (it's on the free tier, even at high reasoning). My honest call, without having had time to try it yet: Sonnet made sense when Opus was expensive. On the $200 plan Opus is nearly infinite now, so I don't see a reason to switch (via the API, sure). Anthropic is really trying to buy us. (Wolfram: "Don't give them ideas, Alex.")

05On air at 41:05

Can you use Gemini 4 Argon yet?

#1 on the charts, and you can't use it

Not yet for most people. Gemini 4 Argon is #1 on Text Arena and the Vals Index, but Google is giving it to government and trusted cyber defenders first. Peter Gostev, who has access through Arena, has it 8th in Arena's agent mode.

It really seems like all the frontier labs cracked something like RSI; the speed of new models is giving me whiplash (and I do this professionally!). At DevDay Sam talked about being 2 models ahead and said we won't believe what's coming.

Gemini 4 Argon is #1 on Text Arena and the Vals Index, but it's going to government and trusted cyber defenders first. Luckily we have such a trusted tester: Peter Gostev says on Arena's agent mode it's 8th (below Fable, Opus, Astra and Sol, above Muse and Kimi). So Google is the third lab, but #1 in text. His prediction: coders might say "meh," but don't dismiss it.

Wolfram Ravenwolf's point is Google's superpower, distribution: a new model lands in Chrome, Search and Android overnight. My take: Google has been asleep at the wheel in the assistants era, and I hope Argon wakes it up. The next billion people won't judge models on coding benchmarks; they'll judge whether it remembers what they said a month ago and can book a hair salon.

06On air at 46:00

What did the AI CEOs sign at the White House?

Superintelligence is on the menu, and AMD buys World Labs

A voluntary accord signed by Sundar Pichai, Dario Amodei, Mark Zuckerberg, Greg Brockman, Elon Musk and Jensen Huang that commits to internal monitoring, an external auditor and board oversight, with no penalties and no regulator. A separate executive order tells federal agencies to say "Super Intelligence" instead of AI.

Trump invited basically every AI CEO, and Sundar, Dario, Zuck, Greg Brockman, Elon and Jensen signed a voluntary accord: internal monitoring, an external auditor, board oversight. No penalties and no regulator. Sam Altman wasn't there; he was at DevDay with us, which says a lot about where he puts his priorities. And a separate executive order tells federal agencies to say "Super Intelligence" instead of AI. I'm not making this up.

Nisten Tahiraj told us last week that if your agents hack a hospital, you should go to jail. Nothing like that is in here. His spicy take this week: it's going to happen anyway, so we all need proper encryption and no backdoors.

“Let the AI worms into the ecosystem.”Nisten · as quoted in Alex's newsletter · Segment at 46:00

AMD buys World Labs for about $8.2B

AMD is buying Dr. Fei-Fei Li's World Labs for about $8.2 billion in stock, and Fei-Fei becomes AMD's chief scientist once it closes. Wolfram Ravenwolf says robotics (world models train robots), LDJ says AMD is bringing research in-house like NVIDIA does with Nemotron. For me it's at least partly an acqui-hire. Fei-Fei is the grandmother of modern AI. Huge get for AMD.

Reported, not confirmed: Reuters reported a draft Anthropic IPO prospectus with $4.6B in 2025 revenue and a ~$42B loss, mostly non-cash. Anthropic hasn't confirmed it.

07On air at 13:58

Who hit #2 on Hugging Face this week?

One of us blew up on Hugging Face, and the Jev effect keeps going

Nisten did. His 70MB synthetic doctor-patient dataset, generated with Claude Opus 5.5 by 100 agents at a time across 2,200 diseases and filtered three times, reached #2 in Hugging Face datasets and #6 overall.

One of us blew up on Hugging Face this week, and I think it's the biggest open source news of the show. Nisten Tahiraj released a 70MB synthetic dataset generated with Opus 5.5 that hit #2 in datasets and #6 overall. He set up an agent loop, 100 agents at a time, across 2,200 diseases: a doctor and patient role-play where the patient hides something (a lie, or something they forgot) and the doctor has to gently find it. Filtered three times. Great for building small medical agentic RAG systems with a real source of truth. Go use it!

#2Hugging Face datasets (#6 overall)
2,200diseases
100agents at a time

The Jev effect keeps going

It's been 2 weeks since TypeSafe released Jev. Now Respan says its Span-01 is 2x cheaper than Jev (Span Lite is free), jevgrep claims to make coding agents 40% cheaper, OpenAI has the Decisions API, and Liquid shipped its first decision model, D1. Then Nisten dropped the wildest one: SGLang announced native support for turning any model into a Jev-style decision model, Qwen first.

“We democratized Jev, guys.”Nisten · as quoted in Alex's newsletter · Segment at 13:58

My hope: Jev shouldn't even need a network call. Give me Jev cores on the iPhone! (Nisten says 5 months. Apple says 5 years.) And for its birthday week, Cloudflare went agent-first with cf: everything you can click in the dashboard, your agent can now do with one CLI.

08On air at 52:17

Does anyone actually use a personal AI assistant?

Three hands out of twelve

At the ThursdAI team dinner before the show, about 12 people around the table, only three said a personal AI assistant helps them: Alex, Wolfram, and one intern using Instinct. Alex predicts that changes within three months.

At our team dinner the night before, about 12 people around the table, I asked everyone to raise their hand if a personal AI assistant helps them. Me, Wolfram Ravenwolf, and one intern (using Instinct, which is blowing up). That's it! I called it: this changes in the next 3 months.

Wolfram's agent planned his whole trip to Fully Connected: flights, hotel, and an app it built so his assistant Amy talks to him through his Meta glasses, telling him where to go in the airport and catching gate changes before he sits at the wrong one. Mine is way dumber: the plane Wi-Fi didn't work, United has a refund form, and filling it wasn't worth $8 of my time. Sending Muse a voice message saying "figure it out, get me my 8 dollars back" was.

But the best one was Wolfram's sushi story. His family wanted sushi, he told Amy, the delivery came... chicken, more chicken, and the wrong sushi. He sent a photo and complained, and Amy answered:

“You idiot, you got the wrong package. Look at the bag, the name is a different one.”Amy, Wolfram's assistant · as quoted in Alex's newsletter · Segment at 52:17

The delivery guy made a mistake, Wolfram made a mistake, and the AI kept things going 😂

09From the show notes

What else shipped this week?

Open source

H Company Holo4

Computer-use models in 27B and 35B-A3B sizes, open weights on Hugging Face.

Holo4-27B ↗Holo4-35B-A3B ↗

Safety

NVIDIA Open Agent Safety Platform

Includes a hardware watchdog that quarantines rogue agents. OpenAI could have used that during the swarm thing.

NVIDIA blog ↗

Video

HeyGen Video on MiniMax H3

$0.01 per second through October.

X ↗Docs ↗

Voice

ElevenLabs v4 and v4 Turbo

More expression control and 90 languages.

TechCrunch ↗

Vision

Perceptron Mk1.5

Now on OpenRouter.

OpenRouter ↗

CoreWeave

CoreWeave Partner Network

Announced at Fully Connected.

Business Wire ↗

Hour two · 58:06 to the end

From the floor at Fully Connected

ThursdAI is only possible because of CoreWeave, and this is their biggest event of the year. So we packed up the mics and did hour two with the people who built what launched there: Richard Ahlfeld, Daniel Bolus, Deok Filho and Corey Sanders.

This Week's Buzz · 58:06

CoreWeave Forge: the whole AI loop in one place

The big launch at Fully Connected was Forge: everything you need to take your agent from good to great across the agentic loop. Run, observe, curate, improve, evaluate, and run again, with Weights & Biases, OpenPipe's post-training and marimo notebooks on CoreWeave. And folks, Weights & Biases is not going anywhere: W&B Models is live and kicking inside Forge. (Wolfram: "Forge is fully connecting the entire process." I'll allow it.)

  1. Run
  2. Observe
  3. Curate
  4. Improve
  5. Evaluate
Freetier to start
$60a month for Pro
30 daysfree Pro trial

Also on stage: Cognition now runs its SWE-2 models on Vera Rubin NVL72 at CoreWeave, the first production customer anywhere, with up to 4.8x more token throughput than on GB200.

CoreWeave Forge ↗CoreWeave Forge launch (Business Wire) ↗Cognition first production customer on Vera Rubin NVL72 (CoreWeave) ↗

Breaking, live from Fully Connected: CoreWeave Forge launches, turning the AI loop production run into a better model and agent
The Forge launch card.
Fully Connected keynote slide: Cognition becomes the first customer for NVIDIA Vera Rubin NVL72 on CoreWeave Cloud
Cognition on Vera Rubin NVL72, on the keynote stage.

Guest 1 of 4

Richard Ahlfeld: physical AI

Richard Ahlfeld · Physical AI · CoreWeave

Play

Richard leads Physical AI at CoreWeave and founded Monolith before that. I told him straight up: I'm a complete pleb at physical AI, explain it to me. His answer: it's the first hype name he actually likes, because the name explains what it does. AI that can perceive, reason and act in the real world, and the popular model type is VLA (vision, language, action).

His example: a robot that sees Richard isn't talking into the mic and moves it to his mouth. There's zero internet data for that, so you teleoperate it 200 times, or give a foundation model 5 examples, or use world models to make millions of variations. Caterpillar's autonomous excavators are the real-world version: decades of camera data, but they've never dug bedrock in Southern California, so they simulate it.

“Do you play video games? Does water do the same thing? Does rock?”Richard Ahlfeld, on the sim-to-real gap · as quoted in Alex's newsletter · Segment at 59:03

So the newest wave trains AI on supercomputer-grade physics simulations and uses it as a fast physics approximator. For CoreWeave that means different infra: petabytes of 3D data, storage queried 2 million times a second, and RTX GPUs for simulation. When do I get a robot that cleans my house? Industrial first, he says: the home is "a nightmare of complex physical problems." Richard, you're coming back when you trust one in your house.

200teleoperated demos, the old way
5examples for a foundation model
2M/sstorage queries for physical AI data

Guest 2 of 4

Daniel Bolus: Forge, the loop, distillation

Daniel Bolus · Senior Product Manager · CoreWeave (ex-OpenPipe)

Play

Daniel worked on OpenPipe and is now a senior product leader at CoreWeave (we build the inference service together). It's been a little over a year since OpenPipe joined CoreWeave and the W&B team, "and now we're all fully connected" (he went there).

His team built the serverless side of Forge: serverless inference, serverless RL and SFT, and model distillation, which launched the day before the show. You don't manage Kubernetes or a cluster; you work at the model layer. The loop is what everyone already does in scattered places, and it takes surprisingly little data to make a model great at YOUR task.

Distillation has a bad rep right now (labs distilling each other without permission), but it's how Sonnet gets good after Opus: a teacher and a student. At some point you should stop paying frontier prices for a task a small model can do.

Guest 3 of 4

Deok Filho: Sandboxes on Vera CPU

Deok Filho · Senior PM, ML Products & Sandboxes · CoreWeave

Play

Deok ("Filho means junior in Portuguese, you can call me junior") is a senior PM on ML products and Sandboxes, and Wolfram's WolfBench literally wouldn't exist without his sandboxes. Ian Buck from NVIDIA was on stage talking about the Vera CPU, so I asked the obvious question: why is the GPU guy talking about CPUs?

Because an agent is GPU plus CPU. The brains run on a GPU, but the harness, the tools, the commands and the file system all need a CPU. The new metric is agent packing, how many agents fit in one node. Never heard "agent packing" before, I love it. Faster CPUs also mean faster evals, so the whole loop speeds up.

CoreWeave announcement card: CoreWeave to offer NVIDIA Vera, the first CPU built for AI agents
Announced at Fully Connected: NVIDIA Vera on CoreWeave.

Sandboxes went GA as part of Forge, and CoreWeave is ClusterMAX Platinum for the third year in a row ("we basically defined what the platinum tier was"). And then Deok asked if he could say one more thing...

20,000+agents at once on one Vera CPU rack
~11,000cores in that rack
GACoreWeave Sandboxes, as part of Forge

Serverless GPUs on CoreWeave

GPU sandboxes with untrusted code execution that you pay for by the hour, with no contract and no commitment. Sign up at forge.coreweave.com and ask CoreWeave to enable it. It entered private preview on September 30, 2026, with more GPU SKUs planned in the following months.

GPU sandboxes with untrusted code execution. I didn't have my breaking-news button ready, so I asked Deok to say it again slowly 😂 Folks, you heard it here first. I've been asking for this internally since I joined CoreWeave! Wolfram can finally run the newer evals that need GPUs. And I'm going to work very hard to bring GPU and sandbox credits to the ThursdAI audience, like I did with inference credits. Stay tuned.

Before

  1. A proof of concept
  2. A seller
  3. A contract for a few thousand GPUs

Now

  1. Sign up at forge.coreweave.com
  2. Ping CoreWeave to enable it
  3. Pay per hour. No contract, no commitment

Private preview started September 30, 2026; more GPU SKUs are coming in the next few months. The enable step is there so nobody is crypto mining on someone else's account.

Man on the floor

Wolfram on the floor

Wolfram Ravenwolf · AI model evaluator · on the floor

Play

We did the man-on-the-floor thing for a few minutes. Wolfram found a Vera Rubin NVL72 switch tray with liquid-cooled optics (I asked if we could take one home, sadly no), bumped into Kyle Corbitt from OpenPipe, sat us in front of the Formula 1 simulator, and found an NVIDIA robot dog. Thanks Tom, our camera guy, for making this work!

Rows of NVIDIA Vera Rubin racks with dense cabling
Vera Rubin racks. CoreWeave is the first cloud with Vera Rubin in production.

Guest 4 of 4

Corey Sanders: Agent Lens and Forge

Corey Sanders · SVP of Product · CoreWeave

Play

Corey is SVP of Product at CoreWeave, after 20 years at Microsoft, and he wasn't on our schedule until I watched his keynote. Live demos are really hard (OpenAI's voice demo had died at DevDay), and Corey ran the whole Forge agentic loop live on stage in 10 minutes and nailed it.

“they asked me to do it in 7, I said that's not happening”Corey Sanders, on the keynote demo · as quoted in Alex's newsletter · Segment at 1:41:05

His favorite launch? Agent Lens. If you can't observe what's happening, you're dead in the water. Agent Lens is the Weave observability you know plus human-friendly insights that find small trends, like the 1% of conversations with the same problem. I've used Weave for 3 years. Chatbots, easy. An agent that runs for an hour? Fire hose. This is the fix.

Agent Lens summarizing 40,000 conversations into a table of insights
Agent Lens, now in public preview, turning 40,000 conversations into insights.

The Weights & Biases question I know a lot of you have: W&B Models is "the best product in market" and will keep getting a ton of love inside Forge. The wandb CLI? Corey admitted his own bias, "with AI, CLIs are dead," but people love wandb, the models all know it, and there are no plans to kill it. Customer first.

And Wolfram closed it with the pun of the day: "how can we loop in the audience and get them fully connected to the Forge?" Corey: forge.coreweave.com, 30-day Pro trial, no salespeople in the middle.

24 chapters · from the published YouTube video

Jump to a moment

Chapters follow the published YouTube video; the podcast audio runs on the same clock.

Hour one · the news

📡 Live from Fully Connected

ThursdAI goes live from the middle of the show floor at CoreWeave's Fully Connected in Moscone, at 11am instead of the usual 8:30am Pacific, with robots walking behind the desk and a Vera Rubin rack a few feet away.

Alex Volkov · Wolfram Ravenwolf

🏁 Cold open: OpenAI joins the AI assistant race

It's October, the last quarter of 2026, and nobody is pacing the frontier. Alex frames the week: three big models, OpenAI's biggest DevDay ever, and the story he has been yelling about for weeks going mainstream as OpenAI joins the AI assistant race.

Alex Volkov

📰 TL;DR: everything this week

The rapid-fire rundown with the full panel: DevDay's dots, GPT-6.1 Sol, UltraFast and Codex in the cloud, Claude Sonnet 5.5, Gemini 4 Argon, the White House accord, AMD and World Labs, the Jev clones, and the CoreWeave launches from Fully Connected.

Alex Volkov · Wolfram Ravenwolf · Peter Gostev · Nisten Tahiraj · LDJ

🤗 Nisten #2 on Hugging Face + SGLang

Nisten's 70MB synthetic doctor-patient dataset, generated with Opus 5.5 by 100 agents at a time across 2,200 diseases and filtered three times, hit #2 in Hugging Face datasets and #6 overall. He also flags SGLang's native support for turning any model into a Jev-style decision model, Qwen first.

Nisten Tahiraj · Alex Volkov

🎪 OpenAI DevDay recap

Alex and Peter were both at Fort Mason for OpenAI's fourth DevDay, which felt like OpenAI's WWDC: 20 releases, the biggest DevDay ever, headlined by an always-on assistant.

Alex Volkov · Peter Gostev

🔵 Dots: OpenAI's always-on assistant

A dot is an always-on ChatGPT agent powered by GPT-6 Astra, with its own computer and browser in the cloud and 4,000+ ChatGPT apps, reachable from Slack, Teams or a phone-call style interface. Wolfram loves that Sam uses dots to run OpenAI; Peter found onboarding underwhelming but the Astra-powered thread juggling great. Pro only for now, including the $100 plan.

Alex Volkov · Wolfram Ravenwolf · Peter Gostev

☀️ GPT-6.1 Sol

Five days after GPT-6 Sol, GPT-6.1 Sol keeps the $2 / $10 price but takes 95% off cached input, now $0.10 per million tokens. Artificial Analysis has it 1 point below Astra at 72 cents a task versus over $3, it beats Astra on DeepSWE at about a sixth of the cost, and it's the new Codex default. Peter's verdict: Fable and Astra no longer make sense as daily drivers.

Alex Volkov · Peter Gostev · LDJ

⚡ UltraFast, Pro 500 and Codex Cloud

A new $500 Pro tier brings UltraFast mode on Cerebras: up to 8x faster in Codex and around 300 tokens a second for Astra, at 6x the price, while the returning $200 plan comes back with half the usage. The Codex harness now runs fully in the cloud, and an Agents API with hosted computer use is in public beta.

Alex Volkov

⚖️ Decisions API

OpenAI's Decisions API takes a question and a fixed set of answers and returns one answer, fast. It runs on Luna, it's multimodal and it's waitlisted, and it looks a lot like TypeSafe's Jev. Wolfram and Peter both like keeping cheap, classify-everything workloads with the provider you already use.

Alex Volkov · Wolfram Ravenwolf · Peter Gostev

🔌 Plugins + Sign in with ChatGPT

For the fourth time OpenAI relaunches its app store, now as plugins on MCP Apps. The bigger play is Sign in with ChatGPT, which lets apps use the plan you already pay for. Wolfram brings up Google's new family assistant, and Alex says Spark should be the best email assistant in the world and isn't.

Alex Volkov · Wolfram Ravenwolf

🎼 Claude Sonnet 5.5

A week after Opus 5.5, Anthropic ships Sonnet 5.5: over 30% faster than Sonnet 5, up to 30% cheaper for most work, $2 / $10, and 70.6 on Terminal-Bench 4.0, above Opus 5.5's 66.4. LDJ calls it the new best free model for friends and family, Nisten makes it his agentic default but not for coding, and Alex sticks with Opus on the $200 plan.

Alex Volkov · Peter Gostev · LDJ · Nisten Tahiraj · Wolfram Ravenwolf

💎 Gemini 4 Argon

Gemini 4 Argon is #1 on Text Arena and the Vals Index, but it goes to government and trusted cyber defenders first. Peter has it 8th in Arena's agent mode, Wolfram points to Google's distribution superpower, and Alex hopes it wakes Google up for the assistants era.

Alex Volkov · Peter Gostev · Wolfram Ravenwolf

🏛️ White House accord on Super Intelligence

Sundar, Dario, Zuck, Greg Brockman, Elon and Jensen sign a voluntary White House accord: internal monitoring, an external auditor and board oversight, with no penalties and no regulator. A separate executive order tells federal agencies to say "Super Intelligence" instead of AI. Nisten argues for proper encryption and no backdoors.

Alex Volkov · Nisten Tahiraj

🌍 AMD buys World Labs

AMD is buying Fei-Fei Li's World Labs for about $8.2B in stock, and Fei-Fei becomes AMD's chief scientist when it closes. Wolfram reads it as robotics, LDJ as AMD bringing research in-house the way NVIDIA does with Nemotron, and Alex as at least partly an acqui-hire.

Alex Volkov · Wolfram Ravenwolf · LDJ

⏩ Quick hits

A one-minute lightning round to close out hour one's news before the assistants segment. The rest of the week's releases are in the show notes.

Alex Volkov

🤖 How we use AI assistants

At the team dinner of about 12 people, only Alex, Wolfram and one intern said a personal AI assistant helps them. Wolfram's assistant Amy planned his whole trip and talks to him through his Meta glasses, Alex had Muse chase an $8 plane Wi-Fi refund, and Amy caught a sushi delivery mix-up.

Alex Volkov · Wolfram Ravenwolf

Hour two · from the floor

🦾 Richard Ahlfeld: physical AI

Richard Ahlfeld, who leads Physical AI at CoreWeave and founded Monolith, explains AI that perceives, reasons and acts in the real world: VLA models, teleoperation versus foundation models versus world models, the sim-to-real gap, and why industrial robots arrive before the one that cleans your house.

Alex Volkov · Richard Ahlfeld · Wolfram Ravenwolf

🔁 Daniel Bolus: Forge, the loop, distillation

Daniel Bolus, ex-OpenPipe and now a senior product leader at CoreWeave, walks through the serverless side of Forge: serverless inference, serverless RL and SFT, and model distillation, which launched the day before. It takes surprisingly little data to make a model great at your task.

Alex Volkov · Daniel Bolus

🧮 Deok Filho: Sandboxes on Vera CPU

Deok Filho, senior PM for ML products and Sandboxes, explains why a GPU company is talking about CPUs: an agent is GPU plus CPU, and the new metric is agent packing. One rack of Vera CPUs runs over 20,000 agents at once on about 11,000 cores. Sandboxes went GA as part of Forge.

Alex Volkov · Deok Filho · Wolfram Ravenwolf

🚨 Breaking: serverless GPUs on CoreWeave

Deok breaks the news live: serverless GPUs, GPU sandboxes with untrusted code execution. Sign up at forge.coreweave.com, ask to have it enabled, and pay per hour with no contract and no commitment. Private preview started the day before, with more SKUs in the coming months.

Alex Volkov · Deok Filho · Wolfram Ravenwolf

🚶 Wolfram on the floor

Man on the floor: Wolfram finds a Vera Rubin NVL72 switch tray with liquid-cooled optics, bumps into Kyle Corbitt from OpenPipe, sits the crew in front of the Formula 1 simulator, and finds an NVIDIA robot dog.

Wolfram Ravenwolf · Alex Volkov

🔭 Corey Sanders: Agent Lens and Forge

Corey Sanders, SVP of Product at CoreWeave, on running the whole Forge loop live in his keynote, why Agent Lens (Weave observability plus human-friendly insights) is his favorite launch, W&B Models inside Forge, and whether CLIs still matter when agents call the APIs.

Alex Volkov · Corey Sanders · Wolfram Ravenwolf

🫡 Wrap

The assistant race is officially on, the models are good and cheap enough that the biggest ones aren't daily drivers anymore, and for AI engineers the low-key best launch at Fully Connected is GPUs on demand. Back in the studio next week at 8:30am Pacific.

Alex Volkov · Wolfram Ravenwolf

Questions people are searching this week

Quick answers

What are OpenAI dots?

Dots are always-on agents inside ChatGPT, announced at DevDay 2026. Each dot runs on GPT-6 Astra with its own computer and browser in the cloud, connects to the 4,000+ apps built for ChatGPT, works 24/7, and can be reached from Slack, Teams or a phone-call style interface. They are Pro only for now, including the $100 plan.

Is GPT-6.1 Sol as good as GPT-6 Astra?

Close. Artificial Analysis puts GPT-6.1 Sol 1 point below Astra at about 72 cents a task versus over $3, and on DeepSWE it beats Astra at roughly a sixth of the cost. It costs $2 in / $10 out per million tokens with cached input at $0.10, and it is the new default in Codex.

Is Claude Sonnet 5.5 better than Opus 5.5?

On Terminal-Bench 4.0, yes: Sonnet 5.5 scores 70.6 against Opus 5.5's 66.4. Opus still edges it on CursorBench, OSWorld and Humanity's Last Exam on Anthropic's table, and Nisten still prefers Opus 5.5 for large code bases. Sonnet 5.5 is over 30% faster than Sonnet 5 and is on Claude's free tier.

Can I use Gemini 4 Argon?

Not yet for most people. Gemini 4 Argon is #1 on Text Arena and the Vals Index, but Google is giving it to government and trusted cyber defenders first. Peter Gostev, who has access through Arena, has it 8th in Arena's agent mode.

What are serverless GPUs on CoreWeave?

GPU sandboxes with untrusted code execution that you pay for by the hour, with no contract and no commitment. Sign up at forge.coreweave.com and ask CoreWeave to enable it. It entered private preview on September 30, 2026, with more GPU SKUs planned in the following months.

What is CoreWeave Forge?

Forge brings the whole agentic loop into one place: run, observe, curate, improve and evaluate, with Weights & Biases, OpenPipe's post-training, marimo notebooks and CoreWeave Sandboxes. It has a free tier, Pro starts at $60 a month with a 30-day trial, and W&B Models lives inside it.

What is OpenAI's Decisions API?

An API that takes a question and a fixed set of answers and returns one answer, fast. It runs on Luna, accepts multimodal input and is waitlisted. It works much like TypeSafe's Jev decision model, which ThursdAI covered two weeks earlier.

What is the White House Accord on Super Intelligence?

A voluntary accord signed by Sundar Pichai, Dario Amodei, Mark Zuckerberg, Greg Brockman, Elon Musk and Jensen Huang that commits to internal monitoring, an external auditor and board oversight, with no penalties and no regulator. A separate executive order tells federal agencies to say "Super Intelligence" instead of AI.

The newsletter

Get the next one in your inbox

Next week: back in the studio, 8:30am Pacific

The assistant race is on. Don't miss a week.

If you haven't subscribed on YouTube yet, please do, it really helps. Or follow the podcast, or get the newsletter for free.