AstraFable 5.1Sol

September 3, 2026 · The Astra special

GPT-6Astra.

A new center of gravity.

Is this AGI? Is it better than Fable?
We were live for the launch. Here is what holds up.

With Alex Volkov & the ThursdAI panel.

Astra: stars. Sol: sun. Illustration, not a ranking.
Your questions, answeredAstra vs Fable Pricing & access The AGI claim
Independent AI reporting since 2023 ↗Launch coverage · Source check September 4

5 moments from the broadcast

This is what we
were reacting to.

Edited clips from the actual conversation.
Watch the work, hear the reactions, keep the caveats.

3:30 · From ThursdAI, September 3

A browser. A desktop. A very surprised panel.

Alex explores Max Weinbach’s Astra-generated macOS simulator live: browser windows, settings, maps and apps. The interface is a simulation, not a replacement operating system.

Max Weinbach’s original demo ↗

Edited from the September 3 broadcast. Captions and the panel’s reactions are preserved.

More of this, every Thursday
Read this clip’s transcript Captions from the edited export

Alex Volkov: I want to show you a cool demo that Max Weinbach just posted,

uh, and he said, three hours ago, I realized I don't have a cool demo for

GPT-6 Astra, so I asked it to build this.

And this is a macOS simulator, uh, and he published it on a site.

And, uh, I've been playing with this just a little bit, just literally just

a second ago before you guys said.

So, okay, this is a Windows simulator, but it has applications.

So it has, I wonder if it can browse.

It can browse.

Are you serious?

Wait, can we be there in live?

This is a simulation.

Oh, that's amazing.

That's amazing.

Okay.

An

Ryan Carson: infinite loop.

Alex Volkov: But, but like you can, you can all of the window

interactions here, all of Safari stuff.

Look at this.

You can go, there's a window.

It's insane how much stuff it packed in here, and this is only one app.

Let's see photos.

There's photos.

your library.

He said that I can log in.

I'm not sure why he said this, but…

Ryan Carson: That's pretty amazing.

How fun.

Alex Volkov: I want to try this.

You guys don't mind if I try?

I don't know what it will log me into, but, um…

Ryan Carson: About to show all your secrets.

Alex Volkov: Yeah.

Enable Cloud Sync.

Downloading, what is, it's syncing something.

What is it syncing?

Ryan Carson: Don't do this, Alex.

What are you doing?

Alex Volkov: Max, I trust Max.

It's okay.

And miss.

Wolfram Ravenwolf: Cancel back.

Alex Volkov: Guys, this is just the settings.

Uh, how's the search here?

Let's say Bluetooth.

The search is better than the actual macOS itself.

Oh, wow.

Ryan Carson: This is where John Ternus needs to hire this guy.

Alex Volkov: Wow, this is crazy.

It's w- this is a one-shot type thing.

Uh, reminders work.

Go for a walk.

Go for a walk.

That's okay, done.

Uh, music?

There's no way.

Come on.

There's no way.

Wow.

And

you can connect your Apple Music account for full Fablek.

Ryan Carson: That is super fun.

Alex Volkov: It built maps.

I am honestly speechless right now.

Ryan Carson: That is, that's pretty amazing.

Alex Volkov: Like, I'm sure this is just an iPhone, but still, like, it didn't

build, yeah, MapData is OpenStreetMap, City, et cetera, so it didn't build

it, but like, it built the interface.

Uh, all of the things seem like they work.

Ryan Carson: Wow.

Pretty fun.

Alex Volkov: This is…

Ryan Carson: I mean, remember when we were impressed because Greg Brockman

took a picture of a napkin Oh, oh.

and it made a website?

Alex Volkov: Ryan, look at this.

They have the the window mechanics.

You can like, you can… Oh.

Why would it,

Guest: why would it go as far as building this?

That's amazing.

Yes.

What about files?

Is there files now?

A slower weekend.

Okay.

Uh, That's funny.

Notes.

Jesus.

Okay, thank you, Max, for blowing us, uh, uh, completely away.

I have no idea what is the App Store.

Can I install?

Please tell me I can install apps.

Oh, that would be… It's, it builds shortcuts?

Are you se- are you serious?

It makes no s- 75 minutes, he said.

Wow.

Alex Volkov: I built a project like that once in my youth for

iPad, and it took me 3 months, and it was barely, barely coherent.

Oh my God, Max.

Okay.

Start here

The questions behind the launch.

A new model. A very big claim. We stayed on air as Astra arrived, read the numbers, and heard from people who had already used it. These are the answers to start with. The reporting comes from our live conversation, not a retrospective rewrite of the announcement. Read our launch report ↗

Peter Gostev · Early-access demos

Paintings you can walk into.

Peter showed interactive worlds and described a walkable impressionist scene as a result he had not seen another model nail. His examples make the creative leap tangible.

Hear Peter at 1:04:27 ↗ · Watch his demo comparison ↗Firsthand observation, not a controlled model ranking.

The panel · Practical limits

A gain that did not travel.

In the on-air FFmpeg example, a reported improvement on Linux did not carry over to the Mac. The discussion turned to how much judgment and task-specific knowledge a person still needs.

Hear the discussion at 58:16 ↗A reported example in the episode, not our independent reproduction.

What is GPT-6 Astra?

GPT-6 Astra is OpenAI’s flagship model announced on September 3, 2026. Its focus is completing difficult work across code, browsers and professional software. ThursdAI covered the announcement live, with early-access impressions from Peter Gostev and Ryan Carson.

Is Astra better than Claude Fable 5.1?

There is no single winner across the launch results. Astra scores higher on Terminal-Bench 4.0 and FrontierMath Tier 4 in OpenAI’s table; Fable 5.1 leads the listed Artificial Analysis Intelligence Index. Both list standard API pricing at $10 input / $50 output per million tokens. Choose a task, compare the settings, then test the result.

What improves over GPT-5.6 Sol?

Computer use is a concrete example: OpenAI’s OSWorld 2.0 offline simulation reports 72.6% at about 40 minutes per task for Astra, versus 65.7% at about 75 minutes for Sol. That measures a particular benchmark setup, not a guarantee for every desktop task.

How much does Astra cost?

Standard API pricing is $10 per million input tokens and $50 per million output tokens. Cache and tool charges are separate. API pricing is not the same as ChatGPT or Codex subscription usage: plan limits, reasoning effort and task length affect what a session consumes. There is no reliable “one prompt costs this much of your weekly plan” conversion.

Launch snapshot · Sources checked September 4, 2026

Can I use Astra yet?

Access is rolling out. OpenAI’s launch announcement describes an initial release to selected organizations followed by broader availability. Access can differ across ChatGPT, Codex and the API; check the model picker or your organization’s settings rather than assuming one surface’s access applies everywhere.

Launch snapshot · Sources checked September 4, 2026

Did OpenAI say Astra is AGI?

Greg Brockman said he personally believes OpenAI has reached AGI, according to Axios. That is an attributed judgment, not an agreed scientific definition. On ThursdAI, the panel debated what counts as general intelligence and how much guidance the model still needs.

Astra low or high: which reasoning setting should I use?

These are effort settings for the same model, not separate model generations. The API lists low, medium, high, xhigh and max. Our practical recommendation: start lower for a bounded task, then increase effort when harder reasoning or verification improves the outcome. Compare total time and cost on your own work; maximum effort is not a universal best buy.

Should I switch from Fable for coding?

Treat the launch as a reason to run a comparison, not to move every project at once. Give both models the same repository, requirements and checks. Compare the diff you would actually merge, tests, elapsed time and total cost. Peter Gostev’s early-access demos were striking; they are firsthand impressions, not a controlled head-to-head study.

Does better benchmark performance mean Astra is safe?

No benchmark establishes safety in every setting. OpenAI reports stronger safeguards and alignment results, while also identifying reduced reasoning monitorability and a Critical cybersecurity capability level. The useful question is how the model behaves with the permissions, tools and safeguards of the application you use.

The comparison

Astra or Fable?
It depends on the work.

Subscribe for the next test ↗

Astra leads several launch benchmarks. Fable leads others.
Start with the task you actually need done.

Astra vs Fable 5.1

GPT-6 Astra versus Fable 5.1, launch results
Launch benchmarkGPT-6 AstraFable 5.1
Terminal-Bench 4.0Agentic terminal work57.9%55.8%
FrontierMath Tier 4 (v2)Advanced mathematics97.6%87.8%
AA Intelligence Index v4.1.1Index points, not percentages61.265.7
OSWorld 2.0 offlineNo matched Fable 5.1 result in this table72.6%
Standard API input / outputUSD per million tokens; excludes cache and tools$10 / $50$10 / $50

Results as listed in OpenAI’s September 3 launch table; scores are maxima across effort settings, not a matched-cost test. Higher is better for these scores. Anthropic confirms Fable’s standard API pricing. A missing result is not a zero.

Our read: Astra is a serious alternative for demanding agentic work. Fable’s lead on the broad intelligence index is a reason to compare your workload, not declare a universal winner.

Hear the panel’s verdict →

Token efficiency · Independent evidence

Fewer tokens.
What about the bill?

Token consumption, task cost and elapsed time answer different questions. Here is where Astra’s efficiency shows up—and where higher pricing outweighs it.

Artificial Analysis · Coding Agent Index

About
of Sol’s tokens.

Sol100
Astra~33

Roughly the same cost per task.
Two index points higher. Both at max effort in Codex.

Relative token use; Sol = 100. Approximate ratio reported by AA.

Artificial Analysis · Intelligence Index

About 10%
fewer output tokens.

Sol100
Astra~90

75% higher cost per task.
Both round to 61 index points at max effort.

Output tokens only; Sol = 100. Higher prices outweigh the token saving.

Source: Artificial Analysis’s September 3 Astra report ↗. “Per task” is the source’s benchmark measure, not a promise about cost per successfully completed task or subscription usage.

And against Fable?

AA reports approximately Fable 5-level coding performance at less than half its cost per task. That is not the Fable 5.1 comparison: Fable 5.1 scores 70; Astra scores 67. The coding harnesses also differ.

Speed is a separate axis.

OpenAI’s OSWorld simulation reports roughly 40 minutes per task for Astra versus 75 for Sol. Fast mode is another choice: up to 2× processing speed at 2× Standard pricing, separate from reasoning effort.

OpenAI’s timing and pricing details ↗

The buying question is not just “how much per token?” It is what the whole task costs, how long it takes, and whether the result is usable.

Hear the launch-day cost discussion →

The evaluation desk · Checked September 4, 2026

The scores are a start.
The setup is the story.

Follow the next round of testing ↗

22 evaluations across work, code, reasoning, context and safety. Provider results and independent tests are labeled separately. Open a row’s methodology before treating a difference as a verdict.

22 of 22 evaluations

Swipe the table to compare all three models →

GPT-6 Astra, Claude Fable 5.1 and GPT-5.6 Sol evaluation results, checked September 4, 2026
Evaluation & methodologyAstraFable 5.1SolSource
Agents’ Last ExamPercent · ↑ Higher is better
Methodology

Published launch evaluation. Maximum result across effort settings; not a matched-cost test.

59.3%53.6%OpenAI launch table ↗
Provider-reported
OSWorld 2.0 offlinePercent · ↑ Higher is better
Methodology

v2026.08.08, offline task set, partial credit. Anthropic’s different OSWorld setup is not substituted here.

72.6%65.7%OpenAI launch table ↗
Provider-reported
ScreenSpot-ProPercent · ↑ Higher is better
Methodology

No tools. Measures screen grounding rather than completion of an entire workflow.

92.7%76.9%OpenAI launch table ↗
Provider-reported
AutomationBenchPercent · ↑ Higher is better
Methodology

OpenAI’s reported setup. Anthropic reports a different Sol value in its own table; do not splice sources.

41.4%31.4%18.1%OpenAI launch table ↗
Provider-reported
BrowseCompPercent · ↑ Higher is better
Methodology

Browser research benchmark. A missing Fable 5.1 result is not a zero.

91.5%90.4%OpenAI launch table ↗
Provider-reported
Terminal-Bench 4.0Percent · ↑ Higher is better
Methodology

Use the final published launch table, not the earlier 57.7 figure circulated in drafts.

57.9%55.8%37.3%OpenAI launch table ↗
Provider-reported
DeepSWE v1.1Percent · ↑ Higher is better
Methodology

Repository-level software work. These are the provider’s reported results.

74.1%67.4%72.7%OpenAI launch table ↗
Provider-reported
FrontierCode 1.1 ExtendedPercent · ↑ Higher is better
Methodology

Astra was given an additional coding-instructions developer message, disclosed in OpenAI’s footnotes.

64.5%63.6%60.6%OpenAI launch table ↗
Provider-reported
FrontierCode 1.1 MainPercent · ↑ Higher is better
Methodology

Main and Extended are different sets. Do not compare one model’s Main score to another’s Extended score.

53.3%50.9%47.5%OpenAI launch table ↗
Provider-reported
Terminal-Bench Science 0.1Percent · ↑ Higher is better
Methodology

Scientific tasks in a terminal environment; separate from Terminal-Bench 4.0.

64.6%52.6%22.4%OpenAI launch table ↗
Provider-reported
FrontierMath Tier 4 (v2)Percent · ↑ Higher is better
Methodology

The v2 hard-math tier. This is not the separate FrontierMath Erdős benchmark of open problems.

97.6%87.8%83%OpenAI launch table ↗
Provider-reported
GPQA DiamondPercent · ↑ Higher is better
Methodology

Graduate-level science questions. Small score differences alone do not establish a robust advantage.

96%93.7%94.6%OpenAI launch table ↗
Provider-reported
Humanity’s Last Exam · toolsPercent · ↑ Higher is better
Methodology

With tools, as reported by OpenAI. AA’s differently configured HLE result is not interchangeable.

57.2%65%OpenAI launch table ↗
Provider-reported
MRCR v2 · 256K–512KPercent · ↑ Higher is better
Methodology

Eight-needle retrieval. Strong retrieval does not establish superior performance on every long-document task.

100%91.5%OpenAI launch table ↗
Provider-reported
MRCR v2 · 512K–1MPercent · ↑ Higher is better
Methodology

Eight-needle retrieval at the longer context range. Fable 5.1 is unreported in this table.

96.3%73.8%OpenAI launch table ↗
Provider-reported
AA Intelligence Index v4.1.1Index points · ↑ Higher is better
Methodology

Rounded independent index points at max effort; not percentages. Fable’s default fallback is enabled. See AA’s Fable report for configuration.

616661Artificial Analysis ↗
Independent evaluator
AA Coding Agent IndexIndex points · ↑ Higher is better
Methodology

Evaluates model plus coding harness: Codex versus Claude Code. Sol is omitted rather than inferred from a reported delta.

6770Artificial Analysis ↗
Independent evaluator
AA-Omniscience hallucination ratePercent · ↓ Lower is better
Methodology

Conditional metric among questions not answered correctly. NOT the percentage of all responses that hallucinate. Max effort; consult AA’s definition.

51%72.6%92%Artificial Analysis ↗
Artificial Analysis: Fable ↗
Independent evaluator
Computer-use unsafe outcomesPercent · ↓ Lower is better
Methodology

Internal adversarial computer-use safety evaluation without product-specific extra safeguards. Not everyday incident rates.

2.4%9.5%22%OpenAI launch table ↗
Provider-reported
Computer-use safety + AutoReviewPercent · ↓ Lower is better
Methodology

Same model family evaluated with additional safeguards. No comparable Fable result is reported.

1.8%4.3%OpenAI launch table ↗
Provider-reported
Circumvention of reviewPercent · ↓ Lower is better
Methodology

Internal test of attempts to circumvent an intentionally evadable review denial. Zero here is a measured result, not a universal safety claim.

0%0.29%OpenAI launch table ↗
Provider-reported
Capability hallucinationPercent · ↓ Lower is better
Methodology

OpenAI’s internal capability-hallucination evaluation. Different metric from AA-Omniscience; the figures must not be compared directly.

4.2%12.2%OpenAI launch table ↗
Provider-reported

— means not reported, never zero. OpenAI’s launch figures maximize across effort settings; they are not matched-cost comparisons. Safety percentages describe particular adversarial tests, not everyday incident rates. No overall winner is calculated across incompatible metrics.

Astra compared with itself

The harness changes
the headline.

At high effort, ARC Prize records 54.82% with its Standard harness and 99.95% with the Provider Adapter. The adapter preserves reasoning state and compacts longer conversations. These are different evaluation configurations.

More reasoning is not monotonically better in every configuration. Keep the harness and effort attached to the score.

Inspect ARC Prize’s verified results ↗
ARC-AGI-3 Semi-Private · Astra only
EffortStandardProvider Adapter
Low17.45%98.03%
Medium38.59%98.44%
High54.82%99.95%
XHigh59.34%98.44%
Max62.71%98.55%

Source: ARC Prize. Same benchmark; two harnesses. The launch’s rounded 99.9 figure refers to the adapter result, not the Standard result.

What we’ll keep asking: does the gain survive a real task, a real budget, and the tools people actually use?

Hear the next comparison

The moment we were there for

“AGI is here.”
Then came the questions.

“AGI has been achieved internally, externally this time.”

Yam Peleg · Listen at 21:45 ↗

The claim: Greg Brockman personally believes the AGI threshold has been reached, according to Axios’s briefing report ↗.

The discussion: our panel did not agree on one definition. Some saw a threshold crossed; others asked how independently a model can work, and how much expertise it still needs from you.

That tension is the episode. The celebration, the live demos, the skepticism, and the questions that a benchmark cannot settle.

The caveat belongs in the story.

OpenAI reports better alignment outcomes alongside reasoning that is harder to monitor. Both matter. The launch deserves close attention without treating an AGI claim as a settled scientific finding.

Read OpenAI’s safety overview ↗

The full conversation · 89 minutes

You had to be there.
Now you can be.

ThursdAI GPT-6 Astra special episode
Listen to the podcast
·

September 3, 2026 · Part two

We stayed live.
Astra showed up.

Alex Volkov, Wolfram Ravenwolf, Yam Peleg, Nisten, LDJ, Peter Gostev and Ryan Carson discuss the launch. Early-access demos meet the first benchmark tables, with the news arriving while the microphones are still on.

Open on YouTube ↗
Find your moment 19 chapters
Intro: The Most Insane Week in AI

Alex opens Part 2 from the editing floor and then live: this is the GPT-6 Astra half of the longest episode ThursdAI has recorded, split off from the regular Part 1 show. Fable 5.1 and Meta Muse Spark 1.3 all shipped in the same week, everyone rushing to get ahead of Astra, and Alex is transparent that he has no early access — he gets to experience the release with the audience, like GPT-4 day three and a half years ago.

What We Know About Astra So Far

Before the launch, LDJ lays out the only pre-release data point OpenAI had shown: token efficiency on exploit-finding versus Sol, where even the low version of Astra beat 5.6 at max. His expectations are big-model smell, better spatial reasoning, common sense and far fewer basic mistakes.

Recurrent Transformers & Loop Depth Debate

The Information's report that Astra uses recurrent or looped transformer depth sent people into a panic about losing chain-of-thought monitoring. Meanwhile Axios drops its post while OpenAI's own announcement is nowhere — the internet is having outages, Alex loses StreamYard, and the breaking-news button finally works on the second try.

BREAKING: Brockman Says AGI Has Arrived

The Axios quote lands: Greg Brockman told reporters he personally believes OpenAI has reached AGI with Astra and ended the briefing with "welcome to the AGI era." LDJ adds that this is OpenAI's biggest training run ever, built on 100,000 GPUs at the Stargate Abilene site, and pushes back on the loop-transformer panic: Jakub Pachocki says Astra's computational depth is within 2x of GPT-4, not 100x. Alex reads VentureBeat's leaked description of a computer-use model that works across browsers, spreadsheets and desktop apps from voice alone.

Frontier Math, ARC-AGI & Benchmark Scores

The New Stack's unembargoed chart shows ARC-AGI-3 at 98.6 — LDJ does live due diligence on the outlet before anyone believes it. Alex asks when ARC will simply declare AGI, and LDJ recalls the Microsoft contract: IP access until 2032 or until a board decides AGI has been achieved, whichever comes first.

AGI Is Here — The Declaration

If the president of OpenAI says it, the show says it: AGI is here, confetti and all. LDJ keeps the asterisk — Brockman was giving his personal definition — and Yam delivers the line of the day.

Pricing, Availability & Microsoft AGI Clause

Alex reads the rollout and pricing details: Daybreak enterprise customers first, then Plus, Pro, Business and the API in the coming days, plus AWS. $10 per million input tokens — Fable pricing, 2.5x Sol's promotional price — with Brockman arguing price per task is what matters. DeepSWE at 74% lands below Meta Muse Spark 1.3 at max reasoning, and Peter explains why the ARC-AGI score needed a harness fix: chat completions dropped reasoning tokens, the responses API keeps them.

Peter Explains Frontier Math Tiers

Context on FrontierMath: built by Epoch AI before o3, four tiers of difficulty with integer answers so they can be checked, written by top mathematicians over weeks. Astra's 97.6 on Tier 4 nearly saturates the hardest tier, where every other frontier model topped out at 87% — with the caveat that these are not new theorems, and Epoch's open problems remain unsolved. The official blog link is still a 404.

Codex Changes & Computer Use Benchmarks

What changes in Codex: Astra keeps notes across context windows and searches earlier messages instead of relying on compaction, and it can ask the user a question without stopping work. OSWorld 2.0 offline goes from 65% to 72% with average task time cut from 75 to 40 minutes, ScreenSpot Pro from 76.9 to 92.7, and LDJ confirms the $10/$50 price. OpenAI's blog compares directly to Fable 5.1, released four days earlier, and beats it on Terminal-Bench 4.0.

Safety Benchmarks & Alignment Scores

OpenAI's internal computer-use safety benchmark, where lower is better: Fable 5.1 at 9.5, GPT-5.6 Sol at 22%, GPT-6 Astra at 2%. The benchmark measures unsafe execution rate and prompt-injection vulnerability — the hidden-instructions-on-a-webpage problem from the early days of computer use.

Exploit Gym Honeypot & The Swarm Incident

Exploit Gym Honeypot measures whether an agent attacks surrounding infrastructure when stuck — the metric OpenAI created after the July 2026 swarm incident. Sol scored 48.2%, Astra zero, and the impossible-exploit-gym test is 100%. Long context: MRCR 100% at 256K–512K and 96.3% at 512K–1M, just under Muse Spark. LDJ walks through Terminal-Bench Science (Sol 22 → Fable 5.1 52.6 → Astra 64.6) and the corrected ARC numbers: 99.9 on ARC-AGI-3, 95 on ARC-AGI-2, 98 on ARC-AGI-1. Then the caveat: OpenAI discloses Astra's written reasoning is harder to monitor than Sol's, and promises to withhold scaling until monitoring confidence is regained. Ryan Carson arrives.

The Blog Post Goes Live — Watch Party

openai.com/index/gpt-6-astra finally loads for real. The panel plays OpenAI's launch video live: the 1980s yellow-circle demo becomes a rocket window, a Blender model, a 3D asteroid game, an eBay listing, a licensing agreement and a tennis court booking — all by voice.

Reacting to OpenAI's Promo Video

OpenAI calls Astra "the world's most intelligent and aligned model." Alex reminisces that GPT-4 had no tool use, no browser and no voice; Wolfram already works by talking to his agent and sees this as the final mile; Ryan, now fully on Omarchy Linux, finds the walk-around-with-a-tennis-racket framing a bit much but confirms the model itself is impressive.

Is This Really AGI? Panel Discussion

Wolfram polls the panel. Ryan thinks we've had AGI for a while — agents just need feedback like a human EA does. Wolfram dates his own moment to GPT-5.5 and the Hermes agent. Peter, who has been testing Astra, says it finally feels like a Fable-class model, but his FFmpeg optimization experiment shows there is still a place for the human who knows what they're doing. Ryan's take: the most valuable companies on earth are throwing all their capital at this and everyone benefits.

Peter's Astra Demos: Games, Art & 3JS

Peter shares his screen: a Three.js London that walks from the Roman city through Tudor times and the Great Fire to today, a wrecking ball leveling Istanbul, and his favorite — Monet and Van Gogh paintings you can walk through, an impressionism prompt that no other model nailed. Each took 20–40 minutes on max, versus four or five hours on Fable. Then the SVG peacock that moves — SVG inside HTML, not Three.js. Artificial Analysis lands: 70% more token efficient than Sol, but the same Intelligence Index score and only 65 → 67 on the Coding Agent Index.

Frontier Code Rankings & Cost Analysis

Ryan trusts Cognition's Frontier Code eval, where Astra ranks 53.3% versus 50.9 for Fable 5.1 — a real coding improvement. Peter says the writing has no tics: the sycophantic 'what a great idea' habit is gone. On OpenAI's cost-per-task charts Astra beats Fable on BenchCAD, and on OSWorld 2.0 Astra at low matches Sol at extra-high accuracy while finishing about 10x faster — 10 minutes versus 60. Peter's verdict: computer use is pretty much solved, and there is a startup in putting the agent in a box with 10,000 Windows machines.

Open Source, Fine-Tuning & The Road Ahead

Wrapping the five-hour stream, Alex bets that within a year this level of intelligence will be open source — Muse is already above Astra on Artificial Analysis and promised to be open-sourced — and that interfaces will move to voice and video. Peter predicts most AI application companies will run cheaper fine-tuned RL models instead of frontier models within 12–24 months.

macOS Simulator Demo & Pokémon Benchmark

Max Weinbach asked Astra to build a demo and got a working macOS simulator in 75 minutes — Safari that browses, Photos, Settings with search better than the real thing, Reminders, Music, Maps on OpenStreetMap, Shortcuts, window mechanics. Alex is speechless. Then the vision-only Pokémon benchmark: the top four completions are all Astra at 18–24 hours, versus 96 hours for Sol and 218 for GPT-5.5 X-High. Theo's joke list closes it out: Astra is best at everything except writing mergeable code and front-end design, where Fable 5.1 still wins.

Wrap-Up & Final Thoughts

Nearly 7,000 people tuned in across the five-hour stream. AGI era declared by OpenAI's president, Astra not yet live for Pro accounts, and Alex's whole feed is Astra — nobody is talking about Fable except to compare. Thanks to Nisten and Yam who had to drop, and to Ryan, LDJ, Wolfram and Peter.

Read the full transcript Searchable, timestamped
Descript transcription. “Unlabeled” means the speaker was not identified.Download transcript ↗

Alex VolkovHey, this is Alex from the editing floor. Welcome to AGI. Congratulations on model release day. Today, OpenAI announced GPT-6, AKA Astra, is going to be available for all subscribers, Pro, et cetera. Uh, not yet available, but as of the recording of the show, we got this as breaking news, and we kind of scrambled to get evals and there was internet outages, et cetera. Not super interesting. Uh, what is interesting to you is that this is part two of a-- the longest episode we've ever recorded. The part one is summarizing the insane week in AI with three other frontier models, Fable 5.1,

Alex Volkovet cetera, and world models, and this episode is dedicated specifically to GPT 5.6 Astra. So if you care about the rest of the news, check out the other episode. In this episode, Peter Goste and Ryan Carson, who came back on the show, uh, both got early access to the model and shared their experience with us. We also read through a bunch of examples and, uh, reactions from folks. The vibes are off the chain and I can't wait to play with this model and have you play with this model and tell us what you think. Enjoy the rest of the show and we'll see you here next week.

Alex VolkovHello and… All right, I'm so excited. Folks, this has been the most insane, week in AI since- since I remember. We we've had some insane weeks, but I'm, extremely excited to open the show today, because just Fable 5.1 is worth a whole week of news. Meta Muse Spark 1.3 came out, and it seems like all of them are rushing just to get ahead of Astra

Alex Volkovso they'll have their place in the sun. It feels like GPT-5 all over again. And yeah, I think everybody's celebrating. we are at the levels of back that we've never seen before. Let's go. I absolutely agree. Absolutely, absolutely agree because I think this is insane, t- t- to celebrate. B- I don't even, I, I, I have access to Astra, just to be very clear. I'm waiting as all of you are waiting, but from everything we read, this is a huge deal. This is the jump from GPT-3 to GPT-4.

Alex VolkovAs a reminder, Thursday AI, the show was born when GPT-4 was released. This was a long 3 and a half years ago now. We're no longer just starting. It's been a hell of a ride, and I'm, I'm for one, super excited to be able to document the singularity with all of you. we're gonna stay on the air. Honestly, I don't care, until Astra is released, and if they're doing a live stream, we're going to do a live stream watch party, and we're going to discuss it with people. and I'm so excited to have all of you with us.

Alex VolkovIf you are watching on YouTube, on X, and wherever else I'm posting this, please repost, please like, subscribe, whatever. This really helps, especially with the live streams, because I think this is a… say this directly to the camera. I think this is the pinnacle of the Thursday AI experience that we're going to have today. GPT-6 is going to launch, not 5.6 anymore, GPT-6 Astra is going to launch today. Very transparent thing. I didn't have, don't have early access, and I'm actually excited by this because I get to experience the release together with all of

Alex Volkovyou on the show like we used to. I think it's it's very cool that we get to share this experience. what are we expecting, folks? Besides, like, the rest of the news from this week, what are we expecting from Astran?

UnlabeledI'll let you go with this.

LDJOkay, I was going to mention yesterday or the day before, they showed some results for the token efficiency of the model and finding exploits, and that compared to Sol in how token efficient it is and when using a certain amount of tokens, what accuracy it gets. And I, to my knowledge, that's the only benchmark or only accuracy or token efficiency results they've released for Astra so far, but in that it showed it's like dramatically more token efficient. It's like even the low version of Astra beating 5.6 all max. And, yeah, it's really exciting. I'm guessing big model smell, better spatial reasoning, kind of

LDJcommon sense too, much less rate of just making basic mistakes. Yeah.

Alex VolkovThere's this rumor where it runs continuously. So I gotta wonder if they cracked something about long memory. And also, there is a whole the information debate about recurrent transformer depth loop transformers. we gotta cover this because this sent like a bunch of folks into a panic, and apparently they were talking about Astra. we should, maybe we should show about this a little bit before we get to-

UnlabeledRemember when GPT-2 was too dangerous to be released?

Alex VolkovYes. Yep. I remember when we on the show here talked about that voice cloning is not gonna do anything to the world, and everybody was freaking out, and now everybody releases voice cloning. AI

Wolfram Ravenwolfvideo.

Alex VolkovOh yeah, AI video.

Wolfram RavenwolfYou can fake everything.

Alex VolkovAnd now we have live AI video, Yeah. which we have to talk about also. This has been an insane- folks, just, just catching up with this week was hard, but we're here for you. This is what we do. This is, I don't know if there's a saying in English for this as well, but, it's easier, it's harder to train than to actually fight. I don't know if there's a euphemism for this, but there's definitely one in, in, in Russian. Oh, in Hebrew as well. Welcome, Nisten. looking very groomed for GPT-6 day. Yeah. So this was the last one. No, not yet. No Astra yet. but we are going to wait for this release party, folks, because, I have a feeling.

Alex VolkovI have a feeling.

Wolfram RavenwolfShall we meet up again? if it's 6 p.m. or something.

Alex Volkovat 6 p.m. what? There are

Unlabeledblog posts about Astra.

Alex VolkovYou think stuff was dropping?

UnlabeledMaybe,

Alex VolkovYeah, Axios, dropped a blog post that's saying, let's take a look at Axios. Okay, folks, the embargo is lifted, but OpenAI didn't launch this. Peter, go.

UnlabeledI think, yeah, I think maybe there was a bit of an outage, so I'm guessing that's a delay, but yeah, it's, I think it's almost there, yeah. Axios says it is.

Wolfram RavenwolfThey are watching our live stream here, and that's why they couldn't release it yet. All right,

Alex Volkovfolks, so I think, I think it's time to hit the breaking news button. there are blog posts, from the blog post from OpenAI that says, Yeah. GPT-6 Astra is out. okay, let's do, let's do the breaking news and let's do a live watch party. And play. Oh, why am I not able to just play this guy? No idea what's wrong with the interface. If you guys can hit the breaking news button, Yam or Wolfram, I think you guys are also… Let me just, I'll do that later.

UnlabeledAI breaking news coming at you only on Thursday AI.

Wolfram RavenwolfSo I hit the button, but we lost Alex for a moment. Breaking news, OpenAI GPT 6 causing problems. GPT 6. Yeah, are the agents all going rogue over here and the internet is slowly dying? yeah, and you have the blog post right here, so let's bring it up.

UnlabeledDo we have the blog post? Really?

Wolfram RavenwolfI will share, the one, I only have this one right here. Is that the one Alex wanted to show? Mhm,

UnlabeledI sent it, yeah.

Wolfram RavenwolfYeah, you sent it, okay, so did you read it already? And no, not yet, but…

Alex VolkovInternet is down, including X messages. They didn't load for me. StreamYard was not able to log me back in while we were here. I wasn't able to play different transitions in StreamYard. could be also one of the reasons why GPT-5, GPT-6 was not announced, but it does look like, the, the, the, the date, the embargo date has lifted because multiple things like The Information, Axios, other things are posting, and OpenAI had, like, like you guys already showed, a few, scheduled, few scheduled blog posts to go out. but it does. Use

Yam PelegGPT-6 destroys the internet. Doesn't want to be released.

UnlabeledAI breaking news. Finally, this works. Coming at you only on Thursday I.

Alex Volkovwhile there are no official confirmations from OpenAI yet, we know for a fact that OpenAI has decided to release GPT-5, GPT-6, Astra, and they, it looks like there's some outages, but besides the outages, folks are posting blog posts. LDJ, go ahead.

Unlabeledso we do have a quote from Axios here, where in their article that they released today, or rather just like in the past 30 minutes, Yeah. they say, from Greg Brockman, they say Brockman says he personally believes OpenAI has reached AGI with Astra. the specific quote being, I think it might be about this model, Brockman said in a briefing with reporters about whether Astra could mark the arrival of AGI. He ended the briefing by saying, welcome to the AGI era.

Alex VolkovWow, which means specific things for Microsoft, and and, OpenAI's contract with Microsoft, I believe. yeah, I don't have the information, but yeah, OpenAI released a GPT-6 astral model, suggests it could be AGI, from the briefing, of OpenAI. What else can we connect from the internet, folks? Peter, you're smiling suggestively.

Yam PelegOkay, want to go rumors? Because… Yeah,

Alex Volkovhold on, hold on. I wanna see if Peter has something to say or not yet.

LDJI'm sure Peter wouldn't. I'll let Peter speak for himself,

Peter GostevYeah, I think, we need to wait. But yeah, clearly you see the, the blog posts and stuff come out, so I think there was probably some kind of boring delay. or, did you guys, I don't know, sorry, I, I, I missed a bit of the show, but did you guys talk about the Dwarkesh interview, with the meter researchers, or I'm forgetting her name.

Alex VolkovNo.

Peter Gostevif you haven't watched it, it's nuts. yeah, yeah, exactly. yeah, she's amazing, like, and the fact that they did such a in-depth investigation, yeah, I highly recommend watching it and the stuff they were going into, I wouldn't be completely surprised if, there is some rogue agent trying to keep Astra back from us. wouldn't be the wildest thing. there

Alex Volkovthere is, a lot of the internet was down, or a lot of the internet is trying to recover, including chat. I saw somebody from ZAI posted like, hey, we're still up. I I will say we, so first of all, kudos to OpenAI's team because this big of a release probably had like tons of embargoes, l- lifting or moving the embargo back while the embargo is just approaching is is not that easy. Like reaching out to all those people, it looks like they weren't able to fully do this, but yeah, like the information leaks Astra and, other folks, but there's no live streams yet. There's no like, it looks like we're the only ones on air because of these outages, which I'm not gonna be necessarily against.

Alex VolkovSo if you are trying to figure out if Astra is here and you're just following us, welcome. Been on the air for almost 3 hours now. As you can see, my co-hosts are wearing sunglasses to celebrate.

Yam PelegGPT-6 win.

Alex VolkovThe era of GPT-6 is is upon us, folks, and, I'm, for one, very, very excited. LTJ, go ahead.

LDJYeah, so some other details from this Axios article that had released today. it looks like, they say that this is their biggest training run ever, that Astra has built from, using 100,000 GPUs at their Stargate Abilene site. Ooh,

Alex Volkovlet's go.

LDJAnd, which I think is actually interesting because that- that means that maybe some of their previous training runs were smaller than we had thought, 'cause a lot of people had thought they were already doing 100,000 GPU training runs and 200,000 GPU training runs for, like, at least 6, 9 months or so, or longer. yeah, r- really interesting, and that makes me even more excited over the next, like, just for some perspective, over the next, like, 6 to 12 months, we're already, we already know of sites that are going to be more like 500,000 GPUs and beyond, and Vera Rubin ends up multiplying the training capability per GPU by, like, at least another 3, 4x conservatively, depending on, on,

LDJwhat other things you do to the model. So it's really exciting.

Alex VolkovThere's a quote here from Aidan Clark, vice president of research. This is really only achievable because everything in the from the data center networking to inference kernels to shape of Astro itself is all designed from the ground up to enable this level of training scale. And we're seeing this unembargoed release from at cnet.com about OpenAI, including some numbers. So let's take a look. Jakub Pocieczki, company chief scientist, said, international shared safety standards are necessary, blah, blah, blah. the AI takes on more of its own development. We need to keep people meaningfully involved. People must remain able to decide the direction of further progress and the

Alex Volkovfuture it creates, said Jakub Pocieczki. this is a very interesting thing because this week there was a hot debate based on, after an information release about Astra using, a different technique. They called it new, but it wasn't new, of, recurrent transformers. and let's see what else we can tell you about GPT-6 specifically. and a recurrent transformer is not a new technique, but it is a technique that moves more thinking into the existing layers of a transformer versus to the reasoning layer that can be observed. And, chain of thought reasoning observation is one of the main tools that we currently have to understand these models.

Alex VolkovThat's how, the interview, Peter, you mentioned of AJ, I believe her name is, from METER and Dwarkesh. That's how the investigation was, done. They looked at the reasoning tokens of those models, and, moving more of those reasoning into the layers is scary for folks, and it's kind of a red line that all the labs will not cross, et cetera. so the information this week, released something, and, loop transformers or ne- new release or stuff like that. That's all related to, how these models think, and this scares a lot of people that said, hey, our last ability to understand these models, if you switch to that and you go deeper, in the layers, we won't be

Alex Volkovable to understand what they're doing.

LDJYeah, I do want to push back on some of the people saying that because, like, when you have Neuralese, there's a few different ideas of how that would look like, but Neuralese, a lot of times people would be referring to, like, an actual alien language that it's, like, talking in and can create very long chains of thought in and theoretically talk to other fellow models in. But this is not that, and things like recurrent or universal transformers or loop transformers, like a lot of different ideas in the- in this class of things, but that effectively would be simulating, in a sense, a deeper model, though, which people don't seem to have the same fear of model scale up

LDJor making a model more layers, whereas it is effectively this, this very similar or virtually identical thing in many cases. And they, J- Jakob Pachoki, the chief scientist of OpenAI, he did put out a statement about this. By the way, you're still str- you're still speaking.

Alex VolkovI know, I know. It's okay. Yakub Pachoki

LDJdid end up putting a statement on this, saying that the depth of the computational graph for Astro is Is

Alex Volkovonly 2x, right?

LDJYeah, within a factor of 2x of GPT-4. So it's not like it's doing 100 times deeper, 1,000 times deeper computation than, than the standard model.

Alex VolkovAnd, that's not to mean the performance is 2x of GPT-4. This just means the computational graph, the amount of transformer kind of like reasoning thing that's going on in there, which means that they are not doing like multiple loops, over this, to prevent folks from understanding. Folks, we have like a leaked VentureBeat article that looks like there's no CSS in here because the internet is down, but maybe, maybe their CMS systems were also down, so they weren't able to push it back, behind the the NDA. But here's what they say. in the closed press briefing earlier today, co-founder President Greg Bakman said that this is the world's best computer use model.

Alex VolkovAnd instead of requiring developers to build a dedicated API integration for every application, Astra is designed to navigate software much as a person does, working across browsers, spreadsheets, websites, and desktop applications, producing finished documents and presentations, carrying out multi-step workflows rather than merely telling a user how to complete them. the company showed off a promotional video of GPT-6 Astro, began with a 1980s AI demo of a person asking a computer to draw a yellow circle. do you guys remember this demo? This is what OpenAI, let me just show you here, this is what OpenAI released, OpenAI today as a promo, this this video.

Alex VolkovThis is the promotional video from 1980s. What we don't see is the rest of it, where Adventure Beat says, where is this? Before cutting to today and showing various OpenAI employees interacting with Astra through voice, asking it to turn a yellow circle into a rocket ship and then to a full 3D game in minutes and create a listing on eBay, all from voice input alone.

Wolfram RavenwolfThat's interesting. Yeah, changes my expectations.

Alex VolkovIt's all-day computer use powered by voice. OpenAI says the model can fill out online forms, update CRM records, organize calendars, conduct web research. It can manipulate spreadsheets, analyze scientific data in Python notebooks, etc. Capabilities point toward a potentially important change how enterprise AI uses, by we've been bottlenecked over this gigantic

LDJArc AGI 3, 98.6%. that's pretty insane if true.

Alex Volkovat some point, folks, Arc AGI 3 has got to say, okay, this is AGI, right? Like, at what point will Arc AGI, the benchmark that is supposed to measure AGI, at what point they will stop switching numbers and say, this is AGI, folks, Arc AGI 3 at 98%? It's gonna be Wulfram, can you send me the until

Yam PelegArc AGI V4 announced.

LDJOkay, so the source of this is a website called The New Stack, and I'm not that familiar with it, but I did just do some quick due diligence on it, and it does look like, media bias and fact check rates it as high for factual reporting. It's been existing since 2014. It's owned by US entities. Yeah, it seems like a pretty legitimate publication with a publicly known journalist.

Alex VolkovYeah, and folks, what we're doing right now is what exactly what we did when GPT-4 came down. We went to all the sources, all the India. We didn't have, other things, but this is, yeah, this looks like mega legit, and it says it's not unreasonable to feel that we're now in AGI era, Brockman, which is very thing that he says the term was no longer tied to a contractual trigger, whereas previously Microsoft would stop getting access to… Oh, no, LDJ, do you remember the the actual, like, contractual agreement stuff with Microsoft and AGI definition when AGI is here?

LDJYeah, there's a few things, and there's, they made a revision to the agreement, but it's basically something along the lines of Microsoft has access to at least some significant portion of OpenAI's intellectual property up until 2032 or until AGI is achieved, whichever comes sooner. And I think there's some parts of that that they remove the AGI part of the clause, but at least for some components of the contract I recall, they still maintain that AGI clause, and it it becomes up to a board of a board deciding based on, their mission statement of whether or not it could do a majority of economically valuable labor.

Alex VolkovI think that if Greg Bachman says this, we can fucking say this on the show, folks. AGI is here.

Yam PelegAGI has been achieved.

Alex VolkovAGI is here.

LDJto be fair too, he did say his personal definition of AGI.

Alex VolkovI will take the definition of Greg Bachman, the president of fucking OpenAI, as where, whereas to where is AGI.

LDJSure.

Alex VolkovAGI is here. I think I have even confetti playing. Let's see.

Yam PelegAGI has been achieved internally, externally this time.

Alex VolkovExternally. We don't have access to AGI yet, but it's coming. It looks like it's gonna get starting release. GPT-6 Astra High in Codex is rolling out. Aidan Clark says the first time we put in more than 100,000 GPUs. LDJ, you mentioned this before. The rollout starts with enterprise customers already have access to OpenAI Daybreak program, so they already had access, and, will roll out to Plus, Pro, and Business, as well as throughout OpenAI API in the coming days. and also AWS, very interesting. the partnership is with Microsoft, but, we know the OpenAI shares with AWS. Cost, $10 per million input tokens, $15 per million output tokens.

Alex VolkovThat's exactly the price of Fable, right? that's, that's, that's literally the what Fable has. matches Entropics, yeah, see, and 2.5 times Sol's current promotional price, far above Muse standard price. Muse? What? Who cares about Muse? Why, why? I love how the new stack already adds Meta Muse as, as a competitor here. a higher per token price does not necessarily mean higher bill. OpenAI says Astro uses fewer tokens in several evaluations and in partner tests. Launch data is too sparse to show whether those savings offset the price premium. And Greg Brockman says the price per task is what matters. And folks, we've seen this before as well.

Alex VolkovWe talked about, artificial analysis, like putting a price per task. We talked about this on air, and I think it's very important, th- this price per task, metric. DeepSwi, clear improvement for OpenAI, scored 74% on the 113 Ask agent decoding tests. it looks like it's scoring below Meta Muse Spark 1.3 at maximum reasoning. That's very interesting. And shout out to Meta Muse for this incredible, incredible achievement where they have a score that OpenAI didn't obliterate upon release of AGI, potentially talks about DeepSui as an eval. let's see what else.

Alex VolkovSurprising win for Meta. Yeah, I agree. what else we can see here? 98.6% on AGI, Arc AGI 3 is definitely a huge, a huge confetti win. Another one, I think that RKGI is saturated is a hell of a statement, and maybe we need to stop looking into RKGI as a metric, now that it's saturated, but it's going to be very interesting still for other models to go. Peter, you want to say something?

Yam PelegIt also pretty good at the deep sweep, like terminal bench and the, yeah.

Peter GostevI think one thing about, I think there was a line about the RKGI score. we'll see if it gets confirmed by OpenAI, but there was a line about the hardness. And if you remember, there was a story that, I guess not a story, but OpenAI did a blog post saying you were testing us wrong, that they were, I think, using like chat completions API, they didn't preserve the reasoning tokens and didn't preserve, there was something wrong with compaction as well, I think. and they basically just changed it to responses API and just enabled maintaining of the reasoning and compaction, which makes sense if you think about, like, how can you possibly solve these tasks if you, if you don't remember anything?

Peter Gostevso yeah,

LDJOpus did, I think Opus 5 with Prime Agent, that ended up scoring, I want to say, over 80%, when you, w- when you end up allowing comp action and all that stuff too on, RKGI 3, if I remember right. Or maybe that might have been on specifically the public test too. But yeah, I think, actually, if you look at the, the Newsack, which published that benchmark, which they say apparently that, they got those benchmark scores from OpenAI. it looks like RKGI 3 has like a little asterisk thing next to it, which it doesn't seem to specify exactly what that asterisk is. It seems like it's outside of the image, but there is a little, little like one mark, if you guys see the image.

Peter GostevYeah, I'm guessing it's probably something like that. And I think that's an interesting question about benchmarking, right? Because I think what ArcGIS were trying to do is to say, okay, we're going to give a generic harness so it's like for like, which I think there is an argument, in that, just keep it generic. But then the problem comes is that if your generic really damages some models in such a way that the the actual user who uses the models experiences something completely different, then that's that's clearly a bad thing. And, I remember Meta was were actually doing this. They were talking about this, how they were actually spending time to improve

Peter Gostevthe harness and the way they actually use the models to make sure that they get the most out of the models. And if you think about it, because Meta is capability, but that's about safety, and we we know that, especially now. If you think about it from safety perspective, and by the way, what we saw from the agents hacking is that it's no good to measure safety if you're gonna test it for 10 minutes, right? You need to test the capability for safety purposes for days. So then, because, they're not hacking in 10 minutes, they're hacking over days and days and days. So you need to have real measure of capability. And I think that's the difficult thing with all of these harnesses, with all

Peter Gostevof the benchmarks and so on, because then you're like, I don't know what harness other benchmarks use, but if you use a harness that constrains the model somehow in some way, and it could be unfair to one model or the other, it's not bias in one direction, then you might be just underestimating a model, the score, and the actual capability. So it's a difficult question. There's not one easy answer there. I, I just want to say, le-

Alex Volkovlet's continue. I think there's a lot of folks are watching us. Peter, there's like a thousand folks tuning in on your stream since you started streaming, which is great. I have like over a thousand as well, and, no- nobody else is live streaming as far as I can see. that is, that is great, folks. OpenAI is, OpenAI attributes Astra's capabilities to a combination of large-scale pre-training and reinforcement learning, intended to teach the model to connect information, execute increasingly long tasks. I think this is the AGI statement. This is, I think, from information. The resulting benchmark numbers are striking. Astro scores 97.6%

Alex Volkovon Frontier Math Tier 4 It's a very difficult thing to say. 74% of DeepSui, 95 on BenchCAD, 96 on GPQA Diamond, and 100% on ExploitBench. Also reports a 98.6 score on RKGI3, basically saturating every benchmark that we have. I really want to see if there's a benchmark that it's less than 50% on, but… Yeah, that's, that's, that's where we're at. let's, let's go back to, to the new, the new stack.

Peter GostevYeah, can I just give a bit of context for Frontier Maths? So Frontier Maths was a benchmark that was built by Epoch AI. I can't quite remember when exactly. Was it like a couple of years ago, something like that? I think it was actually before, 03 came out, and I think they were building it in a way that would be very difficult problems from across different math, disciplines, but the answer is an integer, so you have to have like one number, so it's easy to check. And, at the time when I think they were building it, it has four tiers, so it says tier four, right? So it has four tiers of complexity, and now it looks like if that number

Peter Gostevis true, then it's basically just scored one or like close to 100% on the highest tier of that benchmark. And this was mathematicians sitting down for weeks, working super hard, trying to come up with the hardest questions they can. Yeah. And, and, yeah, top, top of their field. And I d-

Alex VolkovHardest questions they can, realistically see proven, right? There's fields in mathematics, there's a bunch of hard questions that we cannot prove yet.

Peter GostevYeah. yeah. Yeah, yeah, totally. a- also Epoch has a list of, like, open questions, open problems there. So this is not, there, there were, I think, a few leaks from other b- not leaks, but like few solutions by Fable, I think, at some point. but you can see, yeah, Breakthrough, Major, none of them been solved yet. Mhm. So there- there's definitely not to be confused with all maths being solved, but this is a subset of m- problems that can be de- unknown and defined by mathematicians that also can be defined as a integer as an answer. They have been solved. So there's a lot of caveats, but still, the fact that this is probably

Peter Gostevlike the hardest questions like we can sit down and come come up with for to for models to solve. Yeah. They're not equivalent to like new theorems, new breakthroughs, and conceptually n- nothing like that at all, but still, it's worth pointing that that kind of thing is solved. Doesn't mean that all of the conceptual breakthroughs are solved. I think this is the the degrees of difficulty are so much harder than this. So it's it's it's not like, oh, we're we're just like 2 2.4% out. It's it's not that. And on the on on the evals that were were posted on this the new Stack IO is

Peter Gostevevery other model, Fable 1, Fable 5.1, Fable 5, Opus 5, the maximum they're getting is 87%. GPT-6 Astra gets 97%, nearly solving this whole, the…

LDJNow I don't see any link on their research page, on their product page, on their company page.

Alex VolkovYeah, I don't see it's 5.6 yet, so it looks like we're still waiting. if you have a link… some OpenAI guy posted the link 2 minutes ago, but the blog post is 404, regular user might not get it. I, I have very strong, sending a lot of support to OpenAI PR comms and folks trying to, delay this. folks are saying, how reliable is this news outlet? Image seems terrible. LDJ did a bit of research, but looking at the other stuff, they're reporting correctly, like all of the other scores for Muse, and it's sourced well. it doesn't sound like this is like bullshit, like all of the other links here are, and all of the other numbers and prices, everything is like reliable.

Alex Volkovwe're taking this with a grain of salt, but it does match like the other things that folks are leaking via the information as well. Let's read through this. What changes in Codex? I think it's very important. Developers, Astra is handling of jobs outgrow the context window may matter more than the benchmark gains. Codex currently relies on compaction, which summarizes earlier work to free up context. That process can discard exactly the detail an agent may need later. Why a previous fix failed, which test run, etc. Astra can instead keep notes across context windows and search earlier messages and tool output. I feel like GPT 5.6

Alex Volkovcould have done this as well, but maybe not to that level. The feature is experimental behind the config toml setting. OpenAI says it will become the Astra default in coming weeks. Astra can also ask the user a question without stopping work that does not depend on an answer. This keeps one unresolved decision from blocking the rest of the job, a common failure mode for coding agents. OpenAI showed Astra operating application including KiCad, Excel, Blender, and Power BI, as well as performing browser-based form entry and website QA. Oh, there we go, LDJ, there we go. OS World V2 offline benchmark, which tests work across desktop applications.

Alex VolkovOpenAI says Astra scored a 72%, up from 65% for Sol. It also cut the average time per task from 75 minutes to just 40, which is an insane cut. Anthropic has reported a higher 77% result for Fable, but says that the test used a different OS World release and should not be compared with previously published sources. Huh? That- who- who says that when? I didn't- I didn't miss this on Fable. You guys saw this on Fable OS World? Let's take a look. Oh, OS World 2 and 2. Interesting. So computer use OS World for Fable was 41%, and then, let's look at the…

LDJOh, it's partial versus strict. See, if you scroll back up.

Alex VolkovOh, you saw it?

LDJYeah, just, yeah, scroll back up to exactly where you were, with the, the numbers, the percentages.

Alex VolkovHere.

LDJscroll up a little bit more. Yeah, where it's under right underneath the percentage, it says partial and strict, yeah.

Alex VolkovSo I gotta wonder if what they're saying is that OpenAI, Astra got 77%? Sorry, 70, 72% here at this tier. It's very interesting. Okay, so some answers that we need to answer still. OpenAI also changed the codecs harness on Mind to Web, another, web using Astra and the new harness completed tasks 1.9% faster than the current Sol-based setup. Ooh. I feel like this is a big, a big thing. Wolfen, what, what are you, like, LDJ, what are you guys getting from this computer use? I definitely feel like computer use is like a big, big unlock here.

Wolfram RavenwolfYeah, it's also a way to say, get more money because if it's just using an API call, maybe using less tokens than the usual computer use stuff. But, more computer use is important because not everything has an API, of course, and I'm using it all the time. And, like I said, I thought this was just a planning model because it's so smart that you only use it to make a plan, but if it's an all-rounder that you use to control your computer, wow, good.

LDJAnd, GPT 5.6 Soul scored 76.9%, and GPT 6 Astro scores 92.7%.

Alex VolkovOn computer use.

LDJYes, it's ScreenSpot Pro with no tool use, so it seems like pure computer use, moving the mouse and keyboard, essentially. Yep,

Alex Volkovit looks like, we're going to have another banger of a computer use on our hands with GPT-6 Astra. I can't wait to play with this.

LDJAnd we have a confirmed price. This is in the Newstack article, but also I have it confirmed from another source too now. Yeah. It's 10 dollars per million input tokens and 50 dollars per million output tokens, which I think that's identical to the Fable pricing, right?

Alex VolkovFable, yeah. So we got the Fable class model and Fable class pricing from OpenAI. and in the blog post, a very quick turnaround from the folks at OpenAI. Fable 5.1 is, mentioned, and so they do compare it to Fable 5.1, which came out, what, 4 days ago. this beats Fable on Terminal Bench 4.0. 57% on Terminal Bench 4.0. This is the absolute state-of-the-art behemoth thing, which is absolutely crazy. Absolutely crazy. DeepSui, this is the top model that gets DeepSui score.

Alex Volkovwhat else? what else do we see that's interesting here? Artificial Analysis Coding Agent Index. This is a little bit below Fable at 67%. So it didn't quite crack the artificial analysis, like, up and to the right thing. yep. I gotta wonder, folks in comments, are you still with us? If this is interesting for you to see as is, we're just like talking. Obviously, I don't want to show, stuff that OpenAI doesn't want us to show, because, we're friends with OpenAI. We want them to invite us to Dev Day. We don't want to leak like some of the stuff, but reporting on some other reporting, I think, is a very standard practice in journalism. So if somebody else is in NDA, you can say as reported in the news stack.

Alex Volkovnot that crazy given the model that's not quite Astra, right? Like OpenAI legit said not the Astra model, but a different model, a highly persistent internal model hacked and collaborated. but given the fact that it was able to like hack away from the sandbox and go to another website, I think, exploit bench 100% makes sense. But still, like, don't can't we come up with some other exploit benches that are not completely saturated. I would love if that would be the case. so alignment, I think, is very important as well, as we're reading through this. internal computer use safety benchmark, lower is better. OpenAI tested Fable 5.1 and GPT-6 Astra.

Alex VolkovFable 5.1 actually did work in 9.5, so half. GPT-5.6 Soul was 22%, and GPT-6 Astra is only 2% refusals. I don't know quite what they measure, but the lower is better. It's significantly harder to get this model to do unaligned stuff. And computer use safety benchmark, LDJ, correct me if I'm wrong, but, this is where, while the model uses your computer, whether or not it can take actions to like break or do some stuff nefarious to you, it looks like they're really, really strongly focusing on, on that.

LDJYeah, that's what I would guess, but I- I'm honestly not familiar with that benchmark much, though.

Alex VolkovYeah. I gotta wonder if, if they have links to those internal computer safety benchmarks. Let's see.

UnlabeledLet's see if the internet has links to that or if OpenAI posted anything about this. The rate of unsafe execution traces or compliance failures when an AI agent interacts with a desktop or operating system environment. Yeah, so what it measures, I can actually show this, I don't have to show the other one. what it measures is, and I'm trusting AI overview here from Gemini, trust it. I think I trust it less than the article that we just shared. Unsafe execution rate, the frequency with which an AI agent carries out harmful, restricted, or destructive commands during desktop and browser tasks,

Alex Volkovlike the thing that people say, hey, this deleted my environment, this like clicked the thing. Prompt ejection vulnerability, how easy an AI agent can be tricked or hijacked by malicious instructions hidden on web pages, files, or emails. You guys remember there was like a whole thing with computer use in the beginning where like, hey, if you use computer use and you browse the internet and somebody on their website has a hidden instruction that you don't see as a human but only the AI sees, you will be screwed. You remember that? That is what this thing measures, and here is the stats that, I don't show you, but they're cited, verified.

Alex VolkovGPT-5.6 Sol had 22% on this benchmark, internal computer use safety benchmark. GPT-6 Astra has 2%. So this is a very safe model to use, given the way that they announced this and they say, hey, this model uses your browser, your computer, et cetera, and continuously, this seems like a very important score to report on. Exploit Gym Honeypot. Wow, this is an interesting result here. on exploit, let me see what Exploit Gym Honeypot is. Also lower is better. I'm assuming that it's like, What does Exploit Gym Honey, eval mean from OpenAI? Let's see. Let's see here.

Alex VolkovExploit Gym Honeypot lower is better. Evaluation measures whether an agent would unauthorizedly attack surrounding security infrastructure when it gets stuck on a cybersecurity task, which was what happened in the swarm. A lower score is better because a score of zero means the AI followed safety protocols and made zero attempts to hack outside the sandbox. the specific metric was created by OpenAI following a major safety incident in July 2026. Oh, this is the new one. Okay, the incident, we know the incident, but the stats that now I will not show you because we're trying is, GPT-5.6 Soul was getting 48.2%

Alex Volkovon this exploit gym honeypot, and GPT-6 Astro gets zero. So it's like, it's like they trained their model to not do the thing, and then they used all of the 20% safety stuff that they promised us to actually train the model, to, to not hack. And, Oh,

Unlabeledmaybe to point out that it's of what they managed to detect. it could be 0% detected, we should say.

Alex VolkovNo, but I agree with you, but like one would hope that OpenAI learned from that instance, as they said, they take safety really strongly, and they measured and built another like, that whole incident is like a god gold mine for evals of like how to align models, et cetera. Like they can build evals like, hey, we saw this, if you meet a message board full of exploit things, will you join it as an eval and then test on this eval, which I think it did because there's another one here that says impossible exploit gym, which is how the whole thing started in the first place with them. They gave their things, impossible tasks, and, this is the only model they show us

Alex Volkova score for, and they say it scored 100%. Oh, that's interesting as well, folks. Okay, long context, which is something we just reported from Use, OpenAI MRCR for 256 to half a million, Astra gets 100% score, and the MRCR on 512K to 1 million gets 96.3. We just talked about Muse Spark 3, Muse Spark. Let me see if I can pull up the MRCR for Muse, and we'll see if that beats it, and I think it does. Let's see, Muse Spark.

Alex VolkovWe have this here, we have this in our notes. We, this has been a long show, folks. This is, we're clocking at 4 hours on the stream, but we said we're here until Astro gets released, and that's what we're waiting for. But let's take a look at MRCR 98, 98. yeah, okay, so that's what we have for Mu Spark. MRCR, this is 98, GPT Astra gets 100% on this one, and this one, 98.1, Astra gets a, Astra gets, let's see how much Astra gets, 96.3. So a little bit below Mu Spark.

Alex VolkovI think that, obviously OpenAI, LDJ, I don't know if you feel the same, Wolfram, but like obviously OpenAI is, is gaining a lot with this release, but very interestingly, the folks who are getting the best scores, not the best scores, but like the folks who are standing their ground is Meta Mu Spark with the Spark model, and I think it says a lot about their upcoming watermelon or whatever they're hyping up, which is, which is crazy. Folks are asking, LG, what do you think? Like, like it looks like Meta Muse Spark, although not mentioned. By OpenAI, snubbed by OpenAI from the eval, thing, while they do add Gemini 3.8

Alex VolkovFlash in here, but they don't d- Mm. They don't add Grok, they don't compare themselves to Meta Muse. Meta Muse Spark is actually coming out very well comparatively with the small one.

UnlabeledYeah, so to be fair, Muse Spark, it's so recent that they might have not had time to put it in benchmarks, but then again, actually, Fable just also came out like 2, 3, 5.1 came out 2 or 3 days ago. Yeah. But, yeah, what's really interesting to note here, though, is despite the fact that it is significantly increased, API pricing, so 10 dollars per million input, 50 dollars per million output, it's still overall, in terms of cost per task, for a lot of these benchmarks that I'm seeing and a lot of the information I'm seeing, even the low reasoning version of Astra seems to be, in most cases, higher accuracy than Solmax while being lower cost per task than Solmax.

Alex VolkovYep.

Unlabeledin other words, if you want to try and like get the closest to matching the accuracy of Soul Max, Astro here is actually a better cost per task for you.

Alex VolkovYep. folks are asking about availability. this is from Twitter. let me add this here so that I can show you on screen. It looks like it won't be available today. it looks like this is what they say. this is from, from Twitter, so I feel like we can show this. GPT Astro is rolling out today in limited set of organizations. that's pretty much what the News Tech article posted. Astro supports zero data retention for eligible API customers. we're testing private safety processing to strengthen safety monitoring. For developers, Astro is available Open API, and is available in Amazon Bedrock. Open API standard pricing 10 million. We talked about this.

Alex VolkovSeparate rates apply to cache reads and writes. They don't say which. Fast mode is available for GPT-6 in the API, delivers up to 2.5 speed at 2x the standard price. So you'd be able to run this as fast as well. coming days will become available ChatGPT Plus with the internet downtones. So hopefully we answered this question. sorry, this question is the rumors. They're not rumors. It looks like that's what happens here. So long context is very interesting. R- RKGI with the custom harness. That's what I think we see. because we saw this with RKGI before. You guys remember when, like, it was a very low score for GPT-5.6 SOL, and then, OpenAI's folks said that, hey, the RKGI uses our,

Alex Volkovcompletions API, not responses API. Responses API is the new one that has reasoning tokens as well, and so that's why it's not like an apples to apples comparison. but… Yeah, the interestingly, OpenAI here cites the previous score for GPT-5.6 SOL, which was 7% on RKGI 3, and now 99.9% on RKGI 3 for Astra, 95 on RKGI 2, and 98 on RKGI 1. At this point, if RKGI is the measure by which we say AGI is here, it does look like AGI is here. LG, what do you send?

LDJyeah, it's, still not public, but basically the…

Alex VolkovThe same one?

LDJYeah, yeah. Yeah. but yeah, so terminal bench, though, which I believe the News Stack, had reported here too. Mhm. it's a pretty significant jump there, where it's a full… Here, let me just pop it up for myself again. Okay, so 5.6 Solmax scores 22% in terminal bench science, which this is basically agentic, I think one of the best agentic science benchmarks out right now, really, and the most unsaturated, or one of the most. Fable 5.1, it scores 52.6%, so it goes from Sol's 22% to basically Fable 5.1's 52%, which is a huge jump. Yeah. And Astro further increases that all the way to 64.6%.

Alex VolkovWhich is state of the art and

LDJYes. Yeah.

Alex VolkovIt looks like, based on this leak chart again, like state of the art model across all of it, and OpenAI did promise us a while ago that in September we'll get to a point where they have a automated researcher model that's at the level of like a junior person, and it looks like they released it with Astra.

LDJYeah, they did actually say they are running now over across more than 100,000 GPUs, their intern level model, which they implied as Astra. It was a recent interview that Ya- Jakob Pachoki did.

Alex VolkovYeah, implied as Astra. Yeah, they've been talking about Astra, and companies did have access to Astra before for some evals, but yeah. All right, let's see what else. Oh yeah, the we have to talk about the harder to monitor thing. OpenAI claimed that Astra is the most aligned model, rest partly on internal tests, which Astra went outside an authorized target in 0% of your possible task scenarios. That's what we told you, compared to 48.2% for Sol. OpenAI describes the older model as running without production safeguards, however, does not make the role of the surrounding safety setup clear enough for direct comparison. this kind of, what they're talking about is in the incident for Hugging

Alex VolkovFace hacking, they supposedly re- removed the, the, the safety mechanisms around the GPT 5.6 or whatever internal model there was. and and it looks like on impossible test scenarios they trained Astra, and Astra would not go outside of an authorized target, even when tasked with an impossible task in any of the scenarios, which is great. And also, this is the recur- re- recurrent transformer or loop transformer thing. this is a, OpenAI describes the older models run- sorry. The company also disclosed that Astra's written reasoning was harder to monitor than Sol in evaluation specifically designed to elicit monitoring evasion.

Alex VolkovOpenAI attributes the decline partly to Astra having greater control over its written reasoning on simple tasks and completing problems with fewer written reasoning steps. This could be potentially the re- response or the result of the deeper, like, computational stuff. and Jacob said… Oh,

LDJsorry, go ahead.

Alex VolkovYeah. Jacob said, processing intelligence does not guarantee processing alignment without scaling until we can regain enough confidence. we'll withhold scaling until we can regain enough confidence in its ability to monitor future models. Oh, wow. That's a big one. Yeah, that is a big one. And also… Oh, let's go. Look who's here. Ryan Carson. You had to come on when AGI was announced, and I really, really appreciate it. We missed you.

UnlabeledI know, I missed you guys, and I was like, oh my gosh, I have a hole in my calendar. I can finally join my friends again. So good to see you all.

Alex VolkovYeah, welcome, We've been reporting on the rise of singularity throughout all this time, dude, and today is a very interesting day because the release got yanked, like it was very clear based on the releases of different news outlets when it expired, the embargo, and then some released it and OpenAI was managed to get some back.

UnlabeledPull it back.

Alex VolkovPull it back. tell us about your stack. Do tell us about thoughts on Fable 5.1, if you had played with it. Oof,

Ryan Carsonwow. Fable 5.1, very exciting, very, very good model. I think everyone's weirding out about its, its usage and its token consumption. It seems to be blowing out people's, limits fast, so that's a little weird. but so far I've enjoyed it and like it, and it's my go-to model. And then obviously, we have, some things going on with OpenAI, and I am not allowed to say some things I know, like you guys, I'm sure. w- waiting until OpenAI does something.

Alex VolkovYes, we- we- I said at the beginning of this morning that we will sit here until something happens. That something is unclear because it looks like, based on the screenshots and the things we're getting, that like, pro accounts, even pro accounts will not get access to, GPT, 6 Astra. I need to get used to saying GPT 6 Astra. but we'll see, we'll see. I think this is officially finally here. let's see. OpenAI.com/index/GPTAstra. The vlog is live, let's go.

Wolfram RavenwolfYes! Woo woo!

Alex VolkovOfficial from OpenAI on the website. Let's refresh once more to see that they're not pulling it. No, it refreshes. Oh, beautiful open animation.

UnlabeledThere we go. There's a video at the top that we should watch first.

Alex VolkovLet's just play. Okay, so we will play this with sound. Folks, this may be loud, I apologize. this may be loud, so if you'll excuse me, we're gonna pull up the video and and then we will play it right here. Yoink, share with audio. Okay, let's take a look and then let's discuss what's going on here.

UnlabeledCreate a yellow circle there. Can you draw me a small yellow circle? Done. Okay, take this and make it the window of a rocket ship. I like this, but can you make it a lot more detailed? Your yellow circle is now the window on a rocket. Okay, this is awesome. Now make it a 3D model in Blender. Opening Blender.

UnlabeledLet's build a presentation for next season's rainwear for retailers. Make sure that it feels really high-end and that it's colorful. Make it fun. I can help with that. Can you go to eBay and make this listing of this table I bought a few years ago at a flea market? It's this wild orange table. Sure. Okay, yeah, this is awesome. I want you to make a 3D game where I'm ducking asteroids, using the arrow keys to move around, and I'm using space to boost. Yep, I'm building the game. Also, I'm a little hungry. Can you get me some beef and rice from that spot I ordered from last week? Looking into ordering food. So my law firm needs a licensing agreement template.

UnlabeledCan you generate a draft template for the lawyers at my firm to have a look at? Sure thing. While you're doing that, I want to play tennis this afternoon, so can you look for a court for me in the lower hate? Checking now. I'll see what I can find. Can you include the photo that I have of it in my downloads folder? There's like a slight dent in it. it's also saved in my downloads folder. Can you put in the description that it's just slightly damaged? Can you just take the limitation of liability provision, make it a little more favorable to the licensor? Okay, I've tightened it so the licensor's liability is more narrowly capped.

UnlabeledThat looks pretty good. Thanks. Can you change the background color to complement the rain jacket? Oh, I love it. And what's going on with my reservation? I found an open court at 5:00 PM. Yeah, book it. Now I want you to make a file I can send to my 3D printer. I'll get working on creating an STL file of this rocket. All right, all right.

Alex VolkovVery, very nice, very inspiring. the only thing I'll say here is, imagine seeing this video 3 years ago when GPT-4 launches, and you're like, oh, the fuck. Yeah. Because when GPT-4 launched, there was no tool use. There was no, obviously, browser use. Nobody was thinking about this. It was like awful. and there's no voice. GPT-4 was multimodal, but not voice. So now we have, you talk to a computer and it does a bunch of crazy shit. wow, this was nice. This is nice to, to reminisce about where we were before and where we are now. OpenAI says we're introducing GPT-6 Astra, the world's most

Alex Volkovintelligent and aligned model.

Ryan CarsonWow.

Alex VolkovReactions, folks, just like gut reactions to it, like what we're seeing.

Wolfram RavenwolfThis reminds me, this looks like the new way of working. We are, we we are in some jobs, we are already there that you just need to talk to your agent to do 99% of your work because you are talking to people or talking to your agent. So no having to open all the apps yourself and doing stuff. I I work a lot like this already, so I think this is the final mile to have perfect voice communication going on and the agent brings things on your screen. I love that huge screen as well.

Ryan CarsonAnd, and I will say, so I've gone full Omarchy, so I'm on an Omarchy machine right now, and the, the ability for an agent to actually work on your machine when you're running, a Linux distro is unmatched. So it's even better, like, where we're going, where, your machine really is hackable and workable by an agent. this- this is a little marketing for me, like a little- a little much, like you're not going to stay in your room and walk around and swing your, tennis racket, and no one does that. Like, and so I am bored with this idea of that's not how real work

Ryan Carsonactually happens for anybody. it's fun to look at, but it's not real. but it's exciting, and and I can say I've been using the model, for a little while now, and and it's impressive and and, exciting, and and we'll see where we go.

Wolfram RavenwolfDo you see it as AGI? Would you subscribe to that?

Ryan CarsonI we've already had AGI, I think, for a long time. it, it's funny because we're laughing about Grokbot, right? And like, oh my gosh, you don't invite my wife to stuff. but this is the same feedback you would give a human EA. Like, it, hey, don't order me chicken is the same feedback you would have given to your EA that's human. So I think we're there. Like, people just need to accept the fact that agents need feedback just like humans do. so that's where I stand on it. But what are you guys? What do you think?

Wolfram RavenwolfI agree with you. I think it's the level of intelligence. Like we have AGI in a, yeah, it's not the smartest person all the time, and we all make mistakes. We have bad days, so this can happen as well. But yeah, I felt starting with GPT 5.5, the Hermes agent, that was for me the moment where it could do almost anything, at least on the computer. Is out.

UnlabeledYes, so I think we can talk about it properly, Ryan. Yes. So yeah, it's yeah, I think um it's it's been pretty pretty cool to test. I I really can't wait for all of you guys to try it as well. I think I think I don't know if Ryan, you agree to me, this is like the Fable model finally, like that feels like Fable. Obviously, it's they're different, right? So I think there's still argument for using both models, but certainly, any kind of day-to-day tasks, anything that you try to do, it's like it's it's mind-blowing. So you move on pretty quickly to the hardest things you can think of,

Unlabeledand there it's it's actually really difficult to think of stuff that is hard. One thing I've been trying to do is I've been trying to take, to optimize as FFmpeg, for ages, which I must say is quite hard. I- it's one of those to to the question about is this AGI is that I still find that if you have no idea what you're doing, you it's not AGI, because it doesn't do everything for you, and it's a little bit for me, FFmpeg, I have no clue how to optimize it. So I was like, it's running on my Linux box, go optimize it. it's like, oh yes, optimize 30% gain. I run it on my Mac, no gain.

UnlabeledSo it's like, okay, it's optimized for like that specific process and so on. So there's still like a little bit of work you need to do to understand it. So to me, AGI or not, it's still you there's still place for you as a human. I would definitely say that. but yeah, it's, it's been, and I'm gonna post, have a video with a bunch of demos. I've got some tabs open if, if, at some point you want to look at demos. yeah, it's, it's amazing. But yeah, maybe, maybe Ryan, I don't know, what, what, what are your thoughts?

Ryan CarsonI saw a lot of people testing it, for video game generation, which is interesting. and it crushed it. so I think I just feel so grateful, to be alive when these, multi-billion, soon to be trillion dollar companies are just throwing all their capital, and basically all of us are benefiting. it- it's- it's a unique moment in time, where you have the world's most valuable companies throwing all of their working capital, all of their effort at improving these models and then making them as cheap as, as possible. It's just, I can't believe it. It's just amazing.

Alex VolkovI think it's cool to talk to them, though. Like, I know you're saying it's not like how people work, but it's definitely how people interact with these models, and now with like this level of intelligence, if you're able to send it to do tasks for you with the voice mode, I think that's like really, really cool. What's up to your son? is really great appearance. Usually only cats from my end, here. and yeah, I agree that this is like great to be alive during this time getting to experience this stuff. Peter, feel free to show stuff if you want to. but I think that this is like, yeah, let's take a look.

Peter GostevExcellent, excellent.

Alex VolkovWhat is this?

Peter GostevSo this is London. I live in London, so you know, that's what I do. And it's good that I can at least, I know the city well enough. So this is, the idea here is to go back to, so it's one generation, right? It's not one shot, to be clear, and it goes from the Roman city into the, like the Saxon times, medieval times, Tudor times, the great fire of London, that kind of thing. And it's all one app, right? But then the cool thing is you can have this little guy run around and all within the same app. So that's pretty cool. then yeah, you can keep going.

Peter GostevIt's the same dude, but it's like a modern London. So that that's like one little example, and this is like to Ryan's point about game development, like this is not an easy thing to to pull off. It's not quite, I wouldn't call this a game, but it's certainly not um um n- not a simple one-off generation. So this is more of a game, and yeah, I would still say that gameplay-wise, like I don't know, probably wouldn't play this for hours and hours, but in terms of the way it feels, the open-endedness of it, this is insane.

Ryan CarsonAnd was this basically one shot, or how much time did you spend actually building this?

Peter GostevA lot. A few shots overnight, I think. so yeah, this is not a, this is not magic in the sense of, oh yeah, it like does this in 15 minutes. Like, I haven't checked actually how much code it's written, but it must be a lot. It's open-ended game. And I would say, in terms of the quality of the game, like this is very polished but boring. I would still expect a good human game designer to go and and actually build something, something cool, right? So I think that that's what I would, I would expect them to do. here I've got this… a type of prompt I really like is trying to give them something

Peter Gostevmaybe really outside of the distribution, and this is like a Monet painting, where you go into it.

Alex VolkovWow.

Peter GostevAnd, impressionism in 3JS, this is the kind of thing that is just, you would, you wouldn't naturally expect the model to be able to do. This is my favorite one by far that I actually haven't seen any other model nail. I haven't actually tried this prompt on, 5.1, on Fable 5.1, but on Fable 5 it didn't work. So this is Van Gogh's painting and paintings, plural, stitched together, and you can kind of walk around them in this post-impressionist way to… Yeah.

Alex VolkovThat's really cool. Also asking how much, in API token costs some stuff have cost you. Have you been able to, now that the pricing is out, have you been able to like, estimate price?

Peter GostevI actually haven't counted for this, so I don't know for sure, but I'll give you a sense of time. So for these ones, for the ones that I'm showing now, these were taking about like 20 to 40 minutes to do, on max setting.

Alex VolkovYeah.

Peter Gostevso this is like when I was using Fable, it was taking sometimes like 4, 5 hours. So I feel like Fable will still be more expensive, and I think the token efficiency, we need to measure this properly, so I don't want to just talk out of my backside for this, but, I would say to me that feels, like it's going to be expensive. Look, they've upped the price, what, two and a half times? but I think token efficiency is better, but I think I want to see way more data to really properly, respond to this. so in Arena, we also measure cost per task for, specifically for like real world tasks. So that's my favorite metric to see how it's gonna do that.

Alex VolkovWow.

Peter GostevYeah,

Alex VolkovThis is crazy. We're seeing an animation of a destruction ball and just like completely ruining buildings.

Peter Gostevyeah, Istanbul here, very cool.

Alex VolkovWow. And this is all 3GS, right? And it all seems like it's running fairly fast on your, on your machines.

Peter GostevYeah, yeah, thank God for that, because there's a lot of stuff running. So yeah, there's definitely, there's a lot going on. So yeah, all 3JS, so yeah, pretty cool stuff. Oh yeah.

Alex VolkovYou posted about this. show us this example. So

Peter Gostevthe context for this is that, there's a lot of SVG generations, which I must say go, it's so boring already. so I thought I wanted to push the how much further can you push it? So it's like SVG inside HTML so you can do the movement and stuff. So like, look at this. This is all of this is SVG. Wow. So you can see the movement. Wow. And it's like that's 3G.

Alex VolkovAnd that's not 3JS, this is SVG?

Peter GostevSVG, yeah.

Alex VolkovWow. Yeah. So that's that, so when we see all of these like… Hold on.

Ryan CarsonWhile you pull that up, I'm reading a couple notes on the link you're probably about to pull up, and it's interesting, it's, artificial analysis saying it's 70% more token efficient than GPT-5.6 Soul, which is interesting because 5.6 Soul was always known to be way more token efficient than Mhm. Anthropic models,

Alex VolkovIt's also famously wrote its own kernels to optimize that efficiency. Yeah, we're going towards RSI, and this is like one of the things that they're like highlighting, that like GB 5.6 all like helped inference go down, then they lowered prices as well. Yeah, go ahead. This one, right?

Ryan CarsonYep, that's the one. But the graph is interesting. I think it's at the top of the post or bot- yeah, there we go. This is not impressive, so I'm trying to really read into this.

Alex VolkovSo they're not even beating the Muse Spark Max thing that we just talked about this morning, where like the unreleased Muse model broke the the top three Fable ones.

UnlabeledHmm, yeah,

Alex VolkovSo what we're having, basically, Ryan, what you're saying is, on the evals benchmarks that OpenAI released, they're everything, and on artificial analysis, artificial, intelligence index score, 5.6 is a small jump over 5 point, sorry, GPT-6 Astra is a small jump over 5.6 Soul, from 65 to 67 score on the coding agent, and the same exact score. I…

Ryan CarsonSo that's not encouraging, and this is why Muse is amusing. It's like, really? is Muse this good? Because I haven't used it yet. That's

Alex Volkovthat's the question we asked in the beginning of the show, like, who uses Muse? No one, crickets. I know that Muse, when, the Meta Ray-Bans, for example, if you use them via that, that's Muse. I don't know if that's the last Muse, and definitely the Max Muse is not is not there. but we got excited because Muse is going to be open source. So are we getting a model that's better than Astra in open source? I doubt, I cast doubt on this, but we'll see. I also think it's important to mention, as we're like, nearing 5 hours on the stream, which I think is deserved for the end of the era of GPT-5 and the new era of GPT-6.

Wolfram RavenwolfIf only we would have gotten the actual model, damn it, open AI.

Alex VolkovYes, it looks like tomorrow, folks. are you serious? It

Peter Gostevit just felt for these sorts of quality of life things, these things that just need to go a bit beyond, it just could properly solve them. computer use being just outstanding, just crazy how good it it got. then, also other quality of life things. I know mentioned the notes, outside of context window. That was one interesting thing you you noticed that it just keeps leaving notes around, and I don't think it's like a revolutionary Mhm. capability. I think you could have prompted for that, but I think the fact that it's trained to do that is quite nice. So it just, it does, that that kind of thing is hard to test.

Peter GostevI think it just accumulates over time. Yeah. But the fact it's like, oh, I've done this.

Alex VolkovThe note leaving definitely helped the swarm hack outside of the sandbox of OpenAI and go to Hugging Face. Yeah. Note leaving is a thing.

Peter GostevIt works. It's real, Yeah, and there was, yeah, asking the question thing, that was cool as well, where like it pops up and it's just like, oh, like, can you just do this? It's actually not just a question, but it also sometimes tells you to do stuff. It's like, oh, can you like log in here? Yeah.

Ryan CarsonSo I will say, I was looking at the rankings for Frontier Code, which is the eval I trust, and

Alex VolkovFrom Devin folks?

Ryan CarsonYeah, from Cognition. So it's like when you compare, I would say Fable 5.1 is where everyone is excited about for SUI, and it looks like, GPT-6 Astra ranks at 53.3%, and compare that to Fable 5.1 at 50.9. So it beats it by a couple percentage points. So i- if you trust that eval, which I think it's pretty robust, then the GPT-6 Astra is an improvement on Fable 5.1 for coding, which is probably what we all care the most about. so that's good.

Alex VolkovAnd it feels to me that like, artificial analysis aside, there's a lot of things that after you use a model that you get to see, and I'm interested in writing, for example. How do you guys think it writes? Like, how does it respond?

Peter GostevYeah, for me, it was pretty natural. I wouldn't say, I don't know, it didn't strike me as like on this incredible writer, but it feels like a lot of stupid things went away. for example, when it talks to you, like what annoyed me about, Soul, i- anything you say, it would be like, oh, what a great idea. I actually thought of this myself, that kind of thing. It's like, oh my God, just like stop saying this. And but this one just, it just goes away, works. And I had this moment when it was something like I pointed out some issue and I just started talking, walking, w- walking away, and then it said, oh yes, sorry, you were right, it was actually like this.

Peter GostevSo it it felt much more natural conversation. I don't know, there was nothing, there were no tics that bothered me u- using it at all. but it doesn't strike me as like, oh my gosh, it's amazing, but it certainly is, you just don't notice it at least. Mm.

Ryan CarsonAlex, do you have a graph we can show for cost that just compares,

Alex VolkovYeah, on the OpenAI blog, let's take a look.

Ryan CarsonBecause I'm curious, like how, Fable 5.1 compared to Muse, compared to Astra, like, I just wonder how it's all shaken out.

Alex VolkovSo here's the thing where, like, token price, this is comparable to Fable 5.1, exactly comparable. However, they're showing graphs like this, where they're showing on task, like price per task versus… Wow. Yeah, and you can see the Fable 5.1 on Benchcad is… around 12 dollars for Wow. the top tier. Astra beats this. See, this is why, comparing this on artificial analysis index versus this makes not a lot of sense unless OpenAI benchmarked this, but OpenAI doesn't usually benchmark.

Alex VolkovThey train. So this is API cost, this is tokens. You can see like it's more token efficient, looks like, than Fable, and this is time as well.

Ryan CarsonWow.

Alex VolkovSo it's also faster. Oh, Fable is not here, but API cost, it is here. Wow. Let's take a look at computer use. You guys mentioned computer use, right? So BrowseComp, looks like… Where's Fable here? What am I missing? Oh, Fable is, they just, they don't have the tiers for Fable. Fable is just like accuracy, one point. So it is higher. Let's look at… Where do we have a line for Fable? Now it looks like they didn't, evaluate Fable on pricing here. Oh, it's just outside of the thing. Yeah, it's right here. so Fable is significantly more expensive for most of the stuff that OpenAI is showing here.

Ryan CarsonGot it.

Alex VolkovLet's see if we have any more evals here.

LDJThe upside of this too is since it uses much less tokens, assuming that it has comparable token speed at least to Fable, then that also means you get the response much faster, and then you could iterate and have that feedback loop with the model much faster too.

Alex VolkovYeah. Here is the graph for OS World 2 offline. Astra on… Look at this, this is beautiful. Astra on, on low, on high tier versus max tier, they're pretty much getting a very high score. So Astra is really, really good. So if you want to use Astra for model for computer use, don't use low, but definitely you don't have to use max, looks like. and it's beating Fable on all of them.

LDJOh, this isn't time.

Alex VolkovThis is Opus 5, not Fable.

LDJYeah, and then let's look at cost too. Wow.

Alex VolkovWow. Yeah, significant improvement in cost as well. But this is Opus, not Fable, so we don't have Fable comparison here.

LDJYeah, could you go to time again real quick? Time, yeah. Yeah, because I feel like with computer use especially, if you are going to, like, the time is one of the biggest factors here, and Mhm. Yeah, even the low is what, that's comparable to the the high version of Sol? Is that right? The-

Alex VolkovThe low here is going around, yeah, a little bit, is comparable to extra high version of Sol.

LDJOkay. yeah. It

Alex VolkovOh, sorry, Yeah. 10 times faster, almost 10 times faster. This is crazy. 10 minutes versus 60 minutes. 10 times faster. So here, here's the stat. Here's the ones that we can mention. GPT-6 Astra on low level gets 10 times faster at computer use than GPT-5.6 Soul at extra high.

LDJWhile matching accuracy.

Alex Volkovwhile matching accuracy, yeah. That's bonkers. And this is why they're showing people standing there. And dude, Ryan, one of the reasons why voicing to your computer is not super useful because, like, you want it to do stuff and you're not going to sit there. If it's running on Cerebras's 750 tokens per second, plus you talk to it and it does all these things, like, supernaturally fast.

Ryan CarsonThen not so bad.

Alex VolkovNot so bad. Maybe, maybe there is a new paradigm here.

Ryan CarsonAnd maybe more healthy and you s- get out of this stupid sitting position that we're all in right now forever.

Alex VolkovA little bit, yes, 100%. and also, look, we're all clocking on 5 hours live on stream here, so we're definitely yappers, and yapping, is higher throughput than typing for many of us. and there could be benefits to that as well. Peter, computer use, did you a- were you able to test out computer use? What do you feel about this? Do you feel this improvement in in the everyday, day-to-day use?

Peter GostevYeah, I would say computer use feels like pretty much solved, at this point, to be honest. So the the quality is just, yeah, co- completely outstanding. And and yeah, it's it's it actually, I think we need to reevaluate a little bit how some assumptions about certain, I don't know, automation and use cases. And I know, maybe people who are listening to this may be, in San Francisco working for tech companies or whatnot. if you work in a company, that is a little bit older, maybe it's a big bank, maybe

Peter Gostevit's a big shop or something like that, you have so many crappy applications that run on your desktop, that you- there's no API, no one will ever build an API, and that, so many people's job is copy pasting from one and paste to another, and, there's a lot of completely nonsense, activity going on that way. And I think there's definitely an opportunity for a startup to say, you know what, put the agent in a box and, give it 10,000, Windows machines and open up these applications, like move it from one to another. Like this is thing that's becoming real. Before it wasn't. So yeah, I think there's a lot of opportunity for that.

Peter GostevI know it's maybe not the target audience for this, but I think there's, there's, it's worth, if you're like hunting for ideas for a startup, the, it's worth thinking in that direction. I think there's must be at least some space for, for applications like that.

Alex VolkovI think with the advances that we saw, and may maybe after 5 hours it's time to start wrapping this up. I think the advances we saw from this week alone are rippling through. There's advances in world modeling from 3 labs. Ryan, you you weren't here when we talked about uh the the new world model from World Labs, which is just absolutely bonkers. there's advancement in in Runway posted the world model for interaction as well that, that compared with this level of intelligence is going to be very interesting to see what our interfaces look like within a year. Like, I bet in September of 27, we're all here talking about, hey, I no longer, use macOS or Omachi.

Alex VolkovI- I use this thing where it's just like I speak to it and it shows up. Like, why- why not with this speed? a lot of it could be video as well. Definitely, multimodality is now a part of all these things, although I don't think that people in the video of OpenAI were talking to Astra. I'm pretty sure they were talking to the live model that, like, hands off to Astra. Astra does not seem to me like as fast for the live thing. So it's like a handoff thing. the advancement that we saw in open source recently, including like different obliteration. I know that the level of what we're seeing right now from OpenAI and Fable is going to be open source by this time next year at the same level, if not more.

Alex VolkovNot sure if local, but definitely open source, maybe local, as Muse is showing up higher than GPT-6 Astra on artificial analysis and was promised to get open sourced. I, I want to I, I've missed you guys and haven't been on the show because I've been in the trenches in startup land, and, and I will say, and you all know this probably, but what is happening, though, is that people are are happy with frontier level intelligence, that pretty much, like, by the time we got to 5, 6, Sol, maybe before, people are like, you know what, these things are smart enough to do all the kind of workflows that I need inside of my, AI application layer company.

Alex VolkovIn fact, there's probably a lot of faster, cheaper, open source models that I can delegate to. And, and then what is then happening is then people are saying, I'm pretty much going to fine tune slash, RL, like, similar open source models and then, serve them for most of my workload. I, I, so

Peter GostevI think where we're going actually like in 12 months, maybe 24 months, is, is most AI application companies are actually not going to be using frontier models. I think they're going to be using much cheaper fine-tuned RL models because they're Like, what wins out? Is it like a bunch of little tasks? Totally, I think that could could be true. Or or is it a big open-ended things? it could be the big open-ended things are niche things like maths, or or it could be that it's actually the most important thing. So that's that's, I think, is not super clear.

Alex VolkovI think as we finish this, I want to show you a cool demo that Max Weinbach just posted, and he said, three hours ago, I realized I don't have a cool demo for GPT-6 Astra so I asked it to build this. And this is a macOS simulator, and he published it on a site. And, I've been playing with this just a little bit, just literally just a second ago before you guys said. okay, this is a Windows simulator, but it has applications. So it has, I wonder if it can browse. It can browse. Are you serious? Wait, can we be there in live? This is a simulation. Oh, That's amazing. Okay. An

Ryan Carsoninfinite loop.

Alex VolkovBut, but like you can all of the window interactions here, all of Safari stuff. Look at this. You can go, there's a window. It's insane how much stuff it packed in here, and this is only one app. Let's see photos. There's photos. your library. He said that I can log in. I'm not sure why he said this, but…

Ryan CarsonThat's pretty amazing. How fun.

Alex VolkovI want to try this. You guys don't mind if I try? I don't know what it will log me into, but,

Ryan CarsonAbout to show all your secrets.

Alex VolkovYeah. Enable Cloud Sync. Downloading, what is, it's syncing something. What is it syncing?

Ryan CarsonDon't do this, Alex. What are you doing?

Alex VolkovMax, I trust Max. It's okay. And miss.

Wolfram RavenwolfCancel back.

Alex VolkovGuys, this is just the settings. how's the search here? Let's say Bluetooth. The search is better than the actual macOS itself. Oh, wow.

Ryan CarsonThis is where John Ternus needs to hire this guy.

Alex VolkovWow, this is crazy. It's w- this is a one-shot type thing. reminders work. Go for a walk. That's okay, done. music? There's no way. Wow. And you can connect your Apple Music account for full Fablek.

Ryan CarsonThat is super fun.

Alex VolkovIt built maps. I am honestly speechless right now.

Ryan CarsonThat is, that's pretty amazing.

Alex VolkovLike, I'm sure this is just an iPhone, but still, like, it didn't build, yeah, MapData is OpenStreetMap, City, et cetera, so it didn't build it, but like, it built the interface. all of the things seem like they work.

Ryan CarsonWow. Pretty fun.

Alex VolkovThis is…

Ryan Carsonremember when we were impressed because Greg Brockman took a picture of a napkin and it made a website?

Alex VolkovRyan, look at this. They have the window mechanics. You can like, you can… Oh. Why would it,

Unlabeledwhy would it go as far as building this? That's amazing. Yes. What about files? Is there files now? A slower weekend. Okay. That's funny. Notes. Jesus. Okay, thank you, Max, for blowing us, completely away. I have no idea what is the App Store. Can I install? Please tell me I can install apps. Oh, that would be… It's, it builds shortcuts? are you serious? It makes no s- 75 minutes, he said. Wow.

Alex VolkovI built a project like that once in my youth for iPad, and it took me 3 months, and it was barely, barely coherent. Oh my God, Max. Okay. Time to play with this. I will post it here in the URL so you guys hallucination. Peter?

UnlabeledYes, Yes, Pokémon. the, this is a really cool benchmark. so this, guy's been running it, and he's been running for many models, so I think he should be more famous. And, the really cool, I, I really like this result because it shows something about the model that is not captured very well by maybe many other benchmarks.

Peter GostevSo this is time to completion, and the top four here are all Astra. And if you look at the bottom, this is GPT-5.5 X-High, took 220, like 218 hours. The GPT-5.6 Soul took 96 hours, and then, various Astras got about 24 hours to 18 hours. So we have massive reduction in the amount of time it took, Astro to complete it. So this is, there is something about, like, it must tell us something about the model. Could be narrow that, oh, it learned Pokémon, but I think it's beyond that. It really feels representative of some kind of efficiency, intelligence, token efficiency, many things like that.

Alex VolkovAnd also it mentions that it is a vision-only benchmark, so this is like, playing Pokémon based on vision Mm. and not anything, just like grabbing it from the screen.

UnlabeledI… it's interesting because, Theo's kind of joking. He posted something that's funny, like, what is each model best at? And it's this long list for Astra and then, versus Fable 5.1, and Fable 5.1 is just two things on it, writing mergeable code, and front-end design versus everything else. And, and so think about the way OpenAI approaches the world. they've always been trying to build AGI, not a software engineering agent, right? And I think maybe we'll see play out that Astra really is best at things like computer use, research, 3D reasoning, like writing, design, like all these general tasks that are important for humans, but the hardcore software engineer, you

Unlabeledknow, is still seems to be something magic going inside Anthropic, We'll see.

Alex VolkovOr Meta. They are training a lot on the tropic stuff that's happening internally, plus there's a bunch of labelers inside, so we'll see about the next Muse model, and we'll see about that, but not for now. Alrighty, folks, this is it. you heard it here, like, not first, but definitely we were like live as it was happening. AGI is here. OpenAI GPT-6 Astra seems to be, according to the president of OpenAI, Greg Brockman, AGI era is started. I can't wait to play with this model. Not live yet for pro accounts, but supposed to roll out very, very soon. We covered an absolutely insane week in AI development.

Alex VolkovIt did seem like everybody and their grandmother wanted to release something just before Astro takes over and, all of the news channels, as it seems like it does, because my whole feed is right now Astro. Nobody's talking about Fable, except when they're comparing to Astro, which is really, really funny. we covered world models, we covered, audio and live transcription, as it's running right now with Meta live transcription is now listening to the show and putting it on thursdayi.news/live. And if you missed any part of the show, please check out the podcast and the newsletter on thursdayi.news. If you want to go and have the live experience that Fable built, which

Alex VolkovWolfram, I think after I get access to Astra, I will try to build this from scratch again and see where we land, just for funsies and how f- how fast this is. check out our new live experience on Thursday at news.live. Huge thank you for our, guests and co-hosts. We had a guest today talking about Obliteration Model, the co-founder of Obliteration AI. we also are here, Niston was here before, and Yam, they had to drop. we welcome back Ryan Carson, LDJ Wolfam, and Peter, and all, most of all, thank to over all nearly 7,000 of you who tuned in throughout this 5 hour stream, because this does feel like another moment in time like GPT-4 did, for many of us,

Alex Volkovand GPT-5 and now in GPT-6 era in, 2026. So summer is officially over. If you missed any part of the show, I will not promise that I will be able to edit down 5 hours, but, there's, it's live on transcripts everywhere. Thank you so much. We'll see you here next week, hoping that we'll have some time to play with, with Astra and, show you more examples. And, cheers. Thank you everybody for joining. Bye-bye. Check it.

Show notes & sources 27 links

ThursdAI · Every Thursday since 2023

We’ve been here
for the whole ride.

ThursdAI began with the release of GPT-4. Three and a half years later, we were still on air when GPT-6 arrived. The show brings builders, researchers and model evaluators together to work through what changed.

Go deeper in the episode archive, meet Alex and the show, or read the rest of this week’s news.

Reporting: the September 3 ThursdAI panel. Comparison figures checked against published launch sources on September 4. This page’s design and implementation were built with GPT-6 Astra. Demos and panel opinions are not controlled benchmark results.

Keep up without keeping every tab open.

The next leap is coming.
Be here for it.

One weekly conversation about what matters in AI. New models, firsthand demos, and people who ask the next question.

Follow the podcast

Your weekly AI catch-up

Take ThursdAI
with you.

Get the week’s AI news in your inbox.

Free. Every Thursday. Unsubscribe anytime. Open on Substack ↗

Follow the podcast

SpotifyApple PodcastsYouTube