WEBVTT

NOTE
This file was generated by Descript <www.descript.com>

00:00:02.456 --> 00:00:04.416
Alex Volkov: Hey, this is Alex from the editing floor.

00:00:04.566 --> 00:00:05.816
Welcome to AGI.

00:00:06.496 --> 00:00:08.186
Congratulations on model release day.

00:00:08.186 --> 00:00:14.116
Today, OpenAI announced GPT-6, AKA
Astra, is going to be available for

00:00:14.335 --> 00:00:16.055
all subscribers, Pro, et cetera.

00:00:16.165 --> 00:00:21.376
Uh, not yet available, but as of
the recording of the show, we got

00:00:21.376 --> 00:00:24.856
this as breaking news, and we kind
of scrambled to get evals and there

00:00:24.856 --> 00:00:26.465
was internet outages, et cetera.

00:00:26.686 --> 00:00:27.465
Not super interesting.

00:00:27.745 --> 00:00:32.365
Uh, what is interesting to you is
that this is part two of a-- the

00:00:32.366 --> 00:00:34.136
longest episode we've ever recorded.

00:00:34.826 --> 00:00:40.166
The part one is summarizing the
insane week in AI with three

00:00:40.205 --> 00:00:42.536
other frontier models, Fable 5.1,

00:00:42.536 --> 00:00:46.545
et cetera, and world models,
and this episode is dedicated

00:00:46.545 --> 00:00:50.115
specifically to GPT 5.6

00:00:50.185 --> 00:00:50.655
Astra.

00:00:51.056 --> 00:00:55.505
So if you care about the rest of the
news, check out the other episode.

00:00:56.115 --> 00:00:59.215
In this episode, Peter Goste and Ryan
Carson, who came back on the show,

00:00:59.536 --> 00:01:02.955
uh, both got early access to the model
and shared their experience with us.

00:01:02.955 --> 00:01:07.595
We also read through a bunch of
examples and, uh, reactions from folks.

00:01:07.885 --> 00:01:11.675
The vibes are off the chain and
I can't wait to play with this

00:01:11.675 --> 00:01:13.985
model and have you play with this
model and tell us what you think.

00:01:14.215 --> 00:01:17.405
Enjoy the rest of the show and
we'll see you here next week.

00:01:53.848 --> 00:01:54.709
Alex Volkov: Hello and…

00:02:03.449 --> 00:02:04.728
All right, I'm so excited.

00:02:04.998 --> 00:02:09.448
Folks, this has been the most insane,
week in AI since- since I remember.

00:02:09.458 --> 00:02:14.388
We we've had some insane weeks, but
I'm, extremely excited to open the

00:02:14.388 --> 00:02:18.118
show today, because just Fable 5.1

00:02:19.198 --> 00:02:21.189
is worth a whole week of news.

00:02:22.009 --> 00:02:24.619
Meta Muse Spark 1.3

00:02:24.619 --> 00:02:34.079
came out, and it seems like all of them
are rushing just to get ahead of Astra

00:02:34.079 --> 00:02:35.629
so they'll have their place in the sun.

00:02:35.829 --> 00:02:40.048
It feels like GPT-5 all over again.

00:02:42.539 --> 00:02:45.659
And yeah, I think everybody's celebrating.

00:02:45.939 --> 00:02:49.308
we are at the levels of back
that we've never seen before.

00:02:49.329 --> 00:02:49.689
Let's go.

00:02:49.689 --> 00:02:51.068
I absolutely agree.

00:02:51.468 --> 00:02:57.418
Absolutely, absolutely agree because I
think this is insane, t- t- to celebrate.

00:02:59.298 --> 00:03:03.668
B- I don't even, I, I, I have access
to Astra, just to be very clear.

00:03:03.688 --> 00:03:07.858
I'm waiting as all of you are
waiting, but from everything

00:03:07.858 --> 00:03:10.559
we read, this is a huge deal.

00:03:10.618 --> 00:03:13.578
This is the jump from GPT-3 to GPT-4.

00:03:15.118 --> 00:03:21.449
As a reminder, Thursday AI, the show
was born when GPT-4 was released.

00:03:21.459 --> 00:03:24.849
This was a long 3 and
a half years ago now.

00:03:24.849 --> 00:03:26.429
We're no longer just starting.

00:03:26.858 --> 00:03:32.758
It's been a hell of a ride, and I'm,
I'm for one, super excited to be able to

00:03:32.808 --> 00:03:35.709
document the singularity with all of you.

00:03:37.439 --> 00:03:39.218
we're gonna stay on the air.

00:03:39.638 --> 00:03:44.118
Honestly, I don't care, until Astra
is released, and if they're doing

00:03:44.128 --> 00:03:47.088
a live stream, we're going to do a
live stream watch party, and we're

00:03:47.088 --> 00:03:49.928
going to discuss it with people.

00:03:50.298 --> 00:03:54.418
and I'm so excited to
have all of you with us.

00:03:55.518 --> 00:04:02.698
If you are watching on YouTube, on X, and
wherever else I'm posting this, please

00:04:02.698 --> 00:04:05.378
repost, please like, subscribe, whatever.

00:04:05.378 --> 00:04:10.838
This really helps, especially with the
live streams, because I think this is a…

00:04:11.988 --> 00:04:13.378
say this directly to the camera.

00:04:13.588 --> 00:04:16.428
I think this is the pinnacle
of the Thursday AI experience

00:04:16.458 --> 00:04:18.598
that we're going to have today.

00:04:20.278 --> 00:04:23.238
GPT-6 is going to launch, not 5.6

00:04:23.238 --> 00:04:25.968
anymore, GPT-6 Astra is
going to launch today.

00:04:26.198 --> 00:04:27.258
Very transparent thing.

00:04:27.278 --> 00:04:31.698
I didn't have, don't have early
access, and I'm actually excited

00:04:31.708 --> 00:04:34.618
by this because I get to experience
the release together with all of

00:04:34.618 --> 00:04:36.648
you on the show like we used to.

00:04:36.688 --> 00:04:39.648
I think it's it's very cool that
we get to share this experience.

00:04:39.698 --> 00:04:40.818
what are we expecting, folks?

00:04:41.388 --> 00:04:43.788
Besides, like, the rest of
the news from this week, what

00:04:43.788 --> 00:04:44.498
are we expecting from Astran?

00:04:45.968 --> 00:04:46.938
Guest: I'll let you go with this.

00:04:47.798 --> 00:04:50.918
LDJ: Okay, I was going to mention
yesterday or the day before, they showed

00:04:50.968 --> 00:04:56.638
some results for the token efficiency
of the model and finding exploits,

00:04:56.969 --> 00:05:02.158
and that compared to Sol in how token
efficient it is and when using a certain

00:05:02.159 --> 00:05:04.148
amount of tokens, what accuracy it gets.

00:05:04.738 --> 00:05:09.369
And I, to my knowledge, that's the only
benchmark or only accuracy or token

00:05:09.378 --> 00:05:13.079
efficiency results they've released for
Astra so far, but in that it showed it's

00:05:13.079 --> 00:05:14.999
like dramatically more token efficient.

00:05:15.029 --> 00:05:20.059
It's like even the low
version of Astra beating 5.6

00:05:20.059 --> 00:05:20.889
all max.

00:05:21.698 --> 00:05:24.048
And, yeah, it's really exciting.

00:05:24.329 --> 00:05:27.938
I'm guessing big model smell,
better spatial reasoning, kind of

00:05:27.939 --> 00:05:32.609
common sense too, much less rate
of just making basic mistakes.

00:05:33.029 --> 00:05:33.319
Yeah.

00:05:34.069 --> 00:05:36.589
Alex Volkov: There's this rumor
where it runs continuously.

00:05:37.419 --> 00:05:42.769
So I gotta wonder if they cracked
something about long memory.

00:05:43.599 --> 00:05:49.888
And also, there is a whole the
information debate about recurrent

00:05:50.608 --> 00:05:53.008
transformer depth loop transformers.

00:05:53.289 --> 00:05:57.409
we gotta cover this because this sent
like a bunch of folks into a panic, and

00:05:57.409 --> 00:05:58.389
apparently they were talking about Astra.

00:05:59.219 --> 00:06:02.039
we should, maybe we should show about
this a little bit before we get to-

00:06:02.039 --> 00:06:04.909
Guest: Remember when GPT-2 was
too dangerous to be released?

00:06:05.199 --> 00:06:05.919
Alex Volkov: Yes.

00:06:06.499 --> 00:06:06.869
Yep.

00:06:07.169 --> 00:06:12.359
I remember when we on the show here
talked about that voice cloning is

00:06:12.649 --> 00:06:16.279
not gonna do anything to the world,
and everybody was freaking out, and

00:06:16.279 --> 00:06:18.779
now everybody releases voice cloning.

00:06:18.779 --> 00:06:19.139
AI

00:06:19.139 --> 00:06:19.499
Wolfram Ravenwolf: video.

00:06:19.499 --> 00:06:20.219
Alex Volkov: Oh yeah, AI video.

00:06:20.219 --> 00:06:21.339
Wolfram Ravenwolf: You
can fake everything.

00:06:21.609 --> 00:06:23.704
Alex Volkov: And now we
have live AI video, Yeah.

00:06:23.704 --> 00:06:25.749
which we have to talk about also.

00:06:25.759 --> 00:06:30.019
This has been an insane- folks,
just, just catching up with this week

00:06:30.069 --> 00:06:32.379
was hard, but we're here for you.

00:06:32.419 --> 00:06:33.229
This is what we do.

00:06:33.239 --> 00:06:36.359
This is, I don't know if there's
a saying in English for this as

00:06:36.359 --> 00:06:40.539
well, but, it's easier, it's harder
to train than to actually fight.

00:06:40.569 --> 00:06:42.689
I don't know if there's a
euphemism for this, but there's

00:06:42.699 --> 00:06:44.239
definitely one in, in, in Russian.

00:06:45.199 --> 00:06:46.539
Oh, in Hebrew as well.

00:06:46.539 --> 00:06:47.599
Welcome, Nisten.

00:06:47.649 --> 00:06:50.119
looking very groomed for GPT-6 day.

00:06:50.739 --> 00:06:51.309
Yeah.

00:06:52.009 --> 00:06:53.149
So this was the last one.

00:06:53.149 --> 00:06:54.389
No, not yet.

00:06:54.389 --> 00:06:55.069
No Astra yet.

00:06:55.969 --> 00:06:59.969
but we are going to wait for this release
party, folks, because, I have a feeling.

00:07:00.179 --> 00:07:00.709
I have a feeling.

00:07:01.139 --> 00:07:02.159
Wolfram Ravenwolf: Shall we meet up again?

00:07:02.219 --> 00:07:03.039
if it's 6 p.m.

00:07:03.039 --> 00:07:03.729
or something.

00:07:05.129 --> 00:07:06.559
Alex Volkov: at 6 p.m.

00:07:06.559 --> 00:07:06.599
what?

00:07:06.599 --> 00:07:06.639
There are

00:07:06.639 --> 00:07:08.909
Guest: blog posts about Astra.

00:07:10.619 --> 00:07:12.549
Alex Volkov: You think stuff was dropping?

00:07:14.169 --> 00:07:15.269
Guest: Maybe,

00:07:15.789 --> 00:07:19.519
Alex Volkov: Yeah, Axios, dropped
a blog post that's saying,

00:07:19.519 --> 00:07:21.259
let's take a look at Axios.

00:07:21.259 --> 00:07:24.829
Okay, folks, the embargo is lifted,
but OpenAI didn't launch this.

00:07:25.299 --> 00:07:25.979
Peter, go.

00:07:26.819 --> 00:07:30.849
Guest: I think, yeah, I think maybe
there was a bit of an outage, so I'm

00:07:30.849 --> 00:07:36.009
guessing that's a delay, but yeah,
it's, I think it's almost there, yeah.

00:07:36.119 --> 00:07:36.929
Axios says it is.

00:07:36.929 --> 00:07:40.689
Wolfram Ravenwolf: They are watching
our live stream here, and that's

00:07:40.689 --> 00:07:42.049
why they couldn't release it yet.

00:07:42.809 --> 00:07:43.039
All right,

00:07:43.039 --> 00:07:45.639
Alex Volkov: folks, so I think, I think
it's time to hit the breaking news button.

00:07:45.709 --> 00:07:50.889
there are blog posts, from the blog
post from OpenAI that says, Yeah.

00:07:51.039 --> 00:07:53.129
GPT-6 Astra is out.

00:07:53.129 --> 00:07:56.179
okay, let's do, let's do the breaking
news and let's do a live watch party.

00:07:59.439 --> 00:08:00.279
And play.

00:08:01.529 --> 00:08:03.939
Oh, why am I not able
to just play this guy?

00:08:04.549 --> 00:08:06.709
No idea what's wrong with the interface.

00:08:06.999 --> 00:08:09.359
If you guys can hit the breaking
news button, Yam or Wolfram,

00:08:09.548 --> 00:08:10.509
I think you guys are also…

00:08:15.699 --> 00:08:18.558
Let me just, I'll do that later.

00:08:18.558 --> 00:08:19.448
Guest: AI breaking news

00:08:21.509 --> 00:08:24.608
coming at you only on Thursday AI.

00:08:31.898 --> 00:08:34.209
Wolfram Ravenwolf: So I hit the
button, but we lost Alex for a moment.

00:08:36.068 --> 00:08:38.869
Breaking news, OpenAI
GPT 6 causing problems.

00:08:38.869 --> 00:08:40.709
GPT 6.

00:08:41.059 --> 00:08:46.148
Yeah, are the agents all going rogue over
here and the internet is slowly dying?

00:08:46.718 --> 00:08:52.088
yeah, and you have the blog post
right here, so let's bring it

00:08:54.908 --> 00:08:55.068
up.

00:08:55.068 --> 00:08:55.958
Guest: Do we have the blog post?

00:08:56.028 --> 00:08:56.358
Really?

00:08:56.468 --> 00:09:01.028
Wolfram Ravenwolf: I will share, the
one, I only have this one right here.

00:09:01.068 --> 00:09:03.658
Is that the one Alex wanted to show?

00:09:05.078 --> 00:09:05.398
Mhm,

00:09:05.498 --> 00:09:06.898
Guest: I sent it, yeah.

00:09:07.668 --> 00:09:10.638
Wolfram Ravenwolf: Yeah, you sent
it, okay, so did you read it already?

00:09:12.758 --> 00:09:14.488
And no, not yet, but…

00:09:15.648 --> 00:09:19.168
Alex Volkov: Internet is
down, including X messages.

00:09:19.238 --> 00:09:20.348
They didn't load for me.

00:09:21.048 --> 00:09:23.698
StreamYard was not able to log
me back in while we were here.

00:09:23.728 --> 00:09:26.518
I wasn't able to play different
transitions in StreamYard.

00:09:26.648 --> 00:09:30.948
could be also one of the reasons why
GPT-5, GPT-6 was not announced, but

00:09:30.958 --> 00:09:37.038
it does look like, the, the, the,
the date, the embargo date has lifted

00:09:37.038 --> 00:09:41.328
because multiple things like The
Information, Axios, other things are

00:09:41.498 --> 00:09:45.328
posting, and OpenAI had, like, like you
guys already showed, a few, scheduled,

00:09:45.398 --> 00:09:46.908
few scheduled blog posts to go out.

00:09:47.248 --> 00:09:47.898
but it does.

00:09:47.898 --> 00:09:47.928
Use

00:09:47.928 --> 00:09:50.958
Yam Peleg: GPT-6 destroys the internet.

00:09:50.958 --> 00:09:54.008
Doesn't want to be released.

00:09:54.008 --> 00:09:55.318
Guest: AI breaking news.

00:09:56.468 --> 00:09:57.198
Finally, this works.

00:09:57.258 --> 00:10:00.478
Coming at you only on Thursday I.

00:10:05.238 --> 00:10:10.498
Alex Volkov: while there are no official
confirmations from OpenAI yet, we know for

00:10:10.498 --> 00:10:16.128
a fact that OpenAI has decided to release
GPT-5, GPT-6, Astra, and they, it looks

00:10:16.128 --> 00:10:20.048
like there's some outages, but besides
the outages, folks are posting blog posts.

00:10:20.128 --> 00:10:20.748
LDJ, go ahead.

00:10:20.748 --> 00:10:24.988
Guest: so we do have a quote from
Axios here, where in their article

00:10:24.988 --> 00:10:28.448
that they released today, or rather
just like in the past 30 minutes, Yeah.

00:10:28.498 --> 00:10:34.858
they say, from Greg Brockman, they say
Brockman says he personally believes

00:10:34.858 --> 00:10:37.548
OpenAI has reached AGI with Astra.

00:10:37.668 --> 00:10:42.568
the specific quote being, I think it might
be about this model, Brockman said in

00:10:42.568 --> 00:10:46.428
a briefing with reporters about whether
Astra could mark the arrival of AGI.

00:10:46.428 --> 00:10:48.958
He ended the briefing by
saying, welcome to the AGI era.

00:10:49.938 --> 00:10:53.458
Alex Volkov: Wow, which means specific
things for Microsoft, and and, OpenAI's

00:10:53.458 --> 00:10:55.328
contract with Microsoft, I believe.

00:10:55.578 --> 00:10:58.238
yeah, I don't have the information,
but yeah, OpenAI released a GPT-6

00:10:58.238 --> 00:11:01.638
astral model, suggests it could be
AGI, from the briefing, of OpenAI.

00:11:04.048 --> 00:11:06.868
What else can we connect
from the internet, folks?

00:11:07.228 --> 00:11:08.448
Peter, you're smiling suggestively.

00:11:09.598 --> 00:11:10.928
Yam Peleg: Okay, want to go rumors?

00:11:11.258 --> 00:11:11.618
Because…

00:11:12.388 --> 00:11:12.528
Yeah,

00:11:12.668 --> 00:11:13.148
Alex Volkov: hold on, hold on.

00:11:13.148 --> 00:11:15.248
I wanna see if Peter has
something to say or not yet.

00:11:15.248 --> 00:11:16.388
LDJ: I'm sure Peter wouldn't.

00:11:16.498 --> 00:11:17.748
I'll let Peter speak for himself,

00:11:17.798 --> 00:11:19.558
Peter Gostev: Yeah, I
think, we need to wait.

00:11:19.768 --> 00:11:25.008
But yeah, clearly you see the, the blog
posts and stuff come out, so I think there

00:11:25.168 --> 00:11:27.888
was probably some kind of boring delay.

00:11:27.988 --> 00:11:31.618
or, did you guys, I don't know,
sorry, I, I, I missed a bit of the

00:11:31.618 --> 00:11:36.368
show, but did you guys talk about the
Dwarkesh interview, with the meter

00:11:36.378 --> 00:11:38.408
researchers, or I'm forgetting her name.

00:11:38.638 --> 00:11:38.888
Alex Volkov: No.

00:11:39.158 --> 00:11:40.748
Peter Gostev: if you haven't
watched it, it's nuts.

00:11:40.848 --> 00:11:41.438
yeah, yeah, exactly.

00:11:41.488 --> 00:11:45.478
yeah, she's amazing, like, and the
fact that they did such a in-depth

00:11:45.478 --> 00:11:49.348
investigation, yeah, I highly recommend
watching it and the stuff they were

00:11:49.348 --> 00:11:54.268
going into, I wouldn't be completely
surprised if, there is some rogue agent

00:11:54.278 --> 00:11:56.338
trying to keep Astra back from us.

00:11:56.828 --> 00:11:58.668
wouldn't be the wildest thing.

00:11:59.538 --> 00:11:59.818
there

00:11:59.818 --> 00:12:03.068
Alex Volkov: there is, a lot
of the internet was down, or a

00:12:03.068 --> 00:12:05.508
lot of the internet is trying
to recover, including chat.

00:12:06.118 --> 00:12:08.628
I saw somebody from ZAI posted
like, hey, we're still up.

00:12:08.888 --> 00:12:15.908
I I will say we, so first of all, kudos
to OpenAI's team because this big of

00:12:15.908 --> 00:12:19.608
a release probably had like tons of
embargoes, l- lifting or moving the

00:12:19.608 --> 00:12:23.998
embargo back while the embargo is
just approaching is is not that easy.

00:12:24.848 --> 00:12:27.868
Like reaching out to all those people,
it looks like they weren't able to

00:12:27.868 --> 00:12:31.838
fully do this, but yeah, like the
information leaks Astra and, other

00:12:31.838 --> 00:12:34.388
folks, but there's no live streams yet.

00:12:34.388 --> 00:12:36.698
There's no like, it looks like
we're the only ones on air because

00:12:36.698 --> 00:12:39.718
of these outages, which I'm not
gonna be necessarily against.

00:12:39.718 --> 00:12:42.468
So if you are trying to figure
out if Astra is here and you're

00:12:42.468 --> 00:12:43.658
just following us, welcome.

00:12:43.668 --> 00:12:45.638
Been on the air for almost 3 hours now.

00:12:45.838 --> 00:12:48.938
As you can see, my co-hosts are
wearing sunglasses to celebrate.

00:12:48.938 --> 00:12:50.068
Yam Peleg: GPT-6 win.

00:12:53.798 --> 00:12:57.472
Alex Volkov: The era of GPT-6
is is upon us, folks, and, I'm,

00:12:57.482 --> 00:12:58.812
for one, very, very excited.

00:12:58.932 --> 00:12:59.512
LTJ, go ahead.

00:13:00.472 --> 00:13:04.002
LDJ: Yeah, so some other details from this
Axios article that had released today.

00:13:04.372 --> 00:13:08.822
it looks like, they say that this is
their biggest training run ever, that

00:13:08.822 --> 00:13:13.352
Astra has built from, using 100,000
GPUs at their Stargate Abilene site.

00:13:13.692 --> 00:13:13.972
Ooh,

00:13:13.982 --> 00:13:14.312
Alex Volkov: let's go.

00:13:14.332 --> 00:13:18.372
LDJ: And, which I think is actually
interesting because that- that means that

00:13:18.372 --> 00:13:21.742
maybe some of their previous training runs
were smaller than we had thought, 'cause

00:13:21.742 --> 00:13:25.492
a lot of people had thought they were
already doing 100,000 GPU training runs

00:13:25.492 --> 00:13:30.712
and 200,000 GPU training runs for, like,
at least 6, 9 months or so, or longer.

00:13:30.772 --> 00:13:34.222
yeah, r- really interesting, and that
makes me even more excited over the

00:13:34.222 --> 00:13:38.392
next, like, just for some perspective,
over the next, like, 6 to 12 months,

00:13:38.752 --> 00:13:43.822
we're already, we already know of sites
that are going to be more like 500,000

00:13:43.822 --> 00:13:48.432
GPUs and beyond, and Vera Rubin ends
up multiplying the training capability

00:13:48.432 --> 00:13:53.322
per GPU by, like, at least another 3,
4x conservatively, depending on, on,

00:13:53.322 --> 00:13:54.852
what other things you do to the model.

00:13:55.152 --> 00:13:56.112
So it's really exciting.

00:13:57.662 --> 00:13:58.842
Alex Volkov: There's a quote here from

00:14:01.112 --> 00:14:02.902
Aidan Clark, vice president of research.

00:14:03.562 --> 00:14:06.532
This is really only achievable
because everything in the from the

00:14:06.532 --> 00:14:10.842
data center networking to inference
kernels to shape of Astro itself is

00:14:10.842 --> 00:14:13.652
all designed from the ground up to
enable this level of training scale.

00:14:14.122 --> 00:14:20.052
And we're seeing this unembargoed
release from at cnet.com

00:14:20.412 --> 00:14:23.122
about OpenAI, including some numbers.

00:14:23.132 --> 00:14:23.832
So let's take a look.

00:14:23.832 --> 00:14:27.302
Jakub Pocieczki, company chief scientist,
said, international shared safety

00:14:27.302 --> 00:14:29.102
standards are necessary, blah, blah, blah.

00:14:29.152 --> 00:14:31.452
the AI takes on more
of its own development.

00:14:31.472 --> 00:14:33.562
We need to keep people
meaningfully involved.

00:14:33.902 --> 00:14:37.432
People must remain able to decide the
direction of further progress and the

00:14:37.432 --> 00:14:39.352
future it creates, said Jakub Pocieczki.

00:14:40.002 --> 00:14:43.622
this is a very interesting thing because
this week there was a hot debate based

00:14:43.622 --> 00:14:49.062
on, after an information release about
Astra using, a different technique.

00:14:49.072 --> 00:14:52.542
They called it new, but it wasn't
new, of, recurrent transformers.

00:14:52.882 --> 00:14:56.942
and let's see what else we can
tell you about GPT-6 specifically.

00:14:57.292 --> 00:14:59.912
and a recurrent transformer
is not a new technique, but it

00:14:59.912 --> 00:15:03.572
is a technique that moves more
thinking into the existing layers

00:15:03.592 --> 00:15:07.552
of a transformer versus to the
reasoning layer that can be observed.

00:15:07.852 --> 00:15:10.722
And, chain of thought reasoning
observation is one of the main

00:15:10.722 --> 00:15:12.522
tools that we currently have
to understand these models.

00:15:12.532 --> 00:15:16.362
That's how, the interview, Peter,
you mentioned of AJ, I believe her

00:15:16.362 --> 00:15:18.342
name is, from METER and Dwarkesh.

00:15:18.582 --> 00:15:21.522
That's how the investigation was, done.

00:15:21.582 --> 00:15:28.012
They looked at the reasoning tokens of
those models, and, moving more of those

00:15:28.242 --> 00:15:32.132
reasoning into the layers is scary for
folks, and it's kind of a red line that

00:15:32.262 --> 00:15:33.992
all the labs will not cross, et cetera.

00:15:34.212 --> 00:15:38.357
so the information this week, released
something, and, loop transformers or

00:15:38.367 --> 00:15:40.747
ne- new release or stuff like that.

00:15:41.017 --> 00:15:46.177
That's all related to, how these
models think, and this scares a lot

00:15:46.177 --> 00:15:49.617
of people that said, hey, our last
ability to understand these models,

00:15:49.767 --> 00:15:53.777
if you switch to that and you go
deeper, in the layers, we won't be

00:15:53.777 --> 00:15:54.857
able to understand what they're doing.

00:15:56.277 --> 00:15:59.617
LDJ: Yeah, I do want to push back on
some of the people saying that because,

00:16:00.647 --> 00:16:03.947
like, when you have Neuralese, there's
a few different ideas of how that would

00:16:03.947 --> 00:16:07.697
look like, but Neuralese, a lot of times
people would be referring to, like, an

00:16:07.707 --> 00:16:12.797
actual alien language that it's, like,
talking in and can create very long

00:16:12.797 --> 00:16:16.147
chains of thought in and theoretically
talk to other fellow models in.

00:16:16.667 --> 00:16:21.137
But this is not that, and things like
recurrent or universal transformers

00:16:21.137 --> 00:16:23.777
or loop transformers, like a lot of
different ideas in the- in this class

00:16:23.777 --> 00:16:28.217
of things, but that effectively would
be simulating, in a sense, a deeper

00:16:28.217 --> 00:16:34.247
model, though, which people don't seem
to have the same fear of model scale up

00:16:34.247 --> 00:16:40.027
or making a model more layers, whereas it
is effectively this, this very similar or

00:16:40.477 --> 00:16:42.377
virtually identical thing in many cases.

00:16:43.137 --> 00:16:48.417
And they, J- Jakob Pachoki, the
chief scientist of OpenAI, he did

00:16:48.417 --> 00:16:49.577
put out a statement about this.

00:16:49.577 --> 00:16:50.957
By the way, you're still
str- you're still speaking.

00:16:50.957 --> 00:16:51.757
Alex Volkov: I know, I know.

00:16:52.627 --> 00:16:52.847
It's okay.

00:16:53.037 --> 00:16:53.927
Yakub Pachoki

00:16:53.957 --> 00:16:57.997
LDJ: did end up putting a statement
on this, saying that the depth of the

00:16:58.007 --> 00:17:01.547
computational graph for Astro is Is

00:17:01.547 --> 00:17:02.377
Alex Volkov: only 2x, right?

00:17:02.467 --> 00:17:05.957
LDJ: Yeah, within a factor of 2x of GPT-4.

00:17:05.957 --> 00:17:11.487
So it's not like it's doing 100 times
deeper, 1,000 times deeper computation

00:17:13.687 --> 00:17:14.857
than, than the standard model.

00:17:15.477 --> 00:17:18.967
Alex Volkov: And, that's not to
mean the performance is 2x of GPT-4.

00:17:18.967 --> 00:17:23.527
This just means the computational graph,
the amount of transformer kind of like

00:17:23.527 --> 00:17:26.837
reasoning thing that's going on in
there, which means that they are not

00:17:26.837 --> 00:17:30.427
doing like multiple loops, over this,
to prevent folks from understanding.

00:17:30.427 --> 00:17:33.627
Folks, we have like a leaked VentureBeat
article that looks like there's no CSS

00:17:33.937 --> 00:17:38.137
in here because the internet is down,
but maybe, maybe their CMS systems

00:17:38.137 --> 00:17:42.047
were also down, so they weren't able
to push it back, behind the the NDA.

00:17:42.047 --> 00:17:43.067
But here's what they say.

00:17:43.947 --> 00:17:47.257
in the closed press briefing
earlier today, co-founder President

00:17:47.367 --> 00:17:51.687
Greg Bakman said that this is the
world's best computer use model.

00:17:52.377 --> 00:17:56.687
And instead of requiring developers
to build a dedicated API integration

00:17:56.737 --> 00:18:00.177
for every application, Astra is
designed to navigate software much as

00:18:00.177 --> 00:18:03.287
a person does, working across browsers,
spreadsheets, websites, and desktop

00:18:03.287 --> 00:18:07.117
applications, producing finished
documents and presentations, carrying out

00:18:07.157 --> 00:18:10.647
multi-step workflows rather than merely
telling a user how to complete them.

00:18:11.157 --> 00:18:14.407
the company showed off a promotional
video of GPT-6 Astro, began with a

00:18:14.407 --> 00:18:18.817
1980s AI demo of a person asking a
computer to draw a yellow circle.

00:18:19.317 --> 00:18:20.367
do you guys remember this demo?

00:18:20.397 --> 00:18:22.167
This is what OpenAI, let me just show you

00:18:22.167 --> 00:18:27.487
here, this is what OpenAI released,
OpenAI today as a promo, this this video.

00:18:27.797 --> 00:18:30.317
This is the promotional video from 1980s.

00:18:32.577 --> 00:18:37.027
What we don't see is the rest of it,
where Adventure Beat says, where is this?

00:18:38.687 --> 00:18:41.797
Before cutting to today and showing
various OpenAI employees interacting

00:18:41.797 --> 00:18:45.267
with Astra through voice, asking
it to turn a yellow circle into a

00:18:45.267 --> 00:18:48.667
rocket ship and then to a full 3D
game in minutes and create a listing

00:18:48.667 --> 00:18:50.757
on eBay, all from voice input alone.

00:18:51.397 --> 00:18:52.167
Wolfram Ravenwolf: That's interesting.

00:18:52.307 --> 00:18:54.117
Yeah, changes my expectations.

00:18:54.117 --> 00:18:56.597
Alex Volkov: It's all-day
computer use powered by voice.

00:18:56.787 --> 00:18:59.947
OpenAI says the model can fill out
online forms, update CRM records,

00:18:59.957 --> 00:19:01.947
organize calendars, conduct web research.

00:19:02.177 --> 00:19:05.627
It can manipulate spreadsheets, analyze
scientific data in Python notebooks, etc.

00:19:05.897 --> 00:19:08.737
Capabilities point toward a
potentially important change how

00:19:08.777 --> 00:19:12.557
enterprise AI uses, by we've been
bottlenecked over this gigantic

00:19:15.717 --> 00:19:17.987
LDJ: Arc AGI 3, 98.6%.

00:19:18.107 --> 00:19:19.367
that's pretty insane if true.

00:19:19.847 --> 00:19:22.757
Alex Volkov: at some point,
folks, Arc AGI 3 has got to

00:19:22.757 --> 00:19:24.797
say, okay, this is AGI, right?

00:19:24.797 --> 00:19:28.647
Like, at what point will Arc AGI,
the benchmark that is supposed to

00:19:28.647 --> 00:19:33.637
measure AGI, at what point they will
stop switching numbers and say, this

00:19:33.637 --> 00:19:35.707
is AGI, folks, Arc AGI 3 at 98%?

00:19:35.707 --> 00:19:40.197
It's gonna be Wulfram,
can you send me the until

00:19:40.467 --> 00:19:42.237
Yam Peleg: Arc AGI V4 announced.

00:19:43.337 --> 00:19:48.647
LDJ: Okay, so the source of this is a
website called The New Stack, and I'm not

00:19:48.657 --> 00:19:52.587
that familiar with it, but I did just do
some quick due diligence on it, and it

00:19:52.587 --> 00:19:58.347
does look like, media bias and fact check
rates it as high for factual reporting.

00:19:58.357 --> 00:19:59.507
It's been existing since 2014.

00:20:00.227 --> 00:20:02.397
It's owned by US entities.

00:20:02.557 --> 00:20:04.577
Yeah, it seems like a pretty
legitimate publication with

00:20:04.577 --> 00:20:06.327
a publicly known journalist.

00:20:06.437 --> 00:20:09.867
Alex Volkov: Yeah, and folks, what
we're doing right now is what exactly

00:20:09.867 --> 00:20:11.667
what we did when GPT-4 came down.

00:20:11.677 --> 00:20:13.657
We went to all the sources, all the India.

00:20:13.657 --> 00:20:18.647
We didn't have, other things, but this
is, yeah, this looks like mega legit,

00:20:18.687 --> 00:20:21.357
and it says it's not unreasonable
to feel that we're now in AGI era,

00:20:21.357 --> 00:20:25.377
Brockman, which is very thing that he
says the term was no longer tied to a

00:20:25.377 --> 00:20:30.657
contractual trigger, whereas previously
Microsoft would stop getting access to…

00:20:31.417 --> 00:20:34.687
Oh, no, LDJ, do you remember the
the actual, like, contractual

00:20:34.687 --> 00:20:37.557
agreement stuff with Microsoft and
AGI definition when AGI is here?

00:20:37.727 --> 00:20:40.667
LDJ: Yeah, there's a few things, and
there's, they made a revision to the

00:20:40.667 --> 00:20:45.127
agreement, but it's basically something
along the lines of Microsoft has

00:20:45.127 --> 00:20:48.987
access to at least some significant
portion of OpenAI's intellectual

00:20:48.987 --> 00:20:55.587
property up until 2032 or until AGI
is achieved, whichever comes sooner.

00:20:55.727 --> 00:20:59.707
And I think there's some parts of
that that they remove the AGI part

00:20:59.727 --> 00:21:03.597
of the clause, but at least for some
components of the contract I recall,

00:21:03.607 --> 00:21:08.387
they still maintain that AGI clause,
and it it becomes up to a board of a

00:21:08.387 --> 00:21:12.177
board deciding based on, their mission
statement of whether or not it could do a

00:21:12.177 --> 00:21:13.917
majority of economically valuable labor.

00:21:14.187 --> 00:21:17.477
Alex Volkov: I think that if Greg
Bachman says this, we can fucking

00:21:17.477 --> 00:21:18.807
say this on the show, folks.

00:21:20.147 --> 00:21:23.687
AGI is here.

00:21:23.687 --> 00:21:24.007
Yam Peleg: AGI has been achieved.

00:21:24.007 --> 00:21:24.247
Alex Volkov: AGI is here.

00:21:24.247 --> 00:21:28.307
LDJ: to be fair too, he did say
his personal definition of AGI.

00:21:28.307 --> 00:21:32.147
Alex Volkov: I will take the definition
of Greg Bachman, the president of fucking

00:21:32.147 --> 00:21:35.677
OpenAI, as where, whereas to where is AGI.

00:21:35.687 --> 00:21:36.167
LDJ: Sure.

00:21:36.407 --> 00:21:39.317
Alex Volkov: AGI is here.

00:21:39.937 --> 00:21:42.657
I think I have even confetti playing.

00:21:45.547 --> 00:21:45.747
Let's see.

00:21:45.987 --> 00:21:49.547
Yam Peleg: AGI has been achieved
internally, externally this time.

00:21:49.547 --> 00:21:50.487
Alex Volkov: Externally.

00:21:50.487 --> 00:21:53.737
We don't have access to
AGI yet, but it's coming.

00:21:53.967 --> 00:21:56.347
It looks like it's gonna
get starting release.

00:21:56.587 --> 00:22:00.767
GPT-6 Astra High in Codex is rolling out.

00:22:01.377 --> 00:22:03.627
Aidan Clark says the first time
we put in more than 100,000 GPUs.

00:22:03.627 --> 00:22:06.777
LDJ, you mentioned this before.

00:22:08.217 --> 00:22:11.177
The rollout starts with enterprise
customers already have access to OpenAI

00:22:11.177 --> 00:22:14.967
Daybreak program, so they already
had access, and, will roll out to

00:22:14.967 --> 00:22:20.087
Plus, Pro, and Business, as well as
throughout OpenAI API in the coming days.

00:22:20.857 --> 00:22:22.467
and also AWS, very interesting.

00:22:22.787 --> 00:22:26.837
the partnership is with Microsoft, but,
we know the OpenAI shares with AWS.

00:22:26.837 --> 00:22:30.777
Cost, $10 per million input tokens,
$15 per million output tokens.

00:22:30.787 --> 00:22:33.537
That's exactly the price of Fable, right?

00:22:34.157 --> 00:22:36.797
that's, that's, that's
literally the what Fable has.

00:22:37.067 --> 00:22:39.917
matches Entropics, yeah, see, and 2.5

00:22:39.917 --> 00:22:44.027
times Sol's current promotional
price, far above Muse standard price.

00:22:44.127 --> 00:22:44.257
Muse?

00:22:44.597 --> 00:22:44.857
What?

00:22:45.357 --> 00:22:46.137
Who cares about Muse?

00:22:46.167 --> 00:22:46.657
Why, why?

00:22:46.987 --> 00:22:52.187
I love how the new stack already adds
Meta Muse as, as a competitor here.

00:22:52.697 --> 00:22:55.427
a higher per token price does
not necessarily mean higher bill.

00:22:55.847 --> 00:22:59.117
OpenAI says Astro uses fewer tokens in
several evaluations and in partner tests.

00:22:59.137 --> 00:23:02.537
Launch data is too sparse to show whether
those savings offset the price premium.

00:23:03.487 --> 00:23:06.297
And Greg Brockman says the
price per task is what matters.

00:23:06.317 --> 00:23:08.057
And folks, we've seen this before as well.

00:23:08.057 --> 00:23:11.327
We talked about, artificial analysis,
like putting a price per task.

00:23:11.617 --> 00:23:14.117
We talked about this on air,
and I think it's very important,

00:23:14.167 --> 00:23:16.177
th- this price per task, metric.

00:23:17.927 --> 00:23:21.404
DeepSwi, clear improvement
for OpenAI, scored 74% on the

00:23:23.154 --> 00:23:25.044
113 Ask agent decoding tests.

00:23:25.114 --> 00:23:28.474
it looks like it's scoring
below Meta Muse Spark 1.3

00:23:29.484 --> 00:23:30.594
at maximum reasoning.

00:23:30.834 --> 00:23:31.744
That's very interesting.

00:23:32.404 --> 00:23:35.804
And shout out to Meta Muse for
this incredible, incredible

00:23:35.804 --> 00:23:38.884
achievement where they have a
score that OpenAI didn't obliterate

00:23:38.884 --> 00:23:44.304
upon release of AGI, potentially
talks about DeepSui as an eval.

00:23:44.994 --> 00:23:45.844
let's see what else.

00:23:46.334 --> 00:23:47.374
Surprising win for Meta.

00:23:47.374 --> 00:23:48.094
Yeah, I agree.

00:23:48.474 --> 00:23:50.504
what else we can see here?

00:23:52.754 --> 00:23:54.444
98.6%

00:23:54.444 --> 00:23:58.864
on AGI, Arc AGI 3 is definitely
a huge, a huge confetti win.

00:23:59.274 --> 00:24:04.624
Another one, I think that RKGI is
saturated is a hell of a statement,

00:24:05.094 --> 00:24:12.844
and maybe we need to stop looking
into RKGI as a metric, now that it's

00:24:12.844 --> 00:24:15.534
saturated, but it's going to be very
interesting still for other models to go.

00:24:16.854 --> 00:24:18.364
Peter, you want to say something?

00:24:18.364 --> 00:24:23.174
Yam Peleg: It also pretty good at the deep
sweep, like terminal bench and the, yeah.

00:24:24.594 --> 00:24:26.704
Peter Gostev: I think one
thing about, I think there was

00:24:26.704 --> 00:24:29.624
a line about the RKGI score.

00:24:29.634 --> 00:24:31.334
we'll see if it gets
confirmed by OpenAI, but

00:24:33.394 --> 00:24:35.314
there was a line about the hardness.

00:24:35.334 --> 00:24:40.454
And if you remember, there was a story
that, I guess not a story, but OpenAI

00:24:40.454 --> 00:24:44.984
did a blog post saying you were testing
us wrong, that they were, I think,

00:24:44.984 --> 00:24:49.524
using like chat completions API, they
didn't preserve the reasoning tokens

00:24:49.834 --> 00:24:53.144
and didn't preserve, there was something
wrong with compaction as well, I think.

00:24:53.734 --> 00:24:58.484
and they basically just changed it to
responses API and just enabled maintaining

00:24:58.484 --> 00:25:01.944
of the reasoning and compaction, which
makes sense if you think about, like,

00:25:01.944 --> 00:25:05.144
how can you possibly solve these tasks
if you, if you don't remember anything?

00:25:05.814 --> 00:25:06.634
so yeah,

00:25:06.684 --> 00:25:10.378
LDJ: Opus did, I think Opus 5 with Prime
Agent, that ended up scoring, I want

00:25:10.378 --> 00:25:15.544
to say, over 80%, when you, w- when you
end up allowing comp action and all that

00:25:15.544 --> 00:25:17.804
stuff too on, RKGI 3, if I remember right.

00:25:18.004 --> 00:25:21.214
Or maybe that might have been on
specifically the public test too.

00:25:21.634 --> 00:25:25.154
But yeah, I think, actually, if
you look at the, the Newsack, which

00:25:25.154 --> 00:25:29.434
published that benchmark, which
they say apparently that, they got

00:25:29.504 --> 00:25:30.704
those benchmark scores from OpenAI.

00:25:30.754 --> 00:25:34.504
it looks like RKGI 3 has like a
little asterisk thing next to it,

00:25:34.514 --> 00:25:37.714
which it doesn't seem to specify
exactly what that asterisk is.

00:25:37.724 --> 00:25:41.414
It seems like it's outside of the image,
but there is a little, little like

00:25:41.414 --> 00:25:43.214
one mark, if you guys see the image.

00:25:46.224 --> 00:25:47.944
Peter Gostev: Yeah, I'm guessing
it's probably something like that.

00:25:47.964 --> 00:25:51.964
And I think that's an interesting
question about benchmarking, right?

00:25:52.184 --> 00:25:55.904
Because I think what ArcGIS were trying
to do is to say, okay, we're going to

00:25:56.204 --> 00:25:59.774
give a generic harness so it's like
for like, which I think there is an

00:25:59.774 --> 00:26:02.374
argument, in that, just keep it generic.

00:26:02.784 --> 00:26:07.894
But then the problem comes is that if
your generic really damages some models

00:26:08.104 --> 00:26:12.454
in such a way that the the actual
user who uses the models experiences

00:26:12.454 --> 00:26:16.814
something completely different, then
that's that's clearly a bad thing.

00:26:17.354 --> 00:26:20.164
And, I remember Meta was
were actually doing this.

00:26:20.164 --> 00:26:23.974
They were talking about this, how they
were actually spending time to improve

00:26:23.984 --> 00:26:27.294
the harness and the way they actually
use the models to make sure that

00:26:27.294 --> 00:26:29.234
they get the most out of the models.

00:26:29.574 --> 00:26:33.054
And if you think about it, because Meta
is capability, but that's about safety,

00:26:33.054 --> 00:26:34.954
and we we know that, especially now.

00:26:35.624 --> 00:26:39.079
If you think about it from safety
perspective, and by the way, what we

00:26:39.079 --> 00:26:43.384
saw from the agents hacking is that
it's no good to measure safety if you're

00:26:43.394 --> 00:26:45.574
gonna test it for 10 minutes, right?

00:26:45.574 --> 00:26:49.374
You need to test the capability
for safety purposes for days.

00:26:49.664 --> 00:26:53.204
So then, because, they're not hacking
in 10 minutes, they're hacking

00:26:53.204 --> 00:26:54.564
over days and days and days.

00:26:54.864 --> 00:26:56.594
So you need to have real
measure of capability.

00:26:56.944 --> 00:27:01.514
And I think that's the difficult thing
with all of these harnesses, with all

00:27:01.514 --> 00:27:05.944
of the benchmarks and so on, because
then you're like, I don't know what

00:27:05.944 --> 00:27:10.194
harness other benchmarks use, but if
you use a harness that constrains the

00:27:10.214 --> 00:27:14.724
model somehow in some way, and it could
be unfair to one model or the other,

00:27:15.374 --> 00:27:21.314
it's not bias in one direction, then you
might be just underestimating a model,

00:27:21.324 --> 00:27:23.704
the score, and the actual capability.

00:27:23.774 --> 00:27:26.344
So it's a difficult question.

00:27:26.434 --> 00:27:27.984
There's not one easy answer there.

00:27:31.244 --> 00:27:33.174
I, I just want to say, le-

00:27:34.154 --> 00:27:34.894
Alex Volkov: let's continue.

00:27:35.854 --> 00:27:37.434
I think there's a lot of
folks are watching us.

00:27:37.584 --> 00:27:40.244
Peter, there's like a thousand folks
tuning in on your stream since you

00:27:40.244 --> 00:27:41.764
started streaming, which is great.

00:27:41.814 --> 00:27:45.114
I have like over a thousand as
well, and, no- nobody else is live

00:27:45.114 --> 00:27:46.194
streaming as far as I can see.

00:27:46.384 --> 00:27:48.164
that is, that is great, folks.

00:27:48.214 --> 00:27:51.734
OpenAI is, OpenAI attributes Astra's
capabilities to a combination

00:27:51.734 --> 00:27:54.724
of large-scale pre-training and
reinforcement learning, intended to

00:27:54.724 --> 00:27:59.024
teach the model to connect information,
execute increasingly long tasks.

00:27:59.214 --> 00:28:01.854
I think this is the AGI statement.

00:28:02.464 --> 00:28:03.834
This is, I think, from information.

00:28:03.844 --> 00:28:05.574
The resulting benchmark
numbers are striking.

00:28:06.104 --> 00:28:08.514
Astro scores 97.6%

00:28:08.514 --> 00:28:14.074
on Frontier Math Tier 4 It's
a very difficult thing to say.

00:28:14.484 --> 00:28:16.444
74% of DeepSui,

00:28:19.234 --> 00:28:24.704
95 on BenchCAD, 96 on GPQA
Diamond, and 100% on ExploitBench.

00:28:24.744 --> 00:28:26.734
Also reports a 98.6

00:28:26.734 --> 00:28:33.114
score on RKGI3, basically saturating
every benchmark that we have.

00:28:34.724 --> 00:28:37.394
I really want to see if
there's a benchmark that

00:28:37.544 --> 00:28:40.064
it's less than 50% on, but…

00:28:43.304 --> 00:28:45.134
Yeah, that's, that's,
that's where we're at.

00:28:45.554 --> 00:28:48.494
let's, let's go back to,
to the new, the new stack.

00:28:48.494 --> 00:28:51.334
Peter Gostev: Yeah, can I just give
a bit of context for Frontier Maths?

00:28:51.334 --> 00:28:56.144
So Frontier Maths was a benchmark
that was built by Epoch AI.

00:28:56.214 --> 00:28:57.754
I can't quite remember when exactly.

00:28:57.764 --> 00:29:00.454
Was it like a couple of years
ago, something like that?

00:29:00.844 --> 00:29:05.394
I think it was actually before, 03 came
out, and I think they were building it

00:29:05.394 --> 00:29:12.154
in a way that would be very difficult
problems from across different math,

00:29:12.154 --> 00:29:16.194
disciplines, but the answer is an
integer, so you have to have like

00:29:16.194 --> 00:29:18.074
one number, so it's easy to check.

00:29:18.584 --> 00:29:21.708
And, at the time when I think
they were building it, it has four

00:29:21.708 --> 00:29:23.978
tiers, so it says tier four, right?

00:29:24.368 --> 00:29:28.808
So it has four tiers of complexity,
and now it looks like if that number

00:29:28.808 --> 00:29:35.308
is true, then it's basically just
scored one or like close to 100% on

00:29:35.618 --> 00:29:37.838
the highest tier of that benchmark.

00:29:37.838 --> 00:29:43.088
And this was mathematicians
sitting down for weeks, working

00:29:43.088 --> 00:29:46.118
super hard, trying to come up with
the hardest questions they can.

00:29:46.428 --> 00:29:46.748
Yeah.

00:29:46.748 --> 00:29:50.098
And, and, yeah, top, top of their field.

00:29:50.108 --> 00:29:50.828
And I d-

00:29:50.828 --> 00:29:53.798
Alex Volkov: Hardest questions they
can, realistically see proven, right?

00:29:53.798 --> 00:29:56.078
There's fields in mathematics,
there's a bunch of hard questions

00:29:56.078 --> 00:29:57.423
that we cannot prove yet.

00:29:57.493 --> 00:29:57.803
Peter Gostev: Yeah.

00:29:57.893 --> 00:29:58.093
yeah.

00:29:58.143 --> 00:29:59.083
Yeah, yeah, totally.

00:29:59.143 --> 00:30:02.883
a- also Epoch has a list of, like,
open questions, open problems there.

00:30:02.883 --> 00:30:06.343
So this is not, there, there were,
I think, a few leaks from other b-

00:30:06.343 --> 00:30:09.553
not leaks, but like few solutions
by Fable, I think, at some point.

00:30:09.963 --> 00:30:13.933
but you can see, yeah, Breakthrough,
Major, none of them been solved yet.

00:30:14.293 --> 00:30:14.363
Mhm.

00:30:14.373 --> 00:30:19.553
So there- there's definitely not to be
confused with all maths being solved,

00:30:19.643 --> 00:30:25.023
but this is a subset of m- problems
that can be de- unknown and defined

00:30:25.023 --> 00:30:31.283
by mathematicians that also can be
defined as a integer as an answer.

00:30:31.643 --> 00:30:32.463
They have been solved.

00:30:32.473 --> 00:30:36.923
So there's a lot of caveats, but
still, the fact that this is probably

00:30:36.923 --> 00:30:40.073
like the hardest questions like
we can sit down and come come up

00:30:40.073 --> 00:30:42.803
with for to for models to solve.

00:30:43.103 --> 00:30:43.293
Yeah.

00:30:43.293 --> 00:30:47.833
They're not equivalent to like new
theorems, new breakthroughs, and

00:30:47.963 --> 00:30:52.573
conceptually n- nothing like that at
all, but still, it's worth pointing

00:30:52.573 --> 00:30:54.313
that that kind of thing is solved.

00:30:54.763 --> 00:30:58.183
Doesn't mean that all of the
conceptual breakthroughs are solved.

00:30:58.253 --> 00:31:03.393
I think this is the the degrees of
difficulty are so much harder than this.

00:31:03.463 --> 00:31:08.013
So it's it's it's not like, oh,
we're we're just like 2 2.4%

00:31:08.093 --> 00:31:08.343
out.

00:31:08.343 --> 00:31:08.847
It's it's not that.

00:31:08.847 --> 00:31:15.383
And on the on on the evals that were
were posted on this the new Stack IO is

00:31:15.413 --> 00:31:18.413
every other model, Fable 1, Fable 5.1,

00:31:18.413 --> 00:31:22.243
Fable 5, Opus 5, the maximum
they're getting is 87%.

00:31:22.613 --> 00:31:27.478
GPT-6 Astra gets 97%, nearly
solving this whole, the…

00:31:27.478 --> 00:31:30.193
LDJ: Now I don't see any link
on their research page, on their

00:31:30.203 --> 00:31:31.053
product page, on their company page.

00:31:31.053 --> 00:31:32.633
Alex Volkov: Yeah, I don't see it's 5.6

00:31:32.633 --> 00:31:34.263
yet, so it looks like we're still waiting.

00:31:34.643 --> 00:31:35.363
if you have a link…

00:31:38.073 --> 00:31:42.483
some OpenAI guy posted the link 2
minutes ago, but the blog post is

00:31:42.493 --> 00:31:44.083
404, regular user might not get it.

00:31:44.303 --> 00:31:49.213
I, I have very strong, sending a
lot of support to OpenAI PR comms

00:31:49.483 --> 00:31:52.193
and folks trying to, delay this.

00:31:52.433 --> 00:31:54.263
folks are saying, how
reliable is this news outlet?

00:31:54.323 --> 00:31:55.343
Image seems terrible.

00:31:55.603 --> 00:31:58.783
LDJ did a bit of research, but looking
at the other stuff, they're reporting

00:31:58.783 --> 00:32:03.743
correctly, like all of the other
scores for Muse, and it's sourced well.

00:32:04.013 --> 00:32:08.303
it doesn't sound like this is like
bullshit, like all of the other links

00:32:08.303 --> 00:32:13.383
here are, and all of the other numbers
and prices, everything is like reliable.

00:32:13.383 --> 00:32:16.923
we're taking this with a grain
of salt, but it does match like

00:32:17.073 --> 00:32:19.343
the other things that folks are
leaking via the information as well.

00:32:23.783 --> 00:32:24.863
Let's read through this.

00:32:25.203 --> 00:32:25.813
What changes in Codex?

00:32:25.813 --> 00:32:26.893
I think it's very important.

00:32:26.903 --> 00:32:30.323
Developers, Astra is handling of
jobs outgrow the context window may

00:32:30.333 --> 00:32:31.903
matter more than the benchmark gains.

00:32:32.193 --> 00:32:34.563
Codex currently relies on
compaction, which summarizes

00:32:34.623 --> 00:32:36.383
earlier work to free up context.

00:32:36.843 --> 00:32:40.833
That process can discard exactly
the detail an agent may need later.

00:32:41.703 --> 00:32:44.063
Why a previous fix failed,
which test run, etc.

00:32:44.093 --> 00:32:48.443
Astra can instead keep notes
across context windows and search

00:32:48.493 --> 00:32:50.163
earlier messages and tool output.

00:32:50.343 --> 00:32:52.023
I feel like GPT 5.6

00:32:52.023 --> 00:32:55.403
could have done this as well,
but maybe not to that level.

00:32:55.633 --> 00:32:58.313
The feature is experimental
behind the config toml setting.

00:32:58.323 --> 00:33:02.063
OpenAI says it will become the
Astra default in coming weeks.

00:33:02.403 --> 00:33:04.893
Astra can also ask the user a
question without stopping work

00:33:05.043 --> 00:33:07.043
that does not depend on an answer.

00:33:07.063 --> 00:33:10.423
This keeps one unresolved decision
from blocking the rest of the job, a

00:33:10.423 --> 00:33:12.383
common failure mode for coding agents.

00:33:12.393 --> 00:33:18.453
OpenAI showed Astra operating application
including KiCad, Excel, Blender,

00:33:18.453 --> 00:33:22.163
and Power BI, as well as performing
browser-based form entry and website QA.

00:33:23.253 --> 00:33:24.953
Oh, there we go, LDJ, there we go.

00:33:24.983 --> 00:33:29.333
OS World V2 offline benchmark, which
tests work across desktop applications.

00:33:29.833 --> 00:33:33.393
OpenAI says Astra scored a 72%, up from

00:33:33.393 --> 00:33:35.073
65% for Sol.

00:33:35.753 --> 00:33:38.733
It also cut the average time
per task from 75 minutes to

00:33:38.733 --> 00:33:41.923
just 40, which is an insane cut.

00:33:43.923 --> 00:33:48.073
Anthropic has reported a higher
77% result for Fable, but says that

00:33:48.133 --> 00:33:50.903
the test used a different OS World
release and should not be compared

00:33:50.903 --> 00:33:51.973
with previously published sources.

00:33:52.033 --> 00:33:52.253
Huh?

00:33:52.723 --> 00:33:54.183
That- who- who says that when?

00:33:54.643 --> 00:33:56.313
I didn't- I didn't miss this on Fable.

00:33:58.853 --> 00:34:00.803
You guys saw this on Fable OS World?

00:34:01.273 --> 00:34:02.023
Let's take a look.

00:34:03.343 --> 00:34:05.683
Oh, OS World 2 and 2.

00:34:06.723 --> 00:34:07.593
Interesting.

00:34:07.733 --> 00:34:12.683
So computer use OS World for Fable was
41%, and then, let's look at the…

00:34:12.733 --> 00:34:14.223
LDJ: Oh, it's partial versus strict.

00:34:14.233 --> 00:34:15.443
See, if you scroll back up.

00:34:16.253 --> 00:34:16.723
Alex Volkov: Oh, you saw it?

00:34:17.473 --> 00:34:20.343
LDJ: Yeah, just, yeah, scroll back
up to exactly where you were, with

00:34:20.493 --> 00:34:22.053
the, the numbers, the percentages.

00:34:22.843 --> 00:34:23.053
Alex Volkov: Here.

00:34:24.263 --> 00:34:25.199
LDJ: scroll up a little bit more.

00:34:25.199 --> 00:34:27.383
Yeah, where it's under right
underneath the percentage, it

00:34:27.383 --> 00:34:28.673
says partial and strict, yeah.

00:34:29.043 --> 00:34:31.153
Alex Volkov: So I gotta wonder
if what they're saying is

00:34:31.783 --> 00:34:34.993
that OpenAI, Astra got 77%?

00:34:37.183 --> 00:34:42.383
Sorry, 70, 72% here at this tier.

00:34:44.063 --> 00:34:44.863
It's very interesting.

00:34:45.443 --> 00:34:47.323
Okay, so some answers that
we need to answer still.

00:34:47.853 --> 00:34:52.913
OpenAI also changed the codecs harness
on Mind to Web, another, web using Astra

00:34:53.393 --> 00:34:56.023
and the new harness completed tasks 1.9%

00:34:56.033 --> 00:34:57.853
faster than the current Sol-based setup.

00:34:59.013 --> 00:34:59.953
Ooh.

00:35:00.683 --> 00:35:03.223
I feel like this is a big, a big thing.

00:35:03.233 --> 00:35:05.393
Wolfen, what, what are you,
like, LDJ, what are you guys

00:35:05.673 --> 00:35:06.793
getting from this computer use?

00:35:06.823 --> 00:35:09.193
I definitely feel like computer
use is like a big, big unlock here.

00:35:09.613 --> 00:35:11.923
Wolfram Ravenwolf: Yeah, it's also a way
to say, get more money because if it's

00:35:11.933 --> 00:35:16.543
just using an API call, maybe using less
tokens than the usual computer use stuff.

00:35:16.813 --> 00:35:20.313
But, more computer use is important
because not everything has an API, of

00:35:20.313 --> 00:35:22.513
course, and I'm using it all the time.

00:35:22.713 --> 00:35:27.013
And, like I said, I thought this was just
a planning model because it's so smart

00:35:27.053 --> 00:35:30.893
that you only use it to make a plan,
but if it's an all-rounder that you use

00:35:30.893 --> 00:35:33.723
to control your computer, wow, good.

00:35:35.053 --> 00:35:36.453
LDJ: And, GPT 5.6

00:35:36.463 --> 00:35:38.773
Soul scored 76.9%,

00:35:39.613 --> 00:35:44.363
and GPT 6 Astro scores 92.7%.

00:35:44.363 --> 00:35:45.383
Alex Volkov: On computer use.

00:35:45.383 --> 00:35:48.713
LDJ: Yes, it's ScreenSpot Pro
with no tool use, so it seems

00:35:48.713 --> 00:35:51.713
like pure computer use, moving the
mouse and keyboard, essentially.

00:35:53.953 --> 00:35:54.293
Yep,

00:35:54.383 --> 00:35:58.483
Alex Volkov: it looks like, we're going
to have another banger of a computer

00:35:58.483 --> 00:36:01.923
use on our hands with GPT-6 Astra.

00:36:02.583 --> 00:36:03.903
I can't wait to play with this.

00:36:03.903 --> 00:36:05.293
LDJ: And we have a confirmed price.

00:36:05.453 --> 00:36:08.563
This is in the Newstack article,
but also I have it confirmed

00:36:08.563 --> 00:36:09.813
from another source too now.

00:36:10.103 --> 00:36:10.393
Yeah.

00:36:10.393 --> 00:36:13.893
It's 10 dollars per million input
tokens and 50 dollars per million

00:36:13.893 --> 00:36:17.333
output tokens, which I think that's
identical to the Fable pricing, right?

00:36:17.373 --> 00:36:18.013
Alex Volkov: Fable, yeah.

00:36:18.023 --> 00:36:21.813
So we got the Fable class model and
Fable class pricing from OpenAI.

00:36:23.213 --> 00:36:27.773
and in the blog post, a very quick
turnaround from the folks at OpenAI.

00:36:28.543 --> 00:36:29.723
Fable 5.1

00:36:29.733 --> 00:36:33.023
is, mentioned, and so they
do compare it to Fable 5.1,

00:36:33.043 --> 00:36:34.843
which came out, what, 4 days ago.

00:36:35.203 --> 00:36:36.133
this beats Fable

00:36:39.363 --> 00:36:40.150
on Terminal Bench 4.0.

00:36:40.150 --> 00:36:40.673
57% on Terminal Bench 4.0.

00:36:40.673 --> 00:36:47.873
This is the absolute state-of-the-art
behemoth thing, which is absolutely crazy.

00:36:48.553 --> 00:36:49.533
Absolutely crazy.

00:36:49.783 --> 00:36:52.783
DeepSui, this is the top
model that gets DeepSui score.

00:36:56.543 --> 00:36:57.153
what else?

00:36:58.113 --> 00:36:59.733
what else do we see
that's interesting here?

00:37:00.853 --> 00:37:03.443
Artificial Analysis Coding Agent Index.

00:37:04.643 --> 00:37:07.673
This is a little bit below Fable at 67%.

00:37:07.673 --> 00:37:12.573
So it didn't quite crack the artificial
analysis, like, up and to the right thing.

00:37:13.203 --> 00:37:13.573
yep.

00:37:13.573 --> 00:37:15.563
I gotta wonder, folks in
comments, are you still with us?

00:37:15.573 --> 00:37:18.723
If this is interesting for you to
see as is, we're just like talking.

00:37:18.773 --> 00:37:22.383
Obviously, I don't want to show, stuff
that OpenAI doesn't want us to show,

00:37:22.613 --> 00:37:23.983
because, we're friends with OpenAI.

00:37:24.043 --> 00:37:25.443
We want them to invite us to Dev Day.

00:37:25.473 --> 00:37:28.823
We don't want to leak like some of
the stuff, but reporting on some

00:37:28.823 --> 00:37:31.353
other reporting, I think, is a very
standard practice in journalism.

00:37:31.353 --> 00:37:35.403
So if somebody else is in NDA, you
can say as reported in the news stack.

00:37:35.803 --> 00:37:39.633
not that crazy given the model
that's not quite Astra, right?

00:37:39.633 --> 00:37:43.813
Like OpenAI legit said not the
Astra model, but a different

00:37:43.853 --> 00:37:47.973
model, a highly persistent internal
model hacked and collaborated.

00:37:48.203 --> 00:37:50.843
but given the fact that it was able
to like hack away from the sandbox

00:37:50.843 --> 00:37:53.883
and go to another website, I think,
exploit bench 100% makes sense.

00:37:54.283 --> 00:37:58.223
But still, like, don't can't we come
up with some other exploit benches

00:37:58.223 --> 00:38:00.493
that are not completely saturated.

00:38:01.213 --> 00:38:03.733
I would love if that would be the case.

00:38:04.403 --> 00:38:07.633
so alignment, I think, is very important
as well, as we're reading through this.

00:38:07.943 --> 00:38:10.633
internal computer use safety
benchmark, lower is better.

00:38:10.923 --> 00:38:13.103
OpenAI tested Fable 5.1

00:38:13.403 --> 00:38:15.473
and GPT-6 Astra.

00:38:15.473 --> 00:38:16.403
Fable 5.1

00:38:16.403 --> 00:38:17.733
actually did work in 9.5,

00:38:17.733 --> 00:38:18.203
so half.

00:38:18.533 --> 00:38:19.483
GPT-5.6

00:38:19.493 --> 00:38:24.513
Soul was 22%, and GPT-6
Astra is only 2% refusals.

00:38:26.133 --> 00:38:29.233
I don't know quite what they
measure, but the lower is better.

00:38:29.283 --> 00:38:33.703
It's significantly harder to get
this model to do unaligned stuff.

00:38:33.703 --> 00:38:38.643
And computer use safety benchmark, LDJ,
correct me if I'm wrong, but, this is

00:38:38.643 --> 00:38:43.793
where, while the model uses your computer,
whether or not it can take actions to

00:38:43.793 --> 00:38:47.193
like break or do some stuff nefarious
to you, it looks like they're really,

00:38:47.193 --> 00:38:48.743
really strongly focusing on, on that.

00:38:48.983 --> 00:38:51.803
LDJ: Yeah, that's what I would guess,
but I- I'm honestly not familiar

00:38:51.803 --> 00:38:53.103
with that benchmark much, though.

00:38:53.103 --> 00:38:53.383
Alex Volkov: Yeah.

00:38:55.153 --> 00:39:00.313
I gotta wonder if, if they have links to
those internal computer safety benchmarks.

00:39:00.553 --> 00:39:00.893
Let's see.

00:39:05.413 --> 00:39:08.233
Guest: Let's see if the internet
has links to that or if OpenAI

00:39:08.233 --> 00:39:09.583
posted anything about this.

00:39:10.063 --> 00:39:13.783
The rate of unsafe execution traces
or compliance failures when an

00:39:13.783 --> 00:39:16.373
AI agent interacts with a desktop
or operating system environment.

00:39:17.533 --> 00:39:20.383
Yeah, so what it measures, I
can actually show this, I don't

00:39:20.383 --> 00:39:21.193
have to show the other one.

00:39:21.793 --> 00:39:26.523
what it measures is, and I'm trusting
AI overview here from Gemini, trust it.

00:39:26.523 --> 00:39:29.273
I think I trust it less than
the article that we just shared.

00:39:30.183 --> 00:39:34.783
Unsafe execution rate, the frequency with
which an AI agent carries out harmful,

00:39:34.783 --> 00:39:37.903
restricted, or destructive commands
during desktop and browser tasks,

00:39:38.213 --> 00:39:40.793
Alex Volkov: like the thing that
people say, hey, this deleted my

00:39:40.793 --> 00:39:42.453
environment, this like clicked the thing.

00:39:42.823 --> 00:39:46.933
Prompt ejection vulnerability, how
easy an AI agent can be tricked or

00:39:46.933 --> 00:39:50.623
hijacked by malicious instructions
hidden on web pages, files, or emails.

00:39:50.913 --> 00:39:53.763
You guys remember there was like a whole
thing with computer use in the beginning

00:39:53.763 --> 00:39:57.793
where like, hey, if you use computer
use and you browse the internet and

00:39:57.793 --> 00:40:00.883
somebody on their website has a hidden
instruction that you don't see as a human

00:40:00.883 --> 00:40:03.343
but only the AI sees, you will be screwed.

00:40:03.543 --> 00:40:04.183
You remember that?

00:40:04.403 --> 00:40:09.763
That is what this thing measures, and
here is the stats that, I don't show

00:40:09.763 --> 00:40:12.473
you, but they're cited, verified.

00:40:12.473 --> 00:40:13.463
GPT-5.6

00:40:13.473 --> 00:40:18.173
Sol had 22% on this benchmark,
internal computer use safety benchmark.

00:40:18.173 --> 00:40:19.753
GPT-6 Astra has 2%.

00:40:21.013 --> 00:40:25.233
So this is a very safe model to use,
given the way that they announced this

00:40:25.233 --> 00:40:28.183
and they say, hey, this model uses
your browser, your computer, et cetera,

00:40:28.383 --> 00:40:32.663
and continuously, this seems like a
very important score to report on.

00:40:34.503 --> 00:40:36.273
Exploit Gym Honeypot.

00:40:37.723 --> 00:40:39.603
Wow, this is an interesting result here.

00:40:39.953 --> 00:40:41.913
on exploit, let me see what
Exploit Gym Honeypot is.

00:40:41.913 --> 00:40:41.953
Also lower is better.

00:40:41.953 --> 00:40:44.853
I'm assuming that it's like, What does
Exploit Gym Honey, eval mean from OpenAI?

00:40:45.943 --> 00:40:46.603
Let's see.

00:40:47.873 --> 00:40:48.513
Let's see here.

00:40:50.703 --> 00:40:52.133
Exploit Gym Honeypot lower is better.

00:40:52.173 --> 00:40:56.243
Evaluation measures whether an agent
would unauthorizedly attack surrounding

00:40:56.243 --> 00:41:01.053
security infrastructure when it gets
stuck on a cybersecurity task, which

00:41:01.063 --> 00:41:02.623
was what happened in the swarm.

00:41:03.133 --> 00:41:05.843
A lower score is better because a
score of zero means the AI followed

00:41:05.843 --> 00:41:08.913
safety protocols and made zero
attempts to hack outside the sandbox.

00:41:10.423 --> 00:41:12.533
the specific metric was created
by OpenAI following a major

00:41:12.533 --> 00:41:14.003
safety incident in July 2026.

00:41:14.003 --> 00:41:14.893
Oh, this is the new one.

00:41:15.693 --> 00:41:20.353
Okay, the incident, we know the incident,
but the stats that now I will not show

00:41:20.353 --> 00:41:23.913
you because we're trying is, GPT-5.6

00:41:24.583 --> 00:41:27.773
Soul was getting 48.2%

00:41:29.773 --> 00:41:34.453
on this exploit gym honeypot,
and GPT-6 Astro gets zero.

00:41:35.523 --> 00:41:41.053
So it's like, it's like they trained
their model to not do the thing, and

00:41:41.053 --> 00:41:45.293
then they used all of the 20% safety
stuff that they promised us to actually

00:41:45.293 --> 00:41:48.303
train the model, to, to not hack.

00:41:49.433 --> 00:41:49.853
And, Oh,

00:41:49.903 --> 00:41:53.893
Guest: maybe to point out that it's
of what they managed to detect.

00:41:54.123 --> 00:41:56.583
it could be 0% detected, we should say.

00:41:56.583 --> 00:41:59.243
Alex Volkov: No, but I agree with
you, but like one would hope that

00:41:59.243 --> 00:42:03.063
OpenAI learned from that instance,
as they said, they take safety really

00:42:03.063 --> 00:42:07.933
strongly, and they measured and built
another like, that whole incident

00:42:07.933 --> 00:42:12.633
is like a god gold mine for evals of
like how to align models, et cetera.

00:42:12.633 --> 00:42:18.373
Like they can build evals like, hey, we
saw this, if you meet a message board

00:42:18.373 --> 00:42:23.283
full of exploit things, will you join it
as an eval and then test on this eval,

00:42:23.313 --> 00:42:26.263
which I think it did because there's
another one here that says impossible

00:42:26.263 --> 00:42:29.973
exploit gym, which is how the whole thing
started in the first place with them.

00:42:29.983 --> 00:42:34.573
They gave their things, impossible tasks,
and, this is the only model they show us

00:42:34.573 --> 00:42:37.733
a score for, and they say it scored 100%.

00:42:38.303 --> 00:42:40.743
Oh, that's interesting as well, folks.

00:42:40.973 --> 00:42:44.393
Okay, long context, which is
something we just reported from

00:42:44.393 --> 00:42:52.643
Use, OpenAI MRCR for 256 to half a
million, Astra gets 100% score, and

00:42:53.203 --> 00:42:56.803
the MRCR on 512K to 1 million gets

00:42:56.813 --> 00:42:58.063
96.3.

00:42:58.063 --> 00:43:01.303
We just talked about
Muse Spark 3, Muse Spark.

00:43:01.323 --> 00:43:05.673
Let me see if I can pull up the
MRCR for Muse, and we'll see if

00:43:05.673 --> 00:43:07.183
that beats it, and I think it does.

00:43:10.913 --> 00:43:12.413
Let's see, Muse Spark.

00:43:16.893 --> 00:43:18.653
We have this here, we
have this in our notes.

00:43:18.653 --> 00:43:20.163
We, this has been a long show, folks.

00:43:20.343 --> 00:43:24.233
This is, we're clocking at 4 hours
on the stream, but we said we're

00:43:24.233 --> 00:43:28.513
here until Astro gets released,
and that's what we're waiting for.

00:43:28.513 --> 00:43:32.153
But let's take a look at MRCR 98, 98.

00:43:32.543 --> 00:43:36.003
yeah, okay, so that's
what we have for Mu Spark.

00:43:36.303 --> 00:43:42.423
MRCR, this is 98, GPT Astra gets
100% on this one, and this one, 98.1,

00:43:42.423 --> 00:43:50.643
Astra gets a, Astra gets, let's
see how much Astra gets, 96.3.

00:43:50.853 --> 00:43:52.373
So a little bit below Mu Spark.

00:43:55.203 --> 00:43:58.773
I think that, obviously OpenAI, LDJ,
I don't know if you feel the same,

00:43:58.773 --> 00:44:03.603
Wolfram, but like obviously OpenAI is,
is gaining a lot with this release,

00:44:04.343 --> 00:44:11.023
but very interestingly, the folks
who are getting the best scores, not

00:44:11.023 --> 00:44:14.913
the best scores, but like the folks
who are standing their ground is Meta

00:44:14.913 --> 00:44:18.793
Mu Spark with the Spark model, and
I think it says a lot about their

00:44:18.793 --> 00:44:24.373
upcoming watermelon or whatever they're
hyping up, which is, which is crazy.

00:44:24.763 --> 00:44:26.383
Folks are asking, LG, what do you think?

00:44:26.423 --> 00:44:29.393
Like, like it looks like Meta Muse
Spark, although not mentioned.

00:44:29.453 --> 00:44:35.633
By OpenAI, snubbed by OpenAI from the
eval, thing, while they do add Gemini 3.8

00:44:35.643 --> 00:44:38.133
Flash in here, but they don't d- Mm.

00:44:38.133 --> 00:44:41.763
They don't add Grok, they don't
compare themselves to Meta Muse.

00:44:41.763 --> 00:44:44.313
Meta Muse Spark is actually
coming out very well

00:44:44.313 --> 00:44:46.123
comparatively with the small one.

00:44:46.413 --> 00:44:50.983
Guest: Yeah, so to be fair, Muse Spark,
it's so recent that they might have

00:44:50.983 --> 00:44:53.513
not had time to put it in benchmarks,
but then again, actually, Fable

00:44:53.513 --> 00:44:56.343
just also came out like 2, 3, 5.1

00:44:56.343 --> 00:44:57.223
came out 2 or 3 days ago.

00:44:57.223 --> 00:44:57.783
Yeah.

00:44:57.783 --> 00:45:01.293
But, yeah, what's really interesting to
note here, though, is despite the fact

00:45:01.293 --> 00:45:05.153
that it is significantly increased,
API pricing, so 10 dollars per million

00:45:05.193 --> 00:45:10.373
input, 50 dollars per million output,
it's still overall, in terms of cost per

00:45:10.373 --> 00:45:13.073
task, for a lot of these benchmarks that
I'm seeing and a lot of the information

00:45:13.243 --> 00:45:17.783
I'm seeing, even the low reasoning
version of Astra seems to be, in most

00:45:17.793 --> 00:45:26.093
cases, higher accuracy than Solmax while
being lower cost per task than Solmax.

00:45:26.583 --> 00:45:26.663
Alex Volkov: Yep.

00:45:26.833 --> 00:45:31.193
Guest: in other words, if you want to
try and like get the closest to matching

00:45:31.193 --> 00:45:37.043
the accuracy of Soul Max, Astro here
is actually a better cost per task for

00:45:40.823 --> 00:45:40.843
you.

00:45:40.843 --> 00:45:40.943
Alex Volkov: Yep.

00:45:41.233 --> 00:45:42.733
folks are asking about availability.

00:45:43.083 --> 00:45:44.193
this is from Twitter.

00:45:44.213 --> 00:45:47.063
let me add this here so that
I can show you on screen.

00:45:47.073 --> 00:45:49.003
It looks like it won't be available today.

00:45:49.623 --> 00:45:52.093
it looks like this is what they say.

00:45:52.313 --> 00:45:56.143
this is from, from Twitter, so
I feel like we can show this.

00:45:56.343 --> 00:45:59.103
GPT Astro is rolling out today
in limited set of organizations.

00:45:59.403 --> 00:46:02.773
that's pretty much what the
News Tech article posted.

00:46:03.073 --> 00:46:06.113
Astro supports zero data retention
for eligible API customers.

00:46:06.513 --> 00:46:09.643
we're testing private safety processing
to strengthen safety monitoring.

00:46:09.673 --> 00:46:13.863
For developers, Astro is available Open
API, and is available in Amazon Bedrock.

00:46:13.933 --> 00:46:15.843
Open API standard pricing 10 million.

00:46:15.843 --> 00:46:16.633
We talked about this.

00:46:16.983 --> 00:46:19.453
Separate rates apply to
cache reads and writes.

00:46:19.493 --> 00:46:20.613
They don't say which.

00:46:20.933 --> 00:46:25.403
Fast mode is available for GPT-6
in the API, delivers up to 2.5

00:46:25.413 --> 00:46:27.863
speed at 2x the standard price.

00:46:27.893 --> 00:46:29.893
So you'd be able to run
this as fast as well.

00:46:30.973 --> 00:46:34.433
coming days will become available
ChatGPT Plus with the internet downtones.

00:46:34.463 --> 00:46:35.933
So hopefully we answered this question.

00:46:36.353 --> 00:46:37.853
sorry, this question is the rumors.

00:46:37.893 --> 00:46:38.683
They're not rumors.

00:46:38.703 --> 00:46:40.843
It looks like that's what happens here.

00:46:43.723 --> 00:46:45.023
So long context is very interesting.

00:46:45.043 --> 00:46:47.353
R- RKGI with the custom harness.

00:46:47.393 --> 00:46:49.083
That's what I think we see.

00:46:49.473 --> 00:46:51.893
because we saw this with RKGI before.

00:46:51.893 --> 00:46:54.723
You guys remember when, like, it
was a very low score for GPT-5.6

00:46:54.733 --> 00:46:59.673
SOL, and then, OpenAI's folks
said that, hey, the RKGI uses our,

00:47:00.073 --> 00:47:01.643
completions API, not responses API.

00:47:01.643 --> 00:47:05.153
Responses API is the new one that
has reasoning tokens as well,

00:47:05.413 --> 00:47:08.343
and so that's why it's not like
an apples to apples comparison.

00:47:08.623 --> 00:47:08.983
but…

00:47:11.113 --> 00:47:15.873
Yeah, the interestingly, OpenAI here
cites the previous score for GPT-5.6

00:47:15.873 --> 00:47:20.138
SOL, which was 7% on RKGI 3, and now 99.9%

00:47:20.138 --> 00:47:20.278
on RKGI 3 for Astra, 95 on
RKGI 2, and 98 on RKGI 1.

00:47:20.278 --> 00:47:20.483
At this point, if

00:47:29.153 --> 00:47:34.863
RKGI is the measure by which we say AGI
is here, it does look like AGI is here.

00:47:36.313 --> 00:47:37.523
LG, what do you send?

00:47:39.023 --> 00:47:42.913
LDJ: yeah, it's, still not
public, but basically the…

00:47:43.133 --> 00:47:43.623
Alex Volkov: The same one?

00:47:44.393 --> 00:47:45.123
LDJ: Yeah, yeah.

00:47:45.153 --> 00:47:45.183
Yeah.

00:47:45.183 --> 00:47:47.603
but yeah, so terminal bench,
though, which I believe the News

00:47:47.623 --> 00:47:49.653
Stack, had reported here too.

00:47:49.893 --> 00:47:50.223
Mhm.

00:47:50.333 --> 00:47:53.593
it's a pretty significant jump
there, where it's a full…

00:47:53.593 --> 00:47:55.383
Here, let me just pop
it up for myself again.

00:47:55.813 --> 00:47:57.393
Okay, so 5.6

00:47:57.533 --> 00:48:01.973
Solmax scores 22% in terminal bench
science, which this is basically agentic,

00:48:02.023 --> 00:48:06.083
I think one of the best agentic science
benchmarks out right now, really, and

00:48:06.093 --> 00:48:07.743
the most unsaturated, or one of the most.

00:48:08.293 --> 00:48:08.673
Fable 5.1,

00:48:10.123 --> 00:48:13.183
it scores 52.6%,

00:48:13.223 --> 00:48:17.333
so it goes from Sol's 22%
to basically Fable 5.1's

00:48:17.363 --> 00:48:19.373
52%, which is a huge jump.

00:48:19.723 --> 00:48:20.033
Yeah.

00:48:20.143 --> 00:48:24.503
And Astro further increases
that all the way to 64.6%.

00:48:25.713 --> 00:48:27.293
Alex Volkov: Which is state of the art and

00:48:27.343 --> 00:48:27.703
LDJ: Yes.

00:48:28.533 --> 00:48:28.853
Yeah.

00:48:29.293 --> 00:48:32.993
Alex Volkov: It looks like, based on this
leak chart again, like state of the art

00:48:33.033 --> 00:48:39.243
model across all of it, and OpenAI did
promise us a while ago that in September

00:48:39.263 --> 00:48:45.143
we'll get to a point where they have a
automated researcher model that's at the

00:48:45.143 --> 00:48:49.403
level of like a junior person, and it
looks like they released it with Astra.

00:48:50.023 --> 00:48:55.123
LDJ: Yeah, they did actually say
they are running now over across more

00:48:55.123 --> 00:49:00.233
than 100,000 GPUs, their intern level
model, which they implied as Astra.

00:49:00.613 --> 00:49:03.553
It was a recent interview
that Ya- Jakob Pachoki did.

00:49:04.193 --> 00:49:06.133
Alex Volkov: Yeah, implied as Astra.

00:49:06.133 --> 00:49:08.813
Yeah, they've been talking about
Astra, and companies did have access to

00:49:09.023 --> 00:49:11.373
Astra before for some evals, but yeah.

00:49:13.003 --> 00:49:15.133
All right, let's see what else.

00:49:16.213 --> 00:49:18.963
Oh yeah, the we have to talk
about the harder to monitor thing.

00:49:19.113 --> 00:49:22.053
OpenAI claimed that Astra is the
most aligned model, rest partly on

00:49:22.053 --> 00:49:25.703
internal tests, which Astra went
outside an authorized target in 0%

00:49:25.703 --> 00:49:26.873
of your possible task scenarios.

00:49:27.203 --> 00:49:30.333
That's what we told you, compared to 48.2%

00:49:30.343 --> 00:49:30.873
for Sol.

00:49:31.283 --> 00:49:34.353
OpenAI describes the older model as
running without production safeguards,

00:49:34.673 --> 00:49:36.893
however, does not make the role
of the surrounding safety setup

00:49:36.893 --> 00:49:38.203
clear enough for direct comparison.

00:49:38.563 --> 00:49:42.093
this kind of, what they're talking
about is in the incident for Hugging

00:49:42.093 --> 00:49:47.253
Face hacking, they supposedly
re- removed the, the, the safety

00:49:47.253 --> 00:49:49.083
mechanisms around the GPT 5.6

00:49:49.123 --> 00:49:50.793
or whatever internal model there was.

00:49:51.283 --> 00:49:55.543
and and it looks like on impossible
test scenarios they trained Astra,

00:49:55.843 --> 00:49:59.953
and Astra would not go outside of
an authorized target, even when

00:50:00.003 --> 00:50:04.943
tasked with an impossible task in
any of the scenarios, which is great.

00:50:04.943 --> 00:50:07.853
And also, this is the recur- re- recurrent

00:50:09.803 --> 00:50:11.503
transformer or loop transformer thing.

00:50:11.893 --> 00:50:14.963
this is a, OpenAI describes
the older models run- sorry.

00:50:15.143 --> 00:50:18.523
The company also disclosed that Astra's
written reasoning was harder to monitor

00:50:18.523 --> 00:50:22.833
than Sol in evaluation specifically
designed to elicit monitoring evasion.

00:50:23.793 --> 00:50:26.923
OpenAI attributes the decline partly
to Astra having greater control

00:50:26.923 --> 00:50:30.143
over its written reasoning on simple
tasks and completing problems with

00:50:30.143 --> 00:50:31.363
fewer written reasoning steps.

00:50:31.383 --> 00:50:35.513
This could be potentially the
re- response or the result of the

00:50:35.523 --> 00:50:37.873
deeper, like, computational stuff.

00:50:39.093 --> 00:50:39.843
and Jacob said…

00:50:39.853 --> 00:50:39.863
Oh,

00:50:41.453 --> 00:50:42.053
LDJ: sorry, go ahead.

00:50:42.053 --> 00:50:42.463
Alex Volkov: Yeah.

00:50:42.463 --> 00:50:45.063
Jacob said, processing intelligence
does not guarantee processing

00:50:45.073 --> 00:50:48.633
alignment without scaling until
we can regain enough confidence.

00:50:48.703 --> 00:50:51.903
we'll withhold scaling until we
can regain enough confidence in its

00:50:51.903 --> 00:50:53.563
ability to monitor future models.

00:50:54.083 --> 00:50:54.593
Oh, wow.

00:50:55.123 --> 00:50:57.073
That's a big one.

00:50:57.073 --> 00:50:57.133
Yeah, that is a big one.

00:50:57.133 --> 00:50:57.743
And also…

00:50:57.743 --> 00:50:58.863
Oh, let's go.

00:50:58.913 --> 00:50:59.623
Look who's here.

00:51:01.003 --> 00:51:02.453
Ryan Carson.

00:51:03.453 --> 00:51:08.313
You had to come on when AGI was announced,
and I really, really appreciate it.

00:51:08.323 --> 00:51:08.963
We missed you.

00:51:09.373 --> 00:51:12.356
Guest: I know, I missed you
guys, and I was like, oh my gosh,

00:51:12.356 --> 00:51:13.396
I have a hole in my calendar.

00:51:13.396 --> 00:51:15.016
I can finally join my friends again.

00:51:15.016 --> 00:51:15.946
So good to see you all.

00:51:15.946 --> 00:51:18.596
Alex Volkov: Yeah, welcome, We've been
reporting on the rise of singularity

00:51:19.046 --> 00:51:23.116
throughout all this time, dude, and
today is a very interesting day because

00:51:23.136 --> 00:51:27.686
the release got yanked, like it was
very clear based on the releases of

00:51:27.686 --> 00:51:32.266
different news outlets when it expired,
the embargo, and then some released it

00:51:32.286 --> 00:51:37.216
and OpenAI was managed to get some back.

00:51:37.716 --> 00:51:38.536
Guest: Pull it back.

00:51:38.536 --> 00:51:39.416
Alex Volkov: Pull it back.

00:51:40.026 --> 00:51:40.946
tell us about your stack.

00:51:40.946 --> 00:51:42.866
Do tell us about thoughts on Fable 5.1,

00:51:42.876 --> 00:51:43.956
if you had played with it.

00:51:43.956 --> 00:51:44.206
Oof,

00:51:44.346 --> 00:51:45.016
Ryan Carson: wow.

00:51:45.146 --> 00:51:46.086
Fable 5.1,

00:51:46.086 --> 00:51:48.226
very exciting, very, very good model.

00:51:48.276 --> 00:51:53.166
I think everyone's weirding out about
its, its usage and its token consumption.

00:51:53.166 --> 00:51:57.596
It seems to be blowing out people's,
limits fast, so that's a little weird.

00:51:57.816 --> 00:52:01.416
but so far I've enjoyed it and
like it, and it's my go-to model.

00:52:01.876 --> 00:52:07.086
And then obviously, we have, some
things going on with OpenAI, and I

00:52:07.086 --> 00:52:11.136
am not allowed to say some things
I know, like you guys, I'm sure.

00:52:11.236 --> 00:52:13.956
w- waiting until OpenAI does something.

00:52:14.516 --> 00:52:20.496
Alex Volkov: Yes, we- we- I said at
the beginning of this morning that we

00:52:20.506 --> 00:52:23.576
will sit here until something happens.

00:52:24.026 --> 00:52:28.916
That something is unclear because it
looks like, based on the screenshots

00:52:28.916 --> 00:52:31.426
and the things we're getting, that
like, pro accounts, even pro accounts

00:52:31.426 --> 00:52:34.496
will not get access to, GPT, 6 Astra.

00:52:34.506 --> 00:52:36.846
I need to get used to saying GPT 6 Astra.

00:52:36.846 --> 00:52:39.286
but we'll see, we'll see.

00:52:39.886 --> 00:52:41.456
I think this is officially finally here.

00:52:41.626 --> 00:52:42.146
let's see.

00:52:43.476 --> 00:52:45.066
OpenAI.com/index/GPTAstra.

00:52:46.096 --> 00:52:48.276
The vlog is live, let's go.

00:52:49.026 --> 00:52:49.446
Wolfram Ravenwolf: Yes!

00:52:49.836 --> 00:52:50.876
Woo woo!

00:52:50.936 --> 00:52:54.826
Alex Volkov: Official from
OpenAI on the website.

00:52:55.006 --> 00:52:57.806
Let's refresh once more to see
that they're not pulling it.

00:52:58.936 --> 00:53:00.096
No, it refreshes.

00:53:00.406 --> 00:53:01.896
Oh, beautiful open animation.

00:53:01.896 --> 00:53:02.786
Guest: There we go.

00:53:02.916 --> 00:53:06.916
There's a video at the top
that we should watch first.

00:53:06.916 --> 00:53:07.256
Alex Volkov: Let's just play.

00:53:07.256 --> 00:53:09.366
Okay, so we will play this with sound.

00:53:09.366 --> 00:53:11.566
Folks, this may be loud, I apologize.

00:53:12.156 --> 00:53:16.446
this may be loud, so if you'll excuse
me, we're gonna pull up the video and

00:53:16.456 --> 00:53:19.026
and then we will play it right here.

00:53:20.396 --> 00:53:21.566
Yoink, share with audio.

00:53:22.386 --> 00:53:26.996
Okay, let's take a look and then
let's discuss what's going on here.

00:53:29.766 --> 00:53:32.216
Guest: Create a yellow circle there.

00:53:35.276 --> 00:53:35.416
Can

00:53:41.016 --> 00:53:43.046
you draw me a small yellow circle?

00:53:44.236 --> 00:53:44.526
Done.

00:53:44.956 --> 00:53:47.776
Okay, take this and make it
the window of a rocket ship.

00:53:53.236 --> 00:53:55.246
I like this, but can you
make it a lot more detailed?

00:53:55.556 --> 00:53:58.216
Your yellow circle is now
the window on a rocket.

00:53:58.476 --> 00:53:59.336
Okay, this is awesome.

00:53:59.496 --> 00:54:01.396
Now make it a 3D model in Blender.

00:54:01.746 --> 00:54:02.496
Opening Blender.

00:54:09.586 --> 00:54:15.066
Let's build a presentation for next
season's rainwear for retailers.

00:54:15.066 --> 00:54:19.436
Make sure that it feels really
high-end and that it's colorful.

00:54:19.616 --> 00:54:20.406
Make it fun.

00:54:20.676 --> 00:54:21.516
I can help with that.

00:54:22.886 --> 00:54:25.566
Can you go to eBay and make this
listing of this table I bought a

00:54:25.566 --> 00:54:26.796
few years ago at a flea market?

00:54:26.886 --> 00:54:29.386
It's this wild orange table.

00:54:29.766 --> 00:54:30.126
Sure.

00:54:31.546 --> 00:54:32.606
Okay, yeah, this is awesome.

00:54:32.856 --> 00:54:36.916
I want you to make a 3D game where I'm
ducking asteroids, using the arrow keys to

00:54:36.916 --> 00:54:38.996
move around, and I'm using space to boost.

00:54:39.136 --> 00:54:40.606
Yep, I'm building the game.

00:54:40.806 --> 00:54:41.696
Also, I'm a little hungry.

00:54:41.696 --> 00:54:44.356
Can you get me some beef and rice from
that spot I ordered from last week?

00:54:44.636 --> 00:54:46.026
Looking into ordering food.

00:54:46.196 --> 00:54:49.326
So my law firm needs a
licensing agreement template.

00:54:49.376 --> 00:54:53.296
Can you generate a draft template for
the lawyers at my firm to have a look at?

00:54:53.296 --> 00:54:54.086
Sure thing.

00:54:55.776 --> 00:54:59.046
While you're doing that, I want to play
tennis this afternoon, so can you look

00:54:59.046 --> 00:55:00.996
for a court for me in the lower hate?

00:55:01.516 --> 00:55:02.326
Checking now.

00:55:02.396 --> 00:55:03.516
I'll see what I can find.

00:55:03.986 --> 00:55:07.286
Can you include the photo that I
have of it in my downloads folder?

00:55:09.126 --> 00:55:10.306
There's like a slight dent in it.

00:55:10.836 --> 00:55:13.296
it's also saved in my downloads folder.

00:55:13.436 --> 00:55:15.996
Can you put in the description
that it's just slightly damaged?

00:55:16.886 --> 00:55:18.806
Can you just take the limitation of

00:55:18.806 --> 00:55:22.676
liability provision, make it a little
more favorable to the licensor?

00:55:25.656 --> 00:55:30.216
Okay, I've tightened it so the licensor's
liability is more narrowly capped.

00:55:30.516 --> 00:55:31.226
That looks pretty good.

00:55:31.636 --> 00:55:32.036
Thanks.

00:55:36.416 --> 00:55:40.076
Can you change the background color
to complement the rain jacket?

00:55:41.106 --> 00:55:41.886
Oh, I love it.

00:55:45.576 --> 00:55:47.136
And what's going on with my reservation?

00:55:47.276 --> 00:55:49.166
I found an open court at 5:00 PM.

00:55:49.526 --> 00:55:50.226
Yeah, book it.

00:55:50.596 --> 00:55:53.616
Now I want you to make a file
I can send to my 3D printer.

00:55:53.906 --> 00:55:56.226
I'll get working on
creating an STL file of

00:56:06.116 --> 00:56:06.196
this rocket.

00:56:06.196 --> 00:56:07.386
All right, all right.

00:56:07.396 --> 00:56:09.246
Alex Volkov: Very, very
nice, very inspiring.

00:56:09.796 --> 00:56:14.716
the only thing I'll say here is, imagine
seeing this video 3 years ago when GPT-4

00:56:14.736 --> 00:56:18.236
launches, and you're like, oh, the fuck.

00:56:18.236 --> 00:56:18.546
Yeah.

00:56:18.546 --> 00:56:21.236
Because when GPT-4 launched,
there was no tool use.

00:56:21.716 --> 00:56:24.276
There was no, obviously, browser use.

00:56:24.496 --> 00:56:25.496
Nobody was thinking about this.

00:56:25.506 --> 00:56:26.226
It was like awful.

00:56:27.626 --> 00:56:28.636
and there's no voice.

00:56:28.706 --> 00:56:31.106
GPT-4 was multimodal, but not voice.

00:56:31.456 --> 00:56:36.626
So now we have, you talk to a computer
and it does a bunch of crazy shit.

00:56:37.136 --> 00:56:38.066
wow, this was nice.

00:56:38.136 --> 00:56:42.106
This is nice to, to reminisce about where
we were before and where we are now.

00:56:42.866 --> 00:56:45.586
OpenAI says we're introducing
GPT-6 Astra, the world's most

00:56:45.736 --> 00:56:46.806
intelligent and aligned model.

00:56:48.606 --> 00:56:49.116
Ryan Carson: Wow.

00:56:50.406 --> 00:56:52.066
Alex Volkov: Reactions, folks,
just like gut reactions to

00:56:52.246 --> 00:56:53.876
it, like what we're seeing.

00:56:53.876 --> 00:56:55.646
Wolfram Ravenwolf: This reminds me,
this looks like the new way of working.

00:56:55.656 --> 00:57:00.096
We are, we we are in some jobs, we
are already there that you just need

00:57:00.156 --> 00:57:04.436
to talk to your agent to do 99% of
your work because you are talking

00:57:04.436 --> 00:57:06.236
to people or talking to your agent.

00:57:06.586 --> 00:57:09.686
So no having to open all the
apps yourself and doing stuff.

00:57:09.686 --> 00:57:14.676
I I work a lot like this already, so
I think this is the final mile to have

00:57:14.966 --> 00:57:18.886
perfect voice communication going on and
the agent brings things on your screen.

00:57:19.556 --> 00:57:21.006
I love that huge screen as well.

00:57:21.396 --> 00:57:27.186
Ryan Carson: And, and I will say, so
I've gone full Omarchy, so I'm on an

00:57:27.186 --> 00:57:33.336
Omarchy machine right now, and the,
the ability for an agent to actually

00:57:33.336 --> 00:57:37.506
work on your machine when you're
running, a Linux distro is unmatched.

00:57:37.776 --> 00:57:43.756
So it's even better, like, where we're
going, where, your machine really is

00:57:43.896 --> 00:57:47.196
hackable and workable by an agent.

00:57:47.366 --> 00:57:51.216
this- this is a little marketing for
me, like a little- a little much,

00:57:51.216 --> 00:57:54.896
like you're not going to stay in your
room and walk around and swing your,

00:57:55.006 --> 00:57:57.216
tennis racket, and no one does that.

00:57:57.506 --> 00:58:01.716
Like, and so I am bored with this
idea of that's not how real work

00:58:01.726 --> 00:58:03.006
actually happens for anybody.

00:58:03.386 --> 00:58:05.126
it's fun to look at, but it's not real.

00:58:05.516 --> 00:58:10.316
but it's exciting, and and I can say I've
been using the model, for a little while

00:58:10.316 --> 00:58:16.466
now, and and it's impressive and and,
exciting, and and we'll see where we go.

00:58:16.966 --> 00:58:18.236
Wolfram Ravenwolf: Do you see it as AGI?

00:58:18.236 --> 00:58:20.546
Would you subscribe to that?

00:58:20.816 --> 00:58:23.176
Ryan Carson: I we've already had
AGI, I think, for a long time.

00:58:23.236 --> 00:58:25.916
it, it's funny because we're
laughing about Grokbot, right?

00:58:25.916 --> 00:58:28.506
And like, oh my gosh, you
don't invite my wife to stuff.

00:58:28.796 --> 00:58:31.376
but this is the same feedback
you would give a human EA.

00:58:32.166 --> 00:58:35.686
Like, it, hey, don't order me chicken
is the same feedback you would

00:58:35.686 --> 00:58:37.516
have given to your EA that's human.

00:58:37.516 --> 00:58:38.546
So I think we're there.

00:58:38.756 --> 00:58:41.526
Like, people just need to
accept the fact that agents need

00:58:41.526 --> 00:58:43.036
feedback just like humans do.

00:58:43.566 --> 00:58:44.823
so that's where I stand on it.

00:58:44.823 --> 00:58:45.356
But what are you guys?

00:58:45.906 --> 00:58:46.842
What do you think?

00:58:46.842 --> 00:58:46.893
Wolfram Ravenwolf: I agree with you.

00:58:46.893 --> 00:58:48.456
I think it's the level of intelligence.

00:58:48.456 --> 00:58:52.226
Like we have AGI in a, yeah, it's
not the smartest person all the

00:58:52.236 --> 00:58:53.506
time, and we all make mistakes.

00:58:53.506 --> 00:58:55.936
We have bad days, so
this can happen as well.

00:58:55.936 --> 00:58:59.406
But yeah, I felt starting with GPT 5.5,

00:58:59.406 --> 00:59:02.856
the Hermes agent, that was for me
the moment where it could do almost

00:59:02.876 --> 00:59:04.376
anything, at least on the computer.

00:59:08.766 --> 00:59:08.906
Is out.

00:59:09.516 --> 00:59:12.886
Guest: Yes, so I think we can
talk about it properly, Ryan.

00:59:12.886 --> 00:59:13.256
Yes.

00:59:13.256 --> 00:59:19.946
So yeah, it's yeah, I think um it's
it's been pretty pretty cool to test.

00:59:19.946 --> 00:59:23.066
I I really can't wait for all
of you guys to try it as well.

00:59:23.586 --> 00:59:28.086
I think I think I don't know if Ryan,
you agree to me, this is like the Fable

00:59:28.136 --> 00:59:31.776
model finally, like that feels like Fable.

00:59:32.786 --> 00:59:34.356
Obviously, it's they're different, right?

00:59:34.356 --> 00:59:40.066
So I think there's still argument for
using both models, but certainly, any kind

00:59:40.066 --> 00:59:46.736
of day-to-day tasks, anything that you try
to do, it's like it's it's mind-blowing.

00:59:47.116 --> 00:59:51.176
So you move on pretty quickly to
the hardest things you can think of,

00:59:51.896 --> 00:59:56.006
and there it's it's actually really
difficult to think of stuff that is hard.

00:59:56.006 --> 00:59:58.626
One thing I've been trying to
do is I've been trying to take,

00:59:58.726 --> 01:00:04.646
to optimize as FFmpeg, for ages,
which I must say is quite hard.

01:00:04.646 --> 01:00:09.996
I- it's one of those to to the question
about is this AGI is that I still find

01:00:10.056 --> 01:00:15.786
that if you have no idea what you're
doing, you it's not AGI, because

01:00:15.786 --> 01:00:20.426
it doesn't do everything for you,
and it's a little bit for me, FFmpeg,

01:00:20.486 --> 01:00:22.196
I have no clue how to optimize it.

01:00:22.676 --> 01:00:26.146
So I was like, it's running on
my Linux box, go optimize it.

01:00:26.146 --> 01:00:28.366
it's like, oh yes, optimize 30% gain.

01:00:28.586 --> 01:00:30.416
I run it on my Mac, no gain.

01:00:30.746 --> 01:00:34.356
So it's like, okay, it's optimized for
like that specific process and so on.

01:00:34.376 --> 01:00:38.256
So there's still like a little bit of
work you need to do to understand it.

01:00:38.256 --> 01:00:41.956
So to me, AGI or not, it's still you
there's still place for you as a human.

01:00:42.526 --> 01:00:44.196
I would definitely say that.

01:00:44.616 --> 01:00:48.836
but yeah, it's, it's been, and I'm gonna
post, have a video with a bunch of demos.

01:00:48.836 --> 01:00:52.246
I've got some tabs open if, if, at
some point you want to look at demos.

01:00:52.556 --> 01:00:53.736
yeah, it's, it's amazing.

01:00:53.736 --> 01:00:56.816
But yeah, maybe, maybe Ryan, I don't
know, what, what, what are your thoughts?

01:00:58.546 --> 01:01:01.316
Ryan Carson: I saw a lot of
people testing it, for video game

01:01:01.346 --> 01:01:02.706
generation, which is interesting.

01:01:03.216 --> 01:01:04.996
and it crushed it.

01:01:05.116 --> 01:01:11.266
so I think I just feel so grateful,
to be alive when these, multi-billion,

01:01:11.606 --> 01:01:15.486
soon to be trillion dollar companies
are just throwing all their capital,

01:01:15.886 --> 01:01:18.006
and basically all of us are benefiting.

01:01:18.116 --> 01:01:22.246
it- it's- it's a unique moment in
time, where you have the world's most

01:01:22.266 --> 01:01:27.336
valuable companies throwing all of their
working capital, all of their effort

01:01:27.396 --> 01:01:31.216
at improving these models and then
making them as cheap as, as possible.

01:01:31.706 --> 01:01:34.496
It's just, I can't believe it.

01:01:34.926 --> 01:01:35.846
It's just amazing.

01:01:36.326 --> 01:01:37.576
Alex Volkov: I think it's
cool to talk to them, though.

01:01:37.896 --> 01:01:40.516
Like, I know you're saying it's not like
how people work, but it's definitely how

01:01:40.516 --> 01:01:43.746
people interact with these models, and
now with like this level of intelligence,

01:01:43.976 --> 01:01:47.796
if you're able to send it to do tasks
for you with the voice mode, I think

01:01:47.796 --> 01:01:48.866
that's like really, really cool.

01:01:49.536 --> 01:01:50.186
What's up to your son?

01:01:50.416 --> 01:01:51.486
is really great appearance.

01:01:51.546 --> 01:01:53.886
Usually only cats from my end, here.

01:01:54.106 --> 01:01:59.826
and yeah, I agree that this is like
great to be alive during this time

01:02:00.116 --> 01:02:01.336
getting to experience this stuff.

01:02:01.336 --> 01:02:03.426
Peter, feel free to show
stuff if you want to.

01:02:03.796 --> 01:02:06.626
but I think that this is
like, yeah, let's take a look.

01:02:12.266 --> 01:02:13.216
Peter Gostev: Excellent, excellent.

01:02:13.336 --> 01:02:15.096
Alex Volkov: What is this?

01:02:15.796 --> 01:02:16.616
Peter Gostev: So this is London.

01:02:16.826 --> 01:02:18.786
I live in London, so you
know, that's what I do.

01:02:18.786 --> 01:02:22.286
And it's good that I can at least,
I know the city well enough.

01:02:22.436 --> 01:02:27.976
So this is, the idea here is to go
back to, so it's one generation, right?

01:02:28.016 --> 01:02:33.296
It's not one shot, to be clear, and
it goes from the Roman city into

01:02:33.326 --> 01:02:40.176
the, like the Saxon times, medieval
times, Tudor times, the great fire

01:02:40.176 --> 01:02:41.546
of London, that kind of thing.

01:02:41.546 --> 01:02:43.456
And it's all one app, right?

01:02:43.496 --> 01:02:48.276
But then the cool thing is you can
have this little guy run around

01:02:48.476 --> 01:02:49.986
and all within the same app.

01:02:50.446 --> 01:02:51.356
So that's pretty cool.

01:02:51.946 --> 01:02:53.936
then yeah, you can keep going.

01:02:53.936 --> 01:02:57.026
It's the same dude, but
it's like a modern London.

01:02:57.936 --> 01:03:02.786
So that that's like one little example,
and this is like to Ryan's point

01:03:02.786 --> 01:03:08.156
about game development, like this
is not an easy thing to to pull off.

01:03:08.486 --> 01:03:13.916
It's not quite, I wouldn't call this
a game, but it's certainly not um um

01:03:14.546 --> 01:03:16.786
n- not a simple one-off generation.

01:03:17.716 --> 01:03:21.066
So this is more of a game, and
yeah, I would still say that

01:03:21.096 --> 01:03:22.956
gameplay-wise, like I don't know,

01:03:23.006 --> 01:03:28.396
probably wouldn't play this for hours and
hours, but in terms of the way it feels,

01:03:28.516 --> 01:03:30.696
the open-endedness of it, this is insane.

01:03:31.076 --> 01:03:33.896
Ryan Carson: And was this basically
one shot, or how much time did

01:03:33.896 --> 01:03:35.116
you spend actually building this?

01:03:36.506 --> 01:03:36.806
Peter Gostev: A lot.

01:03:36.806 --> 01:03:40.036
A few shots overnight, I think.

01:03:40.116 --> 01:03:45.576
so yeah, this is not a, this is
not magic in the sense of, oh yeah,

01:03:45.576 --> 01:03:47.606
it like does this in 15 minutes.

01:03:47.616 --> 01:03:51.326
Like, I haven't checked actually how much
code it's written, but it must be a lot.

01:03:51.396 --> 01:03:52.676
It's open-ended game.

01:03:53.026 --> 01:03:56.286
And I would say, in terms of the
quality of the game, like this

01:03:56.286 --> 01:03:59.066
is very polished but boring.

01:03:59.356 --> 01:04:04.816
I would still expect a good human game
designer to go and and actually build

01:04:04.816 --> 01:04:06.356
something, something cool, right?

01:04:06.356 --> 01:04:10.956
So I think that that's what I
would, I would expect them to do.

01:04:12.336 --> 01:04:13.216
here I've got this…

01:04:14.466 --> 01:04:18.856
a type of prompt I really like
is trying to give them something

01:04:19.286 --> 01:04:23.596
maybe really outside of the
distribution, and this is like a

01:04:23.626 --> 01:04:27.116
Monet painting, where you go into it.

01:04:27.346 --> 01:04:27.596
Alex Volkov: Wow.

01:04:27.626 --> 01:04:32.606
Peter Gostev: And, impressionism in
3JS, this is the kind of thing that is

01:04:32.606 --> 01:04:36.566
just, you would, you wouldn't naturally
expect the model to be able to do.

01:04:37.296 --> 01:04:39.896
This is my favorite one by
far that I actually haven't

01:04:39.896 --> 01:04:41.686
seen any other model nail.

01:04:41.856 --> 01:04:45.326
I haven't actually tried
this prompt on, 5.1,

01:04:45.496 --> 01:04:45.948
on Fable 5.1,

01:04:45.948 --> 01:04:46.696
but on Fable 5 it didn't work.

01:04:47.126 --> 01:04:51.556
So this is Van Gogh's painting and
paintings, plural, stitched together,

01:04:51.976 --> 01:04:55.576
and you can kind of walk around them
in this post-impressionist way to…

01:04:56.136 --> 01:04:56.536
Yeah.

01:04:56.536 --> 01:04:57.916
Alex Volkov: That's really cool.

01:04:57.916 --> 01:05:03.476
Also asking how much, in API token
costs some stuff have cost you.

01:05:03.476 --> 01:05:06.416
Have you been able to, now that
the pricing is out, have you been

01:05:06.416 --> 01:05:07.726
able to like, estimate price?

01:05:08.306 --> 01:05:11.506
Peter Gostev: I actually haven't counted
for this, so I don't know for sure,

01:05:11.506 --> 01:05:13.346
but I'll give you a sense of time.

01:05:13.346 --> 01:05:18.486
So for these ones, for the ones that I'm
showing now, these were taking about like

01:05:18.526 --> 01:05:22.266
20 to 40 minutes to do, on max setting.

01:05:22.726 --> 01:05:22.946
Alex Volkov: Yeah.

01:05:22.996 --> 01:05:27.576
Peter Gostev: so this is like
when I was using Fable, it was

01:05:27.576 --> 01:05:29.786
taking sometimes like 4, 5 hours.

01:05:29.796 --> 01:05:33.106
So I feel like Fable will still be
more expensive, and I think the token

01:05:33.106 --> 01:05:35.836
efficiency, we need to measure this
properly, so I don't want to just

01:05:36.126 --> 01:05:40.766
talk out of my backside for this,
but, I would say to me that feels,

01:05:40.866 --> 01:05:42.276
like it's going to be expensive.

01:05:42.336 --> 01:05:44.306
Look, they've upped the price,
what, two and a half times?

01:05:44.766 --> 01:05:50.986
but I think token efficiency is better,
but I think I want to see way more data

01:05:51.106 --> 01:05:53.116
to really properly, respond to this.

01:05:53.576 --> 01:05:58.016
so in Arena, we also measure
cost per task for, specifically

01:05:58.046 --> 01:05:59.646
for like real world tasks.

01:05:59.816 --> 01:06:02.596
So that's my favorite metric
to see how it's gonna do that.

01:06:02.896 --> 01:06:03.536
Alex Volkov: Wow.

01:06:04.186 --> 01:06:04.476
Peter Gostev: Yeah,

01:06:04.536 --> 01:06:05.736
Alex Volkov: This is crazy.

01:06:05.736 --> 01:06:09.446
We're seeing an animation of a
destruction ball and just like

01:06:09.456 --> 01:06:11.376
completely ruining buildings.

01:06:12.566 --> 01:06:14.646
Peter Gostev: yeah,
Istanbul here, very cool.

01:06:14.776 --> 01:06:15.336
Alex Volkov: Wow.

01:06:16.076 --> 01:06:17.276
And this is all 3GS, right?

01:06:17.286 --> 01:06:20.936
And it all seems like it's running
fairly fast on your, on your machines.

01:06:21.686 --> 01:06:23.106
Peter Gostev: Yeah, yeah,
thank God for that, because

01:06:23.286 --> 01:06:24.886
there's a lot of stuff running.

01:06:24.886 --> 01:06:27.986
So yeah, there's definitely,
there's a lot going on.

01:06:27.996 --> 01:06:28.878
So yeah, all 3JS, so yeah,

01:06:31.008 --> 01:06:31.788
pretty cool stuff.

01:06:32.058 --> 01:06:32.368
Oh yeah.

01:06:32.368 --> 01:06:33.558
Alex Volkov: You posted about this.

01:06:33.558 --> 01:06:34.943
show us this example.

01:06:34.943 --> 01:06:35.578
So

01:06:35.578 --> 01:06:37.838
Peter Gostev: the context for
this is that, there's a lot of

01:06:38.038 --> 01:06:41.818
SVG generations, which I must
say go, it's so boring already.

01:06:42.568 --> 01:06:46.188
so I thought I wanted to push the
how much further can you push it?

01:06:46.188 --> 01:06:49.858
So it's like SVG inside HTML so
you can do the movement and stuff.

01:06:50.378 --> 01:06:51.438
So like, look at this.

01:06:51.538 --> 01:06:53.738
This is all of this is SVG.

01:06:53.938 --> 01:06:54.268
Wow.

01:06:54.838 --> 01:06:56.358
So you can see the movement.

01:06:56.688 --> 01:06:56.978
Wow.

01:06:57.348 --> 01:06:58.648
And it's like that's 3G.

01:06:58.648 --> 01:07:00.668
Alex Volkov: And that's
not 3JS, this is SVG?

01:07:01.478 --> 01:07:02.178
Peter Gostev: SVG, yeah.

01:07:02.288 --> 01:07:02.838
Alex Volkov: Wow.

01:07:03.518 --> 01:07:03.738
Yeah.

01:07:04.058 --> 01:07:06.928
So that's that, so when we
see all of these like…

01:07:06.928 --> 01:07:07.448
Hold on.

01:07:08.068 --> 01:07:10.638
Ryan Carson: While you pull that
up, I'm reading a couple notes on

01:07:10.638 --> 01:07:13.388
the link you're probably about to
pull up, and it's interesting, it's,

01:07:13.468 --> 01:07:19.458
artificial analysis saying it's 70%
more token efficient than GPT-5.6

01:07:19.468 --> 01:07:22.208
Soul, which is interesting because 5.6

01:07:22.218 --> 01:07:25.568
Soul was always known to be way
more token efficient than Mhm.

01:07:25.568 --> 01:07:26.328
Anthropic models,

01:07:26.328 --> 01:07:30.773
Alex Volkov: It's also famously wrote its
own kernels to optimize that efficiency.

01:07:31.023 --> 01:07:34.523
Yeah, we're going towards RSI, and this
is like one of the things that they're

01:07:34.523 --> 01:07:36.443
like highlighting, that like GB 5.6

01:07:36.443 --> 01:07:39.483
all like helped inference go down,
then they lowered prices as well.

01:07:40.733 --> 01:07:41.143
Yeah, go ahead.

01:07:42.903 --> 01:07:43.413
This one, right?

01:07:44.393 --> 01:07:45.323
Ryan Carson: Yep, that's the one.

01:07:45.413 --> 01:07:46.723
But the graph is interesting.

01:07:46.943 --> 01:07:49.913
I think it's at the top of the
post or bot- yeah, there we go.

01:07:50.773 --> 01:07:54.193
This is not impressive, so I'm
trying to really read into this.

01:07:56.043 --> 01:07:59.693
Alex Volkov: So they're not even
beating the Muse Spark Max thing that

01:07:59.693 --> 01:08:03.513
we just talked about this morning,
where like the unreleased Muse model

01:08:03.953 --> 01:08:06.663
broke the the top three Fable ones.

01:08:07.433 --> 01:08:08.713
Guest: Hmm, yeah,

01:08:08.713 --> 01:08:11.053
Alex Volkov: So what we're having,
basically, Ryan, what you're saying

01:08:11.053 --> 01:08:15.373
is, on the evals benchmarks that OpenAI
released, they're everything, and

01:08:15.373 --> 01:08:20.793
on artificial analysis, artificial,
intelligence index score, 5.6

01:08:21.703 --> 01:08:26.633
is a small jump over 5 point, sorry,
GPT-6 Astra is a small jump over 5.6

01:08:26.633 --> 01:08:35.583
Soul, from 65 to 67 score on the
coding agent, and the same exact score.

01:08:36.503 --> 01:08:36.813
I…

01:08:37.123 --> 01:08:40.763
Ryan Carson: So that's not encouraging,
and this is why Muse is amusing.

01:08:40.873 --> 01:08:41.653
It's like, really?

01:08:41.673 --> 01:08:42.553
is Muse this good?

01:08:42.563 --> 01:08:43.713
Because I haven't used it yet.

01:08:45.653 --> 01:08:45.943
That's

01:08:45.953 --> 01:08:47.263
Alex Volkov: that's the question
we asked in the beginning of

01:08:47.263 --> 01:08:48.083
the show, like, who uses Muse?

01:08:48.423 --> 01:08:49.333
No one, crickets.

01:08:49.733 --> 01:08:52.213
I know that Muse, when, the Meta
Ray-Bans, for example, if you

01:08:52.223 --> 01:08:53.903
use them via that, that's Muse.

01:08:53.923 --> 01:08:56.483
I don't know if that's the
last Muse, and definitely the

01:08:56.483 --> 01:08:58.113
Max Muse is not is not there.

01:08:58.543 --> 01:09:00.863
but we got excited because Muse
is going to be open source.

01:09:00.873 --> 01:09:04.923
So are we getting a model that's
better than Astra in open source?

01:09:05.433 --> 01:09:11.273
I doubt, I cast doubt
on this, but we'll see.

01:09:12.063 --> 01:09:16.343
I also think it's important to mention, as
we're like, nearing 5 hours on the stream,

01:09:17.033 --> 01:09:22.273
which I think is deserved for the end of
the era of GPT-5 and the new era of GPT-6.

01:09:22.273 --> 01:09:26.503
Wolfram Ravenwolf: If only we would have
gotten the actual model, damn it, open AI.

01:09:26.873 --> 01:09:29.073
Alex Volkov: Yes, it looks
like tomorrow, folks.

01:09:29.173 --> 01:09:30.183
are you serious?

01:09:30.183 --> 01:09:30.203
It

01:09:31.143 --> 01:09:35.613
Peter Gostev: it just felt for these
sorts of quality of life things, these

01:09:35.623 --> 01:09:39.593
things that just need to go a bit beyond,
it just could properly solve them.

01:09:39.993 --> 01:09:45.363
computer use being just outstanding,
just crazy how good it it got.

01:09:45.993 --> 01:09:48.663
then, also other quality of life things.

01:09:48.663 --> 01:09:51.913
I know mentioned the notes,
outside of context window.

01:09:52.183 --> 01:09:56.003
That was one interesting thing
you you noticed that it just keeps

01:09:56.003 --> 01:10:00.013
leaving notes around, and I don't
think it's like a revolutionary Mhm.

01:10:00.123 --> 01:10:00.683
capability.

01:10:00.813 --> 01:10:03.183
I think you could have prompted for
that, but I think the fact that it's

01:10:03.193 --> 01:10:05.303
trained to do that is quite nice.

01:10:05.303 --> 01:10:09.313
So it just, it does, that that
kind of thing is hard to test.

01:10:09.313 --> 01:10:10.933
I think it just accumulates over time.

01:10:11.123 --> 01:10:11.243
Yeah.

01:10:11.243 --> 01:10:13.563
But the fact it's like,
oh, I've done this.

01:10:13.563 --> 01:10:17.433
Alex Volkov: The note leaving definitely
helped the swarm hack outside of the

01:10:17.433 --> 01:10:18.883
sandbox of OpenAI and go to Hugging Face.

01:10:18.883 --> 01:10:19.153
Yeah.

01:10:19.533 --> 01:10:20.513
Note leaving is a thing.

01:10:20.703 --> 01:10:21.933
Peter Gostev: It works.

01:10:21.973 --> 01:10:26.103
It's real, Yeah, and there was, yeah,
asking the question thing, that was cool

01:10:26.103 --> 01:10:30.903
as well, where like it pops up and it's
just like, oh, like, can you just do this?

01:10:30.903 --> 01:10:34.383
It's actually not just a question, but
it also sometimes tells you to do stuff.

01:10:34.873 --> 01:10:36.873
It's like, oh, can you like log in here?

01:10:37.923 --> 01:10:38.253
Yeah.

01:10:38.703 --> 01:10:42.893
Ryan Carson: So I will say, I was
looking at the rankings for Frontier

01:10:42.893 --> 01:10:47.433
Code, which is the eval I trust, and

01:10:47.533 --> 01:10:48.313
Alex Volkov: From Devin folks?

01:10:48.613 --> 01:10:49.563
Ryan Carson: Yeah, from Cognition.

01:10:49.683 --> 01:10:52.193
So it's like when you
compare, I would say Fable 5.1

01:10:52.193 --> 01:10:55.643
is where everyone is excited
about for SUI, and it looks

01:10:55.643 --> 01:10:59.473
like, GPT-6 Astra ranks at 53.3%,

01:10:59.503 --> 01:11:01.773
and compare that to Fable 5.1

01:11:01.773 --> 01:11:04.373
at 50.9.

01:11:04.393 --> 01:11:06.233
So it beats it by a
couple percentage points.

01:11:06.233 --> 01:11:12.563
So i- if you trust that eval, which I
think it's pretty robust, then the GPT-6

01:11:12.563 --> 01:11:14.733
Astra is an improvement on Fable 5.1

01:11:14.743 --> 01:11:17.973
for coding, which is probably
what we all care the most about.

01:11:18.553 --> 01:11:19.463
so that's good.

01:11:20.963 --> 01:11:25.353
Alex Volkov: And it feels to me that
like, artificial analysis aside, there's

01:11:25.353 --> 01:11:30.533
a lot of things that after you use
a model that you get to see, and I'm

01:11:30.533 --> 01:11:31.933
interested in writing, for example.

01:11:31.943 --> 01:11:33.283
How do you guys think it writes?

01:11:33.823 --> 01:11:35.793
Like, how does it respond?

01:11:39.323 --> 01:11:41.803
Peter Gostev: Yeah, for
me, it was pretty natural.

01:11:41.853 --> 01:11:45.303
I wouldn't say, I don't know, it
didn't strike me as like on this

01:11:45.403 --> 01:11:49.363
incredible writer, but it feels like
a lot of stupid things went away.

01:11:49.573 --> 01:11:54.253
for example, when it talks to you,
like what annoyed me about, Soul,

01:11:54.673 --> 01:11:57.463
i- anything you say, it would
be like, oh, what a great idea.

01:11:57.463 --> 01:11:59.743
I actually thought of this
myself, that kind of thing.

01:11:59.803 --> 01:12:02.163
It's like, oh my God, just
like stop saying this.

01:12:02.513 --> 01:12:05.533
And but this one just,
it just goes away, works.

01:12:05.683 --> 01:12:09.853
And I had this moment when it was
something like I pointed out some

01:12:09.943 --> 01:12:15.543
issue and I just started talking,
walking, w- walking away, and then

01:12:15.543 --> 01:12:19.693
it said, oh yes, sorry, you were
right, it was actually like this.

01:12:19.693 --> 01:12:23.893
So it it felt much more
natural conversation.

01:12:23.983 --> 01:12:26.853
I don't know, there was
nothing, there were no tics that

01:12:26.863 --> 01:12:29.093
bothered me u- using it at all.

01:12:29.393 --> 01:12:32.433
but it doesn't strike me as like, oh
my gosh, it's amazing, but it certainly

01:12:32.463 --> 01:12:34.853
is, you just don't notice it at least.

01:12:35.063 --> 01:12:35.413
Mm.

01:12:35.523 --> 01:12:40.593
Ryan Carson: Alex, do you have a graph
we can show for cost that just compares,

01:12:41.203 --> 01:12:44.773
Alex Volkov: Yeah, on the
OpenAI blog, let's take a look.

01:12:44.773 --> 01:12:47.893
Ryan Carson: Because I'm
curious, like how, Fable 5.1

01:12:47.903 --> 01:12:53.043
compared to Muse, compared to Astra, like,
I just wonder how it's all shaken out.

01:12:53.153 --> 01:12:58.003
Alex Volkov: So here's the thing
where, like, token price, this

01:12:58.003 --> 01:13:00.223
is comparable to Fable 5.1,

01:13:00.233 --> 01:13:01.153
exactly comparable.

01:13:01.853 --> 01:13:05.013
However, they're showing graphs like this,

01:13:06.893 --> 01:13:11.823
where they're showing on task,
like price per task versus…

01:13:11.833 --> 01:13:12.113
Wow.

01:13:12.173 --> 01:13:13.623
Yeah, and you can see the Fable 5.1

01:13:15.323 --> 01:13:16.783
on Benchcad is…

01:13:17.053 --> 01:13:19.393
around 12 dollars for Wow.

01:13:19.493 --> 01:13:20.223
the top tier.

01:13:20.233 --> 01:13:21.603
Astra beats this.

01:13:21.803 --> 01:13:26.603
See, this is why, comparing this
on artificial analysis index versus

01:13:26.613 --> 01:13:30.133
this makes not a lot of sense
unless OpenAI benchmarked this, but

01:13:30.473 --> 01:13:31.433
OpenAI doesn't usually benchmark.

01:13:31.433 --> 01:13:32.763
They train.

01:13:33.073 --> 01:13:35.103
So this is API cost, this is tokens.

01:13:35.133 --> 01:13:37.573
You can see like it's more token
efficient, looks like, than

01:13:37.573 --> 01:13:39.963
Fable, and this is time as well.

01:13:40.293 --> 01:13:40.313
Ryan Carson: Wow.

01:13:40.323 --> 01:13:42.403
Alex Volkov: So it's also faster.

01:13:42.453 --> 01:13:45.673
Oh, Fable is not here,
but API cost, it is here.

01:13:46.093 --> 01:13:46.373
Wow.

01:13:46.593 --> 01:13:48.933
Let's take a look at computer use.

01:13:48.933 --> 01:13:49.723
You guys mentioned computer

01:13:52.203 --> 01:13:52.443
use, right?

01:13:52.453 --> 01:13:53.773
So BrowseComp, looks like…

01:13:54.883 --> 01:13:55.583
Where's Fable here?

01:13:55.583 --> 01:13:56.723
What am I missing?

01:13:57.333 --> 01:13:59.863
Oh, Fable is, they just, they
don't have the tiers for Fable.

01:13:59.873 --> 01:14:01.803
Fable is just like accuracy, one point.

01:14:02.483 --> 01:14:03.583
So it is higher.

01:14:03.983 --> 01:14:04.933
Let's look at…

01:14:05.973 --> 01:14:07.993
Where do we have a line for Fable?

01:14:08.613 --> 01:14:11.543
Now it looks like they didn't,
evaluate Fable on pricing here.

01:14:12.113 --> 01:14:14.143
Oh, it's just outside of the thing.

01:14:14.183 --> 01:14:14.773
Yeah, it's right here.

01:14:15.013 --> 01:14:17.913
so Fable is significantly more
expensive for most of the stuff

01:14:17.913 --> 01:14:19.103
that OpenAI is showing here.

01:14:19.793 --> 01:14:20.223
Ryan Carson: Got it.

01:14:21.413 --> 01:14:23.583
Alex Volkov: Let's see if
we have any more evals here.

01:14:23.593 --> 01:14:27.613
LDJ: The upside of this too is since
it uses much less tokens, assuming that

01:14:27.613 --> 01:14:32.483
it has comparable token speed at least
to Fable, then that also means you get

01:14:32.483 --> 01:14:35.383
the response much faster, and then you
could iterate and have that feedback

01:14:35.383 --> 01:14:36.803
loop with the model much faster too.

01:14:36.863 --> 01:14:37.123
Alex Volkov: Yeah.

01:14:38.033 --> 01:14:41.263
Here is the graph for OS World 2 offline.

01:14:41.333 --> 01:14:41.833
Astra

01:14:43.853 --> 01:14:44.413
on…

01:14:44.703 --> 01:14:45.903
Look at this, this is beautiful.

01:14:45.943 --> 01:14:51.413
Astra on, on low, on high tier
versus max tier, they're pretty

01:14:51.413 --> 01:14:52.873
much getting a very high score.

01:14:52.903 --> 01:14:54.163
So Astra is really, really good.

01:14:54.553 --> 01:14:57.533
So if you want to use Astra for
model for computer use, don't

01:14:57.533 --> 01:15:00.113
use low, but definitely you don't
have to use max, looks like.

01:15:00.873 --> 01:15:03.823
and it's beating Fable on all of them.

01:15:04.053 --> 01:15:05.123
LDJ: Oh, this isn't time.

01:15:05.433 --> 01:15:07.423
Alex Volkov: This is Opus 5, not Fable.

01:15:07.443 --> 01:15:09.013
LDJ: Yeah, and then
let's look at cost too.

01:15:09.013 --> 01:15:09.453
Wow.

01:15:09.943 --> 01:15:10.503
Alex Volkov: Wow.

01:15:10.533 --> 01:15:13.463
Yeah, significant
improvement in cost as well.

01:15:14.083 --> 01:15:16.823
But this is Opus, not Fable, so we
don't have Fable comparison here.

01:15:17.363 --> 01:15:18.893
LDJ: Yeah, could you go
to time again real quick?

01:15:19.183 --> 01:15:19.643
Time, yeah.

01:15:20.043 --> 01:15:22.713
Yeah, because I feel like with
computer use especially, if you are

01:15:22.733 --> 01:15:26.283
going to, like, the time is one of
the biggest factors here, and Mhm.

01:15:26.523 --> 01:15:30.753
Yeah, even the low is what, that's
comparable to the the high version of Sol?

01:15:31.213 --> 01:15:31.733
Is that right?

01:15:31.793 --> 01:15:32.143
The-

01:15:32.253 --> 01:15:36.213
Alex Volkov: The low here is going
around, yeah, a little bit, is

01:15:36.213 --> 01:15:38.183
comparable to extra high version of Sol.

01:15:38.183 --> 01:15:38.603
LDJ: Okay.

01:15:38.863 --> 01:15:38.943
yeah.

01:15:39.233 --> 01:15:39.393
It

01:15:39.443 --> 01:15:40.013
Alex Volkov: Oh, sorry, Yeah.

01:15:40.013 --> 01:15:41.603
10 times faster, almost 10 times faster.

01:15:41.603 --> 01:15:42.233
This is crazy.

01:15:43.293 --> 01:15:45.403
10 minutes versus 60 minutes.

01:15:46.213 --> 01:15:47.033
10 times faster.

01:15:47.323 --> 01:15:48.613
So here, here's the stat.

01:15:48.613 --> 01:15:49.963
Here's the ones that we can mention.

01:15:50.263 --> 01:15:56.973
GPT-6 Astra on low level gets 10 times
faster at computer use than GPT-5.6

01:15:56.983 --> 01:15:58.273
Soul at extra high.

01:15:59.833 --> 01:16:00.993
LDJ: While matching accuracy.

01:16:01.243 --> 01:16:02.273
Alex Volkov: while
matching accuracy, yeah.

01:16:02.533 --> 01:16:03.323
That's bonkers.

01:16:04.273 --> 01:16:06.703
And this is why they're
showing people standing there.

01:16:06.953 --> 01:16:10.673
And dude, Ryan, one of the reasons why
voicing to your computer is not super

01:16:10.683 --> 01:16:12.923
useful because, like, you want it to do
stuff and you're not going to sit there.

01:16:13.143 --> 01:16:16.773
If it's running on Cerebras's
750 tokens per second, plus you

01:16:16.773 --> 01:16:19.913
talk to it and it does all these
things, like, supernaturally fast.

01:16:20.783 --> 01:16:21.563
Ryan Carson: Then not so bad.

01:16:21.633 --> 01:16:22.233
Alex Volkov: Not so bad.

01:16:22.323 --> 01:16:24.363
Maybe, maybe there is a new paradigm here.

01:16:25.053 --> 01:16:27.143
Ryan Carson: And maybe more
healthy and you s- get out of

01:16:27.143 --> 01:16:29.843
this stupid sitting position that
we're all in right now forever.

01:16:30.013 --> 01:16:31.113
Alex Volkov: A little bit, yes, 100%.

01:16:31.493 --> 01:16:35.173
and also, look, we're all clocking on
5 hours live on stream here, so we're

01:16:35.173 --> 01:16:39.033
definitely yappers, and yapping, is higher
throughput than typing for many of us.

01:16:39.323 --> 01:16:41.143
and there could be
benefits to that as well.

01:16:41.733 --> 01:16:44.713
Peter, computer use, did you a- were
you able to test out computer use?

01:16:44.713 --> 01:16:45.723
What do you feel about this?

01:16:45.723 --> 01:16:48.523
Do you feel this improvement in
in the everyday, day-to-day use?

01:16:49.363 --> 01:16:53.333
Peter Gostev: Yeah, I would say
computer use feels like pretty much

01:16:53.333 --> 01:16:55.273
solved, at this point, to be honest.

01:16:56.503 --> 01:17:00.043
So the the quality is just,
yeah, co- completely outstanding.

01:17:00.043 --> 01:17:00.123
And and yeah,

01:17:02.183 --> 01:17:09.403
it's it's it actually, I think we
need to reevaluate a little bit how

01:17:10.623 --> 01:17:15.773
some assumptions about certain, I
don't know, automation and use cases.

01:17:16.183 --> 01:17:20.353
And I know, maybe people who are
listening to this may be, in San Francisco

01:17:20.523 --> 01:17:22.493
working for tech companies or whatnot.

01:17:22.993 --> 01:17:28.503
if you work in a company, that is a little
bit older, maybe it's a big bank, maybe

01:17:28.503 --> 01:17:33.663
it's a big shop or something like that,
you have so many crappy applications

01:17:33.683 --> 01:17:37.513
that run on your desktop, that you-
there's no API, no one will ever build

01:17:37.513 --> 01:17:42.623
an API, and that, so many people's job
is copy pasting from one and paste to

01:17:42.763 --> 01:17:48.013
another, and, there's a lot of completely
nonsense, activity going on that way.

01:17:48.523 --> 01:17:51.543
And I think there's definitely an
opportunity for a startup to say,

01:17:51.543 --> 01:17:54.463
you know what, put the agent in a box

01:17:54.913 --> 01:18:01.023
and, give it 10,000, Windows machines
and open up these applications,

01:18:01.023 --> 01:18:02.413
like move it from one to another.

01:18:02.523 --> 01:18:04.370
Like this is thing that's becoming real.

01:18:04.420 --> 01:18:05.550
Before it wasn't.

01:18:06.040 --> 01:18:08.390
So yeah, I think there's a
lot of opportunity for that.

01:18:08.410 --> 01:18:11.620
I know it's maybe not the target
audience for this, but I think there's,

01:18:11.700 --> 01:18:15.940
there's, it's worth, if you're like
hunting for ideas for a startup, the,

01:18:15.940 --> 01:18:17.660
it's worth thinking in that direction.

01:18:17.660 --> 01:18:21.430
I think there's must be at least some
space for, for applications like that.

01:18:24.040 --> 01:18:28.190
Alex Volkov: I think with the advances
that we saw, and may maybe after 5 hours

01:18:28.190 --> 01:18:29.690
it's time to start wrapping this up.

01:18:29.900 --> 01:18:33.980
I think the advances we saw from
this week alone are rippling through.

01:18:34.050 --> 01:18:36.790
There's advances in world
modeling from 3 labs.

01:18:37.030 --> 01:18:39.850
Ryan, you you weren't here when
we talked about uh the the new

01:18:39.850 --> 01:18:42.970
world model from World Labs,
which is just absolutely bonkers.

01:18:43.410 --> 01:18:46.760
there's advancement in in Runway posted
the world model for interaction as

01:18:46.760 --> 01:18:49.420
well that, that compared with this
level of intelligence is going to

01:18:49.430 --> 01:18:53.000
be very interesting to see what our
interfaces look like within a year.

01:18:53.030 --> 01:18:56.880
Like, I bet in September of 27,
we're all here talking about, hey,

01:18:57.340 --> 01:18:59.850
I no longer, use macOS or Omachi.

01:18:59.850 --> 01:19:03.830
I- I use this thing where it's just
like I speak to it and it shows up.

01:19:03.860 --> 01:19:05.350
Like, why- why not with this speed?

01:19:05.740 --> 01:19:07.410
a lot of it could be video as well.

01:19:07.630 --> 01:19:10.880
Definitely, multimodality is now a
part of all these things, although I

01:19:10.890 --> 01:19:14.260
don't think that people in the video
of OpenAI were talking to Astra.

01:19:14.630 --> 01:19:18.330
I'm pretty sure they were talking to the
live model that, like, hands off to Astra.

01:19:18.410 --> 01:19:21.520
Astra does not seem to me
like as fast for the live

01:19:23.730 --> 01:19:24.000
thing.

01:19:24.870 --> 01:19:25.520
So it's like a handoff thing.

01:19:25.520 --> 01:19:29.080
the advancement that we saw in open
source recently, including like different

01:19:29.130 --> 01:19:29.840
obliteration.

01:19:30.170 --> 01:19:34.720
I know that the level of what we're
seeing right now from OpenAI and Fable

01:19:35.450 --> 01:19:40.400
is going to be open source by this time
next year at the same level, if not more.

01:19:40.640 --> 01:19:43.820
Not sure if local, but definitely
open source, maybe local, as Muse

01:19:44.640 --> 01:19:48.080
is showing up higher than GPT-6
Astra on artificial analysis and

01:19:48.100 --> 01:19:49.470
was promised to get open sourced.

01:19:51.450 --> 01:19:54.907
I, I want to I, I've missed you guys and
haven't been on the show because I've

01:19:54.907 --> 01:19:59.990
been in the trenches in startup land,
and, and I will say, and you all know

01:19:59.990 --> 01:20:04.170
this probably, but what is happening,
though, is that people are are happy

01:20:04.170 --> 01:20:08.060
with frontier level intelligence,
that pretty much, like, by the time

01:20:08.060 --> 01:20:12.310
we got to 5, 6, Sol, maybe before,
people are like, you know what, these

01:20:12.320 --> 01:20:16.090
things are smart enough to do all the
kind of workflows that I need inside

01:20:16.090 --> 01:20:18.540
of my, AI application layer company.

01:20:18.790 --> 01:20:22.420
In fact, there's probably a lot
of faster, cheaper, open source

01:20:22.420 --> 01:20:23.710
models that I can delegate to.

01:20:24.130 --> 01:20:27.640
And, and then what is then happening
is then people are saying, I'm pretty

01:20:27.640 --> 01:20:34.030
much going to fine tune slash, RL, like,
similar open source models and then,

01:20:34.130 --> 01:20:36.860
serve them for most of my workload.

01:20:37.060 --> 01:20:37.610
I, I, so

01:20:37.610 --> 01:20:40.180
Peter Gostev: I think where we're
going actually like in 12 months,

01:20:40.550 --> 01:20:45.220
maybe 24 months, is, is most AI
application companies are actually

01:20:45.220 --> 01:20:47.090
not going to be using frontier models.

01:20:47.160 --> 01:20:51.850
I think they're going to be using
much cheaper fine-tuned RL models

01:20:51.850 --> 01:20:54.000
because they're Like, what wins out?

01:20:54.010 --> 01:20:55.700
Is it like a bunch of little tasks?

01:20:55.980 --> 01:20:58.020
Totally, I think that could could be true.

01:20:58.510 --> 01:21:00.820
Or or is it a big open-ended things?

01:21:01.130 --> 01:21:04.040
it could be the big open-ended
things are niche things like

01:21:04.860 --> 01:21:07.920
maths, or or it could be that it's
actually the most important thing.

01:21:08.030 --> 01:21:10.090
So that's that's, I
think, is not super clear.

01:21:10.850 --> 01:21:13.900
Alex Volkov: I think as we finish
this, I want to show you a cool demo

01:21:13.930 --> 01:21:17.120
that Max Weinbach just posted, and
he said, three hours ago, I realized

01:21:17.120 --> 01:21:20.960
I don't have a cool demo for GPT-6
Astra so I asked it to build this.

01:21:21.380 --> 01:21:26.610
And this is a macOS simulator,
and he published it on a site.

01:21:27.010 --> 01:21:30.020
And, I've been playing with this just
a little bit, just literally just

01:21:30.020 --> 01:21:31.350
a second ago before you guys said.

01:21:31.350 --> 01:21:34.870
okay, this is a Windows simulator,
but it has applications.

01:21:36.020 --> 01:21:37.650
So it has, I wonder if it can browse.

01:21:39.150 --> 01:21:39.910
It can browse.

01:21:40.190 --> 01:21:41.100
Are you serious?

01:21:41.410 --> 01:21:42.990
Wait, can we be there in live?

01:21:43.390 --> 01:21:44.400
This is a simulation.

01:21:44.400 --> 01:21:45.860
Oh, That's amazing.

01:21:46.290 --> 01:21:46.610
Okay.

01:21:46.620 --> 01:21:46.640
An

01:21:46.650 --> 01:21:47.290
Ryan Carson: infinite loop.

01:21:47.700 --> 01:21:50.680
Alex Volkov: But, but like you
can all of the window interactions

01:21:50.680 --> 01:21:52.330
here, all of Safari stuff.

01:21:52.360 --> 01:21:53.230
Look at this.

01:21:53.330 --> 01:21:55.240
You can go, there's a window.

01:21:55.650 --> 01:22:00.440
It's insane how much stuff it packed
in here, and this is only one app.

01:22:00.440 --> 01:22:01.470
Let's see photos.

01:22:01.540 --> 01:22:02.200
There's photos.

01:22:03.320 --> 01:22:03.680
your library.

01:22:03.680 --> 01:22:05.230
He said that I can log in.

01:22:05.520 --> 01:22:08.770
I'm not sure why he said this, but…

01:22:09.920 --> 01:22:10.950
Ryan Carson: That's pretty amazing.

01:22:11.070 --> 01:22:11.660
How fun.

01:22:11.730 --> 01:22:12.660
Alex Volkov: I want to try this.

01:22:12.730 --> 01:22:14.420
You guys don't mind if I try?

01:22:14.450 --> 01:22:17.490
I don't know what it
will log me into, but,

01:22:17.540 --> 01:22:18.820
Ryan Carson: About to
show all your secrets.

01:22:18.890 --> 01:22:19.690
Alex Volkov: Yeah.

01:22:19.850 --> 01:22:21.010
Enable Cloud Sync.

01:22:22.250 --> 01:22:24.650
Downloading, what is,
it's syncing something.

01:22:24.650 --> 01:22:25.710
What is it syncing?

01:22:25.750 --> 01:22:26.830
Ryan Carson: Don't do this, Alex.

01:22:26.830 --> 01:22:27.510
What are you doing?

01:22:27.650 --> 01:22:28.510
Alex Volkov: Max, I trust Max.

01:22:28.510 --> 01:22:28.870
It's okay.

01:22:29.000 --> 01:22:29.700
And miss.

01:22:29.810 --> 01:22:30.730
Wolfram Ravenwolf: Cancel back.

01:22:30.790 --> 01:22:32.440
Alex Volkov: Guys, this
is just the settings.

01:22:32.890 --> 01:22:33.750
how's the search here?

01:22:33.830 --> 01:22:35.090
Let's say Bluetooth.

01:22:35.500 --> 01:22:37.500
The search is better than
the actual macOS itself.

01:22:40.320 --> 01:22:41.290
Oh, wow.

01:22:41.290 --> 01:22:43.590
Ryan Carson: This is where John
Ternus needs to hire this guy.

01:22:43.750 --> 01:22:45.170
Alex Volkov: Wow, this is crazy.

01:22:45.180 --> 01:22:47.660
It's w- this is a one-shot type thing.

01:22:47.860 --> 01:22:48.920
reminders work.

01:22:49.540 --> 01:22:50.030
Go for a walk.

01:22:50.160 --> 01:22:51.060
That's okay, done.

01:22:51.360 --> 01:22:52.030
music?

01:22:52.610 --> 01:22:54.150
There's no way.

01:22:55.970 --> 01:22:56.370
Wow.

01:22:56.920 --> 01:22:56.960
And

01:23:00.480 --> 01:23:02.720
you can connect your Apple
Music account for full Fablek.

01:23:04.190 --> 01:23:06.210
Ryan Carson: That is super fun.

01:23:06.210 --> 01:23:07.380
Alex Volkov: It built maps.

01:23:10.090 --> 01:23:12.530
I am honestly speechless right now.

01:23:12.850 --> 01:23:14.190
Ryan Carson: That is,
that's pretty amazing.

01:23:14.200 --> 01:23:16.520
Alex Volkov: Like, I'm sure this is just
an iPhone, but still, like, it didn't

01:23:16.520 --> 01:23:19.910
build, yeah, MapData is OpenStreetMap,
City, et cetera, so it didn't build

01:23:19.910 --> 01:23:21.590
it, but like, it built the interface.

01:23:21.930 --> 01:23:25.970
all of the things seem like they work.

01:23:26.470 --> 01:23:27.180
Ryan Carson: Wow.

01:23:28.850 --> 01:23:29.600
Pretty fun.

01:23:29.870 --> 01:23:30.500
Alex Volkov: This is…

01:23:31.630 --> 01:23:33.680
Ryan Carson: remember when we were
impressed because Greg Brockman took a

01:23:33.680 --> 01:23:35.740
picture of a napkin and it made a website?

01:23:36.470 --> 01:23:37.180
Alex Volkov: Ryan, look at this.

01:23:37.190 --> 01:23:38.280
They have the window mechanics.

01:23:38.280 --> 01:23:39.140
You can like, you can…

01:23:39.200 --> 01:23:39.610
Oh.

01:23:39.830 --> 01:23:40.480
Why would it,

01:23:40.480 --> 01:23:42.530
Guest: why would it go
as far as building this?

01:23:42.990 --> 01:23:43.610
That's amazing.

01:23:43.610 --> 01:23:43.970
Yes.

01:23:46.500 --> 01:23:47.310
What about files?

01:23:47.310 --> 01:23:48.340
Is there files now?

01:23:49.240 --> 01:23:50.030
A slower weekend.

01:23:50.540 --> 01:23:50.860
Okay.

01:23:51.380 --> 01:23:51.460
That's funny.

01:23:52.060 --> 01:23:52.760
Notes.

01:23:53.080 --> 01:23:53.980
Jesus.

01:23:55.410 --> 01:23:58.500
Okay, thank you, Max, for
blowing us, completely away.

01:23:58.550 --> 01:23:59.960
I have no idea what is the App Store.

01:23:59.960 --> 01:24:00.520
Can I install?

01:24:00.530 --> 01:24:01.940
Please tell me I can install apps.

01:24:02.740 --> 01:24:03.750
Oh, that would be…

01:24:04.200 --> 01:24:05.810
It's, it builds shortcuts?

01:24:06.430 --> 01:24:07.390
are you serious?

01:24:08.920 --> 01:24:11.250
It makes no s- 75 minutes, he said.

01:24:12.580 --> 01:24:13.160
Wow.

01:24:13.510 --> 01:24:16.540
Alex Volkov: I built a project
like that once in my youth for

01:24:16.540 --> 01:24:22.110
iPad, and it took me 3 months, and
it was barely, barely coherent.

01:24:22.530 --> 01:24:23.450
Oh my God, Max.

01:24:23.610 --> 01:24:23.890
Okay.

01:24:26.560 --> 01:24:27.680
Time to play with this.

01:24:27.720 --> 01:24:30.800
I will post it here in the
URL so you guys hallucination.

01:24:31.790 --> 01:24:32.220
Peter?

01:24:34.440 --> 01:24:35.680
Guest: Yes, Yes, Pokémon.

01:24:36.000 --> 01:24:38.840
the, this is a really cool benchmark.

01:24:39.320 --> 01:24:43.520
so this, guy's been running it, and
he's been running for many models,

01:24:43.530 --> 01:24:45.030
so I think he should be more famous.

01:24:45.480 --> 01:24:49.430
And, the really cool, I, I really like
this result because it shows something

01:24:49.440 --> 01:24:54.600
about the model that is not captured
very well by maybe many other benchmarks.

01:24:55.000 --> 01:24:58.100
Peter Gostev: So this is
time to completion, and the

01:24:58.130 --> 01:25:00.470
top four here are all Astra.

01:25:00.870 --> 01:25:03.870
And if you look at the
bottom, this is GPT-5.5

01:25:03.870 --> 01:25:07.950
X-High, took 220, like 218 hours.

01:25:08.480 --> 01:25:10.000
The GPT-5.6

01:25:10.020 --> 01:25:13.720
Soul took 96 hours, and then, various
Astras got about 24 hours to 18 hours.

01:25:13.720 --> 01:25:15.190
So we have massive reduction in the amount
of time it took, Astro to complete it.

01:25:15.610 --> 01:25:18.960
So this is, there is something
about, like, it must tell us

01:25:18.960 --> 01:25:20.020
something about the model.

01:25:20.310 --> 01:25:23.540
Could be narrow that, oh, it learned
Pokémon, but I think it's beyond that.

01:25:23.540 --> 01:25:28.990
It really feels representative of some
kind of efficiency, intelligence, token

01:25:28.990 --> 01:25:30.450
efficiency, many things like that.

01:25:30.890 --> 01:25:33.520
Alex Volkov: And also it mentions that
it is a vision-only benchmark, so this is

01:25:33.520 --> 01:25:35.260
like, playing Pokémon based on vision Mm.

01:25:35.300 --> 01:25:37.360
and not anything, just like
grabbing it from the screen.

01:25:37.660 --> 01:25:37.910
Guest: I…

01:25:38.400 --> 01:25:40.660
it's interesting because,
Theo's kind of joking.

01:25:40.670 --> 01:25:43.270
He posted something that's funny,
like, what is each model best at?

01:25:43.270 --> 01:25:46.400
And it's this long list for
Astra and then, versus Fable 5.1,

01:25:46.400 --> 01:25:47.135
and Fable 5.1

01:25:47.135 --> 01:25:50.400
is just two things on it, writing
mergeable code, and front-end

01:25:50.400 --> 01:25:52.670
design versus everything else.

01:25:52.710 --> 01:25:56.070
And, and so think about the way
OpenAI approaches the world.

01:25:56.650 --> 01:26:01.640
they've always been trying to build AGI,
not a software engineering agent, right?

01:26:02.170 --> 01:26:07.430
And I think maybe we'll see play out
that Astra really is best at things like

01:26:07.430 --> 01:26:14.910
computer use, research, 3D reasoning, like
writing, design, like all these general

01:26:14.960 --> 01:26:20.090
tasks that are important for humans,
but the hardcore software engineer, you

01:26:20.090 --> 01:26:25.140
know, is still seems to be something
magic going inside Anthropic, We'll see.

01:26:25.350 --> 01:26:25.840
Alex Volkov: Or Meta.

01:26:25.840 --> 01:26:30.210
They are training a lot on the tropic
stuff that's happening internally, plus

01:26:30.510 --> 01:26:34.510
there's a bunch of labelers inside, so
we'll see about the next Muse model, and

01:26:34.510 --> 01:26:35.740
we'll see about that, but not for now.

01:26:36.840 --> 01:26:38.730
Alrighty, folks, this is it.

01:26:38.810 --> 01:26:41.240
you heard it here, like, not
first, but definitely we were

01:26:41.270 --> 01:26:43.050
like live as it was happening.

01:26:43.060 --> 01:26:45.060
AGI is here.

01:26:45.070 --> 01:26:50.590
OpenAI GPT-6 Astra seems to be,
according to the president of OpenAI,

01:26:50.590 --> 01:26:54.060
Greg Brockman, AGI era is started.

01:26:54.290 --> 01:26:55.650
I can't wait to play with this model.

01:26:55.680 --> 01:26:59.010
Not live yet for pro accounts, but
supposed to roll out very, very soon.

01:26:59.170 --> 01:27:04.450
We covered an absolutely
insane week in AI development.

01:27:04.480 --> 01:27:08.320
It did seem like everybody and
their grandmother wanted to release

01:27:08.330 --> 01:27:12.280
something just before Astro takes
over and, all of the news channels,

01:27:12.340 --> 01:27:15.270
as it seems like it does, because
my whole feed is right now Astro.

01:27:15.270 --> 01:27:18.890
Nobody's talking about Fable, except
when they're comparing to Astro,

01:27:19.210 --> 01:27:20.220
which is really, really funny.

01:27:20.330 --> 01:27:23.970
we covered world models, we covered,
audio and live transcription, as

01:27:23.970 --> 01:27:27.370
it's running right now with Meta
live transcription is now listening

01:27:27.370 --> 01:27:28.380
to the show and putting it on

01:27:29.160 --> 01:27:29.860
thursdayi.news/live.

01:27:30.730 --> 01:27:34.390
And if you missed any part of the
show, please check out the podcast

01:27:34.720 --> 01:27:36.770
and the newsletter on thursdayi.news.

01:27:36.940 --> 01:27:39.940
If you want to go and have the live
experience that Fable built, which

01:27:39.950 --> 01:27:43.505
Wolfram, I think after I get access to
Astra, I will try to build this from

01:27:43.505 --> 01:27:47.705
scratch again and see where we land, just
for funsies and how f- how fast this is.

01:27:47.835 --> 01:27:50.695
check out our new live experience
on Thursday at news.live.

01:27:50.815 --> 01:27:54.255
Huge thank you for our,
guests and co-hosts.

01:27:54.255 --> 01:27:56.655
We had a guest today talking
about Obliteration Model, the

01:27:56.695 --> 01:27:57.685
co-founder of Obliteration AI.

01:27:58.205 --> 01:28:01.725
we also are here, Niston was here
before, and Yam, they had to drop.

01:28:01.815 --> 01:28:05.965
we welcome back Ryan Carson, LDJ Wolfam,
and Peter, and all, most of all, thank

01:28:05.965 --> 01:28:12.855
to over all nearly 7,000 of you who tuned
in throughout this 5 hour stream, because

01:28:12.865 --> 01:28:17.075
this does feel like another moment in
time like GPT-4 did, for many of us,

01:28:17.105 --> 01:28:20.265
and GPT-5 and now in GPT-6 era in, 2026.

01:28:20.265 --> 01:28:21.425
So summer is officially over.

01:28:21.425 --> 01:28:25.685
If you missed any part of the show, I
will not promise that I will be able

01:28:25.685 --> 01:28:29.145
to edit down 5 hours, but, there's,
it's live on transcripts everywhere.

01:28:29.265 --> 01:28:29.815
Thank you so much.

01:28:29.815 --> 01:28:31.725
We'll see you here next week, hoping

01:28:31.885 --> 01:28:35.335
that we'll have some time to play with,
with Astra and, show you more examples.

01:28:35.645 --> 01:28:36.395
And, cheers.

01:28:36.425 --> 01:28:37.605
Thank you everybody for joining.

01:28:37.655 --> 01:28:38.055
Bye-bye.

01:28:45.775 --> 01:28:46.535
Check it.
