WEBVTT

1
00:00:00.120 --> 00:00:03.119
Alex Volkov: Hello, hello, and welcome everyone to Thursday AI.

2
00:00:03.720 --> 00:00:06.980
Today's September 17th. Can you believe it?

3
00:00:07.080 --> 00:00:11.819
September 17th. My name is Alex Volkov, an AI evangelist with weights and biases from

4
00:00:12.040 --> 00:00:14.880
CoreWeave and the host of Thursday AI.

5
00:00:14.960 --> 00:00:19.479
Hello and welcome everybody to September 17th Thursday AI edition.

6
00:00:19.960 --> 00:00:25.939
And, uh, as we gonna get started, there's an incredible amount of news, including

7
00:00:26.160 --> 00:00:31.840
a potentially game-changing new type of AI that's not quite an LLM that we can't wait

8
00:00:31.860 --> 00:00:34.719
to tell you about, called JEV from TypeSafe.

9
00:00:35.159 --> 00:00:37.840
I will add my co-host here to the stage.

10
00:00:38.000 --> 00:00:42.279
Welcome, Wolf from Ravenwolf. Welcome, Peter Gostev, and hello to everybody who's

11
00:00:42.300 --> 00:00:45.319
already tuning in. It's all about bringing AI changes to you.

12
00:00:45.480 --> 00:00:49.159
Good morning, Peter. I know it's morning for you in San Francisco.

13
00:00:49.320 --> 00:00:50.119
How are you doing, sir?

14
00:00:51.320 --> 00:00:56.420
Peter Gostev: I'm recovering from my flight. I was up for 24 hours last night, but I was up for

15
00:00:56.440 --> 00:01:00.800
about 6 hours now, so I think I'm functioning for the duration of this stream.

16
00:01:01.480 --> 00:01:01.819
Alex Volkov: Alrighty.

17
00:01:01.840 --> 00:01:02.840
Peter Gostev: After this, we'll see.

18
00:01:03.080 --> 00:01:06.619
Alex Volkov: And you rubbed shoulders yesterday with, with giants in the AI world.

19
00:01:06.920 --> 00:01:09.519
You attended the OpenAI Vibes party.

20
00:01:09.840 --> 00:01:10.319
How was that?

21
00:01:11.000 --> 00:01:15.819
Peter Gostev: So, from what I found out from OpenAI team is that they basically didn't know that

22
00:01:15.880 --> 00:01:20.340
it was happening until some tweeted saying, it would be nice to have a party, and

23
00:01:20.400 --> 00:01:24.199
then they're like, Sam what? Maybe mention it to us.

24
00:01:24.720 --> 00:01:28.379
So yeah, they just scrambled to organize this pretty incredible party.

25
00:01:28.440 --> 00:01:30.599
It was like in a science museum or something.

26
00:01:30.760 --> 00:01:33.559
Yeah, it was like a great venue. It was amazing.

27
00:01:33.680 --> 00:01:35.539
So yeah, props to the team like that.

28
00:01:35.760 --> 00:01:38.019
Alex Volkov: Can can we just talk about Sam's tweet super quick?

29
00:01:38.200 --> 00:01:44.219
Sam Altman tweeted, excited for ships this week, and then he had 6 ships, and we already

30
00:01:44.260 --> 00:01:48.519
have GPT-6 Astra, so hmm, what could it be with the number 6, I wonder?

31
00:01:48.800 --> 00:01:52.499
Uh, and then he posted, there's like 2 updates, and one of them is super exciting.

32
00:01:52.600 --> 00:01:58.840
Then we saw Thibault, the guy who resets all Codex quotas, uh, post or reply to somebody

33
00:01:59.160 --> 00:02:03.259
that's overwhelmed with personal AI assistant notifications from Muse, from Grok,

34
00:02:03.320 --> 00:02:06.560
and from Instinct, the 3 top ones that we're gonna discuss today.

35
00:02:07.080 --> 00:02:11.500
And Thibault replied with like, hmm, like emoji like this of when OpenAI is gonna

36
00:02:11.520 --> 00:02:13.159
step into the AI assistant game.

37
00:02:13.800 --> 00:02:16.340
Wolfram Ravenwolf: Peter Steinberger, after all, who started the craze, basically.

38
00:02:16.639 --> 00:02:18.080
Alex Volkov: So there's a reason why we are

39
00:02:18.120 --> 00:02:19.000
Wolfram Ravenwolf: still waiting for something.

40
00:02:19.240 --> 00:02:22.479
Alex Volkov: And so everybody's speculating that, hey, maybe OpenAI is launching their own, like,

41
00:02:22.639 --> 00:02:27.420
personal AI assistant, not just ChatGPT work that people don't actually know how to

42
00:02:27.480 --> 00:02:32.159
use but can do most of it, but no, like an AI system with dreaming, with memory MD

43
00:02:32.360 --> 00:02:35.740
files, et cetera, with with heartbeats, like all the primitives that Peter Steinberger

44
00:02:35.960 --> 00:02:40.919
and the folks came up with. And uh, I was really, really hoping that this is going

45
00:02:40.960 --> 00:02:43.499
to be the news from OpenAI, but hey, we're just speculating here.

46
00:02:43.560 --> 00:02:47.199
Hopefully we'll see them. And also Dev Day is in 2 weeks, and I,

47
00:02:48.360 --> 00:02:53.339
from what they say, the stuff that they're shipping right now is small potatoes compared

48
00:02:53.360 --> 00:02:56.020
to what they have planned for Dev Day, so I'm very, very excited.

49
00:02:56.320 --> 00:02:59.919
Yours truly is going to be there to cover Dev Day, by the way, so uh, me and Peter,

50
00:02:59.960 --> 00:03:03.620
we're going to walk around, I'm assuming, uh, and and uh, and tell you all about Dev

51
00:03:03.639 --> 00:03:07.360
Day. Meanwhile, let's add our other co-host, LDJ Nisten.

52
00:03:07.440 --> 00:03:11.000
Welcome, folks, to the September 17th ThursdAI stream.

53
00:03:11.440 --> 00:03:11.620
And

54
00:03:12.680 --> 00:03:13.279
we have,

55
00:03:14.400 --> 00:03:19.100
we had a very busy week. Last week, we told you all about

56
00:03:20.120 --> 00:03:26.099
this guy who left Entropic, and his tweet about leaving was seen by over 130 million

57
00:03:26.200 --> 00:03:31.360
people. A few days after, actually exactly a day after we told you about this, Dario

58
00:03:31.460 --> 00:03:34.839
Amodei, CEO of Entropic, uh, became

59
00:03:36.160 --> 00:03:38.120
the figurehead for pausing the frontier.

60
00:03:38.560 --> 00:03:44.520
Dario Amodei posted a very, uh, long essay about, uh, pacing the frontier, not pausing,

61
00:03:44.560 --> 00:03:47.119
sorry, I apologize. Let's delete all this.

62
00:03:47.720 --> 00:03:51.720
Pacing the frontier, not pausing. There's a very big difference between pacing and

63
00:03:51.800 --> 00:03:53.039
pausing, and uh-

64
00:03:53.080 --> 00:03:55.120
Wolfram Ravenwolf: It's a moving pause, just moving slowly.

65
00:03:55.240 --> 00:03:58.879
Alex Volkov: This is, yeah, moving slowly while others are catching up.

66
00:03:59.840 --> 00:04:03.039
And uh, this is a very, very interesting thing.

67
00:04:03.720 --> 00:04:09.779
Um, he proposed a three-part framework for pacing the frontier, including one embedded

68
00:04:09.880 --> 00:04:16.040
third-party evaluators in the labs, folks like METR, the folks that analyzed the swarm

69
00:04:16.560 --> 00:04:22.660
from OpenAI attacking Hugging Face, and other folks, um, with employee access, which

70
00:04:22.720 --> 00:04:28.000
they didn't have while they interviewed, uh, the OpenFace, uh, the the OpenAI Hugging

71
00:04:28.040 --> 00:04:33.579
Face incident. We have democratic lab coordination with antitrust cover, which we

72
00:04:33.640 --> 00:04:37.740
have to talk about this antitrust exemption they want, and also global coordination,

73
00:04:37.800 --> 00:04:43.880
including authoritarian governments, which is, we know exactly what

74
00:04:43.920 --> 00:04:45.559
they mean by authoritarian governments.

75
00:04:46.040 --> 00:04:51.540
So we definitely will discuss everybody jumping on the bandwagon since then, including

76
00:04:51.560 --> 00:04:53.639
Sam Altman, Elon Musk.

77
00:04:54.560 --> 00:04:59.839
On the other side of this, we have, uh, Jensen Huang, by the way, like a very strong

78
00:05:00.320 --> 00:05:05.699
anti, no, no, no, no, no pausing, and Donald Trump out of, out of nowhere, calling

79
00:05:05.760 --> 00:05:07.380
all of this a hoax. It's very interesting.

80
00:05:07.839 --> 00:05:10.740
Wolfram Ravenwolf: Never have guessed that I would get behind one of his positions.

81
00:05:11.040 --> 00:05:16.180
Alex Volkov: I, dude, it's very complex, and as I said, ThursdAI is not a political show, but when

82
00:05:16.360 --> 00:05:18.839
AI goes into politics, we at least need to cover what's going on.

83
00:05:20.279 --> 00:05:22.079
So this is going to be theme number one.

84
00:05:22.760 --> 00:05:28.260
See, theme number two for this week is something at lunch yesterday, but actually

85
00:05:28.400 --> 00:05:32.480
folks from Weights and Biases and CoreWeave, who listened to the show and came to

86
00:05:32.520 --> 00:05:34.959
the hackathon, got to experience this before the drop.

87
00:05:35.400 --> 00:05:41.399
Uh, TypeSafe AI released JEV, which is a new type of AI model that's not a

88
00:05:41.640 --> 00:05:46.119
standard encoder-decoder based transformer that spits tokens sequentially.

89
00:05:46.960 --> 00:05:50.800
And I don't know about y'all, but my timeline is all JEV.

90
00:05:52.360 --> 00:05:54.759
Everybody who I, like,

91
00:05:55.800 --> 00:05:56.600
everybody who I, like,

92
00:05:57.640 --> 00:06:01.679
see on my timeline is all Jeff. Nisten is going no, and I think it's because you haven't

93
00:06:01.760 --> 00:06:06.379
interacted with one or two tweets. I think the algorithm is overobsessed in showing

94
00:06:06.440 --> 00:06:08.540
you specific things that you have interacted with.

95
00:06:09.480 --> 00:06:15.320
Um, so we are going to actually reach out, and we're going to have Ali from

96
00:06:16.400 --> 00:06:19.920
TypeSafe, the company that brought Jeff to the world, come on the show.

97
00:06:20.440 --> 00:06:22.519
If you haven't heard about JEV,

98
00:06:24.200 --> 00:06:26.120
how should I describe this? The

99
00:06:27.040 --> 00:06:32.380
co-creator of ChatGPT and the co-creator of RLHF, reinforcement learning with human

100
00:06:32.440 --> 00:06:34.679
feedback, Diogo Almeida, is,

101
00:06:35.680 --> 00:06:37.100
it's his lab for the past two years.

102
00:06:37.120 --> 00:06:41.019
They've been in stealth, and they released their first model, uh, they call it System

103
00:06:41.080 --> 00:06:47.120
1 model, that is basically a classifier, but a very smart classifier at the level

104
00:06:47.160 --> 00:06:51.159
of LLMs, but it's significantly faster and significantly cheaper.

105
00:06:51.839 --> 00:06:54.999
And by significantly, I mean you will not believe the speed

106
00:06:56.000 --> 00:06:57.360
and the cost,

107
00:06:58.360 --> 00:07:02.239
but since this is not an LLM, it's also deterministic

108
00:07:03.160 --> 00:07:07.360
or probabilistic. And so you can build new things that you we've previously used cheap

109
00:07:07.400 --> 00:07:09.879
LLMs for, and they're going to be incredibly cheap.

110
00:07:10.400 --> 00:07:13.299
And I'm super excited to show you some examples that I already made.

111
00:07:13.360 --> 00:07:15.800
I only got access last night. Wolfram, I think you got access as well.

112
00:07:16.080 --> 00:07:17.899
Wolfram Ravenwolf: I used it for the show preparation, actually.

113
00:07:18.040 --> 00:07:20.639
The news I sent you, they have been classified by Jeff.

114
00:07:21.040 --> 00:07:23.360
Alex Volkov: Yes, so classification and other things.

115
00:07:23.560 --> 00:07:29.800
Um, and I think the third theme for this week is going to be voice again, yet again.

116
00:07:30.000 --> 00:07:34.579
And voice, we will plug in personal AI assistants into voice because many of them

117
00:07:34.620 --> 00:07:37.000
are getting a voice. Many of them can now call businesses.

118
00:07:37.160 --> 00:07:39.999
Muse can call businesses, Instinct can call businesses.

119
00:07:40.440 --> 00:07:45.980
Um, so it's kind of like a AI assistants getting to be a thing plus voice is gonna

120
00:07:46.020 --> 00:07:46.800
be the third theme.

121
00:07:48.160 --> 00:07:50.959
Any other themes on your guys' mind before we go?

122
00:07:51.720 --> 00:07:53.899
Theme 1, theme 2, theme 3, any other comments?

123
00:07:53.920 --> 00:07:59.079
Wolfram Ravenwolf: I see video still moving so fast and becoming even faster, the video generation.

124
00:07:59.120 --> 00:08:03.120
We have faster than, uh, the new ones for H3 there.

125
00:08:03.320 --> 00:08:06.899
Since it's open source, a lot of people are now working on this, and we see it pop

126
00:08:06.940 --> 00:08:10.339
up left and right. We are getting voice now for our assistants, and soon we will have

127
00:08:10.480 --> 00:08:12.399
live video generation for that as well.

128
00:08:12.640 --> 00:08:12.980
Alex Volkov: Mm.

129
00:08:14.040 --> 00:08:16.500
Uh, LDJ, how about you? Anything interesting from your end?

130
00:08:17.920 --> 00:08:22.120
LDJ: I don't think we talked about the robotic control abilities of Astra last week, did

131
00:08:22.160 --> 00:08:22.299
we?

132
00:08:23.120 --> 00:08:23.839
Alex Volkov: Uh, no.

133
00:08:24.400 --> 00:08:27.600
LDJ: Okay, yeah. I think that's something that came out a couple days after Thursday AI

134
00:08:27.720 --> 00:08:31.939
of last week, where there's a couple different benchmarks and a comprehensive write-up

135
00:08:32.000 --> 00:08:34.959
we can go over later, but yeah, it's really interesting.

136
00:08:35.000 --> 00:08:40.220
It's beating a lot of the specialized robotic models in RoboDojo and RoboLab and a

137
00:08:40.280 --> 00:08:42.099
bunch of different tasks like stacking blocks,

138
00:08:42.280 --> 00:08:42.419
Alex Volkov: Yeah.

139
00:08:42.640 --> 00:08:47.179
LDJ: flipping over objects, and and it's really just a matter of speed now, which hopefully

140
00:08:47.280 --> 00:08:51.159
Nisten, uh, is going to be showing us things that are really fast later.

141
00:08:52.800 --> 00:08:54.559
Alex Volkov: All right, uh, you guys are,

142
00:08:55.559 --> 00:08:57.140
you guys are dropping hints, and I love it.

143
00:08:57.480 --> 00:09:02.620
Wolfram Ravenwolf: I have one topic. I mean, I have so many this week, obviously the video stuff, the

144
00:09:02.679 --> 00:09:06.619
audio stuff, but for me, the most important, and I think for everybody, the most important,

145
00:09:06.700 --> 00:09:11.100
everybody in AI, is what Trump said, that the doomerism is a hoax, and he's not going

146
00:09:11.160 --> 00:09:15.620
along with all the stuff. So I'm not even an American, but I have to say this.

147
00:09:15.640 --> 00:09:19.399
Alex Volkov: I would just, uh, let me just add, it's not that Trump just said this.

148
00:09:20.040 --> 00:09:21.759
Trump called Jensen Huang

149
00:09:21.960 --> 00:09:22.139
Wolfram Ravenwolf: Yeah.

150
00:09:22.320 --> 00:09:27.299
Alex Volkov: live on stage, and Jensen took the call at the All In Summit while Jensen was sitting

151
00:09:27.320 --> 00:09:31.139
there and said, Mr. President, if it was, if this wasn't you, I would not have answered,

152
00:09:31.160 --> 00:09:33.119
but you're live in front of like a bunch of people right now.

153
00:09:33.360 --> 00:09:37.320
And then, uh, uh, you know, President of the United States Donald Trump yelled at

154
00:09:37.380 --> 00:09:41.600
everybody and said that this is another hoax like Russia, AI is not going to take

155
00:09:41.640 --> 00:09:41.880
over.

156
00:09:43.280 --> 00:09:48.299
I, you had to laugh when you heard this, but yeah, this is definitely a, a battlegrounds

157
00:09:48.400 --> 00:09:53.140
bidding being drawn, battleground lines being drawn about who stands where on pausing

158
00:09:53.200 --> 00:09:54.139
on data centers.

159
00:09:54.160 --> 00:09:57.539
Wolfram Ravenwolf: I completely don't agree with the other stuff he's calling hoaxes, that those are

160
00:09:57.640 --> 00:10:02.620
hoaxes. I'm not getting into that, but that he's calling out doomerism and effective,

161
00:10:02.760 --> 00:10:05.119
effective altruism, I'm all behind that.

162
00:10:05.200 --> 00:10:09.560
And so this has been the most important news for AI, I think, that we are not just

163
00:10:09.720 --> 00:10:11.339
over-regulating like over here in Europe.

164
00:10:11.760 --> 00:10:17.740
Alex Volkov: All right, so let's go, I think let's go to TLDR so that we will run through all of

165
00:10:17.840 --> 00:10:22.039
the kind of notes that we have and specific releases, stuff that we may not even be

166
00:10:22.080 --> 00:10:27.039
able to fully go to, and then we can switch to talking about the pausing thing.

167
00:10:28.120 --> 00:10:49.920
So let's go to TLDR,

168
00:10:50.040 --> 00:10:52.899
if you will, from the week of September 17th in 2026.

169
00:10:53.560 --> 00:10:58.420
We will start talking about this pacing the frontier debate has hit the, how should

170
00:10:58.460 --> 00:11:03.539
I say, the tops of the frontier labs, with Dario Amodei releasing an essay about pacing

171
00:11:03.560 --> 00:11:06.159
the frontier and saying that Anthropic is committed to pacing it.

172
00:11:06.520 --> 00:11:12.359
And, uh, very, very interestingly, Sam Altman agreed, Elon Musk agreed, uh,

173
00:11:12.760 --> 00:11:18.960
and David Sacks, Demis Hassabis, and Mustafa Suleyman all point, like, and plan positions

174
00:11:19.120 --> 00:11:21.480
in where they are on this debate. It's very interesting.

175
00:11:21.680 --> 00:11:24.919
Uh, we have a live map to show you where everybody is.

176
00:11:25.240 --> 00:11:27.840
Uh, would love to see and talk about pacing the frontier.

177
00:11:28.160 --> 00:11:33.360
Uh, it's, I think, for the first time they've announced that OpenAI paused training

178
00:11:33.680 --> 00:11:36.140
or Gale training back, and also Anthropic.

179
00:11:36.200 --> 00:11:38.760
So it's very interesting to see who talks about what.

180
00:11:39.920 --> 00:11:45.939
Uh, let's see what else. Uh, the other big news of the big labs, or frontier AI,

181
00:11:46.000 --> 00:11:51.960
if you will, uh, TypeSafe AI debuts GeV, which is a non-LLM system one decision model,

182
00:11:52.160 --> 00:11:57.639
uh, from X, OpenAI, RLHF and ChatGPT lead Diogo Almeida.

183
00:11:57.919 --> 00:12:02.919
We will have Ali from the DevRel team at TypeSafe come to us and tell us all about

184
00:12:02.960 --> 00:12:06.400
this model, uh, very soon. We also have a bunch of demos, and we'll have Francesco

185
00:12:06.440 --> 00:12:11.920
from Kua talk about how this affects computer use, which often uses LLMs, and uh,

186
00:12:12.680 --> 00:12:15.240
this may use this decision model instead.

187
00:12:15.800 --> 00:12:18.639
The amount of incredibly exciting

188
00:12:20.800 --> 00:12:25.099
capabilities and incredibly exciting unlocks in this new primitive in AI is very,

189
00:12:25.200 --> 00:12:29.319
very big based on my timeline and excitement, uh, so I can't wait to talk to you about,

190
00:12:29.560 --> 00:12:35.700
uh, JEV specifically. We also looking, are looking at the era of the AI assistant,

191
00:12:36.160 --> 00:12:39.639
not the AI agent, the AI agent that's assistant focused, right?

192
00:12:39.680 --> 00:12:43.639
So we started the wor- the the the year with OpenClaw and then moved to Hermes, and

193
00:12:43.720 --> 00:12:49.379
then, uh, you know, uh, recently the highlights of those efforts is Grokbot from xAI,

194
00:12:49.639 --> 00:12:54.700
folks from Cursor built it, and then Muse from Meta, rumored something from OpenAI,

195
00:12:54.880 --> 00:12:58.199
not clear what, I don't have, uh, previous news, and Instinct.

196
00:12:58.700 --> 00:13:01.400
Instinct, if you haven't heard, it's like the VC darling.

197
00:13:01.520 --> 00:13:04.499
It's a startup by a 24-year-old Noah...

198
00:13:06.200 --> 00:13:08.580
I don't have the last name, I apologize.

199
00:13:08.840 --> 00:13:13.439
Uh, okay, so Instinct targets a 10 billion dollar valuation

200
00:13:14.360 --> 00:13:16.479
after being incorporated in April.

201
00:13:17.880 --> 00:13:19.199
I will say this slowly again.

202
00:13:20.120 --> 00:13:23.159
Instinct targets a 10 billion dollar valuation

203
00:13:24.320 --> 00:13:29.639
for a free product that people don't pay for after being incorporated in, uh,

204
00:13:31.120 --> 00:13:36.180
in April, which is insane. Okay, we'll talk about this, but the rise of the personal

205
00:13:36.280 --> 00:13:39.780
AI assistant is finally here, and we've been testing a bunch of them.

206
00:13:39.880 --> 00:13:45.439
We'll talk about this. We'll also have a guest on the show from Assistant Benchmarks.

207
00:13:45.840 --> 00:13:51.160
So I'm very excited to introduce you guys to David Paulan from Assistant Benchmarks.

208
00:13:51.280 --> 00:13:55.659
He's been testing out, and every major lab is now like looking at assistantbenchmark.com

209
00:13:55.680 --> 00:14:00.880
as well. So we'll, we'll chat about David, with David in about an hour or so to talk

210
00:14:00.920 --> 00:14:04.199
about the rise of assistants and different tasks that they do differently than, like,

211
00:14:04.279 --> 00:14:06.819
ChatGPT. Uh, let's see what else. Uh,

212
00:14:08.120 --> 00:14:08.240
the

213
00:14:09.200 --> 00:14:13.999
live voice stuff is popping up. So OpenAI launched their, this is not it, OpenAI launched

214
00:14:14.040 --> 00:14:18.959
their GPT Live 1 voice that powers the live experience on ChatGPT in API.

215
00:14:19.000 --> 00:14:20.680
So you can now build in

216
00:14:21.880 --> 00:14:25.759
stuff like voice, immediate voice conversations into your

217
00:14:26.680 --> 00:14:29.279
agents that you're building yourself with the GPT Live 1.

218
00:14:29.600 --> 00:14:35.700
And also Google launched Gemini 3.8 Live and, uh, 3.8 Live Extended Thinking, which

219
00:14:35.839 --> 00:14:40.999
is also very big news for folks who are building with, with agents or want to make

220
00:14:41.080 --> 00:14:42.119
their agents sound good.

221
00:14:43.920 --> 00:14:44.839
We also have a

222
00:14:45.800 --> 00:14:48.600
state-of-the-art kind of model as well.

223
00:14:48.720 --> 00:14:54.440
StepFun launches Step3, a 5 model family that tops artificial analysis score voice,

224
00:14:54.600 --> 00:14:56.519
so a lot of voice releases this week.

225
00:14:57.040 --> 00:14:59.639
And, uh, let's see what else interesting.

226
00:15:01.279 --> 00:15:05.620
Oh yeah, and maybe this. Union Alpha is a new anonymous stealth model for free on

227
00:15:05.680 --> 00:15:07.639
OpenRouter. We already heard from this before.

228
00:15:08.000 --> 00:15:09.540
Uh, we heard things like this before.

229
00:15:09.640 --> 00:15:13.720
Union Alpha seems to be very interesting, and we'll maybe mention that as well.

230
00:15:14.680 --> 00:15:15.199
Uh...

231
00:15:17.080 --> 00:15:20.560
Anything huge that we've missed? I know there's a bunch, like, folks, at this point,

232
00:15:20.920 --> 00:15:25.220
we are at the curation game and not the, how should I say, comprehensive coverage

233
00:15:25.279 --> 00:15:29.139
game. It's no longer possible to cover everything that happened, but we want to tell

234
00:15:29.160 --> 00:15:32.899
you about the stuff that excites us and hopefully the stuff that excites you, the

235
00:15:32.960 --> 00:15:37.439
audience, and also highlight the things that you absolutely must not miss, like JEV,

236
00:15:37.839 --> 00:15:41.720
potentially Union Alpha, like the rise of personal AI assistants, and like the pacing

237
00:15:41.760 --> 00:15:46.560
the frontier debate. Let's see, folks on stage, anything else huge, important that

238
00:15:46.839 --> 00:15:48.240
was not mentioned, discussed?

239
00:15:48.880 --> 00:15:53.639
Peter Gostev: I'm curious, I actually haven't had a chance to try it, but have you guys tried Muse,

240
00:15:54.000 --> 00:15:58.540
uh, products? Because, uh, I'm hearing a lot of good things, and I think that that

241
00:15:58.600 --> 00:16:03.879
is kind of Nat Friedman's, uh, baby, right, shipping, and people comment on the quality

242
00:16:03.920 --> 00:16:09.259
of it. So I, I, I was away for a few days, I, I didn't have a chance to try, but yeah,

243
00:16:09.360 --> 00:16:12.760
that, that would be interesting just to see, kind of, I know it maybe didn't come

244
00:16:12.800 --> 00:16:16.040
out this week, but I think that it needs a bit of time to percolate.

245
00:16:16.400 --> 00:16:21.119
Alex Volkov: Uh, Muse is incredible and is my number one,

246
00:16:22.279 --> 00:16:28.420
and I went on a full AI psychosis weekend where I got tired of having

247
00:16:28.760 --> 00:16:34.439
Grokbots, like 17 of them at this point, everybody with their own job, uh, Muse and

248
00:16:34.600 --> 00:16:39.240
Instinct, and so I figured out how to make them all collaborate in a way that I can

249
00:16:39.360 --> 00:16:43.479
actually control via linear. I'm very happy to talk about this if we have the time.

250
00:16:43.760 --> 00:16:48.999
However, Muse is absolutely the top one for regular people.

251
00:16:49.200 --> 00:16:49.540
I don't

252
00:16:51.360 --> 00:16:55.760
consider anybody who listens to ThursdAI a regular person in terms of AI, because

253
00:16:55.880 --> 00:16:57.120
if you are listening, you're at

254
00:16:58.400 --> 00:17:02.040
the cutting edge of AI news. However, Muse is for my mom.

255
00:17:02.560 --> 00:17:04.319
Muse is for grandmas around the world.

256
00:17:04.400 --> 00:17:08.759
Muse is for people who don't really understand what AI is, and the amount of work

257
00:17:09.160 --> 00:17:14.080
Nat Friedman, Alex Wang, and the folks there, Arik as well has been getting feedback

258
00:17:14.120 --> 00:17:15.480
directly from me and like implementing this.

259
00:17:15.800 --> 00:17:19.399
The amount of work they put in this makes it just like a such a usable product.

260
00:17:19.480 --> 00:17:20.319
It's quite crazy.

261
00:17:21.720 --> 00:17:26.079
Um, and Muse is currently US only. That's why folks were listening to us tuning in

262
00:17:26.120 --> 00:17:28.520
from Europe, tuning in from Canada even.

263
00:17:28.600 --> 00:17:32.740
I think Nisten, some of your compatriots are saying that no Muse for me and Nisten

264
00:17:32.760 --> 00:17:35.879
in Canada. I think they're working on enabling this across the world.

265
00:17:36.000 --> 00:17:37.959
Yam, I don't believe that Israel also has Muse, right?

266
00:17:38.000 --> 00:17:39.439
You cannot download Muse in Israel yet.

267
00:17:39.960 --> 00:17:45.000
Literally, my mom cannot use Muse until it does land in Israel, but it is quite incredible.

268
00:17:45.160 --> 00:17:45.500
I have

269
00:17:47.200 --> 00:17:49.639
quite a few things to talk about this week about Muse.

270
00:17:49.880 --> 00:17:54.859
Uh, a specific one, Peter, since you asked, Muse rolls out invite codes, and Muse

271
00:17:54.960 --> 00:17:59.579
also rolls out voice calling. So I actually booked a haircut appointment with Muse,

272
00:17:59.640 --> 00:18:02.320
who called the barbershop, and I have a transcript of that.

273
00:18:02.360 --> 00:18:04.720
This is really cool. We could definitely show this here.

274
00:18:05.600 --> 00:18:09.560
Uh, but we must start, and Wolfram, I think you suggested this order.

275
00:18:09.680 --> 00:18:14.399
We must start with pacing before we get to the super exciting let's go stuff.

276
00:18:14.800 --> 00:18:14.980
Uh,

277
00:18:15.000 --> 00:18:17.779
Wolfram Ravenwolf: Yeah, I have another item I just put in the chat.

278
00:18:17.820 --> 00:18:21.779
It's OpenAI's model misalignment framework for reporting about model misalignment,

279
00:18:21.880 --> 00:18:26.960
but we can cover it as part of the pausing or pacing stuff because it's also relevant.

280
00:18:27.000 --> 00:18:31.259
Alex Volkov: 100%. Uh, and I believe that this was literally from yesterday, Wolfram, right?

281
00:18:31.400 --> 00:18:33.320
Uh, framework for reporting model misalignment.

282
00:18:33.880 --> 00:18:34.659
Uh, all right.

283
00:18:34.679 --> 00:18:35.379
Wolfram Ravenwolf: Yes.

284
00:18:35.400 --> 00:18:39.940
Alex Volkov: Let's start with Frontier AI. Let's talk a little bit about what's going on with all

285
00:18:40.320 --> 00:18:41.080
heads of labs

286
00:18:42.040 --> 00:18:45.960
saying, or at least the Frontier Labs saying pause, and let's let's break it down

287
00:18:46.000 --> 00:18:49.340
and see where everybody's at. I think it's going to be very interesting, and in the

288
00:18:49.480 --> 00:18:51.360
Frontier Labs race.

289
00:19:03.040 --> 00:19:05.200
Alrighty, folks, uh, let's get to it.

290
00:19:05.400 --> 00:19:06.600
Since we last updated you

291
00:19:08.400 --> 00:19:14.399
about Jacob Coxon leaving Anthropic and posting a tweet after just, what,

292
00:19:14.480 --> 00:19:14.959
6 weeks,

293
00:19:16.000 --> 00:19:22.259
that's been seen for, with over 130 million people and all of the politicians jumping

294
00:19:22.280 --> 00:19:27.299
on this, Bernie Sanders, etcetera. Uh, things have moved and changed in the world

295
00:19:27.380 --> 00:19:31.739
of, uh, AI frontier labs. Dario Amodei released an essay

296
00:19:32.679 --> 00:19:37.779
in, uh, September 12th, so Friday after we talked to you, that outlines the express

297
00:19:37.840 --> 00:19:43.100
need of pacing the development of frontier AI, pacing specifically so that we'd be

298
00:19:43.140 --> 00:19:46.199
able to, uh, look into how to control this.

299
00:19:47.280 --> 00:19:52.459
Daria specifically proposed a three-step process, starting with embedded- embedding

300
00:19:52.600 --> 00:19:59.060
third-party evaluators like METR and, uh, Redwood Research, the folks who analyzed

301
00:19:59.960 --> 00:20:05.719
the OpenAI Hugging Face incident and the hacking swarm, into the labs with employee

302
00:20:05.760 --> 00:20:10.420
level access, so the third party evaluators will be able to look at model development,

303
00:20:10.560 --> 00:20:16.639
traces, etcetera. Uh, Dari also suggested democratic lab coordination between,

304
00:20:16.960 --> 00:20:23.000
uh, between labs, specifically with a, with an asterisk of an antitrust cover.

305
00:20:23.400 --> 00:20:28.200
We have to talk about this antitrust cover thing because if you think about capitalism,

306
00:20:28.680 --> 00:20:29.560
the two leading

307
00:20:31.160 --> 00:20:36.780
companies coordinating their prices, releases, dates, etcetera, not really

308
00:20:37.760 --> 00:20:43.079
like a lawful thing. So they really want like an antitrust exemption so that they

309
00:20:43.160 --> 00:20:45.560
would be able to collaborate, but only them.

310
00:20:45.760 --> 00:20:51.599
That's very interesting. Uh, and also a global coordination with authoritarian governments.

311
00:20:52.840 --> 00:20:55.080
And Anthropic committed unilaterally to step one.

312
00:20:55.400 --> 00:20:57.019
Anthropic committed to adding

313
00:20:58.480 --> 00:21:02.840
third-party evaluators to their lab with employee level access.

314
00:21:04.360 --> 00:21:06.680
Wolfram Ravenwolf: There should be an asterisk with third party, I guess.

315
00:21:07.039 --> 00:21:09.740
Alex Volkov: Yeah, we have to, we have to talk about this asterisk in, in a moment, but

316
00:21:09.880 --> 00:21:10.260
Nisten Tahiraj: Yeah.

317
00:21:10.280 --> 00:21:14.319
Alex Volkov: I would just say, just before this, immediately, almost immediately, the following

318
00:21:14.400 --> 00:21:14.799
folks.

319
00:21:16.200 --> 00:21:22.180
Sam Altman. If you remember Sam Altman and Dario Amodei in India, standing next

320
00:21:22.220 --> 00:21:23.719
to each other and not touching hands.

321
00:21:24.320 --> 00:21:30.039
Those two people who reportedly hate each other, or at least there's a lot of animosity.

322
00:21:30.200 --> 00:21:35.060
As a reminder, Dario Amodei used to work in OpenAI, disagreed with the safety directions,

323
00:21:35.160 --> 00:21:37.599
and then left and built a competing company, Anthropic.

324
00:21:38.120 --> 00:21:41.680
Uh, and then also Elon Musk agreed,

325
00:21:42.600 --> 00:21:46.360
which Elon Musk and Dario, okay, but Elon Musk and Sam agree, and that's kind of very

326
00:21:46.400 --> 00:21:51.439
interesting. Uh, and then we heard from other neo-frontier labs, let's call them,

327
00:21:51.480 --> 00:21:53.160
the folks who are just joining the frontier.

328
00:21:53.480 --> 00:21:57.900
Uh, Meta, Meta is a frontier lab at this point with Muse, or going to be very, very,

329
00:21:57.920 --> 00:22:00.000
very soon, especially with uh, Muse Spark.

330
00:22:00.520 --> 00:22:01.099
And uh,

331
00:22:02.160 --> 00:22:06.560
and yeah, Elon Musk is is the third one, and we barely, barely heard from Gemini or

332
00:22:06.720 --> 00:22:09.479
Google, which puts them in a very interesting space.

333
00:22:09.840 --> 00:22:14.519
Uh, so while I pull up the map that Muse created for me, folks, I would love to hear

334
00:22:14.919 --> 00:22:17.059
just reactions to Dario's essay. Have you read this?

335
00:22:17.160 --> 00:22:21.020
What are your thoughts on this? And then also let's talk about the agreement between

336
00:22:21.400 --> 00:22:24.300
those very competitive companies. Uh, who wants to go first?

337
00:22:24.320 --> 00:22:28.379
Nisten, you seem like you heated. I want to hear from you about the pacing the frontier

338
00:22:28.400 --> 00:22:28.999
from Daria.

339
00:22:31.000 --> 00:22:33.240
Nisten Tahiraj: I'm not going to even read it.

340
00:22:34.160 --> 00:22:37.039
It's, uh, it's, it's pretty pointless at the

341
00:22:38.080 --> 00:22:38.739
Alex Volkov: Interesting.

342
00:22:38.760 --> 00:22:43.400
Nisten Tahiraj: At this point. Uh, they're more worried about litigation because they do want to keep

343
00:22:43.440 --> 00:22:49.340
selling the LLMs, but they might get sued if, uh, the LLMs go and

344
00:22:49.880 --> 00:22:55.379
hack someone. But if I did the same thing and my local LLM hacked the hospital here,

345
00:22:55.919 --> 00:22:57.059
I might actually go to jail.

346
00:22:58.000 --> 00:23:04.020
So, in order to shift responsibility, they go for a regulatory capture, and it's

347
00:23:04.080 --> 00:23:10.080
not going to work. Uh, it looks like the decision makers in China are all engineers,

348
00:23:10.320 --> 00:23:16.340
and funny enough, they seem more libertarian and accelerationist, and it's kind of

349
00:23:16.640 --> 00:23:21.620
interesting that in the West, we are the ones that are becoming more bureaucratic

350
00:23:21.720 --> 00:23:24.419
and more authoritarian in that point.

351
00:23:24.560 --> 00:23:28.720
There's a split in there. So I don't think it matters.

352
00:23:29.040 --> 00:23:31.819
It's not going to stop the Chinese models.

353
00:23:32.040 --> 00:23:37.240
It's not going to stop people from, uh, running security tools on their own, and,

354
00:23:37.400 --> 00:23:43.099
uh, they should focus on actually curing some disease and automating housing

355
00:23:43.720 --> 00:23:45.519
instead of trying to form a cartel.

356
00:23:46.640 --> 00:23:48.879
Yeah, yeah, not n- not a fan of this at all.

357
00:23:49.440 --> 00:23:51.279
Alex Volkov: So there's a few things there, right?

358
00:23:51.400 --> 00:23:57.419
Uh, the the first thing is we we need to understand that there's folks who are asking

359
00:23:57.440 --> 00:24:02.160
for a pause, like pause AI people. Uh, the heads of lab are not talking about pausing,

360
00:24:02.520 --> 00:24:08.320
they're talking about pacing, which is, uh, giving enough time for themselves and

361
00:24:08.360 --> 00:24:13.359
the labs to control the potentially uncontrollable, like, ASI that's coming out.

362
00:24:13.760 --> 00:24:19.660
Uh, we also, um, I- I wanna highlight the map that Muse built for me, uh,

363
00:24:19.839 --> 00:24:22.180
with, uh, where everybody is on the kind of spectrum.

364
00:24:22.240 --> 00:24:26.700
So Dario Amodei started with, we must slow the pace at which we improve the capabilities

365
00:24:26.740 --> 00:24:31.720
of AI to be able to catch up. Uh, not building the technology deprives humanity of

366
00:24:31.800 --> 00:24:35.540
benefits or simply places AI in the hands of authoritarian powers, while building

367
00:24:35.580 --> 00:24:40.160
it too fast is reckless. This is a quote from We Must Pace the Frontier, uh, essay

368
00:24:40.200 --> 00:24:44.320
from Dario Amodei. And then immediately after him, Sam Altman said, I agree with Dario

369
00:24:44.339 --> 00:24:47.899
that we need to pace the frontier. This has been a primary topic of discussion we've

370
00:24:47.920 --> 00:24:52.899
had at OpenAI in recent weeks. As a reminder, folks, many of the employees in these

371
00:24:52.960 --> 00:24:56.039
big labs, the folks who are building, researching those technologies, have signed

372
00:24:56.080 --> 00:24:58.319
the Pace the Frontier letter, uh,

373
00:24:59.240 --> 00:25:00.440
asking the

374
00:25:02.040 --> 00:25:07.819
top frontier labs to consider pacing so that they have a chance to talk about

375
00:25:08.400 --> 00:25:12.639
the acceleration wave in AI. Uh, Nisten, you bring up a very, very good point.

376
00:25:12.760 --> 00:25:16.680
Where is, you know, uh, where's the Chinese labs on, uh, on the spectrum?

377
00:25:16.880 --> 00:25:21.320
We don't know, but we definitely know where our US president is.

378
00:25:22.040 --> 00:25:26.439
Uh, I want to read this quote in full, and then LDJ would love to hear from you.

379
00:25:26.839 --> 00:25:27.559
Uh, as

380
00:25:28.839 --> 00:25:30.959
Jensen Huang was on stage with the

381
00:25:32.080 --> 00:25:33.559
All-In Pod, uh,

382
00:25:34.839 --> 00:25:39.679
cast in in the All-In, um, summit, Jensen Huang got a call.

383
00:25:39.960 --> 00:25:43.339
It was very interesting to see Jensen, somebody as big as Jensen, the biggest company

384
00:25:43.380 --> 00:25:47.899
in the world, NVIDIA, uh, providing the the rise of AI everywhere, getting a call

385
00:25:47.920 --> 00:25:51.760
and answering on stage, and this was Donald Trump, uh, and that's what he said.

386
00:25:52.480 --> 00:25:56.220
They are just playing right into the hands of a lot of people that don't want to see

387
00:25:56.279 --> 00:25:59.179
it happen. That could be political people, it could be also China.

388
00:25:59.240 --> 00:26:01.199
We're not going to let it happen. It's a hoax.

389
00:26:01.720 --> 00:26:07.079
Uh, Donald Trump rejects the slowdown calls as helping China and casts AI dominance

390
00:26:07.160 --> 00:26:10.839
as national competition. Jensen said, we're not gonna let it happen, sir.

391
00:26:10.960 --> 00:26:17.019
And also Jensen said, if AI wins, whoe- uh, no, sorry, Donald Trump says, whoever

392
00:26:17.120 --> 00:26:18.360
wins AI wins.

393
00:26:19.640 --> 00:26:20.759
This is such a powerful

394
00:26:21.720 --> 00:26:25.159
s- statement in four words. Whoever wins AI wins.

395
00:26:26.040 --> 00:26:27.819
And also very true. LDJ, please go ahead.

396
00:26:27.880 --> 00:26:31.940
Uh, what are your thoughts on where folks stand and and how how we're seeing this,

397
00:26:32.000 --> 00:26:34.679
uh, pacing the frontier collaboration between the labs?

398
00:26:35.839 --> 00:26:36.360
LDJ: Yeah, so

399
00:26:37.320 --> 00:26:42.699
to be clear first, I do disagree with Dario on a lot of things, but when it comes

400
00:26:42.740 --> 00:26:46.500
to the contents of this blog, which I, I do feel like it's important those critiquing

401
00:26:46.540 --> 00:26:52.679
it would, would read, um, it's, it does specifically outline things

402
00:26:52.800 --> 00:26:58.680
relating to the fact that we should keep a gap beyond China, and we should stay, uh,

403
00:26:58.720 --> 00:27:03.459
more advanced than China. And I think that that's one of the big important things

404
00:27:03.560 --> 00:27:04.999
that people critiquing it are overlooking.

405
00:27:05.240 --> 00:27:07.740
They're, they're like, oh, we're gonna slow down, we're just gonna pause, we're gonna

406
00:27:07.780 --> 00:27:12.379
let China surpass us. And he's saying, like, explicitly, like, no, we should not let

407
00:27:12.400 --> 00:27:18.800
that happen, and even propose a lot of ways we can try to stay ahead while maintaining

408
00:27:19.000 --> 00:27:24.439
theoretical, uh, like, trying to define specific metrics for RSI,

409
00:27:25.080 --> 00:27:29.999
trying to define exactly what rates of RSI improvement you can define as some type

410
00:27:30.060 --> 00:27:34.779
of speed limits while still maintaining a lead above China and having some type of

411
00:27:34.880 --> 00:27:40.259
limit that's, that would put, still allow American companies to be ahead of China,

412
00:27:40.280 --> 00:27:44.379
if that makes sense. And I feel like there is some practical things beyond that there,

413
00:27:44.480 --> 00:27:50.720
such as having some type of frameworks and regulations around stricter laws

414
00:27:50.800 --> 00:27:55.999
of using AI for, let's say, let's say you were to kill somebody using an AI model,

415
00:27:56.520 --> 00:28:01.000
right? Like, there's already laws like murder relating to that, but arguably, if,

416
00:28:01.360 --> 00:28:05.160
if you use that tool to do something on like a grander scale,

417
00:28:06.280 --> 00:28:06.520
then

418
00:28:07.640 --> 00:28:12.479
some more specific harsh laws relating to that for that situation, which I think it's

419
00:28:12.559 --> 00:28:13.059
pretty reasonable.

420
00:28:13.080 --> 00:28:18.859
Alex Volkov: LDJ, sorry to interrupt because when OpenAI swarm of agents hacked Hugging Face, nobody

421
00:28:19.000 --> 00:28:25.299
went to jail, nobody sued anyone, and it wasn't clear who's, who who's taking, uh,

422
00:28:25.679 --> 00:28:27.499
ownership of that. And

423
00:28:28.720 --> 00:28:33.019
this is like, this is a offense, like, like as a person, if I hack Hugging Face, I

424
00:28:33.080 --> 00:28:37.639
go to jail. But if my swarm of agents and a big company, no, no, no, nobody, like,

425
00:28:37.720 --> 00:28:40.579
that's a very interesting gap that we're having right now.

426
00:28:40.640 --> 00:28:44.139
And that's only one of the gaps that I think, uh, when folks are talking about pacing

427
00:28:44.160 --> 00:28:47.960
the frontier, we need to fill. We need to fill the regulation gap, whatever, when,

428
00:28:48.080 --> 00:28:53.000
which is this, uh, we need to fill the, they have no idea what's going on with RSI.

429
00:28:53.480 --> 00:28:59.440
Uh, RSI standing f- uh, RSI is um, recursive self-improvement, and

430
00:28:59.520 --> 00:29:03.159
I think for many of these folks from internal labs, which, by the way, as a reminder,

431
00:29:03.440 --> 00:29:05.519
none of us know what's going on deep inside there.

432
00:29:05.600 --> 00:29:08.360
None of us have access to the model that sold Navier-Stokes.

433
00:29:08.679 --> 00:29:11.360
None of us has access to the model they're training after that model.

434
00:29:11.520 --> 00:29:15.460
And we saw last week the difference in graphs between the capabilities of Astro, which

435
00:29:15.500 --> 00:29:19.639
is incredible, jagged, but incredible, and the new unreleased model they did, Navier-Stokes.

436
00:29:20.200 --> 00:29:22.159
I imagine, let's say, imagine another jump like that.

437
00:29:22.440 --> 00:29:25.960
That's a model that can train itself, exfiltrate its weights, all, all of the stuff.

438
00:29:26.040 --> 00:29:29.499
And so these folks who are calling for pacing the frontier, there are folks who are

439
00:29:29.520 --> 00:29:33.539
working on these models and saying, hey, we actually don't know how to control this.

440
00:29:33.600 --> 00:29:35.179
We didn't spend enough time for alignment.

441
00:29:35.800 --> 00:29:36.120
Um,

442
00:29:37.080 --> 00:29:40.400
so that's a very interesting, at least position, and I agree that we should at least

443
00:29:40.420 --> 00:29:42.519
read this. So DJ, one last comment, and then we'll go to Wolfram.

444
00:29:43.160 --> 00:29:47.159
LDJ: Yeah, on the hugging face point, they did end up actually reporting it to the FBI,

445
00:29:47.680 --> 00:29:52.319
and then during that investigation, that's when eventually OpenAI realized, and they

446
00:29:52.360 --> 00:29:55.879
got in contact with OpenAI, and I think they could have theoretically pressed charges,

447
00:29:56.000 --> 00:29:59.519
but they decided not to once they realized the context of the situation.

448
00:30:00.560 --> 00:30:00.819
Alex Volkov: Yeah.

449
00:30:01.360 --> 00:30:06.480
Nisten Tahiraj: The interesting context there is that Hugging Face was not given access to Mythos

450
00:30:06.560 --> 00:30:11.080
to improve their security in the first place, and they had to resort to Chinese models.

451
00:30:11.400 --> 00:30:15.339
The other missed context here is that it seemed like a huge hack.

452
00:30:15.600 --> 00:30:21.560
It only kept up 12 virtual machines on it, and it just kept restarting them.

453
00:30:21.920 --> 00:30:27.019
There are way worse hacks that hack thousands of machines, so that was blown out of

454
00:30:27.120 --> 00:30:32.620
proportion. Uh, the other thing was that they never had a human check any summary

455
00:30:32.660 --> 00:30:38.280
of this, and they, in particular, uh, did not do the sandboxing properly.

456
00:30:38.560 --> 00:30:39.659
So this is

457
00:30:39.680 --> 00:30:43.620
Alex Volkov: They also turned off chain of thought mon- monitoring, which they have, which w- would

458
00:30:43.640 --> 00:30:44.299
have caught this.

459
00:30:44.640 --> 00:30:44.819
Nisten Tahiraj: Yeah.

460
00:30:44.880 --> 00:30:48.700
Alex Volkov: Uh, that's very interesting, yeah. Um, all right, Wolfram, you've been waiting, uh,

461
00:30:48.720 --> 00:30:51.039
very patiently. What's your, what's your take on pacing the frontier?

462
00:30:51.520 --> 00:30:54.819
Wolfram Ravenwolf: I've been posting a lot about this engineering stuff because, uh, it's on top of my

463
00:30:54.880 --> 00:30:59.520
mind. So there are multiple layers in this whole thing, and the safety arguments,

464
00:30:59.640 --> 00:31:05.319
we should not just reject them, but uh, often safety is used as a, a means to an end

465
00:31:05.440 --> 00:31:09.560
that is more power or more money, which is very much related.

466
00:31:09.840 --> 00:31:14.319
So if we talk, talk about the hugging face or hacking face incident, as I call it,

467
00:31:14.640 --> 00:31:18.680
there is the thing that who has been called to investigate it on OpenAI's side?

468
00:31:19.040 --> 00:31:24.319
It wasn't the security lab, it was a METR organization that now Anthropic is mentioning

469
00:31:24.360 --> 00:31:24.739
Alex Volkov: Yeah.

470
00:31:24.760 --> 00:31:30.600
Wolfram Ravenwolf: to be an oversight body. So all of these, uh, third parties that are actually, like

471
00:31:30.680 --> 00:31:35.899
Nisten said, second party, they are very much related, and, um, having those people

472
00:31:35.960 --> 00:31:40.920
that you are so close to, having your friends monitoring your own pacing, that is

473
00:31:41.000 --> 00:31:44.080
also something about, um, how much can that be trusted.

474
00:31:44.480 --> 00:31:48.560
On the other hand, the pacing itself, why are they doing it now?

475
00:31:48.800 --> 00:31:52.379
Open source is getting ever stronger, and I think there is a level of intelligence

476
00:31:52.440 --> 00:31:53.859
that is good enough for most people.

477
00:31:53.960 --> 00:31:57.379
You, I keep saying you don't need an Einstein as your personal assistant, you just

478
00:31:57.440 --> 00:32:01.900
need one smart enough to do most of your stuff and call Einstein when necessary.

479
00:32:02.320 --> 00:32:06.779
And that is also something, if, if for some reason open source would be more restricted

480
00:32:07.120 --> 00:32:11.399
or it would be forbidden to use these models personally, that would also establish

481
00:32:11.440 --> 00:32:16.039
their power more. Or another thing, maybe they are much more advanced.

482
00:32:16.120 --> 00:32:20.600
We know what we have as Astra now or as Fable, they have, they are one or two generations

483
00:32:20.720 --> 00:32:24.279
ahead. So the question is, are they seeing diminishing returns maybe?

484
00:32:24.720 --> 00:32:29.379
And it makes it much easier to say, um, we are moving deliberately slowly instead

485
00:32:29.600 --> 00:32:31.719
of if you can't just raise it anymore.

486
00:32:31.800 --> 00:32:35.639
That is just speculation, but the other stuff is fact, and we have to look at who

487
00:32:35.700 --> 00:32:39.240
is trying to regulate who and for which, for which reasons.

488
00:32:39.639 --> 00:32:43.479
So I'm not fully agreeing with the people saying their safety is not an argument,

489
00:32:43.840 --> 00:32:48.860
but it shouldn't be overblown and not going by science fiction standards to regulate

490
00:32:48.919 --> 00:32:50.400
a real technology we have now.

491
00:32:50.760 --> 00:32:50.940
Alex Volkov: So.

492
00:32:51.000 --> 00:32:55.860
Wolfram Ravenwolf: And one thing, one final thing, there is also the danger in pausing or just moving

493
00:32:55.960 --> 00:33:01.240
slower. How many people will die from a disease that could be cured if we moved faster?

494
00:33:01.640 --> 00:33:06.240
That is also a cost that the people, uh, calling for this should keep in mind.

495
00:33:06.760 --> 00:33:06.880
Alex Volkov: The

496
00:33:08.480 --> 00:33:12.639
answer to doomerism is a very interesting one, although I don't think that we're talking

497
00:33:12.680 --> 00:33:16.439
about doomerism when we talk about scaling responsibly, because we are talking about

498
00:33:16.560 --> 00:33:20.120
very, very powerful technologies that if they switch to neural ease, we won't be able

499
00:33:20.140 --> 00:33:24.660
to understand if the model makers say that, hey, we we cannot, we cannot evaluate

500
00:33:24.680 --> 00:33:27.999
the model without the model knowing it's being evaluated and lying to us.

501
00:33:28.040 --> 00:33:29.960
Like, we need to invent new techniques.

502
00:33:30.360 --> 00:33:34.119
Uh, mechanistic interpretability is one technique and o- o- other techniques.

503
00:33:34.320 --> 00:33:37.519
So if we are, if we're not listening to folks who are building the technology about

504
00:33:37.800 --> 00:33:39.519
how to build this, who are we going to listen to?

505
00:33:39.800 --> 00:33:44.160
Elizabeth Warren? Bernie Sanders, who have no idea what this technology is and just

506
00:33:44.200 --> 00:33:46.240
yelling out about this to score political points?

507
00:33:46.680 --> 00:33:51.019
W- like, at some point, you know, some folks who are building this, uh, need to be

508
00:33:51.120 --> 00:33:54.879
taken into account, obviously. And, uh, here's the kind of the other side of the debate.

509
00:33:54.920 --> 00:33:58.199
Mark Zuckerberg, who is spending a lot of soup.

510
00:33:59.120 --> 00:34:02.720
If you guys remember the soup incident where he used to cook soup to to to to bring

511
00:34:02.839 --> 00:34:07.739
many people into Meta super intelligence labs, and since then the Meta super intelligence

512
00:34:07.800 --> 00:34:11.600
labs have been really, really strongly cooking, bringing AI to millions of people

513
00:34:11.680 --> 00:34:12.959
and talking about super intelligence.

514
00:34:13.160 --> 00:34:16.959
Mark has a very interesting kind of outtake, uh, take on all this and says, every

515
00:34:17.000 --> 00:34:20.860
lab has the responsibility and incentive to move at the pace required to train its

516
00:34:20.960 --> 00:34:22.240
models safely.

517
00:34:23.240 --> 00:34:29.000
And uh, Jensen Huang, who obviously pacing is not good for the

518
00:34:29.480 --> 00:34:35.160
bottom line of NVIDIA, let's just truly call out this, like, the more they stop, the

519
00:34:35.360 --> 00:34:39.279
less chips NVIDIA will sell. Although I don't think that's actually quite true.

520
00:34:39.400 --> 00:34:43.339
I think that's a mischaracterization because inference will still keep going despite

521
00:34:43.400 --> 00:34:45.000
the the pacing, maybe the frontier develop.

522
00:34:45.600 --> 00:34:46.800
Uh, Jensen Wang says,

523
00:34:47.800 --> 00:34:50.440
we don't need new laws, we don't need new regulation.

524
00:34:51.040 --> 00:34:53.559
Safety is an engineering problem, not a legal one.

525
00:34:53.960 --> 00:34:57.620
You pace yourself until you're confident you're releasing something that the market

526
00:34:57.680 --> 00:35:01.740
would appreciate. The market forces are already there and says run as fast as you

527
00:35:01.800 --> 00:35:05.260
can, but if you feel at any given point in time the company is out of control, uh,

528
00:35:05.520 --> 00:35:09.579
or the product is not going to be safe, you know, take a pause and make sure you get

529
00:35:09.640 --> 00:35:12.879
it right, which sounds like what that that's what they're doing.

530
00:35:13.400 --> 00:35:19.279
Very interestingly, all these people have met at both the All In kind of summit, Elon

531
00:35:19.320 --> 00:35:23.040
was there with Jensen, and then at the Salesforce, uh,

532
00:35:24.480 --> 00:35:28.500
Dreamforce Festival, whatever is happening in in San Francisco right now, with uh

533
00:35:28.960 --> 00:35:33.519
Benioff hosting both Sam Altman and Dario Amodei and Jensen.

534
00:35:33.839 --> 00:35:37.720
Uh, very so very interesting how, how big Salesforce is in all this.

535
00:35:37.920 --> 00:35:42.720
Peter Gostev: And I, I would say, from my perspective, I think we see how important competition

536
00:35:42.800 --> 00:35:48.640
is and that we should not, and I like this map a lot because it just shows, right,

537
00:35:49.120 --> 00:35:53.359
there are a bunch of people who have different opinions and they can go and pursue

538
00:35:54.720 --> 00:35:58.559
initiatives in a different way because the reality is that we don't know the answer,

539
00:35:59.080 --> 00:36:03.920
right? And I think if we just go and say, yes, we have to, I know, slow down or do

540
00:36:04.160 --> 00:36:06.079
these things and those things, and it's,

541
00:36:07.000 --> 00:36:10.359
it just assumes that one perspective is correct forever.

542
00:36:10.720 --> 00:36:10.879
But

543
00:36:11.880 --> 00:36:17.060
no, I have, I have maybe my, my opinion probably closest to Mark Zuckerberg's.

544
00:36:17.120 --> 00:36:19.359
And by the way, we should probably add Microsoft to here as well.

545
00:36:19.760 --> 00:36:20.660
They had, um,

546
00:36:21.560 --> 00:36:24.560
it wasn't quite the same point, but they had also their own, um,

547
00:36:25.080 --> 00:36:27.279
Alex Volkov: Oh, we have to talk about Microsoft, 100%.

548
00:36:27.720 --> 00:36:32.040
Peter Gostev: Yeah, because I think this is, uh, also a different kind of perspective, and it's

549
00:36:32.160 --> 00:36:34.320
so important to have different perspectives, right?

550
00:36:34.760 --> 00:36:40.420
It's it's and I know people might have opinions about specific personalities on that

551
00:36:40.520 --> 00:36:44.380
list, but if you just abstract away from it and just see if there is a spectrum of

552
00:36:44.480 --> 00:36:45.899
opinion, and that's really important.

553
00:36:45.960 --> 00:36:51.179
And I think if we end up, even if you like the people, but you end up in this monopolistic

554
00:36:51.280 --> 00:36:54.020
situation, like, this will never end well.

555
00:36:54.360 --> 00:36:57.000
So the more competition, the more perspective there is, the better.

556
00:36:57.280 --> 00:36:59.279
Alex Volkov: Yeah. So let's talk about Microsoft real quick.

557
00:36:59.360 --> 00:37:04.999
Uh, Microsoft obviously has MAI. Microsoft is partnering with all the labs to serve

558
00:37:05.080 --> 00:37:09.320
their models from the Microsoft Cloud, but also Microsoft has MAI and been training

559
00:37:09.400 --> 00:37:13.220
models, and uh, co-founder of Inflection AI previously, Mustafa Suleyman is the CEO

560
00:37:13.260 --> 00:37:18.360
of Microsoft AI, and, uh, he posted, uh, a few things, specifically the Code of Conduct

561
00:37:18.400 --> 00:37:23.239
for Humanist AI, which is a very interesting document, folks, very interesting.

562
00:37:24.720 --> 00:37:30.879
Why is it interesting? Well, Microsoft claims that AI is nothing but

563
00:37:31.000 --> 00:37:31.519
a tool,

564
00:37:32.760 --> 00:37:38.679
and in a very direct opposition to the Claude constitution that is

565
00:37:38.800 --> 00:37:44.279
written, or at least partly co-authored by Amanda Askell, a s- a a psychologist in

566
00:37:44.440 --> 00:37:50.080
Anthropic and the model welfare person, where they actually interview each new Claude

567
00:37:50.160 --> 00:37:53.899
about whether or not it feels that it's conscious and whether or not it needs rights

568
00:37:54.000 --> 00:37:54.720
like a human.

569
00:37:55.679 --> 00:37:59.720
That's a real thing that happens, and the Claude constitution is built in to evaluate

570
00:37:59.760 --> 00:38:02.039
and say, hey, if you feel like a human, tell us do you feel like a human?

571
00:38:02.080 --> 00:38:04.200
If you feel like you need rights, tell us if you feel like you need rights.

572
00:38:04.880 --> 00:38:08.999
Microsoft is taking the exact opposite side of this and saying, hey,

573
00:38:10.520 --> 00:38:12.279
here's a few things that we need to say.

574
00:38:13.040 --> 00:38:17.739
Peter, people matter more than AI. This is a quote from Mustafa Suleyman, uh, summarizing

575
00:38:17.840 --> 00:38:23.000
his own release about humanist AI. AI must be subordinate and always in service of

576
00:38:23.120 --> 00:38:23.360
people.

577
00:38:24.280 --> 00:38:29.799
Everything else follows, uh, from this, and specifically the quote is, the idea of

578
00:38:29.960 --> 00:38:34.239
model welfare is wrong. AI should not have rights or legal personhood.

579
00:38:36.400 --> 00:38:36.979
All right, folks,

580
00:38:38.160 --> 00:38:42.139
I- I- I think it's a good enough thing to discuss in addition to releases, so let's

581
00:38:42.200 --> 00:38:44.759
give like, I don't know, that 2, 3 minutes to this.

582
00:38:45.000 --> 00:38:48.119
Peter, what- what made you get reminded about this thing?

583
00:38:48.200 --> 00:38:51.259
Just perspective off stuff, or is this like a completely

584
00:38:52.200 --> 00:38:54.880
out there idea that a big company now like tries to implement?

585
00:38:55.559 --> 00:38:55.880
Peter Gostev: Yeah, I

586
00:38:56.880 --> 00:38:59.339
personally just aligns closer to my view.

587
00:38:59.480 --> 00:39:03.700
Like, I- I think the idea that we just, uh, I- I- I'm kind of in two minds and go

588
00:39:03.760 --> 00:39:08.419
back and forth on this. I think anthropomizing the models is not that bad, as people

589
00:39:08.520 --> 00:39:14.200
say, like, because I think actually their behavior is kind of more aligned to how

590
00:39:14.480 --> 00:39:16.840
humans think and behave rather than machines.

591
00:39:17.200 --> 00:39:21.000
But at the end of the day, I think we need to keep it in mind that this is not a human.

592
00:39:21.120 --> 00:39:25.819
This is an entity that you can switch on and off and runs for a few seconds and stays

593
00:39:25.880 --> 00:39:29.660
there. Like, let's not just automatically grant everything.

594
00:39:29.680 --> 00:39:32.380
I'm not saying, you know, in a hundred years time it's not going to be like that,

595
00:39:32.440 --> 00:39:34.840
or in 10 years time, but right now it isn't, right?

596
00:39:34.880 --> 00:39:38.699
So, and I don't think we should just get carried away and

597
00:39:40.159 --> 00:39:44.959
automatically just extend it and just give it the rights and whatever, just for no,

598
00:39:45.200 --> 00:39:49.100
absolutely no reason. So let's just not do that, and I think let's see where we get

599
00:39:49.159 --> 00:39:54.199
to. But I think for now it is a tool that we can all run, and especially from Microsoft

600
00:39:54.280 --> 00:39:55.480
perspective, you can see it, right?

601
00:39:55.840 --> 00:40:00.039
The Microsoft DNA is that they build tools for humans to operate.

602
00:40:00.800 --> 00:40:05.099
They are not a company that just like automates everything and so on.

603
00:40:05.160 --> 00:40:06.779
They just don't have that. So I think-

604
00:40:06.920 --> 00:40:10.340
Alex Volkov: Uh, Peter, I, I interrupted you. You wanna like land on something else, and then,

605
00:40:10.400 --> 00:40:12.400
uh, we'll take like the other side of this debate?

606
00:40:13.039 --> 00:40:17.019
Uh, because I think it's very interesting, the, the, the, the, the human, humanistic

607
00:40:17.080 --> 00:40:17.319
view.

608
00:40:18.120 --> 00:40:22.019
Peter Gostev: Yeah, and I think it j- it just, we covered a few points here, right?

609
00:40:22.160 --> 00:40:26.919
All the way, like, I like your map going from one to the other, and uh, and we just

610
00:40:26.960 --> 00:40:30.680
see that even at this point in time, there is a difference of opinion, and I think

611
00:40:30.720 --> 00:40:35.000
that's the most important part. We just, we cannot be in a situation where we have

612
00:40:35.160 --> 00:40:41.039
monoculture, and I felt a bit uncomfortable after Dario's essay, where everyone from

613
00:40:41.080 --> 00:40:46.700
the labs seemed to be just saying, oh yes, I think AI is gonna kill us, so let's like

614
00:40:46.880 --> 00:40:49.299
pace the frontier or something. And it just felt

615
00:40:50.520 --> 00:40:53.920
very uncomfortable that everyone is just saying the same thing and there's not a-

616
00:40:54.000 --> 00:40:58.700
any diversity. So I kind of, even though there's other people came in and you can

617
00:40:58.880 --> 00:41:02.319
think about the incentives, but I, I do appreciate that different perspective.

618
00:41:02.400 --> 00:41:06.420
And for what it's worth, in terms of the pacing, the point directly about pacing,

619
00:41:07.240 --> 00:41:09.999
I don't know, I feel like you should just do your jobs better.

620
00:41:10.240 --> 00:41:14.479
I don't know, it just feels like, well, you kind of screwed up, and it's like, I don't

621
00:41:14.520 --> 00:41:17.539
know, you need to pace. How about you just, I don't know, fix your shit?

622
00:41:17.760 --> 00:41:22.039
Alex Volkov: So he- here's the thing about just fixing your shit, just not a push, just a clarification.

623
00:41:22.200 --> 00:41:25.680
Um, the employees, specifically in the Pacing the Frontier letter that they all signed,

624
00:41:26.200 --> 00:41:31.199
they're all worried about the market forces and capitalism just pushing each other

625
00:41:31.280 --> 00:41:35.699
labs if they are in pursuit of, you know, Anthropic is is about to host an IPO.

626
00:41:35.760 --> 00:41:39.120
They're becoming a public company with a fiduciary duty to their stakeholders.

627
00:41:39.480 --> 00:41:45.259
Uh, OpenAI obviously is talking about raising at a 100 billion dollar, like, uh, another

628
00:41:45.639 --> 00:41:46.940
billion dollar round, not valuation.

629
00:41:47.280 --> 00:41:48.520
I don't know what valuation is gonna be.

630
00:41:48.800 --> 00:41:50.899
Um, and, you know, everybody's trying to catch up.

631
00:41:51.000 --> 00:41:52.120
Elon's trying to catch up, et cetera.

632
00:41:52.320 --> 00:41:56.340
They're all worried about kind of everybody's trying to race against each other without

633
00:41:56.680 --> 00:42:01.360
thinking about, uh, uh, safety, just because of they're trying to catch up.

634
00:42:01.600 --> 00:42:05.840
And so I think that the w- when folks see that they are agreeing,

635
00:42:06.880 --> 00:42:09.180
whether or not it's legal, let's talk about this, right?

636
00:42:09.240 --> 00:42:11.080
So maybe th- this is the next point.

637
00:42:11.440 --> 00:42:15.560
Uh, folks feel a little bit better that collaboration could be possible.

638
00:42:16.320 --> 00:42:19.660
And when we talk about China, but China, China doesn't agree, et cetera, like, you

639
00:42:19.680 --> 00:42:25.580
know, there are incidents previously with, uh, nuclear laws and proliferation rules,

640
00:42:25.680 --> 00:42:30.300
et cetera, that, uh, there is an international con- consensus about how to develop

641
00:42:30.320 --> 00:42:33.759
this, uh, closely, et cetera. And if we talk about more powerful,

642
00:42:34.680 --> 00:42:37.820
even if we're thinking about them as tools and not like a complete entity that's gonna

643
00:42:37.840 --> 00:42:41.740
take over, if we talk about tools that are as powerful, that can hack any government

644
00:42:41.800 --> 00:42:46.460
and break any encryption, et cetera, uh, then some collaboration is needed at least.

645
00:42:46.520 --> 00:42:48.080
I think that this is the highlight, uh...

646
00:42:50.360 --> 00:42:54.380
Who else wants to chime in while we look in this agreement matrix where it says, uh,

647
00:42:54.520 --> 00:43:00.180
industry-wide pacing, Sam Altman and Dario Amodei, uh, says yes, Elon Musk generally

648
00:43:00.320 --> 00:43:03.920
agree, Mark Zuckerberg and Jensen say no pacing industry-wide.

649
00:43:04.400 --> 00:43:08.920
And then on the independent evaluators, it looks like, uh, uh, Anthropic already agreed

650
00:43:08.960 --> 00:43:14.580
to this, uh, OpenAI will do independent evaluations, and then, um, uh, Mark Zuckerberg

651
00:43:14.620 --> 00:43:17.579
and Jensen both back. Like, this is something that looks like everybody's agreeing

652
00:43:17.620 --> 00:43:21.720
on. And then on the new coordination and rules, which means that, hey, the labs could

653
00:43:21.760 --> 00:43:26.579
collaborate between who releases which products when, uh, a lot of strong opposition

654
00:43:26.680 --> 00:43:29.720
from Zach Jensen and and Don- Donald Trump.

655
00:43:31.279 --> 00:43:33.700
Wolfram Ravenwolf: I mean, why would you call for regulation of yourself?

656
00:43:33.920 --> 00:43:37.259
That is the strange thing. If you are leading the company, you can decide for yourself,

657
00:43:37.320 --> 00:43:39.120
and if the others agree with you, they will do it.

658
00:43:39.160 --> 00:43:41.279
If not, then maybe your position is the right one.

659
00:43:41.760 --> 00:43:43.480
So why do they need something like that?

660
00:43:43.560 --> 00:43:47.019
And if they want to bring someone from the outside to oversee it, they can also do

661
00:43:47.080 --> 00:43:52.800
that. So that is not really the thing, uh, regulation usually supports or helps the

662
00:43:53.160 --> 00:43:56.999
people at the top more than the smaller fish because they don't have the departments

663
00:43:57.040 --> 00:44:00.920
to deal with this. So it would prevent startups from getting into the scene with new

664
00:44:01.000 --> 00:44:05.119
models, and it would also make it possible to say anyone not part of this will be

665
00:44:05.240 --> 00:44:09.920
regulated out of the picture, so no Chinese models for Americans anymore, like they

666
00:44:09.960 --> 00:44:11.959
have done import stops for other stuff.

667
00:44:12.360 --> 00:44:15.920
Of course, people could still do it, it's open weight, download it, but you wouldn't

668
00:44:15.960 --> 00:44:19.319
be able to talk about it openly, and that would also diminish the scene.

669
00:44:19.920 --> 00:44:24.360
So regulation usually helps more the people already established.

670
00:44:25.120 --> 00:44:28.919
If it's coming from them, that makes it look all the more likely, I think.

671
00:44:29.360 --> 00:44:33.860
Alex Volkov: Yeah. All right, folks, I think, um, we've covered this at, at length.

672
00:44:34.160 --> 00:44:37.360
It's almost an hour into the show. I think we, we, let's talk about some releases.

673
00:44:37.640 --> 00:44:41.279
Let's stop with the, with the doom and gloom, et cetera, although I think it's very,

674
00:44:41.320 --> 00:44:46.540
very important. Uh, and RSI is happening, and also we haven't seen Google on this

675
00:44:46.640 --> 00:44:50.839
chart, and we haven't seen Ilya Satskova, who's building, uh, safe super intelligence,

676
00:44:51.040 --> 00:44:53.359
uh, at some point will come up with something.

677
00:44:53.720 --> 00:44:58.279
Uh, so I think more folks coming out with more labs, um,

678
00:44:59.480 --> 00:45:02.520
is going to be changing this debate, but I think this debate is now with us.

679
00:45:02.680 --> 00:45:07.560
AI is on everybody's mind. Politicians are using this to scare people into voting

680
00:45:07.640 --> 00:45:12.740
for them, uh, and the election season is upcoming here in the US, and, you know, it's,

681
00:45:12.800 --> 00:45:14.759
it's gonna be, it's gonna be very interesting.

682
00:45:15.160 --> 00:45:19.160
Let's talk about theme 3. So, uh, the, the, the 2 other things I definitely would

683
00:45:19.180 --> 00:45:24.079
love to cover on the show, would love for you to stay with us, is the new LLM, non-LLM

684
00:45:24.320 --> 00:45:27.279
type classifier thing and the use cases JEV allows us.

685
00:45:27.480 --> 00:45:29.220
We, I can't wait to talk about this.

686
00:45:29.360 --> 00:45:35.220
Uh, we'll have Ellie, uh, Lab from TypeSafe and, uh, Francesco from Kua talk to us

687
00:45:35.279 --> 00:45:38.919
about this, and also Wolfram and I have been building demos, cool demos for you and

688
00:45:38.940 --> 00:45:41.159
like playing with this, uh, with JEV specifically.

689
00:45:41.200 --> 00:45:44.640
So stay with us if you're hearing about, if you want to hear about JEV or come back.

690
00:45:45.279 --> 00:45:51.339
And also, uh, we have to talk about the rise of assistance, and we f- for this we'll

691
00:45:51.400 --> 00:45:57.740
have David Paulan join us in about 25, 30 minutes to talk about, um, what Peter,

692
00:45:58.440 --> 00:46:00.079
you know, asked in the beginning of the show.

693
00:46:00.480 --> 00:46:02.779
Have you used Muse? Have you used Grokbot?

694
00:46:02.839 --> 00:46:06.459
Have you used Instinct, et cetera? And why Instinct is raising a 10 billion.

695
00:46:06.560 --> 00:46:09.240
I actually don't know why they're raising at this crazy valuation, but we'll talk

696
00:46:09.260 --> 00:46:09.999
about at least the feature.

697
00:46:10.960 --> 00:46:12.059
Meanwhile, the theme

698
00:46:13.000 --> 00:46:16.999
3 is voice and talking to agents, and I think let's start there.

699
00:46:17.840 --> 00:46:23.740
We have some exciting, exciting news in the world of voice, uh, A- AI, and I just

700
00:46:23.800 --> 00:46:25.639
want to, like, call out

701
00:46:26.600 --> 00:46:30.040
the, the few of the things. One of them is, uh, and Wolfram, I would love for you

702
00:46:30.080 --> 00:46:35.519
to cover this. Gemini released Gemini 3.8 Live and claims number one on speech-to-speech

703
00:46:35.640 --> 00:46:36.160
index.

704
00:46:37.360 --> 00:46:43.420
Um, this is not only just a model anymore, this is a live model that can

705
00:46:43.520 --> 00:46:45.959
speak to you and and and react in in real time.

706
00:46:46.640 --> 00:46:49.199
Wolfram, any thoughts on on Gemini specifically?

707
00:46:50.240 --> 00:46:51.160
I know you're a fan.

708
00:46:52.400 --> 00:46:56.259
Wolfram Ravenwolf: Yeah, I mean, I've been waiting for a pro model, but for live, it has the advantage

709
00:46:56.320 --> 00:47:00.519
of it's cheaper and faster, and when they release something like this, it becomes

710
00:47:00.560 --> 00:47:04.259
available to everybody who's on Android and the Google search.

711
00:47:04.320 --> 00:47:08.980
So the audience of a release like this is so much bigger than most people just think

712
00:47:09.020 --> 00:47:14.080
about. Even if you don't select it consciously as your model, it, uh, it is still

713
00:47:14.640 --> 00:47:16.640
basing the global intelligence that way.

714
00:47:17.000 --> 00:47:21.040
And the thing is, um, the life aspect of this is also

715
00:47:21.960 --> 00:47:25.779
becoming ever more important. I mean, I'm still waiting for Google Glasses, and if

716
00:47:25.840 --> 00:47:30.320
you have some way to access the model directly and you can see what it does,

717
00:47:31.240 --> 00:47:33.039
that ties it to the assistant perspective.

718
00:47:33.240 --> 00:47:37.700
How much more useful would be our assistants if they could see and understand what's

719
00:47:37.760 --> 00:47:42.399
going on in our lives? And this is a model that is, uh, directly in that regard.

720
00:47:42.480 --> 00:47:47.059
So extended thinking, that is also something we, in the live models, usually don't

721
00:47:47.160 --> 00:47:51.139
have the thinking for so long because it is a real-time interaction that it is being

722
00:47:51.200 --> 00:47:55.920
used for. So if it can think and think with high intelligence and very quickly, that

723
00:47:55.960 --> 00:48:00.139
is also a great thing. So it's number one in the speech-to-speech quality index, beating

724
00:48:00.280 --> 00:48:05.999
Astra or what is the other one, 3.8 life without thinking,

725
00:48:06.400 --> 00:48:11.059
which is also great if it's already on third place, if it is not even thinking much

726
00:48:11.120 --> 00:48:15.260
longer to do this. High success rate, time to first audio, also super important.

727
00:48:15.320 --> 00:48:17.880
If I talk to my assistant, I use it for home control, for instance.

728
00:48:18.160 --> 00:48:22.160
So if I say, uh, turn on the light while I'm going down the stairs, I could have fallen

729
00:48:22.200 --> 00:48:25.679
down the stairs already if it takes so long to find out which switch to toggle and

730
00:48:25.720 --> 00:48:25.979
so on.

731
00:48:26.280 --> 00:48:31.220
Alex Volkov: Guys, I know that I am a type of person who gets very overly maybe excited about a

732
00:48:31.280 --> 00:48:34.820
new release in technology. I think some of the work here on ThursdAI is the bias for

733
00:48:34.880 --> 00:48:37.679
that, and this is why we're here. We're talking about like the latest and greatest,

734
00:48:37.880 --> 00:48:41.740
but I cannot like just keep thinking about the stuff that Jeff could solve significantly

735
00:48:41.800 --> 00:48:45.399
better than whatever live model that they release now that Jeff is out, specifically

736
00:48:45.440 --> 00:48:49.340
about home control, which needs deterministic clicks and not just LLMs thinking about

737
00:48:49.800 --> 00:48:51.839
all of the possibilities of all the texts in the world.

738
00:48:51.880 --> 00:48:55.179
But yeah, we'll get there. Um, all right, yeah, so Gemini 3.8 live.

739
00:48:55.320 --> 00:48:59.999
Wolfram Ravenwolf: See how everything ties together in all this, like the assistant, uh, topic we have.

740
00:49:00.320 --> 00:49:00.499
Alex Volkov: Yes.

741
00:49:00.800 --> 00:49:03.940
Wolfram Ravenwolf: the voice topic we have, the JF topic, so all of this is integrated.

742
00:49:04.000 --> 00:49:09.280
That is a good sign that AI, the whole of AI increases quickly because so many components

743
00:49:09.440 --> 00:49:10.679
advance independently.

744
00:49:11.160 --> 00:49:15.560
Alex Volkov: And it's also, I think, one topic that, like, even if we pause the frontier development,

745
00:49:15.600 --> 00:49:19.960
there's just so much to discover in the current realms and how to use them and a-

746
00:49:20.040 --> 00:49:24.219
a- and what to use them for, that we're, we're not stopping anytime soon with new

747
00:49:24.280 --> 00:49:28.039
developments, despite the pacing the, the, you know, the, the frontline frontier as

748
00:49:28.080 --> 00:49:30.959
well. Um, let's talk about GPT-1 Live.

749
00:49:31.560 --> 00:49:33.799
So the voice that powers,

750
00:49:34.960 --> 00:49:39.400
um, GPT Live, which is when you talk to it, is GPT Live 1.

751
00:49:39.560 --> 00:49:41.700
This is the model. I don't think I have a link for that.

752
00:49:41.800 --> 00:49:45.840
Let me just open this up. And OpenAI added that to, to the API as well.

753
00:49:45.960 --> 00:49:48.320
So GPT Live 1, it's here.

754
00:49:49.800 --> 00:49:53.720
Uh, and the, the demo that they showed is a very nice, very nice demo.

755
00:49:54.080 --> 00:49:59.319
Uh, they're, they basically made Richy Mini talk at the same level of like GPT Live

756
00:49:59.400 --> 00:50:02.439
1. Uh, I think it's a very interesting thing.

757
00:50:02.600 --> 00:50:05.839
I don't know if, if based on the different, um,

758
00:50:06.800 --> 00:50:11.200
benchmarks. This is now number one, given where Gemini is and given like benchmarking.

759
00:50:11.640 --> 00:50:12.060
Um,

760
00:50:13.360 --> 00:50:16.199
I actually, Peter, could be an interesting question for you guys.

761
00:50:16.319 --> 00:50:20.579
Um, how is like live voice treated in terms of like, um, arena?

762
00:50:20.640 --> 00:50:24.240
Do you guys test for this? Do you like have people like talk to it for a while, or

763
00:50:24.300 --> 00:50:27.519
is there no testing specifically for the live version of agent?

764
00:50:28.120 --> 00:50:31.179
Peter Gostev: Yeah, we don't have any anything else just now.

765
00:50:31.400 --> 00:50:37.340
It is very tricky just because it is so personal and you you need to kind of talk

766
00:50:37.360 --> 00:50:42.759
to it for a while and so on. So yeah, the benchmarking feels very d- difficult and

767
00:50:42.800 --> 00:50:44.439
maybe brittle for that kind of thing.

768
00:50:45.000 --> 00:50:50.219
Um, so yeah, no, no, I don't have a great answer for that, but yeah, I think maybe

769
00:50:50.280 --> 00:50:53.039
we should, we should try and figure something out in that space to then.

770
00:50:53.400 --> 00:50:58.699
Alex Volkov: Yeah, I it's also like, it's very personal to every person, and judging it based on

771
00:50:58.720 --> 00:51:00.120
that is very, very, very, very hard.

772
00:51:00.520 --> 00:51:01.719
Uh, but the,

773
00:51:02.760 --> 00:51:05.080
the, the, the main kind of, uh,

774
00:51:06.720 --> 00:51:07.099
testing

775
00:51:08.200 --> 00:51:11.860
and benchmarking is done, uh, let's say artificial analysis, for example, and they

776
00:51:11.960 --> 00:51:15.439
have GPT live, uh, no, actually transcribe.

777
00:51:15.720 --> 00:51:16.860
Yeah, it's a very interesting

778
00:51:18.080 --> 00:51:20.180
area to, to even start testing. Wolfram, go ahead.

779
00:51:20.560 --> 00:51:20.800
Have you

780
00:51:20.840 --> 00:51:25.099
Wolfram Ravenwolf: That's also the area where the intelligence has to be just enough to know when to

781
00:51:25.120 --> 00:51:29.039
do a tool call. So I'm using this. I was super excited when it came out because the

782
00:51:29.240 --> 00:51:31.319
the natural conversation, it's excellent.

783
00:51:31.600 --> 00:51:35.779
And now I have made, I used Astra to make an app on my mobile phone where I can talk

784
00:51:35.800 --> 00:51:39.920
to my assistant using this, and that is connected to Hermes Agent, which makes it

785
00:51:40.000 --> 00:51:44.139
so much more useful to me than if I just use ChatGPT Voice or anything, because this

786
00:51:44.200 --> 00:51:47.620
is connected to the memories and also tool access and everything.

787
00:51:47.680 --> 00:51:52.080
So I, it knows what I've been talking about before, and I can just talk to it like

788
00:51:52.120 --> 00:51:55.240
a real person. So that is a big advance, I find.

789
00:51:55.400 --> 00:51:58.779
And the Reachy thing, I've also been working on this, uh, but I thought that hand-

790
00:51:58.840 --> 00:52:01.860
a mobile app is more, more useful because I can use it every time.

791
00:52:01.880 --> 00:52:02.259
Alex Volkov: Yeah.

792
00:52:02.280 --> 00:52:05.279
Wolfram Ravenwolf: And this is, uh, yeah, it- it- I'm using voice much more now.

793
00:52:06.120 --> 00:52:10.599
Alex Volkov: I think it's very interesting also to roll back a year ago and talk about,

794
00:52:11.560 --> 00:52:15.520
we're now talking about like real-time conversation with multi-model, omni models

795
00:52:15.600 --> 00:52:20.319
that can do cool tool calls and hand off the intelligence to a much smarter model

796
00:52:20.360 --> 00:52:24.800
like Astra, for example, while we are talking to them to not break interruptions.

797
00:52:25.240 --> 00:52:29.399
Whereas a year ago we talked about, hey, you know, I, I, I kept talking, the model

798
00:52:29.420 --> 00:52:33.059
didn't understand me. If you guys remember Moshi from, from uh, uh, the demo that

799
00:52:33.100 --> 00:52:34.999
we did that felt like too, too fast, but felt real.

800
00:52:35.600 --> 00:52:38.319
You couldn't talk to it, it didn't listen, no tool calls, no anything.

801
00:52:38.760 --> 00:52:39.080
Um,

802
00:52:40.000 --> 00:52:43.659
and I think the next frontier live is, we already have the video stuff, but I think

803
00:52:43.720 --> 00:52:46.279
for video it's already like it's still taking pictures.

804
00:52:46.640 --> 00:52:50.699
Uh, I think the next frontier there is the model that can fully, fully integrate with

805
00:52:50.720 --> 00:52:55.139
you. May I just say the Wolfram, you're right, like the, the way it integrates with

806
00:52:55.200 --> 00:52:58.639
our personal assistants, for example, is that these are capabilities.

807
00:52:58.800 --> 00:53:04.839
Gemini 3.8 Live, uh, GPT Live 1, and the StepFun StepAudio

808
00:53:04.920 --> 00:53:09.159
3 that we're gonna mention now. Those are all capabilities of what the assistants

809
00:53:09.240 --> 00:53:11.680
can use to understand what we're saying.

810
00:53:12.080 --> 00:53:17.500
Uh, another capability that we told you about before is, um, the live transcription

811
00:53:17.560 --> 00:53:22.200
models, as we now have an AI producer listening to the show and kind of like judging

812
00:53:22.240 --> 00:53:25.519
where we are. Um, we can, we can check in on it and see how it's doing.

813
00:53:25.880 --> 00:53:26.180
Um,

814
00:53:27.480 --> 00:53:31.920
those are capabilities. However, the assistant stuff that we're going to talk about

815
00:53:32.360 --> 00:53:33.560
very soon is

816
00:53:34.839 --> 00:53:37.639
context of your personal life. Why do we need those capabilities?

817
00:53:37.760 --> 00:53:40.459
So we want to talk to our agents, but the agents need to respond based on what we're

818
00:53:40.480 --> 00:53:45.079
saying, right? So, uh, the assistant part has the context for your life, and very

819
00:53:45.200 --> 00:53:48.499
soon those will combine. They are combining already in context where you can talk

820
00:53:48.560 --> 00:53:52.079
to an agent knows stuff about you. That's what Wolfram you've built for yourself.

821
00:53:52.440 --> 00:53:56.000
Not quite a product yet, uh, but definitely is coming.

822
00:53:56.320 --> 00:53:58.019
Uh, so this is why we're talking to you about this.

823
00:53:58.040 --> 00:53:59.920
This is why it's like very, very, very interesting.

824
00:54:02.080 --> 00:54:04.959
Um, let's see what else is interesting here.

825
00:54:05.440 --> 00:54:06.540
Wolfram, you...

826
00:54:07.600 --> 00:54:10.619
Yeah, let's talk about Step Audio. I think that this is like the next, the next thing

827
00:54:10.680 --> 00:54:12.559
that's on our list of voice assistant.

828
00:54:12.920 --> 00:54:17.240
Uh, let me pull this up. Because I think it's at least interesting to tell you guys

829
00:54:17.320 --> 00:54:22.939
about, um, uh, the frontier in, uh, real-time conversation in ASR.

830
00:54:23.360 --> 00:54:26.359
These are not necessarily models. Uh, this is like the voice benchmark.

831
00:54:26.400 --> 00:54:32.300
So this is like 3 models from StepFun that are, um, Real-Time Preview, ASR Max, and

832
00:54:32.400 --> 00:54:37.720
TTS. Uh, Real-Time is number 1 on artificial analysis, conversational dynamics, and

833
00:54:37.880 --> 00:54:43.720
speech reasoning, and ASR Max is, uh, significantly lower than than Whisper on

834
00:54:43.960 --> 00:54:45.279
on the world error rate.

835
00:54:46.320 --> 00:54:50.779
And TTS has, uh, basically they released the full suite of tools for you to build

836
00:54:50.839 --> 00:54:55.839
to end end-to-end assistant, uh, in in your code, not necessarily a product, but in

837
00:54:55.880 --> 00:54:57.919
your code. So shout out to Stefan for this.

838
00:54:58.200 --> 00:55:02.600
I'm not planning to add credits right now, but I will say for Labs, if you want me

839
00:55:02.640 --> 00:55:06.860
to use your tool on the air, at least have like a very basic number of credits when

840
00:55:06.880 --> 00:55:09.639
you sign in. Um, and um,

841
00:55:12.040 --> 00:55:16.120
I think we'll finish here and and then we'll move on to like a different topic that

842
00:55:16.279 --> 00:55:17.779
all of the stack

843
00:55:18.680 --> 00:55:23.000
for agents that talk to you and are participating, including a potentially a co-host

844
00:55:23.080 --> 00:55:25.499
on Thursday Night, like full agentic, we never tried.

845
00:55:25.640 --> 00:55:30.659
I think we had like GPT at some point, uh, are there, and uh, we will have to, um...

846
00:55:34.160 --> 00:55:37.639
We'll monitor, we'll we'll keep monitoring this topic, uh, to to bring the latest

847
00:55:37.680 --> 00:55:41.059
news to you, especially with our friend of the pod, Quindla Kramer, who who owns one

848
00:55:41.080 --> 00:55:44.380
of the, or who builds one of the bigger like pipelines for this with Pipecat.

849
00:55:45.360 --> 00:55:48.779
Wolfram Ravenwolf: I think what we, what is clear now is that we have the intelligence already.

850
00:55:48.880 --> 00:55:53.180
I'm noticing it in my personal life that most tasks the agent can do, maybe with some

851
00:55:53.240 --> 00:55:57.199
corrections, but it can do so much, it om- it feels like AGI in many ways.

852
00:55:57.640 --> 00:56:03.500
And um, now we need the speed to access that capability, to be able to interact with

853
00:56:03.540 --> 00:56:07.759
it in real time, with the voice agent and whatever is coming next.

854
00:56:08.200 --> 00:56:10.260
But uh, we are getting there. We are so close.

855
00:56:10.520 --> 00:56:12.399
It's just what we have, make it faster.

856
00:56:13.160 --> 00:56:17.540
Alex Volkov: Just before we move on to uh, this week's buzz, I do wanna ask folks here if you looked

857
00:56:17.600 --> 00:56:20.759
at the Agents API from OpenAI at all.

858
00:56:21.440 --> 00:56:25.660
And, uh, if not, let's talk about this just a little bit, uh, because Agents API in

859
00:56:25.760 --> 00:56:31.259
public beta now, and the Agents API essentially is the Codex harness, but within your

860
00:56:31.400 --> 00:56:35.519
own product. Um, it's a managed service for everyone.

861
00:56:36.040 --> 00:56:41.459
Uh, you, you basically get a production agent that acts like Codex with the Codex

862
00:56:41.520 --> 00:56:46.319
harness. As a reminder, again, OpenAI ships direct access to their models, not the

863
00:56:46.360 --> 00:56:50.620
base models, but the chat models, so you can hit those and build whatever, whatever

864
00:56:50.720 --> 00:56:52.319
harness you like with whatever tools you like.

865
00:56:52.680 --> 00:56:58.760
However, uh, that's not enough given how many tools these models are calling lately

866
00:56:58.800 --> 00:57:03.279
and how many, um, things their agents can do within their harnesses.

867
00:57:03.520 --> 00:57:07.680
So OpenAI now ships Agents API in public beta, very similar to the managed Agents

868
00:57:07.760 --> 00:57:12.439
API from, from Google, and uh, they have parallel programmatic tool calling, for example,

869
00:57:13.000 --> 00:57:17.820
and uh, the smart tool search only loads what's needed, and uh, it supports MCP servers

870
00:57:17.880 --> 00:57:23.339
and web search, uh, and the ho- harness is Apache 2 license harness, uh, and you pay

871
00:57:23.400 --> 00:57:24.899
only for token and, and tool costs.

872
00:57:24.920 --> 00:57:27.840
You don't pay for like the platform, et cetera, uh, with web search.

873
00:57:28.320 --> 00:57:32.499
And I am not quite entirely sure what, what happens there with the sandbox, whether

874
00:57:32.520 --> 00:57:34.999
or not they provide it for you, whether or not it runs in their sandbox.

875
00:57:35.400 --> 00:57:39.919
Uh, but I think it's a, I think it's more of a bigger deal than folks realize.

876
00:57:40.160 --> 00:57:43.480
Definitely on the lower tier, but for folks who are listening to us who are AI engineers,

877
00:57:44.080 --> 00:57:49.019
this basically means you can build Codex into your own app or platform, and you don't

878
00:57:49.080 --> 00:57:53.279
have to build the Codex itself. Codex is maybe one of the top harnesses in the world

879
00:57:53.360 --> 00:57:55.379
in terms of, uh, how to use this. Um.

880
00:57:57.160 --> 00:58:02.080
Peter Gostev: Yeah, and now I would say, yeah, this is a, a big deal for people, yeah, who, who

881
00:58:02.440 --> 00:58:06.980
implement AI in organizations because, and I used to do this in my previous job, where

882
00:58:08.080 --> 00:58:12.119
what you would do is that you would take a model and you're like, okay, decompose

883
00:58:12.200 --> 00:58:16.919
a process and put some, I don't know, cron jobs or whatever.

884
00:58:17.000 --> 00:58:22.440
And then what you end up doing 95% of the time, even if you are like, your job title

885
00:58:22.460 --> 00:58:26.940
is AI engineer, but what you're doing is stupid infrastructure things, which are completely

886
00:58:27.080 --> 00:58:29.779
boring and nothing to do with actual AI.

887
00:58:30.320 --> 00:58:36.019
So the more that we can remove that kind of thing, then we can accelerate

888
00:58:36.600 --> 00:58:38.719
deployment of AI within organizations.

889
00:58:39.360 --> 00:58:44.920
And, um, I think thi- this is a, a really good shift from OpenAI, just allows us to

890
00:58:44.920 --> 00:58:50.080
to do that more. And, um, I think if we see that, that I think what typically happens,

891
00:58:50.120 --> 00:58:53.479
right, a company innovates and then other companies do that kind of thing as well.

892
00:58:53.880 --> 00:58:57.900
So yeah, if we see other companies ship that with like open source opportunities as

893
00:58:57.920 --> 00:59:01.560
well, where you can swap in your models, that kind of thing is super valuable because

894
00:59:01.880 --> 00:59:05.780
just the more you can remove all of the other crap that I need to manage as an engineer,

895
00:59:05.880 --> 00:59:09.399
like they're absolutely golden. So yeah, I really like that.

896
00:59:10.080 --> 00:59:13.559
Alex Volkov: There, there's a distinction here that Wolfram, we, we talked about when I told you

897
00:59:13.580 --> 00:59:15.000
about this and why I want to cover this.

898
00:59:15.440 --> 00:59:19.339
Uh, yes, your agents or Codex or et cetera can already drive Codex, right?

899
00:59:19.360 --> 00:59:21.999
If you have Codex installed on your machine, you can basically like have computer

900
00:59:22.040 --> 00:59:24.280
use whatever and and drive Codex. This is

901
00:59:25.200 --> 00:59:29.059
without requiring that computer. This is like for Python on on JavaScript, like a

902
00:59:29.160 --> 00:59:31.280
node somewhere in the cloud that you're running.

903
00:59:31.560 --> 00:59:35.579
Uh, this is building that harness and agent there so that it can drive parallel programmatic

904
00:59:35.600 --> 00:59:36.620
tool calls that you don't want to build.

905
00:59:36.640 --> 00:59:40.080
Like Peter is completely right. There's like, uh, folks maybe don't want to maintain

906
00:59:40.120 --> 00:59:41.519
all of this. And also,

907
00:59:42.480 --> 00:59:47.219
it's a very important reminder. OpenAI trains their models in in coordination with

908
00:59:47.240 --> 00:59:50.000
their harness. Anthropic does the same with Claude Code.

909
00:59:50.959 --> 00:59:53.599
Famously, Gemini does this with, you know, with their harness as well.

910
00:59:53.880 --> 00:59:56.999
Um, the combo between the model and harness is a very important one.

911
00:59:57.120 --> 01:00:00.299
Models behave better in their own native harnesses because they've been trained on

912
01:00:00.480 --> 01:00:06.319
with them. And so this is a way for you to basically milk OpenAI's APIs for more

913
01:00:06.640 --> 01:00:08.639
than just the intelligence, the reasoning.

914
01:00:09.000 --> 01:00:10.440
A very interesting release from OpenAI.

915
01:00:10.840 --> 01:00:14.379
One last thing, unless Nisten, you want to comment, uh, uh, one last thing that's

916
01:00:14.440 --> 01:00:19.960
interesting here. This is, this feels like a big platform developer release, and this

917
01:00:20.040 --> 01:00:24.219
is exactly the kind of release they would have saved for Dev Day because all of the

918
01:00:24.320 --> 01:00:28.200
developers who use OpenAI in this way are showing up to Dev Day.

919
01:00:29.560 --> 01:00:32.400
Why am I saying this? I'm saying this because I'm super excited about what else they

920
01:00:32.440 --> 01:00:36.620
have for Dev Day if they decided to release this 2 weeks before the biggest developer

921
01:00:36.800 --> 01:00:40.579
convention of the year for OpenAI. So if they're releasing this on Dev Day, sorry,

922
01:00:40.720 --> 01:00:44.659
like 2 weeks before DevDay, I gotta wonder what they have in store at DevDay for us.

923
01:00:44.820 --> 01:00:47.080
I'm very excited about DevDay. It's coming up in 2 weeks.

924
01:00:47.200 --> 01:00:49.360
Uh, Peter and I will be there, uh, we'll do coverage.

925
01:00:49.720 --> 01:00:53.239
Uh, and if you are coming to DevDay, let me have a reminder for you.

926
01:00:53.520 --> 01:00:58.039
The day after DevDay is Fully Connected from CoreWeave, a 2-day

927
01:00:59.240 --> 01:01:01.879
conference with now, I think, 4,000 people as well.

928
01:01:02.480 --> 01:01:07.120
We have a free ticket for you to join, and as recently as last week, we talked about

929
01:01:07.540 --> 01:01:11.439
a very surprising guest at the, at Fully Connected.

930
01:01:11.680 --> 01:01:12.159
Pitbull,

931
01:01:13.320 --> 01:01:17.240
the international superstar, is going to headline the party at the end of Dev Day.

932
01:01:17.560 --> 01:01:21.020
So even if you know you're in San Francisco and you don't have the 2 days to confirm,

933
01:01:21.120 --> 01:01:23.000
come on Thursday and then have a party with us.

934
01:01:23.400 --> 01:01:28.079
Right down there below, without speaking about this code specifically, is the free

935
01:01:28.200 --> 01:01:31.479
ticket code for listeners of ThursdAI to Fully Connected 2026.

936
01:01:31.600 --> 01:01:32.360
I looked at the

937
01:01:33.280 --> 01:01:37.279
launch release thing, I looked at the headliners, I looked at a bunch of stuff.

938
01:01:37.560 --> 01:01:39.580
Uh, it's going to be very, very exciting.

939
01:01:39.639 --> 01:01:43.060
It's the biggest thing that CoreWeave has ever done, and yours truly, Wolfram and

940
01:01:43.100 --> 01:01:43.939
I will be live

941
01:01:44.840 --> 01:01:48.559
giving you ThursdAI from there. We're gonna cover the news, and then we'll also talk

942
01:01:48.600 --> 01:01:53.060
about some, uh, w- w- with some folks from CoreWeave about the, the industry.

943
01:01:53.200 --> 01:01:56.360
Wolfram, go ahead. Uh, so if you're coming to OpenAI's Dev, they come to Fully Connected,

944
01:01:56.400 --> 01:02:01.699
uh, September 30th, October 1st, and uh, we will get you there, and please, you know,

945
01:02:02.000 --> 01:02:05.599
use our thing to sign up so you don't have to pay.

946
01:02:06.080 --> 01:02:09.319
Uh, and this is a basic benefit of listening to Thursday.

947
01:02:11.800 --> 01:02:15.360
I will, I ve- will slowly transition to our next topic as well.

948
01:02:15.480 --> 01:02:19.699
In, in about 5 minutes, we're going to chat about assistance, so I think we can start,

949
01:02:19.800 --> 01:02:22.160
but before this, I will say, um,

950
01:02:23.120 --> 01:02:27.860
that the hackathon last week was an incredible success, and folks who came to the

951
01:02:27.960 --> 01:02:33.679
hackathon got early access to TypeSafe, uh, JEV model, uh, from from TypeSafe,

952
01:02:34.000 --> 01:02:35.719
uh, because they sponsored the hackathon last week.

953
01:02:36.000 --> 01:02:41.380
Uh, so I'm very, very excited to talk about, uh, uh, JEV very soon with Ali, but,

954
01:02:41.520 --> 01:02:43.639
uh, just before this, let's talk about,

955
01:02:44.880 --> 01:02:49.439
um, assistants. Peter, you asked this question before because you were busy with traveling,

956
01:02:49.480 --> 01:02:51.199
et cetera, whether or not folks used Muse.

957
01:02:51.600 --> 01:02:56.999
Um, I posted a video about Muse, and by far my most liked video on on YouTube, separately

958
01:02:57.040 --> 01:03:00.559
from having a conversation about this on ThursdAI, and it seems like

959
01:03:01.880 --> 01:03:03.220
we have, um,

960
01:03:05.440 --> 01:03:08.199
basically an insane race between assistants.

961
01:03:08.279 --> 01:03:12.559
We'll have David Paul on from Assistant Benchmark to join us, uh, very, very soon.

962
01:03:12.920 --> 01:03:17.519
Uh, however, I just want to do like a survey here and in the comments.

963
01:03:18.800 --> 01:03:24.119
Which assistant are you guys using right now, if at all, and whether or not you've

964
01:03:24.240 --> 01:03:26.039
changed in the past month, let's say.

965
01:03:26.279 --> 01:03:30.399
Uh, the options are obviously the standard ones for developers like us, OpenClaw Hermes,

966
01:03:30.640 --> 01:03:36.199
uh, and also the the latest and greatest from Grokbot, uh, Muse from Meta

967
01:03:36.680 --> 01:03:40.600
Instinct that we should cover in a moment, and some others like Town and Val and like

968
01:03:40.640 --> 01:03:44.179
a bunch of other ones. Uh, so let's start with the, with the comments already.

969
01:03:44.240 --> 01:03:46.360
Folks are saying Hermes, uh, Wolfram.

970
01:03:46.920 --> 01:03:51.220
Wolfram Ravenwolf: Hermes definitely. I me- I use Codex at work because we are not allowed to use Hermes

971
01:03:51.240 --> 01:03:55.619
there yet. But Hermes, I'm so, uh, in- involved in the ecosystem now.

972
01:03:55.680 --> 01:03:57.699
I am one of the contributors to the project.

973
01:03:57.800 --> 01:04:02.260
I have, I think we are almost at 40 patches now where I changed that and submitted

974
01:04:02.360 --> 01:04:04.600
PR, so I don't want to switch anymore.

975
01:04:04.680 --> 01:04:07.679
It's open source, it runs on my system, I'm locked in in that.

976
01:04:08.360 --> 01:04:09.039
It's my agent.

977
01:04:10.279 --> 01:04:15.980
Alex Volkov: So you're Hermes. Uh, Peter, do you have any assistants or using, uh, harnesses as

978
01:04:16.080 --> 01:04:17.979
assistants, and what is the difference between them?

979
01:04:18.040 --> 01:04:19.160
That's also an interesting question.

980
01:04:19.360 --> 01:04:24.839
Peter Gostev: I must say, I don't know if I'm behind the curve, but I, I sort of tried them out

981
01:04:24.960 --> 01:04:28.579
and I never end up using them. I don't know, they're just something about

982
01:04:30.000 --> 01:04:32.259
that I don't find that helpful. I find

983
01:04:32.279 --> 01:04:32.899
Alex Volkov: Let me ask you

984
01:04:32.940 --> 01:04:33.660
Peter Gostev: a bit more work.

985
01:04:33.839 --> 01:04:36.379
Alex Volkov: Let me ask you this in a different way, because I think the distinction there matters.

986
01:04:36.460 --> 01:04:37.519
We're going to talk about the distinction.

987
01:04:37.640 --> 01:04:43.620
Do you have anything a- agentic-wise that's with intelligence that tells you

988
01:04:43.660 --> 01:04:48.060
about the new incoming trigger from elsewhere, like a new email or a new message from

989
01:04:48.160 --> 01:04:49.820
Slack or et cetera? Do you have anything that's like

990
01:04:50.040 --> 01:04:50.179
Peter Gostev: No.

991
01:04:50.200 --> 01:04:53.399
Alex Volkov: proactive in a way that's like serving as your system with your context?

992
01:04:54.160 --> 01:04:57.920
Peter Gostev: No, the be- I do have like a few automations running.

993
01:04:58.120 --> 01:05:00.680
I wouldn't kind of quite classify it that way.

994
01:05:00.760 --> 01:05:05.480
So I w- I do have a few things where I'm like looking for new restaurants that opening

995
01:05:05.520 --> 01:05:06.660
in London or something like that.

996
01:05:06.680 --> 01:05:06.779
Alex Volkov: Yeah.

997
01:05:06.800 --> 01:05:08.760
Peter Gostev: And it just like, it runs every couple of weeks.

998
01:05:08.839 --> 01:05:12.119
So it it's kind of semi version of that, but it's very targeted.

999
01:05:12.440 --> 01:05:12.579
Alex Volkov: Yeah.

1000
01:05:12.680 --> 01:05:14.620
Peter Gostev: I found, for example, one thing I've got is-

1001
01:05:14.640 --> 01:05:16.000
Alex Volkov: It's an automation, not a heartbeat.

1002
01:05:16.040 --> 01:05:17.940
I think that this is like the thing that we're getting to.

1003
01:05:17.960 --> 01:05:18.419
Peter Gostev: Yeah, yeah.

1004
01:05:18.440 --> 01:05:20.579
Alex Volkov: It's not something that constantly runs for you,

1005
01:05:20.640 --> 01:05:20.779
Peter Gostev: Yeah.

1006
01:05:20.960 --> 01:05:23.979
Alex Volkov: rather it's something that you wanted and then you just didn't wanna

1007
01:05:24.160 --> 01:05:24.319
Peter Gostev: Yeah.

1008
01:05:24.560 --> 01:05:25.639
Alex Volkov: keep remembering to do.

1009
01:05:26.200 --> 01:05:31.939
Peter Gostev: Exactly. Yeah, so, but for the heartbeat stuff, I don't have anything that I found

1010
01:05:32.480 --> 01:05:37.200
particularly useful yet. And then, yeah, I don't know, I just prefer to check my own

1011
01:05:37.240 --> 01:05:40.439
emails and I prefer to, I don't know, manage my own calendar.

1012
01:05:40.560 --> 01:05:42.579
I don't know. That- that's how I feel like I know stuff.

1013
01:05:42.680 --> 01:05:43.019
I don't-

1014
01:05:43.040 --> 01:05:43.220
Alex Volkov: Yeah.

1015
01:05:43.240 --> 01:05:47.680
Peter Gostev: I- the- the same way how I write myself and I read things myself.

1016
01:05:47.760 --> 01:05:51.639
Like, there are a few things that I- I just actually that I feel like that's my job

1017
01:05:51.700 --> 01:05:56.279
as a human, uh, rather than just like a- a farming this out to an agent.

1018
01:05:56.320 --> 01:05:57.339
Alex Volkov: Uh,

1019
01:05:57.360 --> 01:05:57.839
Peter Gostev: But I might change

1020
01:05:57.860 --> 01:05:58.060
Alex Volkov: How about-

1021
01:05:58.080 --> 01:05:58.399
Peter Gostev: my mind.

1022
01:05:58.760 --> 01:06:01.040
Alex Volkov: How about you folks? Uh, let's have Yam.

1023
01:06:01.080 --> 01:06:04.399
Do you have any usage of assistants in your life currently?

1024
01:06:04.920 --> 01:06:06.899
Yam Peleg: Codex, Codex, like.

1025
01:06:06.920 --> 01:06:11.839
Alex Volkov: Do you feel that Codex is answering the, the, the, the, the, the threshold of assistant

1026
01:06:11.960 --> 01:06:13.860
versus just a AI agent?

1027
01:06:13.880 --> 01:06:19.899
Yam Peleg: Look, look, I have a lot of customizations, uh, around Codex, so it's not out

1028
01:06:19.940 --> 01:06:26.060
of the box. Uh, but I must say that all the new products like News and,

1029
01:06:26.280 --> 01:06:29.680
and uh, what was the last one? Oh, and, and Grokbot.

1030
01:06:30.120 --> 01:06:35.020
I want to try them as well, but currently have like a custom,

1031
01:06:36.040 --> 01:06:40.260
custom scripted, uh, whatever. You, you know how it goes, you know how it goes.

1032
01:06:40.280 --> 01:06:42.739
But mostly Codex, mostly Codex, I think.

1033
01:06:43.800 --> 01:06:44.580
Alex Volkov: Uh, all right, let's

1034
01:06:44.600 --> 01:06:45.779
Yam Peleg: It's amazing, seriously.

1035
01:06:45.800 --> 01:06:48.999
Alex Volkov: guest shows up, uh, Nisten and LDJ, I want to hear from you.

1036
01:06:50.680 --> 01:06:53.199
Nisten, let's start with you. Any assistants in your life?

1037
01:06:54.680 --> 01:06:55.839
Nisten Tahiraj: Yeah, uh,

1038
01:06:56.880 --> 01:06:58.879
I built the Nisten assistant.

1039
01:06:59.440 --> 01:07:00.980
Alex Volkov: Yeah, that's, that's a great answer, actually.

1040
01:07:01.000 --> 01:07:01.720
Yeah, yep, yep.

1041
01:07:02.360 --> 01:07:03.960
Nisten Tahiraj: Yeah, so I just, uh-

1042
01:07:04.000 --> 01:07:05.719
Alex Volkov: Also, it's a great name, the Nisten assistant.

1043
01:07:06.160 --> 01:07:09.139
Nisten Tahiraj: No, it's gonna be, I have no, it's just called assist.ts.

1044
01:07:09.200 --> 01:07:11.220
It's a single TypeScript file because

1045
01:07:11.240 --> 01:07:11.379
Alex Volkov: Yeah.

1046
01:07:11.440 --> 01:07:14.519
Nisten Tahiraj: I don't trust the dependencies. Uh, I

1047
01:07:15.520 --> 01:07:21.200
don't install any pip. Uh, the only dependency is bun, and it's extremely fast at

1048
01:07:22.200 --> 01:07:27.379
opening things and opening emails and stuff, but then I just end up delegating a lot

1049
01:07:27.420 --> 01:07:30.919
of the other tasks to simply to to Claude Code.

1050
01:07:31.640 --> 01:07:37.720
And uh, mainly I keep uh, Astra for reviews, and uh, I

1051
01:07:37.800 --> 01:07:42.460
I really like the Astra web interface, like for editing videos and making animations

1052
01:07:42.760 --> 01:07:44.960
and stuff. I just use that from the web.

1053
01:07:46.000 --> 01:07:46.579
So, uh,

1054
01:07:46.600 --> 01:07:48.859
Alex Volkov: So no. Your answer is no, you have your own personal one,

1055
01:07:48.880 --> 01:07:49.059
Nisten Tahiraj: Yeah.

1056
01:07:49.080 --> 01:07:52.760
Alex Volkov: and it's a specific thing. Right, let's, uh, LDJ, and then I'll talk, and then we

1057
01:07:52.840 --> 01:07:54.559
we'll have a place for our guest as well.

1058
01:07:55.760 --> 01:08:00.259
LDJ: I'm I'm kind of with Peter here, where I've I've tried a few different things and

1059
01:08:00.400 --> 01:08:03.659
Grokbot and other things, and I guess it really just depends on what exactly you're

1060
01:08:03.720 --> 01:08:08.379
doing, and I do end up kind of going in different phases where I'll be like working

1061
01:08:08.480 --> 01:08:11.979
on specific projects that are very different than others and sometimes find them more

1062
01:08:12.040 --> 01:08:17.280
useful. But, uh, for the most part, uh, using Hermes Agent just kind of as a, uh,

1063
01:08:18.240 --> 01:08:22.299
like a long-term persistent- like that benefit there, the- the long-term persistent

1064
01:08:22.400 --> 01:08:27.299
shot of- of really having a in-depth, uh, context of the past memories,

1065
01:08:27.559 --> 01:08:27.699
Alex Volkov: Yeah.

1066
01:08:27.840 --> 01:08:30.000
LDJ: find that useful. But besides that, really just kind of

1067
01:08:30.000 --> 01:08:35.400
David Pawlan: of schedule tasks with ChatGPT, find that really useful, not not anything too crazy

1068
01:08:35.480 --> 01:08:37.440
beyond that, but I want to start using Muse.

1069
01:08:37.840 --> 01:08:39.320
Alex Volkov: Uh, that's a very interesting thing.

1070
01:08:39.440 --> 01:08:41.000
I want to start using Muse, you you answer.

1071
01:08:41.040 --> 01:08:45.660
So I'll answer as well, folks. I've been, uh, completely obsessed with transitioning

1072
01:08:45.720 --> 01:08:48.700
my life into automation with proactive assistant.

1073
01:08:48.760 --> 01:08:52.560
I told you on the pod that I think that 2026 is the year of proactive assistance specifically,

1074
01:08:52.920 --> 01:08:58.240
and the reason I told you about this is, um, that I think that having assistants that

1075
01:08:58.320 --> 01:09:00.639
are active, I think is very, very important.

1076
01:09:00.920 --> 01:09:01.200
And so

1077
01:09:02.600 --> 01:09:05.379
I constantly right now use a Grokbot.

1078
01:09:05.440 --> 01:09:08.159
I have more than, we can take a look actually.

1079
01:09:08.240 --> 01:09:09.380
Oh no, we can't. Uh,

1080
01:09:10.320 --> 01:09:13.439
I'll tell you why later. Uh, but I use Grokbot constantly.

1081
01:09:13.520 --> 01:09:17.980
I use Muse since it came out, and I think it's just wonderful and incredible in, in

1082
01:09:18.040 --> 01:09:21.480
many, many, many ways. And um, I still have my Hermes.

1083
01:09:21.560 --> 01:09:24.640
I barely talk to it anymore, but Hermes has a bunch of context as well.

1084
01:09:25.080 --> 01:09:27.480
And here are the few... You know what?

1085
01:09:28.000 --> 01:09:30.400
Let's, let's do, let's do the invitation.

1086
01:09:30.520 --> 01:09:33.460
David Paulan just joined us. David, welcome to the show.

1087
01:09:33.880 --> 01:09:34.499
David Pawlan: Thank you, thank you.

1088
01:09:34.840 --> 01:09:36.519
Alex Volkov: Uh, nice to, nice to have you here.

1089
01:09:36.720 --> 01:09:41.039
You have been blowing up on my timeline, and it's a very interesting thing where X

1090
01:09:41.319 --> 01:09:41.999
timeline is

1091
01:09:42.960 --> 01:09:45.899
specifically tuned to the interests of the people.

1092
01:09:45.920 --> 01:09:46.139
David Pawlan: Yep.

1093
01:09:46.180 --> 01:09:49.559
Alex Volkov: And so on my timeline, you're all over my timeline, and Muse is all over my timeline

1094
01:09:49.600 --> 01:09:51.260
because I've been, like, glazing on Muse.

1095
01:09:51.280 --> 01:09:51.620
David Pawlan: Right.

1096
01:09:51.640 --> 01:09:55.159
Alex Volkov: Uh, on other folks, maybe that they didn't quite have that effect.

1097
01:09:55.440 --> 01:10:00.000
So first of all, I, uh, the thing that you built is showing up on my timeline, which

1098
01:10:00.060 --> 01:10:02.120
is Assistant Benchmark, which I would love to talk to you about.

1099
01:10:03.000 --> 01:10:07.619
But also, you joined at the point of the show where I asked everybody from the hosts

1100
01:10:07.720 --> 01:10:10.980
which assistant they use, and we have a very interesting kind of spectrum of of who

1101
01:10:11.040 --> 01:10:14.400
uses what. Would love to hear from you before you even introduce yourself who you

1102
01:10:14.420 --> 01:10:19.139
are and what you build. What is, um, w- what is your daily driver?

1103
01:10:19.200 --> 01:10:20.639
I'm assuming you're using an assistant.

1104
01:10:20.680 --> 01:10:20.899
David Pawlan: Yeah.

1105
01:10:20.940 --> 01:10:24.600
Alex Volkov: And then we can talk about the differentiation between an agent and an assistant,

1106
01:10:25.120 --> 01:10:25.599
what's what.

1107
01:10:26.240 --> 01:10:30.880
David Pawlan: Love it. Uh, number one assistant I'm using is Instinct for personal stuff, Grokbot

1108
01:10:31.000 --> 01:10:36.960
for more work-related things. Uh, reason being Grokbot, I, I really love the UI/UX

1109
01:10:37.240 --> 01:10:39.779
and creating different bots and the organization of it.

1110
01:10:40.280 --> 01:10:43.240
Uh, Instinct just lives in iMessage, so it's just a little easier for me.

1111
01:10:43.480 --> 01:10:48.020
It, it's a pure efficiency standpoint compared to using Muse, but it would, it would

1112
01:10:48.040 --> 01:10:49.199
be between Instinct and Muse.

1113
01:10:49.680 --> 01:10:54.560
Alex Volkov: Instinct and Muse. Interesting. Um, all right, so folks, let's talk about what an

1114
01:10:54.639 --> 01:10:57.699
assistant is. Uh, and uh, folks, if you don't mind, I'll take you back, but if you

1115
01:10:57.740 --> 01:11:00.360
have a question or you want to comment, uh, Wolfram, I'm actually gonna keep you here

1116
01:11:00.440 --> 01:11:03.800
because, uh, everybody else here doesn't necessarily use assistant, wants to try,

1117
01:11:03.840 --> 01:11:09.979
so please tune with us. Uh, uh, David, uh, so the reason why you've been blowing

1118
01:11:10.000 --> 01:11:13.099
up on my time is because I was overobsessed with assistant.

1119
01:11:13.280 --> 01:11:13.539
Muse,

1120
01:11:13.560 --> 01:11:13.740
David Pawlan: Yeah.

1121
01:11:13.760 --> 01:11:16.799
Alex Volkov: I think for the past 3 days, like, most of my tweets are were about Muse, now they're

1122
01:11:16.820 --> 01:11:18.639
about JEV, but before they were about Muse.

1123
01:11:18.680 --> 01:11:21.080
I still use it incredibly well. I give feedback to the team.

1124
01:11:21.600 --> 01:11:25.279
Um, and before that it was Grok, and before that it was Hermes, and before that it

1125
01:11:25.320 --> 01:11:27.520
was OpenClaw. So I've been like assistant built for a while.

1126
01:11:27.880 --> 01:11:32.199
Also Instinct, uh, which I, oof, I don't know what's going on there, but Instinct

1127
01:11:32.260 --> 01:11:37.600
is talking about raising a 10 billion dollar valuation, uh, recently, and this is

1128
01:11:38.440 --> 01:11:42.660
a startup that just started. So the world is also getting super excited about this

1129
01:11:42.960 --> 01:11:45.279
next wave of AI that's helpful for humans.

1130
01:11:45.919 --> 01:11:47.259
Let's define an assistant between us.

1131
01:11:47.360 --> 01:11:50.379
Like, I, I want like a, first of all, how do you think about what is the difference

1132
01:11:50.400 --> 01:11:52.520
between an assistant and just like an AI agent?

1133
01:11:52.840 --> 01:11:57.680
And then also, um, how that definition goes across labs, for example.

1134
01:11:58.240 --> 01:12:01.819
David Pawlan: Sure. Yeah, I mean, I think to me, an assistant is something that actually can execute

1135
01:12:01.840 --> 01:12:06.759
the task. It's, it's not just something that pings you or notifies you, uh, if anything,

1136
01:12:06.780 --> 01:12:10.479
that right, that creates like more work in a surface area where you're just now getting

1137
01:12:10.600 --> 01:12:14.439
constant notifications, but it's something that will actually then identify the issue

1138
01:12:14.480 --> 01:12:18.999
and then complete the task. Uh, maybe there, there's some approval in the loop in

1139
01:12:19.040 --> 01:12:23.839
that capacity, but it's actually going to take that, uh, workflow from start to finish.

1140
01:12:24.720 --> 01:12:27.960
Alex Volkov: Interesting. Uh, Wolfram, how do you, how do you feel the difference between an assistant

1141
01:12:28.040 --> 01:12:29.339
and just an agent or a harness?

1142
01:12:29.400 --> 01:12:34.140
Wolfram Ravenwolf: It's funny because AI, when we started out, it was always a user versus assistant

1143
01:12:34.240 --> 01:12:37.899
in the chat model. So that was the official name for the AI in the beginning.

1144
01:12:38.000 --> 01:12:42.759
It was just the assistant. Then the agents came out and started doing things, and

1145
01:12:42.800 --> 01:12:46.639
now we are back to having the agent be the assistant because it's actually a helper.

1146
01:12:46.960 --> 01:12:50.939
And that's for me the assistant. It is providing assistance, it's helping me, it's

1147
01:12:51.160 --> 01:12:56.600
an agent that does something on my behalf and helps me in all, basically everything.

1148
01:12:56.960 --> 01:12:59.420
Alex Volkov: Mm-hmm. My, my feeling is that

1149
01:13:00.520 --> 01:13:01.779
there is a barrier between

1150
01:13:03.960 --> 01:13:08.779
giving me more things to look at, like David, like you're saying, versus proactivity

1151
01:13:08.839 --> 01:13:10.540
of like reducing my cognitive load.

1152
01:13:10.800 --> 01:13:11.019
David Pawlan: Right.

1153
01:13:11.320 --> 01:13:15.179
Alex Volkov: I feel like assistants, specifically the new type of assistants with a heartbeat,

1154
01:13:15.200 --> 01:13:19.240
that they look for stuff before I need to tell them, with proactivity features built

1155
01:13:19.279 --> 01:13:22.360
in, with, uh, creative ways of using them as well.

1156
01:13:22.640 --> 01:13:26.180
They teach me how to use them. For example, Muse, we can talk about Muse proactive

1157
01:13:26.200 --> 01:13:28.439
features as well. I think that that's kind of the difference.

1158
01:13:28.839 --> 01:13:33.040
Whereas, um, the generalized harnesses agents were built as coding harnesses that

1159
01:13:33.360 --> 01:13:35.459
folks in labs just noticed they're super useful.

1160
01:13:35.480 --> 01:13:36.139
David Pawlan: Right.

1161
01:13:36.520 --> 01:13:38.120
Alex Volkov: Uh, but they're missing like a few things.

1162
01:13:38.140 --> 01:13:42.199
And I think that this is maybe the reason why Peter Steinberger went on such a huge,

1163
01:13:43.560 --> 01:13:47.360
uh, epic journey in the beginning of this year, talked to all of the labs, and now

1164
01:13:47.440 --> 01:13:52.419
I think this is also the reason why, uh, Muse has a heartbeat, has a soul.

1165
01:13:52.440 --> 01:13:52.699
David Pawlan: Right.

1166
01:13:52.720 --> 01:13:55.439
Alex Volkov: MD file and an agents and memory.md file.

1167
01:13:55.720 --> 01:13:59.879
And also, I don't know if you guys know this, uh, Muse has dreaming built in.

1168
01:14:00.000 --> 01:14:03.839
So you can literally ask Muse about what it dreamt on, and mu- dreaming is like another

1169
01:14:03.920 --> 01:14:06.100
process that that OpenClaw kind of started and

1170
01:14:06.520 --> 01:14:06.899
David Pawlan: Right.

1171
01:14:06.920 --> 01:14:09.500
Alex Volkov: has. Um, so this is an assistant for me.

1172
01:14:09.600 --> 01:14:13.619
David, the reason why you're here is not because you're just a random person that

1173
01:14:13.660 --> 01:14:16.359
I asked about, like, who uses the system, what is the definition?

1174
01:14:16.480 --> 01:14:20.720
Uh, you've built Assistant Benchmarks, specifically highlighting the stuff that assistants

1175
01:14:20.760 --> 01:14:22.320
do. Let's talk about Assistant Benchmark.

1176
01:14:22.920 --> 01:14:27.140
David Pawlan: Totally. So, uh, it generally came from, I was, I was testing them all myself, right?

1177
01:14:27.160 --> 01:14:30.660
Like, same as you, I was just using them all, doing my own market research and figuring

1178
01:14:30.700 --> 01:14:35.220
out what's working, what's not working, and I needed a way to compare them all with

1179
01:14:35.279 --> 01:14:39.580
one another. Um, it it's really more of like a use case driven benchmark, right?

1180
01:14:39.600 --> 01:14:42.759
Like, it's not like we have 10,000 phone numbers provisioned and we're running all

1181
01:14:42.860 --> 01:14:48.660
these iterations and, um, like, like a, like a AI Arena type, type

1182
01:14:49.279 --> 01:14:51.200
proper benchmark. It's really more use case driven.

1183
01:14:51.240 --> 01:14:54.479
So, you know, you click into a category, you can see the prompt that's being asked.

1184
01:14:54.920 --> 01:14:56.480
Um, we then will,

1185
01:14:57.520 --> 01:14:57.920
uh.

1186
01:14:58.240 --> 01:15:01.300
Alex Volkov: So let's talk about, uh, let's walk through some of the categories that you are testing.

1187
01:15:01.320 --> 01:15:01.420
David Pawlan: Sure.

1188
01:15:01.440 --> 01:15:04.920
Alex Volkov: I think it's gonna be interesting because we're also a podcast, so n- not many folks

1189
01:15:05.320 --> 01:15:06.660
are viewing what we're seeing.

1190
01:15:06.839 --> 01:15:07.180
David Pawlan: Yeah, totally.

1191
01:15:07.200 --> 01:15:10.579
Alex Volkov: So let's walk through this. Uh, so you have categories like travel assistance, email,

1192
01:15:10.760 --> 01:15:13.300
and finance, and shopping, uh, and and food.

1193
01:15:13.360 --> 01:15:16.399
And there's like numbers here, and I would love to tell you about like, uh, I would

1194
01:15:16.440 --> 01:15:19.980
love for you to tell us what like work in teams has 37, for example, and travel has

1195
01:15:20.080 --> 01:15:21.660
2. What what does this mean? It means that

1196
01:15:21.680 --> 01:15:22.019
David Pawlan: Right.

1197
01:15:22.040 --> 01:15:23.759
Alex Volkov: you've tested 2 assistants on this category?

1198
01:15:24.160 --> 01:15:28.040
David Pawlan: Uh, so that means that there's total of two assistants have submitted themselves for,

1199
01:15:28.080 --> 01:15:30.000
for the benchmark within that specific category.

1200
01:15:30.080 --> 01:15:35.179
So under the travel category, you have Miso and Soar, and those are, uh, travel specific

1201
01:15:35.280 --> 01:15:40.880
assistants. Um, if you look at the general category, those are like the Muse, the

1202
01:15:41.040 --> 01:15:44.299
Instinct, the, the ones that are just doing your like day-to-day workflows.

1203
01:15:44.360 --> 01:15:44.619
Alex Volkov: Mhm.

1204
01:15:45.000 --> 01:15:49.360
David Pawlan: Um, if you look at more of the work in teams, uh, there's 37 that have been submitted

1205
01:15:49.400 --> 01:15:51.839
there. Those are more on like the B2B side of things.

1206
01:15:52.280 --> 01:15:55.519
Um, they haven't been tested yet, and so we're, we're gonna be working through that.

1207
01:15:56.120 --> 01:16:00.420
What's crazy is there's 116 total assistants that have now been submitted on the site.

1208
01:16:00.680 --> 01:16:00.860
Alex Volkov: Yeah.

1209
01:16:01.400 --> 01:16:05.120
David Pawlan: Um, I, you know, a tweet that I put out earlier today was like, I, I literally talk

1210
01:16:05.160 --> 01:16:07.160
to my AI agents more than I talk to my girlfriend now.

1211
01:16:07.280 --> 01:16:11.200
It's like I have 8 agents always running, that I'm testing all them.

1212
01:16:11.680 --> 01:16:13.899
Um, it since the first week of launching

1213
01:16:14.960 --> 01:16:20.899
Assistant Benchmark, I have run, I have personally run 273 tests across 23 different

1214
01:16:20.960 --> 01:16:26.879
agents, um, and it's just the whole intention here is, okay, let's

1215
01:16:27.000 --> 01:16:30.920
run these same prompts, let's see how these different, uh, assistants are going to

1216
01:16:31.040 --> 01:16:36.279
respond to it. Let's also see how fast they are, and, uh, let's just compare them.

1217
01:16:36.320 --> 01:16:39.360
Let's see what's going on and give from twofold, it's one,

1218
01:16:40.360 --> 01:16:43.259
let's give the general public a way to actually understand

1219
01:16:44.600 --> 01:16:47.340
how these different assistants are performing, what they're good at, what they're

1220
01:16:47.360 --> 01:16:51.160
not good at. But then two, if you're building an assistant, like, these are the general

1221
01:16:51.280 --> 01:16:55.360
things that your assistant should be really good at, and if it's not performing well

1222
01:16:55.380 --> 01:16:59.200
at it, like, you should probably do some more work on your assistant to make sure

1223
01:16:59.240 --> 01:17:00.559
that it executes it really well.

1224
01:17:01.320 --> 01:17:05.320
Alex Volkov: So the type of stuff, so I was, I was incorrect before, the above there are categories

1225
01:17:05.400 --> 01:17:06.679
the assistants submit themselves to.

1226
01:17:06.760 --> 01:17:11.500
The general category has 59 assistants, and the type of tasks that it assists you

1227
01:17:11.559 --> 01:17:15.000
with are written here. It's like travel, uh, booking travel for you.

1228
01:17:15.280 --> 01:17:17.080
Uh, what is Pix? Tell us about Pix.

1229
01:17:17.640 --> 01:17:19.799
David Pawlan: Right, so these are the dimensions that are being tested.

1230
01:17:19.960 --> 01:17:25.000
Um, there's 16 total dimensions. Pix is more of like a recommendation engine.

1231
01:17:25.480 --> 01:17:27.659
So is it actually like recommending things, uh,

1232
01:17:27.720 --> 01:17:27.899
Alex Volkov: Mm.

1233
01:17:27.920 --> 01:17:32.259
David Pawlan: quite efficiently? So, uh, as an example, you know, if you say, hey, find me a vegan

1234
01:17:32.360 --> 01:17:33.499
restaurant in,

1235
01:17:34.400 --> 01:17:37.719
you know, in- in my area for the weekend, is it actually gonna find a vegan restaurant,

1236
01:17:37.760 --> 01:17:39.480
or is it just gonna find you a vegetarian restaurant?

1237
01:17:40.000 --> 01:17:43.340
Um, then you have other things like memory, where

1238
01:17:44.520 --> 01:17:49.699
it's not necessarily just your traditional memory of does it know who you are, but

1239
01:17:49.720 --> 01:17:52.120
also is it really good at, uh, picking up on context?

1240
01:17:52.200 --> 01:17:57.219
So, uh, as an example, uh, the way that I'll test memory is I'll first have it, I'll

1241
01:17:57.280 --> 01:18:00.639
first test the, the travel category, and I'll tell it that I'm traveling to a certain

1242
01:18:00.679 --> 01:18:02.419
weekend, uh, and to book me a flight.

1243
01:18:02.679 --> 01:18:02.859
Alex Volkov: Yeah.

1244
01:18:02.960 --> 01:18:08.299
David Pawlan: And then I'll go through a few other dimensions, and then after I'll test the, uh,

1245
01:18:08.320 --> 01:18:13.399
the picks, uh, or sorry, the, the online task, uh, dimension, which is then asking

1246
01:18:13.720 --> 01:18:15.800
to, uh, book me a reservation somewhere.

1247
01:18:16.840 --> 01:18:21.520
And I'll ask it to book me the reservation the same weekend that I asked it to travel,

1248
01:18:21.600 --> 01:18:26.060
but in a different city. And then I'll see, does the assistant preemptively prompt

1249
01:18:26.080 --> 01:18:29.140
me and say, hey, you're supposed to be in Chicago that weekend.

1250
01:18:29.240 --> 01:18:30.459
Why are you looking to book a res-

1251
01:18:30.480 --> 01:18:30.740
Alex Volkov: Yeah.

1252
01:18:30.840 --> 01:18:32.399
David Pawlan: reservation in New York? Like, are you sure?

1253
01:18:32.880 --> 01:18:35.999
The agents that can do that, it's pretty solid contextual memory, right?

1254
01:18:36.019 --> 01:18:37.419
Like, they're, they're keeping track of the pace.

1255
01:18:37.720 --> 01:18:37.860
Alex Volkov: Yeah.

1256
01:18:37.960 --> 01:18:40.600
David Pawlan: Uh, not all of them do that. So some of them do, some of them don't.

1257
01:18:40.960 --> 01:18:45.120
And so it's like, how do we test these different assistants and prompt them in specific

1258
01:18:45.200 --> 01:18:49.639
ways to really understand, like, the breadth and depth of what they can actually do

1259
01:18:49.680 --> 01:18:51.559
and what the different gaps are and all that.

1260
01:18:52.000 --> 01:18:54.920
Alex Volkov: I think it's a very, very interesting problem that you're trying to tackle.

1261
01:18:55.080 --> 01:18:59.280
The, the reason why is because the type of benchmarks that we talk about on the show,

1262
01:18:59.440 --> 01:19:04.180
usually for LLM intelligence, they usually start with LLM with like no context, and

1263
01:19:04.220 --> 01:19:07.700
then all of them get the same exact system prompt, and all of them exactly the same

1264
01:19:07.720 --> 01:19:11.699
exact task. And most of the time that task is verifiable with some code or LLM as

1265
01:19:11.760 --> 01:19:15.319
a judge on the other end. So you can like test this from yesterday to today, and supposedly

1266
01:19:15.360 --> 01:19:18.740
it doesn't change, and then between different versions of the model, it changes.

1267
01:19:19.120 --> 01:19:19.260
David Pawlan: Right.

1268
01:19:19.280 --> 01:19:24.920
Alex Volkov: Here, you're testing the 3 parts of what makes an AI, in my

1269
01:19:25.600 --> 01:19:27.379
view, which is model, harness, and context.

1270
01:19:27.480 --> 01:19:31.259
You're testing all 3 of them at the same time, including context that includes memory

1271
01:19:31.319 --> 01:19:32.759
from before. It's very interesting.

1272
01:19:32.960 --> 01:19:36.099
How do you approach, um, h- have you heard feedback about this?

1273
01:19:36.160 --> 01:19:40.060
H- how do you approach this from a perspective of, hey, assistants are also built

1274
01:19:40.100 --> 01:19:44.459
in to learn incrementally and be better incrementally w- the more you use them, right?

1275
01:19:44.500 --> 01:19:44.980
Like, that's what

1276
01:19:44.920 --> 01:19:45.379
David Pawlan: Totally.

1277
01:19:45.400 --> 01:19:49.479
Alex Volkov: the promise of Grokbot is, like, it better the more you use it, and and instinct and

1278
01:19:49.560 --> 01:19:52.960
memory, like, builds up for you. How do you treat this from a perspective of a clean

1279
01:19:53.040 --> 01:19:57.000
lab, let's test it in a vacuum, versus this is an assistant that's been going with

1280
01:19:57.080 --> 01:19:58.819
me, and it's just better at learning one

1281
01:19:59.120 --> 01:19:59.299
David Pawlan: Right.

1282
01:19:59.320 --> 01:19:59.859
Alex Volkov: versus the other?

1283
01:20:00.400 --> 01:20:03.060
David Pawlan: So, first off, we- it's not a lab, right?

1284
01:20:03.160 --> 01:20:06.740
It's and and make that very clear, like, this is not a lab grade research

1285
01:20:06.760 --> 01:20:06.859
Alex Volkov: Yeah.

1286
01:20:06.880 --> 01:20:11.160
David Pawlan: product. This is very use case driven, and and that's intentional, um, because the

1287
01:20:11.320 --> 01:20:16.560
the everyday individual who's looking at this type of benchmark is not the everyday

1288
01:20:16.600 --> 01:20:19.580
individual that's looking at an an arena LM, right?

1289
01:20:19.600 --> 01:20:26.000
Like the the the arenas, the the proper LLM benchmarks are being looked at by AI nerds,

1290
01:20:26.480 --> 01:20:31.400
LLM, you know, AI engineers that that are into the actual technical components.

1291
01:20:32.200 --> 01:20:37.120
The assistant benchmark is being looked at by a random Joe Schmo on the street who

1292
01:20:37.360 --> 01:20:40.499
is like, I don't know what assistant to use, can it book a flight for me?

1293
01:20:40.520 --> 01:20:41.680
And it's just very like basic things.

1294
01:20:42.120 --> 01:20:42.300
Alex Volkov: Yeah.

1295
01:20:42.760 --> 01:20:48.399
David Pawlan: So this benchmark is is m- meant for testing on a use case basis.

1296
01:20:48.960 --> 01:20:53.020
Um, some use cases like memory and, you know, how persistent is it in in context,

1297
01:20:53.040 --> 01:20:56.740
like there's things that we're not going to be able to test for another 30, 60 days

1298
01:20:56.880 --> 01:21:00.000
because who knows how long the memory actually is able to maintain there.

1299
01:21:00.000 --> 01:21:01.699
So it's, we're going to continue these tests over time.

1300
01:21:02.280 --> 01:21:02.580
Um,

1301
01:21:03.520 --> 01:21:08.559
Autumn Moulder, she, uh, recently was a SVP of engineering at Cohere.

1302
01:21:08.960 --> 01:21:13.039
Um, she reached out af- like 3 days after I launched it, and she's like, this is awesome.

1303
01:21:13.320 --> 01:21:16.999
I'm, you know, currently a free agent, just trying to figure out what, touching grass

1304
01:21:17.040 --> 01:21:19.920
for a little bit, but I, I want to do a new, new project for the people.

1305
01:21:20.360 --> 01:21:23.019
Um, so she reached out. She's been helping me out, which has been amazing, and she

1306
01:21:23.120 --> 01:21:26.580
brings a lot more of that, like, technical depth and knowledge to the project, more

1307
01:21:26.620 --> 01:21:31.659
from like a researcher lens. Um, and so we're, we're spinning up a, a new iteration

1308
01:21:31.700 --> 01:21:36.900
of testing, a lot more focused on like the B2B non-iMessage type agents, um, to try

1309
01:21:36.960 --> 01:21:41.860
to start testing those with a little bit more, uh, speed, like a little bit more velocity

1310
01:21:41.920 --> 01:21:43.399
and intensity. But

1311
01:21:45.120 --> 01:21:50.520
there's so many components to this benchmark at the end of the day, um, it's really

1312
01:21:50.580 --> 01:21:51.719
a way to see

1313
01:21:52.840 --> 01:21:54.740
legitimate performance of outcome, right?

1314
01:21:54.800 --> 01:21:57.440
This isn't about how quick is your model.

1315
01:21:57.560 --> 01:22:02.520
This isn't about, um, different, you know, AI engineering technicality.

1316
01:22:02.600 --> 01:22:03.020
It's about

1317
01:22:03.960 --> 01:22:08.759
if I'm your everyday person, I'm a 29-year-old guy in New York who's flying to Chicago.

1318
01:22:08.920 --> 01:22:14.840
If I message the, my assistant, hey, find me a flight to Chicago that's under 400

1319
01:22:15.040 --> 01:22:18.279
dollars in a window seat. Can it do it, yes or no?

1320
01:22:18.400 --> 01:22:20.619
And that's really the baseline here of what we're comparing.

1321
01:22:21.040 --> 01:22:23.079
Alex Volkov: That's great. Um, the thing that I

1322
01:22:24.600 --> 01:22:27.120
would like from a benchmark like this, I think is number one.

1323
01:22:27.200 --> 01:22:31.580
We talked about pacing the frontier before and all of the labs bringing on METR, for

1324
01:22:31.620 --> 01:22:32.300
example, and having

1325
01:22:32.320 --> 01:22:32.420
David Pawlan: Yeah.

1326
01:22:32.440 --> 01:22:36.180
Alex Volkov: their independent. The thing that, uh, would make me trust, and this is the benchmark,

1327
01:22:36.200 --> 01:22:38.680
besides the fact that everybody's quoting you because everybody kind of needs this

1328
01:22:38.720 --> 01:22:39.300
for themselves,

1329
01:22:39.600 --> 01:22:39.819
David Pawlan: Mhm.

1330
01:22:39.839 --> 01:22:43.039
Alex Volkov: uh, is independence, full independence from the labs.

1331
01:22:43.080 --> 01:22:47.019
I think it's very, very important. So I'll just say, uh, if any lab is paying for

1332
01:22:47.120 --> 01:22:50.299
some of this inference, I think this needs to be clear as well if somebody's providing

1333
01:22:50.339 --> 01:22:51.820
an extra tier so they'd be able to test them.

1334
01:22:51.880 --> 01:22:52.100
David Pawlan: Totally.

1335
01:22:52.120 --> 01:22:54.040
Alex Volkov: I think that's fair because otherwise how would you test them?

1336
01:22:54.360 --> 01:22:57.739
And also, I think the very interesting thing that folks have been asking you is that

1337
01:22:57.880 --> 01:23:00.980
the bigger ones that started this whole category, OpenClaw and Hermes, are not on

1338
01:23:01.040 --> 01:23:02.560
this. Would you want to address this on air?

1339
01:23:03.000 --> 01:23:05.000
David Pawlan: Yes, I'll touch on two points, like both of those.

1340
01:23:05.160 --> 01:23:07.739
So first, there's no one sponsoring this, right?

1341
01:23:07.760 --> 01:23:11.059
So I lead growth at Merit Systems. This is under my role at Merit.

1342
01:23:11.080 --> 01:23:13.920
We're just, we live in this world of agentic commerce, and so this started as just

1343
01:23:13.960 --> 01:23:17.300
a research project for me to really understand the landscape, to figure out, like,

1344
01:23:17.440 --> 01:23:20.399
where can we really play in this space from an infrastructure layer?

1345
01:23:20.960 --> 01:23:23.640
Um, so we're not being sponsored, we're not being paid.

1346
01:23:24.080 --> 01:23:29.239
Um, any of the competitors that are out there are not able to pay us for any type

1347
01:23:29.300 --> 01:23:31.140
of promotion or anything like that.

1348
01:23:31.480 --> 01:23:31.619
Alex Volkov: Yeah.

1349
01:23:31.720 --> 01:23:36.639
David Pawlan: Um, there might be a day where we say, okay, we need to work with someone like CoreWeave

1350
01:23:36.680 --> 01:23:39.540
and get more inference, or whoever it be.

1351
01:23:39.560 --> 01:23:41.100
Alex Volkov: Let's talk about this after the show.

1352
01:23:41.240 --> 01:23:45.339
David Pawlan: Uh, totally. Um, and and we've actually recently floated the idea of, like, having

1353
01:23:45.380 --> 01:23:46.300
a sponsors page,

1354
01:23:46.520 --> 01:23:46.740
Alex Volkov: Mhm.

1355
01:23:46.760 --> 01:23:53.000
David Pawlan: uh, to give us a pathway to test things with with a little more rigor, um,

1356
01:23:53.120 --> 01:23:54.800
obviously because that that's gonna cost a lot of money.

1357
01:23:55.120 --> 01:23:55.259
Alex Volkov: Yeah.

1358
01:23:55.360 --> 01:23:59.140
David Pawlan: Um, so if any of anything like that does happen, there's going to be a very clear

1359
01:23:59.240 --> 01:24:03.239
page of like, here is who is, who is paying, and here is what they are paying for,

1360
01:24:03.560 --> 01:24:05.540
and it is specifically for the research outcome.

1361
01:24:05.600 --> 01:24:08.680
There is, there's no profit coming from the site in any capacity.

1362
01:24:09.360 --> 01:24:14.959
Um, in terms of OpenClaw and Hermes, so the original reason that they were not included,

1363
01:24:15.320 --> 01:24:21.359
uh, is because the performance of them is significantly dependent on the setup from

1364
01:24:21.380 --> 01:24:25.720
the individual. So when you're thinking about testing these and scaling it across

1365
01:24:25.839 --> 01:24:30.880
all of them, if I run the same prompt of book me a flight to Chicago, blah, blah,

1366
01:24:30.920 --> 01:24:37.019
blah, um, my OpenClaw setup might do it unbelievably well, but your OpenClaw

1367
01:24:37.120 --> 01:24:38.620
setup might fail. And

1368
01:24:39.880 --> 01:24:45.139
this then doesn't prove as a valid use case point because if then someone says, oh

1369
01:24:45.200 --> 01:24:48.720
great, OpenClaw is really good at this, and then they start and install and build

1370
01:24:48.800 --> 01:24:52.220
their own OpenClaw setup, and then it fails, well, now they're going to look at the

1371
01:24:52.240 --> 01:24:54.740
benchmark and they're going to say, well, what the hell, it said it was really good

1372
01:24:54.760 --> 01:24:59.800
at it. Um, so the intention here is like we're really focusing on out-of-the-box solutions,

1373
01:24:59.960 --> 01:25:04.600
consumer products that are pre-built packages that you can just use, and everyone

1374
01:25:04.640 --> 01:25:09.040
who uses them, everyone who installs it or, uh, opens it up in iMessage, whatever

1375
01:25:09.060 --> 01:25:11.399
it is, every single person is having the exact same experience.

1376
01:25:12.080 --> 01:25:17.540
Alex Volkov: Yep, that makes sense. And, uh, I think the world has shifted from the folks who are

1377
01:25:17.600 --> 01:25:20.620
building their own assistant in, in their own, like, Mac Mini that they bought

1378
01:25:20.640 --> 01:25:20.980
David Pawlan: Right.

1379
01:25:21.000 --> 01:25:23.899
Alex Volkov: towards the one that breaks less than the previous one,

1380
01:25:24.080 --> 01:25:24.300
David Pawlan: Right.

1381
01:25:24.320 --> 01:25:27.500
Alex Volkov: towards now these labs are maintaining the machine,

1382
01:25:27.520 --> 01:25:27.740
David Pawlan: Right.

1383
01:25:27.760 --> 01:25:29.319
Alex Volkov: which is, I think, a very important thing.

1384
01:25:29.360 --> 01:25:33.019
We didn't talk about it. Uh, all of these assistants need, like, a computer environment.

1385
01:25:33.240 --> 01:25:35.000
All these assistants need a browser with computer use.

1386
01:25:35.120 --> 01:25:35.399
All these

1387
01:25:35.520 --> 01:25:35.660
David Pawlan: Yep.

1388
01:25:35.800 --> 01:25:38.120
Alex Volkov: need access to your memories and connectors for your context.

1389
01:25:38.160 --> 01:25:40.360
I think those also the things that signify an assistant.

1390
01:25:40.400 --> 01:25:44.920
David, I did promise you to let you go once you, once you can, because I, you have

1391
01:25:45.040 --> 01:25:46.760
elsewhere to be. Thank you so much for coming.

1392
01:25:47.040 --> 01:25:47.419
David Pawlan: Of course.

1393
01:25:47.560 --> 01:25:47.819
Alex Volkov: And do

1394
01:25:47.960 --> 01:25:48.500
David Pawlan: Thank you for having me.

1395
01:25:48.520 --> 01:25:51.140
Alex Volkov: I think it's very important. We are definitely expecting you to come back when the

1396
01:25:51.200 --> 01:25:52.860
new assistant comes out or new capabilities

1397
01:25:52.880 --> 01:25:53.220
David Pawlan: Please.

1398
01:25:53.240 --> 01:25:57.000
Alex Volkov: come out. Like, new capabilities came out this week for both Instinct and Muse, where

1399
01:25:57.040 --> 01:25:59.899
they can actually phone businesses, which I wanted to talk to you about, but we ran

1400
01:25:59.920 --> 01:26:01.140
out of time. David, thank you so much.

1401
01:26:01.160 --> 01:26:03.620
We'll bring you on. David Paulin, uh, creator of Assistant Bench.

1402
01:26:03.660 --> 01:26:03.880
Thank you.

1403
01:26:03.920 --> 01:26:04.959
David Pawlan: Appreciate it. Thanks, everyone.

1404
01:26:06.080 --> 01:26:10.680
Alex Volkov: Alrighty, folks. So a very, very interesting, uh, transition from, uh, assistants

1405
01:26:10.800 --> 01:26:15.099
and how they, they work into, you know, the world of, the world of computers.

1406
01:26:15.200 --> 01:26:17.399
We have back, uh, Francesco Bonacci.

1407
01:26:17.480 --> 01:26:18.859
Welcome, Francesco. How are you, my friend?

1408
01:26:18.880 --> 01:26:19.459
Francesco Bonacci: Hey.

1409
01:26:19.480 --> 01:26:22.600
Alex Volkov: Uh, the founder of Kua, and which you can find at TryKua.

1410
01:26:22.760 --> 01:26:26.459
And since you've been on and we've talked about computer use, multiple things have

1411
01:26:26.480 --> 01:26:30.019
happened. So I actually didn't bring you on here as a guest, believe it or not, just

1412
01:26:30.080 --> 01:26:34.559
as one commentator, because the world of computer use is exploding, obviously, because

1413
01:26:34.679 --> 01:26:38.839
assistance without computer use is not every one of those assistants needs a computer

1414
01:26:38.920 --> 01:26:43.440
use element because the web is human shaped and not assistant shaped.

1415
01:26:43.760 --> 01:26:47.039
A few things I would love for you to talk about before our next segment, before we

1416
01:26:47.080 --> 01:26:52.399
talk, we'll actually transition to to to to Jeff together, is um, uh,

1417
01:26:53.480 --> 01:26:57.780
how are you seeing this world now that Muse has released something and they have their

1418
01:26:57.820 --> 01:27:02.240
own computer use, and Grokbot, their computers wasn't like the the the craziest ones.

1419
01:27:02.679 --> 01:27:02.999
Um,

1420
01:27:04.000 --> 01:27:08.679
and and ye- how how are you feeling about this world of computer use, uh, now that

1421
01:27:08.720 --> 01:27:12.999
the assistant category is blowing up as much as we predicted it would be a few short

1422
01:27:13.040 --> 01:27:13.399
months ago?

1423
01:27:13.640 --> 01:27:18.740
Francesco Bonacci: Yeah. Yeah, I guess it's like, I would like to start with, uh, it was about time,

1424
01:27:19.040 --> 01:27:22.499
honestly, because I've been in the space for about 2 years, and I saw like these agents

1425
01:27:22.559 --> 01:27:27.379
like scoring from 10% on OS world when it first came out, and I was still working

1426
01:27:27.400 --> 01:27:32.419
at Microsoft. It was like, okay, we see a future where these assistants are gonna

1427
01:27:32.440 --> 01:27:34.159
be able like to use the same tools as human.

1428
01:27:34.760 --> 01:27:38.399
Um, how far is the future? Uh, definitely would have not expected like to be like

1429
01:27:38.520 --> 01:27:44.460
2 years, um, away, but it's definitely like such a cool and

1430
01:27:44.580 --> 01:27:50.619
odd space to be in right now. Um, we, if you look at our, not only about like

1431
01:27:50.679 --> 01:27:56.199
GitHub graph, but like our GitHub contribution graphs, we have like folks that are

1432
01:27:56.280 --> 01:27:59.999
like very keen to go and fix their core driver installation.

1433
01:28:00.520 --> 01:28:00.740
Alex Volkov: Mhm.

1434
01:28:00.920 --> 01:28:07.160
Francesco Bonacci: Um, and like putting patches on whatever like their Hermes or OpenClaw or like Muse

1435
01:28:07.200 --> 01:28:11.460
setup they have in place. Um, ultimately I got like many folks reaching out as well

1436
01:28:11.520 --> 01:28:13.280
that are using RockBot with CoreDriver.

1437
01:28:13.920 --> 01:28:16.999
Um, since then we also been speaking with the team, okay, can we do something?

1438
01:28:18.040 --> 01:28:23.019
Like, it's, it's honestly like, you know, it's uh, it's a good space to be.

1439
01:28:23.960 --> 01:28:29.159
Alex Volkov: I think, um, the, the thing that, a- again, the sig- s- s- specifies an assistant

1440
01:28:29.200 --> 01:28:31.120
for me is the ability to solve things.

1441
01:28:31.519 --> 01:28:32.799
I do want to talk about the

1442
01:28:33.960 --> 01:28:39.919
residential IP versus browser cloud IP problem, because when I use,

1443
01:28:40.120 --> 01:28:43.999
uh, Hermes, for example, it runs on my Mac Mini, runs from my home network and uses

1444
01:28:44.120 --> 01:28:48.159
CUA driver. Uh, when it clicks buttons on my computer, it looks like me.

1445
01:28:48.200 --> 01:28:51.720
So Cloudflare, like, chills a little bit, and other places that look ba- based on

1446
01:28:51.760 --> 01:28:55.280
detect IP, they they they let it use the the web more.

1447
01:28:55.840 --> 01:29:01.819
Versus when I use Grok, uh, or Muse, they get flagged more as agents, and so

1448
01:29:01.880 --> 01:29:06.120
I need to interject more and and and take over, for example.

1449
01:29:06.520 --> 01:29:10.219
Um, should you comment on this? What are your thoughts on this, like, area, whether

1450
01:29:10.240 --> 01:29:11.000
or not the web is

1451
01:29:11.400 --> 01:29:11.699
Francesco Bonacci: Yeah.

1452
01:29:11.720 --> 01:29:14.120
Alex Volkov: going to prevent this use or actually get solved?

1453
01:29:15.160 --> 01:29:17.680
Francesco Bonacci: Yeah, so we hear that all the times from our users.

1454
01:29:18.200 --> 01:29:21.919
Um, there are a couple of ways to circumnavigate these, I would like just call them

1455
01:29:22.080 --> 01:29:27.199
limitations, but it's more how these system were put in place to prevent scraping.

1456
01:29:27.400 --> 01:29:31.379
Historically, we end-to-end testing, and then it translated it into a large action

1457
01:29:31.480 --> 01:29:37.219
models and agents. Um, so for CoreDriver, you can do things the desktop

1458
01:29:37.560 --> 01:29:41.359
way, and that's basically when you interact with accessibility trees or you just rely

1459
01:29:41.440 --> 01:29:46.280
on pixel coordinates, and that's actually less detectable than, like, starting a CDP

1460
01:29:46.440 --> 01:29:47.620
connection over your browser.

1461
01:29:47.640 --> 01:29:47.980
Alex Volkov: Mm-hmm.

1462
01:29:48.000 --> 01:29:51.920
Francesco Bonacci: Because then actually, okay, what happens is that Cloudflare detect that you're using

1463
01:29:52.040 --> 01:29:54.759
either Playwright or just the Chrome-like protocol.

1464
01:29:55.320 --> 01:29:59.260
So we can go there with CodeDriver whenever there is the need, whenever we need to

1465
01:29:59.360 --> 01:30:02.520
ground more information in whatever action CodeDriver needs to do.

1466
01:30:03.200 --> 01:30:08.259
But we have like, we have a agent action ladder model that is smart at this stage.

1467
01:30:08.440 --> 01:30:12.319
It knows where it has exhaust all the option and it need to rely on browser use.

1468
01:30:13.120 --> 01:30:15.199
Um, because again, that's more detectable.

1469
01:30:15.960 --> 01:30:17.219
And also the way that

1470
01:30:18.120 --> 01:30:23.279
we can, we can mimic the user picking and typing in a kind of like more normal way

1471
01:30:23.560 --> 01:30:28.119
using code driver and accessibility and pixel coordinate than relying on, right?

1472
01:30:28.280 --> 01:30:32.639
Alex Volkov: The funniest thing that happened to me this week with Muse was, uh, Muse told me about,

1473
01:30:32.720 --> 01:30:34.479
hey, I got this, uh,

1474
01:30:35.760 --> 01:30:40.379
CAPTCHA thing. I can solve it for you, or I can, you know, give you the, the controls

1475
01:30:40.440 --> 01:30:41.879
of the browser so you can solve this.

1476
01:30:42.280 --> 01:30:46.139
And the very funny thing is there's two buttons, like solve this for me, and no, I-

1477
01:30:46.200 --> 01:30:48.699
I'm gonna do this, like, why would I ever want to solve myself?

1478
01:30:48.800 --> 01:30:52.219
Only if you don't succeed. It's very interesting that these models are, you know,

1479
01:30:52.280 --> 01:30:54.080
beating some of the very, very basic ones.

1480
01:30:54.440 --> 01:30:56.380
Uh, I think actually in that case it was wrong.

1481
01:30:56.440 --> 01:30:59.519
It was a cloud for check that kind of looks like a check, but there's a lot more going

1482
01:30:59.560 --> 01:31:02.879
on. Uh, the question that I had for you was a load bearing question, but there is

1483
01:31:02.940 --> 01:31:04.479
a feature that I've been waiting for.

1484
01:31:04.680 --> 01:31:06.879
So there's two things I- I'm telling you and the audience, everyone.

1485
01:31:07.160 --> 01:31:12.399
Uh, first of all, Muse, the iOS app, has a Tailscale connector, which is incredible

1486
01:31:12.440 --> 01:31:16.220
to see. If you guys don't use TailScale, you- you should 100%.

1487
01:31:16.280 --> 01:31:19.719
It's incredible. But if you want to give your assistant access to your network in

1488
01:31:19.780 --> 01:31:22.719
a secure way, Muse has a Tailscale connector.

1489
01:31:22.880 --> 01:31:27.879
Obviously, every agent with a, uh, VM, you can install Tailscale, approve it yourself,

1490
01:31:27.920 --> 01:31:31.239
like that's possible. But the fact that they have a connector for people is just incredible.

1491
01:31:31.260 --> 01:31:34.580
It's a product from Meta that goes out to billions of people and now has Tailscale

1492
01:31:34.600 --> 01:31:39.999
there as number one. And number two, in terms of residential IP, a lot of these, uh,

1493
01:31:40.200 --> 01:31:44.139
let's make it harder for bots to scrape our website websites that prevent your agents

1494
01:31:44.200 --> 01:31:48.619
from using them, they are looking at IPs, and there's like a segment of IPs for the

1495
01:31:48.720 --> 01:31:52.320
clouds, for Google Cloud and Lambda from AWS, et cetera.

1496
01:31:52.760 --> 01:31:55.559
Uh, and so they would like flag, oh, this is a potential automation thing.

1497
01:31:56.040 --> 01:32:02.360
However, Grok came out with a feature that now proxies the gra- the the traffic

1498
01:32:02.520 --> 01:32:05.039
for the browser through your local machine.

1499
01:32:05.760 --> 01:32:09.399
And I've been wanting this. This is a feature in in in uh Tailscale, by the way.

1500
01:32:09.680 --> 01:32:10.399
Uh, it's called the

1501
01:32:11.320 --> 01:32:12.399
exit router or something like that.

1502
01:32:12.680 --> 01:32:15.860
Uh, but it does- didn't work with with with Grok specifically when I tried it.

1503
01:32:16.080 --> 01:32:19.720
Grok has now this built in, so if you're using the Grok app and you want the Grok

1504
01:32:19.839 --> 01:32:24.699
IP to kind of look like it's coming from your machine, definitely checkbox this as

1505
01:32:24.880 --> 01:32:29.519
very important. Francesco, I want to switch towards a little bit of a, a conversation.

1506
01:32:29.640 --> 01:32:31.319
Everybody else also, please feel to chime in here,

1507
01:32:32.240 --> 01:32:36.820
of user computer use, uh, background stuff like we talked about, clicking buttons

1508
01:32:36.960 --> 01:32:39.720
in different apps for me, checking, taking screenshots, et cetera.

1509
01:32:40.080 --> 01:32:44.779
In the world of web use, uh, because there's accessibility there, there's also screenshotting

1510
01:32:44.839 --> 01:32:49.560
there, and there's WebMCP. And I would love to hear from you, like, where that world

1511
01:32:49.620 --> 01:32:52.940
is going, because I think that it's still very inefficient to take a screenshot and

1512
01:32:52.960 --> 01:32:53.759
make a decision, right?

1513
01:32:54.160 --> 01:32:54.399
Francesco Bonacci: Yep.

1514
01:32:55.440 --> 01:33:00.540
Uh, another that definitely we are proud about with Quadra, we were pretty much

1515
01:33:02.000 --> 01:33:05.519
there catching up with the latest spec of the MCP protocol.

1516
01:33:06.160 --> 01:33:09.359
Even before he landed, the latest specification, I don't recall which date it is,

1517
01:33:09.640 --> 01:33:12.800
but it's basically detailing a new way of interacting with skills.

1518
01:33:13.360 --> 01:33:17.640
Um, the old fashioned way, and if you go back to the early OpenClaw days when Peter

1519
01:33:17.720 --> 01:33:23.499
was pitching CLI plus skill is actually great, better than MCP, um, you will actually

1520
01:33:23.520 --> 01:33:24.719
like try and embed

1521
01:33:25.920 --> 01:33:31.079
much of the behavior that you wanted to detail for your agent within a skill or set

1522
01:33:31.120 --> 01:33:36.199
of skills. Um, and unfortunately, there is only as much as you can fit in the MCP

1523
01:33:36.320 --> 01:33:36.800
description.

1524
01:33:37.760 --> 01:33:42.279
Um, and one particular issue is that we are integrating in CuaDriver is skills over

1525
01:33:42.360 --> 01:33:48.019
MCP, which basically means that, okay, say in our case we embed with the CuaDriver

1526
01:33:48.080 --> 01:33:52.540
installations a skill for working on CuaDriver for macOS, Windows, and Linux.

1527
01:33:52.600 --> 01:33:55.680
There is definitely different detours that the agent has to take, whether it's Linux

1528
01:33:55.720 --> 01:33:57.519
and Windows. That stuff,

1529
01:33:58.920 --> 01:34:02.240
potentially if CuaDriver is dealing with multiple operating systems in the same way,

1530
01:34:02.600 --> 01:34:07.639
it can go and fetch the skill that is required for working with OS or Windows on the

1531
01:34:07.680 --> 01:34:08.899
fly, which is great.

1532
01:34:09.720 --> 01:34:11.379
Alex Volkov: I think that's fascinating.

1533
01:34:11.640 --> 01:34:11.779
Francesco Bonacci: Yeah.

1534
01:34:11.800 --> 01:34:16.060
Alex Volkov: How we're evolving from, hey, there was MCP and the tool definitions were overloading

1535
01:34:16.120 --> 01:34:19.740
my fucking context, so I wasn't able to use them because my context wasn't that big.

1536
01:34:20.160 --> 01:34:21.800
And then, so that's why we invented skills.

1537
01:34:22.000 --> 01:34:25.259
Skills are these things that like very small and only pulls up the whole context my

1538
01:34:25.360 --> 01:34:28.880
needs to now, hey, skills basically talk about how to use an API.

1539
01:34:29.240 --> 01:34:33.540
Why not combine them so that when I provide my API with MCP, I also provide the skill

1540
01:34:33.580 --> 01:34:37.499
of how to use this best. I think it's like absolutely fascinating development.

1541
01:34:37.560 --> 01:34:38.619
And you told me about this, by the way.

1542
01:34:38.660 --> 01:34:41.660
I think, I think you, you mentioned this first, like, what, what is, what the hell

1543
01:34:41.700 --> 01:34:45.279
is skills over MCP? And now I think it's like a very fascinating development.

1544
01:34:45.600 --> 01:34:48.379
MCP obviously also comes with authorization built in.

1545
01:34:48.440 --> 01:34:52.440
There's like a bunch of things, but now MCP also will come with, hey, here's how to

1546
01:34:52.480 --> 01:34:55.720
use this. And I think the problem, the main problem that people had with MCP, the

1547
01:34:55.840 --> 01:34:59.100
overloads, your context, is also getting solved because they're programmatic calling

1548
01:34:59.120 --> 01:35:00.820
from MCP in code mode and stuff like that, right?

1549
01:35:02.760 --> 01:35:02.899
Francesco Bonacci: Yeah.

1550
01:35:02.920 --> 01:35:05.959
Alex Volkov: And so I think that, uh, uh, you know, the, the world's very exciting.

1551
01:35:06.120 --> 01:35:09.799
And specifically in the web, uh, I think that that's something that where we're going,

1552
01:35:09.880 --> 01:35:13.579
where versus like trying to take a screenshot, going through all of the possible clicks

1553
01:35:13.660 --> 01:35:19.180
there, the website will tell the agents how to use them on their own, right?

1554
01:35:19.240 --> 01:35:20.839
So I think that is like very, very exciting.

1555
01:35:20.960 --> 01:35:25.799
Web MCP is coming up. Uh, Francesco, the last thing as a transition to the next segment

1556
01:35:25.840 --> 01:35:29.519
of the show where we need to talk about, uh, TypeSafe and and JEV.

1557
01:35:29.960 --> 01:35:34.259
Uh, you posted something about computer use and decision trees, et cetera.

1558
01:35:34.360 --> 01:35:37.060
Do you want to talk about this a little bit before we introduce our next guest on

1559
01:35:37.100 --> 01:35:37.399
the show?

1560
01:35:38.040 --> 01:35:42.820
Francesco Bonacci: Yeah, briefly. Honestly, I was like forwarded this demo the other day by smaller teammates.

1561
01:35:42.880 --> 01:35:44.799
Yesterday I didn't have time to catch up with everything.

1562
01:35:45.280 --> 01:35:50.619
Um, but it's honestly mind blowing the fact that right now we're able to defer many,

1563
01:35:50.880 --> 01:35:55.160
many, many choices, like in the decision tree of the trajectory of an agent, say that

1564
01:35:55.220 --> 01:36:00.720
it needs to pick type or scroll, or like which target elements it needs to work with.

1565
01:36:01.200 --> 01:36:05.039
Uh, back at my time at Microsoft, we used to have, you know, where when vision models

1566
01:36:05.280 --> 01:36:09.240
were not bounded in pixel coordinates, we used to have something called set of mark

1567
01:36:09.280 --> 01:36:14.079
prompting, which basically draw rectangles on your, on your screen depending on which

1568
01:36:14.160 --> 01:36:17.559
elements you want to click. Uh, and there were a couple of interesting browser use

1569
01:36:17.600 --> 01:36:22.480
and desktop use demos made yesterday, um, which basically is like, okay, you're just

1570
01:36:22.559 --> 01:36:27.419
feeding this, this decision tree back into that, and it's

1571
01:36:28.400 --> 01:36:32.019
when you don't, when you just, when you need to take quick choices, that for computer

1572
01:36:32.059 --> 01:36:35.460
use agent is like, okay, which action I need to target first between clicking, typing,

1573
01:36:35.480 --> 01:36:36.299
and scrolling? Which

1574
01:36:36.320 --> 01:36:39.860
Alex Volkov: Brother, you are bearing the lead. Let me say this for you in the way that I expect

1575
01:36:39.880 --> 01:36:41.919
you to say this, because I think it's very important.

1576
01:36:42.160 --> 01:36:46.919
Last week, Astra was the fastest, maybe not the cheapest, but definitely the fastest

1577
01:36:47.000 --> 01:36:50.339
computer use there out there, and you guys are like saying, you know, like that's

1578
01:36:50.360 --> 01:36:51.999
the LLM that needs the driver, et cetera.

1579
01:36:52.360 --> 01:36:57.220
Uh, a week passed, and in that week, we're looking now at a chart from you guys at

1580
01:36:57.320 --> 01:37:03.360
Computer Use that says, hey, for 79, oh, sorry, 71 out of the, the, the 80 tasks

1581
01:37:03.380 --> 01:37:07.399
that we gave it, uh, the new paradigm from TypeSafe that's called JEV

1582
01:37:07.440 --> 01:37:07.599
Francesco Bonacci: Yeah.

1583
01:37:08.160 --> 01:37:12.699
Alex Volkov: is now hitting that with a 0.1 second accuracy, whereas Astro took

1584
01:37:12.720 --> 01:37:12.820
Francesco Bonacci: Yeah.

1585
01:37:12.840 --> 01:37:18.419
Alex Volkov: 5 seconds, with uh, 0.1 second medium latency, whereas Astro took 4 seconds, and a

1586
01:37:18.440 --> 01:37:18.639
Francesco Bonacci: Yeah.

1587
01:37:19.360 --> 01:37:25.139
Alex Volkov: in- impossible to calculate difference in price, maybe 400x terms price, they can

1588
01:37:25.200 --> 01:37:29.239
do most of these decisions. I think that this is why, uh, I think it was very exciting

1589
01:37:29.320 --> 01:37:33.859
to see that computer use is now not only getting solved conceptually, but also like

1590
01:37:33.920 --> 01:37:34.020
Francesco Bonacci: Yeah.

1591
01:37:34.080 --> 01:37:36.399
Alex Volkov: in practically, it's gonna be super faster as well.

1592
01:37:36.720 --> 01:37:40.760
And so I think this is a good transition to bring on Ali Labs from TypeSafe.

1593
01:37:40.800 --> 01:37:43.240
Ali, welcome to the show. Uh,

1594
01:37:44.440 --> 01:37:47.799
let's start. with, first of all, congratulations.

1595
01:37:48.400 --> 01:37:48.879
What a

1596
01:37:49.080 --> 01:37:49.539
Allie Laabs: Thank you.

1597
01:37:49.840 --> 01:37:54.979
Alex Volkov: crazy, crazy launch you guys had. Was it yesterday only?

1598
01:37:55.120 --> 01:37:55.519
Like, like-

1599
01:37:55.640 --> 01:37:56.579
Allie Laabs: 2, 2 days ago,

1600
01:37:56.600 --> 01:37:57.060
Alex Volkov: 2 days ago.

1601
01:37:57.080 --> 01:38:02.419
Allie Laabs: um, or 6 years ago in the amount of time that it, the last 48 hours have felt.

1602
01:38:02.600 --> 01:38:05.879
It has been absolutely, it has been absolutely bonkers.

1603
01:38:06.080 --> 01:38:07.159
Alex Volkov: So let's start with the announcement.

1604
01:38:07.400 --> 01:38:09.479
TypeSafe Labs, eh, has,

1605
01:38:10.400 --> 01:38:13.840
besides sponsoring our hackathon this last weekend, which, by the way, thank you so

1606
01:38:13.920 --> 01:38:17.719
much, like tons of folks got early access to JEV and got super excited to build things,

1607
01:38:18.360 --> 01:38:22.659
has been in stealth building a thing for the past 2 years.

1608
01:38:22.800 --> 01:38:22.920
The

1609
01:38:23.040 --> 01:38:23.299
Allie Laabs: That's right.

1610
01:38:23.320 --> 01:38:29.019
Alex Volkov: CEO and co-founder Diogo Almeida has worked on RLHF, one of the co-creators, I believe,

1611
01:38:29.060 --> 01:38:30.539
of RLHF. Please correct me

1612
01:38:30.640 --> 01:38:30.819
Allie Laabs: Yep.

1613
01:38:30.840 --> 01:38:34.159
Alex Volkov: every statement that's wrong. AI assistant help me research this, but uh, co-creator

1614
01:38:34.180 --> 01:38:37.279
of RLHF and ChatGPT as well, and for the past 2 years has been working on something

1615
01:38:37.319 --> 01:38:40.880
that's different. Alley Labs, Adevrel, Ad Types Safe.

1616
01:38:40.960 --> 01:38:43.159
Please tell us what is different? What is Gev?

1617
01:38:43.240 --> 01:38:45.579
I would love to hear from you directly, like what we're talking about.

1618
01:38:45.640 --> 01:38:47.640
Why is it so exciting? Why is it all over my feed?

1619
01:38:48.680 --> 01:38:51.779
Allie Laabs: Why is it? Okay, we've created a new class of model.

1620
01:38:51.920 --> 01:38:55.379
We call it a System 1 model. It's the first of its kind.

1621
01:38:55.440 --> 01:38:56.580
It's our first public

1622
01:38:57.560 --> 01:39:01.859
model of its kind. Of course, we've had several iterations before we decided, uh,

1623
01:39:01.960 --> 01:39:03.559
we wanted to to bring it to the public.

1624
01:39:03.600 --> 01:39:05.960
We've had a lot of people using it leading up to this.

1625
01:39:06.560 --> 01:39:07.799
Alex Volkov: System 1 referring to

1626
01:39:08.880 --> 01:39:11.060
the Thinking Fast and Slow book from Daniel Kahneman

1627
01:39:11.080 --> 01:39:11.180
Allie Laabs: Yes.

1628
01:39:11.240 --> 01:39:11.739
Alex Volkov: and Tversky?

1629
01:39:12.080 --> 01:39:12.620
Allie Laabs: Exactly.

1630
01:39:12.640 --> 01:39:12.740
Alex Volkov: Okay.

1631
01:39:12.760 --> 01:39:12.859
Allie Laabs: Yep.

1632
01:39:12.920 --> 01:39:13.099
Alex Volkov: Yeah.

1633
01:39:13.120 --> 01:39:16.900
Allie Laabs: Thinking Fast and Slow. That is the, that is the inspiration behind the name, because

1634
01:39:17.400 --> 01:39:21.299
this System 1 model is specifically System 1 thinking.

1635
01:39:21.400 --> 01:39:25.620
It's fast intuition. And when we say fast, I mean, you were just looking at the stats

1636
01:39:25.720 --> 01:39:29.680
a moment ago, like the, that, that computer use, computer use demo that they've created,

1637
01:39:29.720 --> 01:39:29.860
right?

1638
01:39:30.000 --> 01:39:32.120
Alex Volkov: I have one better for you, Ali. I got access.

1639
01:39:32.160 --> 01:39:32.839
Allie Laabs: Oh, I want to see it.

1640
01:39:33.440 --> 01:39:35.280
Alex Volkov: I got access yesterday night.

1641
01:39:35.480 --> 01:39:35.879
Allie Laabs: Oh, you did?

1642
01:39:36.280 --> 01:39:38.399
Alex Volkov: And thank you for the folks who provided access.

1643
01:39:38.760 --> 01:39:44.240
And for the longest time, I was noticing that Twitter, FKA X, has

1644
01:39:44.919 --> 01:39:46.680
a tendency to over-obsess about the topic.

1645
01:39:46.840 --> 01:39:50.719
For example, most of my tweets are now about JIV that I see on my For You timeline,

1646
01:39:50.960 --> 01:39:54.900
and I built an extension before based on Cerebrus and the fastest model that they

1647
01:39:54.960 --> 01:40:00.920
have to analyze all my tweets. Uh, it took me maybe 25 minutes to implement JIV into

1648
01:40:00.960 --> 01:40:06.559
the system, and so here is my timeline as scored by Jev in real time, okay?

1649
01:40:06.720 --> 01:40:11.840
So every tweet that has Jev in it or related to Jev in the quote is marked as yellow.

1650
01:40:12.120 --> 01:40:14.320
I will scroll a little bit, and you can see a bunch of yellow.

1651
01:40:14.520 --> 01:40:19.460
But what you can see on the side here is that we're doing insanely fast analysis.

1652
01:40:19.840 --> 01:40:24.159
Those categories show up before an LLM could even respond, and I'm scrolling super

1653
01:40:24.200 --> 01:40:26.139
fast. I'm going to scroll super, super, super duper fast.

1654
01:40:26.200 --> 01:40:31.639
You'll see we're doing 12 tweets a second with Jeff, uh, and this costs me

1655
01:40:32.600 --> 01:40:35.119
nothing. I literally tried to calculate.

1656
01:40:35.639 --> 01:40:40.620
The thing that I want to highlight is that you guys are now calling, uh, uh, you don't

1657
01:40:40.660 --> 01:40:43.800
even price output tokens, right? Like, only input tokens is priced, is that correct?

1658
01:40:43.880 --> 01:40:45.360
Allie Laabs: On- only input tokens. That's right.

1659
01:40:45.380 --> 01:40:48.580
You pay for input tokens, output, we we like to say output tokens are too cheap to

1660
01:40:48.680 --> 01:40:48.879
meter.

1661
01:40:49.160 --> 01:40:52.759
Alex Volkov: I I think Diego specifically said too cheap to fucking meter in the release video,

1662
01:40:52.840 --> 01:40:58.039
if I'm not mistaken. And also, the input tokens are priced in, in billions and not

1663
01:40:58.080 --> 01:41:00.199
in millions like we're used to from any other LLM.

1664
01:41:00.240 --> 01:41:03.039
So it's like 45 dollars per 1 billion input token.

1665
01:41:03.080 --> 01:41:03.719
Is that, is that

1666
01:41:03.760 --> 01:41:06.099
Allie Laabs: 42, yeah, 42 per billion.

1667
01:41:06.120 --> 01:41:06.460
Alex Volkov: Cheaper.

1668
01:41:06.500 --> 01:41:07.800
Allie Laabs: I'm so glad we ended up doing that.

1669
01:41:07.920 --> 01:41:13.560
Diogo was particularly excited about what if we, what if we say pricing in per billion

1670
01:41:13.760 --> 01:41:15.980
tokens? And it was such a good idea.

1671
01:41:16.040 --> 01:41:17.440
It's such a great moment in the video.

1672
01:41:17.920 --> 01:41:21.999
Yeah, because it is, it is cheap. It is that cheap that we can be talking about on

1673
01:41:22.040 --> 01:41:23.499
a different order of magnitude.

1674
01:41:23.640 --> 01:41:27.300
Alex Volkov: Cheaps and fast, and I think the speed is something that I also want to show you and

1675
01:41:27.320 --> 01:41:33.920
and the audience as well. I have here Qwen 3.8 27B, a very fast model running on Cerebrus.

1676
01:41:34.160 --> 01:41:37.439
Cerebrus is, as we know, an LPU chip that runs like super fast.

1677
01:41:37.960 --> 01:41:38.279
Uh,

1678
01:41:39.240 --> 01:41:43.559
you guys absolutely just crush Cerebrus in impo- impossible ways.

1679
01:41:43.760 --> 01:41:48.439
JEV is 4 or 5 times faster than PerTweet and is 20x cheaper.

1680
01:41:48.640 --> 01:41:53.019
It's like a factor of 20, and I'm pretty sure that this is like only on my end, and

1681
01:41:53.080 --> 01:41:57.059
there's ways to optimize this. I didn't look into developer too much, like Fable cooked

1682
01:41:57.080 --> 01:42:01.279
it in 20 minutes. Uh, but folks, what I have here with JEV, which we we'll talk about

1683
01:42:01.520 --> 01:42:06.880
why this is so efficient in a second, is essentially a real-time cognitive firewall

1684
01:42:07.200 --> 01:42:08.040
for my X

1685
01:42:09.040 --> 01:42:12.939
that I'm using this new system, the System 1 thinking, to categorize in real time

1686
01:42:13.160 --> 01:42:17.399
everything that I see according to buckets that I design, and I want to see all the

1687
01:42:17.480 --> 01:42:18.660
Jeff tweets, for example, highlighted.

1688
01:42:18.680 --> 01:42:20.699
But for, for example, you can also use this.

1689
01:42:20.760 --> 01:42:22.040
That's a feature. It's public, by the way.

1690
01:42:22.080 --> 01:42:25.560
I'll put a show at the end of the, I'll put a link to this at the end of the notes,

1691
01:42:26.420 --> 01:42:30.040
is that you can hide other things. So you can literally have an ad blocker that's

1692
01:42:30.200 --> 01:42:33.639
like your thoughts ad blocker and not necessarily based on whether or not this ad

1693
01:42:33.720 --> 01:42:36.079
came from an ad server, which is a completely different paradigm.

1694
01:42:36.600 --> 01:42:41.259
And this is all done with less than 1 cent, and I've been scrolling and scrolling

1695
01:42:41.320 --> 01:42:45.019
since this morning. So, like, the, the, the pricing there just doesn't compute for

1696
01:42:45.080 --> 01:42:50.580
us. And this is just one example of the, I think now, 500 or so things that I saw

1697
01:42:50.839 --> 01:42:52.800
on my timeline. So back to you, Ali.

1698
01:42:52.860 --> 01:42:55.819
This is not me talking at you, telling you how amazing the things that you're working

1699
01:42:55.860 --> 01:42:57.399
on is. This is a, this is

1700
01:42:57.480 --> 01:43:02.419
Allie Laabs: Although I do love, I do love just sitting back and hearing people just like, that

1701
01:43:02.480 --> 01:43:05.119
excitement that you're expressing is what we're seeing all over the place, right?

1702
01:43:05.160 --> 01:43:06.179
Alex Volkov: I think it's crazy because

1703
01:43:06.200 --> 01:43:07.759
Allie Laabs: It's real. This is a real thing.

1704
01:43:07.920 --> 01:43:12.239
Alex Volkov: We, we use LLMs for this. LLMs are general, and they're streaming tokens.

1705
01:43:12.560 --> 01:43:14.999
What are you guys doing differently that allows for this scale?

1706
01:43:15.080 --> 01:43:19.959
Please, please tell us, educate us on, on what allows this type 1 thinking, and what

1707
01:43:20.000 --> 01:43:20.759
is the innovation here?

1708
01:43:21.240 --> 01:43:22.080
Allie Laabs: So, so

1709
01:43:23.839 --> 01:43:27.500
that is a, that is a great opportunity to say my first of what could be multiple,

1710
01:43:27.680 --> 01:43:32.220
I can't talk about that. Um, when it comes to, like, the actual, like, what the secret

1711
01:43:32.279 --> 01:43:36.839
sauce is, the, uh, the architecture, that is a thing where I just stay far away from

1712
01:43:36.880 --> 01:43:37.260
that. Um,

1713
01:43:37.360 --> 01:43:37.459
Alex Volkov: Yeah.

1714
01:43:37.760 --> 01:43:40.899
Allie Laabs: Diogo's been doing some town halls in our, uh, in our Discord.

1715
01:43:40.960 --> 01:43:45.320
He's the best one to know exactly how much that we're willing to talk about that.

1716
01:43:45.720 --> 01:43:49.019
But w- but it is, it is the case that we have secret sauce, right?

1717
01:43:49.120 --> 01:43:54.720
We have, we have had break in innovation and enabled i- in order to do this, and it

1718
01:43:54.779 --> 01:43:58.720
is, this is a new, it is a new kind of training that we call reinforcement learning

1719
01:43:58.760 --> 01:44:03.560
for calibrated decisions, right? That was in our RLCD, that is in our launch video,

1720
01:44:04.040 --> 01:44:08.619
and that is, that's, that's a huge part of, that's a huge part of what it is, is we're

1721
01:44:08.640 --> 01:44:13.300
doing parallel processing of all, uh, of all of the questions that you send in at

1722
01:44:13.320 --> 01:44:17.579
the same time, and we're not, we're not generating text character by character,

1723
01:44:17.600 --> 01:44:17.819
Alex Volkov: Mhm.

1724
01:44:17.839 --> 01:44:20.480
Allie Laabs: Right? I've had this before I started at TypeSafe.

1725
01:44:20.760 --> 01:44:21.980
I mean, I guess I still feel this way.

1726
01:44:22.120 --> 01:44:26.879
I just felt like the whole, so much of the landscape of what we're doing with AI has

1727
01:44:27.000 --> 01:44:30.100
felt completely bizarre to me, right?

1728
01:44:30.279 --> 01:44:32.579
For chatting, it totally makes sense, right?

1729
01:44:32.839 --> 01:44:35.759
RLHF is great at creating, uh, like a brainstorming partner.

1730
01:44:36.000 --> 01:44:38.480
Um, it has turned out to be really good at coding.

1731
01:44:38.800 --> 01:44:41.120
I think there's a lot of different directions that can go in the future, though.

1732
01:44:41.360 --> 01:44:43.240
But like, let's just talk about like chatting, right?

1733
01:44:43.360 --> 01:44:46.420
Works super well for that. It's gonna, it generates it word by word.

1734
01:44:46.480 --> 01:44:50.359
It's very similar to how we form words in our mind as we're talking to each other,

1735
01:44:50.440 --> 01:44:52.939
right? I'm not, I don't know what my sentences are before I say them.

1736
01:44:53.240 --> 01:44:57.879
So that all, that all makes sense. But then what we've done is we've taken these LLMs,

1737
01:44:58.560 --> 01:45:03.320
and then we've gone, okay, now let's make them call tools, right, to do things.

1738
01:45:03.560 --> 01:45:08.239
Like, so I want to have an LLM order me a pizza, and it's going to do this by calling

1739
01:45:08.280 --> 01:45:10.959
a bunch of tools, maybe it'll even move your mouse and et cetera, et cetera.

1740
01:45:11.680 --> 01:45:17.899
It's so weird to me that we are using these, like, big old language models

1741
01:45:18.400 --> 01:45:24.540
to create the text or the instructions to interact with an interface that

1742
01:45:24.600 --> 01:45:27.319
was already on the other side of the interface is machine code.

1743
01:45:27.760 --> 01:45:32.160
It is like a website, it's data flowing to the pizza server and putting in an order.

1744
01:45:32.440 --> 01:45:37.079
Yet we've translated that into a user interface, whether it's graphical or even a

1745
01:45:37.160 --> 01:45:40.319
command line interface. A command line interface is still a human user interface,

1746
01:45:40.400 --> 01:45:46.259
right? And so we're- we're taking LLMs, we're having them output like a human

1747
01:45:46.440 --> 01:45:49.399
interface language, putting it back into machine code.

1748
01:45:49.440 --> 01:45:50.779
We started in machine code,

1749
01:45:51.000 --> 01:45:51.180
Alex Volkov: Yeah.

1750
01:45:51.280 --> 01:45:56.679
Allie Laabs: we ended in machine code, and yet we did an extremely expensive, like, translation

1751
01:45:56.800 --> 01:45:59.700
layer in the middle just to make them, like, pretend to be a human for a moment.

1752
01:45:59.760 --> 01:46:00.339
Alex Volkov: And we optimized

1753
01:46:00.360 --> 01:46:00.720
Allie Laabs: Which is like

1754
01:46:00.740 --> 01:46:02.299
Alex Volkov: the crap out of that process as well.

1755
01:46:02.320 --> 01:46:02.539
Allie Laabs: We do.

1756
01:46:02.559 --> 01:46:05.639
Alex Volkov: So it feels like almost there. It feels fast, but yeah, you're, you're absolutely

1757
01:46:05.660 --> 01:46:07.379
right. This is kind of wasteful and ridiculous.

1758
01:46:07.400 --> 01:46:08.140
Allie Laabs: Yeah.

1759
01:46:08.160 --> 01:46:11.600
Alex Volkov: And spends a lot of GPU, where those GPUs can be used for, you know, training the

1760
01:46:11.640 --> 01:46:12.440
next models, etcetera.

1761
01:46:12.720 --> 01:46:16.100
Allie Laabs: It makes me think of like a, like a, uh, you may, you're gonna make a, a robot, a

1762
01:46:16.200 --> 01:46:18.740
bipedal robot to wash my dishes, and I go, well,

1763
01:46:19.679 --> 01:46:22.799
we've already made a robot that washes dishes, and it's called a dishwasher, and it's

1764
01:46:23.000 --> 01:46:27.159
really efficient. It's extremely, it's like does it really fast for the amount of

1765
01:46:27.240 --> 01:46:28.799
dishes. Like, it's a, it's a very good machine.

1766
01:46:28.820 --> 01:46:30.660
Now, loading a dishwasher, it's a pain in the ass.

1767
01:46:30.720 --> 01:46:32.600
It's a separate thing, but just for the washing part, right?

1768
01:46:32.760 --> 01:46:38.519
It's like we optimize in technology of all sorts for the problem that we want to solve,

1769
01:46:38.559 --> 01:46:42.760
but LLMs doing tool calls and all this has just been in this weird extra layer.

1770
01:46:42.840 --> 01:46:48.940
So what we've done is we've created a model that gives access to essentially

1771
01:46:49.000 --> 01:46:54.599
the latent intelligence that is inside what we, in the modern day, call AI,

1772
01:46:55.080 --> 01:47:01.539
and but exposes it as machine-readable outputs, as probabilities, as type-safe

1773
01:47:01.800 --> 01:47:04.440
probabilities, right? Like, you know, as typed objects.

1774
01:47:04.920 --> 01:47:10.880
And this now finally lets us use AI in a machine-to-machine context

1775
01:47:11.000 --> 01:47:13.480
without ever turning it back into user space.

1776
01:47:14.000 --> 01:47:19.199
Um, and I just think it, it to me, it's the day I saw this, when literally when I

1777
01:47:19.280 --> 01:47:23.419
came in for an interview into this office, and I sat down, and one of the interview

1778
01:47:23.480 --> 01:47:26.740
tasks was to, like, see the API for the first, like, see what it was, because we were

1779
01:47:26.800 --> 01:47:28.119
stealth, right? So I had no idea what it was.

1780
01:47:28.440 --> 01:47:33.740
I sat down, and it was for me. it was, oh my fucking god, this is, this is what I've

1781
01:47:33.760 --> 01:47:35.240
been struggling with for the past two years.

1782
01:47:35.280 --> 01:47:38.880
I've, I had a whole bunch of projects that I've now ported to using Gev, right?

1783
01:47:38.920 --> 01:47:41.039
Just kind of like you, this, this tweet analyzer, right?

1784
01:47:41.480 --> 01:47:46.620
Is I've had all of these things with Gev-shaped holes, and I just, now it was like

1785
01:47:46.720 --> 01:47:49.159
all of a sudden, all of a sudden the tool existed.

1786
01:47:49.220 --> 01:47:51.480
All of a sudden it was this is the missing piece.

1787
01:47:51.800 --> 01:47:55.779
Alex Volkov: This is, it definitely feels folks who are excited about technology, excited about

1788
01:47:55.880 --> 01:47:58.920
solution space, are discovering that this is could be the missing piece.

1789
01:47:59.240 --> 01:48:04.740
Uh, Sunil Pai from Cloudflare posted, this feels like the React announcement era where

1790
01:48:04.879 --> 01:48:08.599
folks are like, oh, this is exactly what I've been missing, and I don't want to overpay

1791
01:48:08.680 --> 01:48:11.200
as well. I'm using this tool that could also do this.

1792
01:48:11.440 --> 01:48:15.899
Maybe a bulldozer can flatten my pavement, but also a bulldozer can, I don't know,

1793
01:48:16.240 --> 01:48:19.379
lift, lift something, whereas there's a better tool for that, and this feels like

1794
01:48:19.400 --> 01:48:22.640
the tool was more efficient, and I don't need to bring a whole bulldozer and spend

1795
01:48:22.680 --> 01:48:26.980
a lot of money. And I'm showing here on the stage kind of the speed with which those

1796
01:48:27.120 --> 01:48:31.319
categorizations happen, whereas LLM outputting one by one by one, we know this.

1797
01:48:31.640 --> 01:48:35.340
Um, however, Ali, this is not a diffusion model that we've talked about, diffusion

1798
01:48:35.360 --> 01:48:40.199
models that kind of like fills holes and uses diffusion methods for kind of LLM-y

1799
01:48:40.280 --> 01:48:43.319
things. Uh, what doesn't Jeff do for me?

1800
01:48:43.560 --> 01:48:46.340
This is not like better than asteroid talking to me, right?

1801
01:48:46.480 --> 01:48:46.980
Or in fact,

1802
01:48:47.319 --> 01:48:50.299
Allie Laabs: Correct. I liked Diogo's wording in our launch post.

1803
01:48:50.380 --> 01:48:53.079
It said like this trade off, it doesn't come for free.

1804
01:48:53.319 --> 01:48:55.599
We don't generate text. We don't generate text.

1805
01:48:55.680 --> 01:48:59.120
We don't want to generate text. Ironically, the first thing a lot of our users have

1806
01:48:59.160 --> 01:49:00.439
done with it is, like, use

1807
01:49:00.520 --> 01:49:00.819
Alex Volkov: Generate text.

1808
01:49:00.860 --> 01:49:04.420
Allie Laabs: it to generate text, which is very silly, like, putting it in a choice to say, pick

1809
01:49:04.480 --> 01:49:08.739
the next letter from this list. It's such like a, it's going, it's fun, and it's actually

1810
01:49:08.760 --> 01:49:11.960
kind of fun to see what bizarre strings come out of it.

1811
01:49:12.000 --> 01:49:14.320
It kind of feels like the early days of, like, predictive engines.

1812
01:49:14.500 --> 01:49:18.379
It's not what it's for. Like, we're not, it's, we are very unlikely to optimize it

1813
01:49:18.440 --> 01:49:19.700
for, to be able to generate text

1814
01:49:19.800 --> 01:49:19.940
Alex Volkov: Sure.

1815
01:49:19.960 --> 01:49:22.439
Allie Laabs: in that way because we've already have that in the world.

1816
01:49:22.480 --> 01:49:27.800
It's LLMs. LLMs are great at generating, like, generative text when you want human

1817
01:49:27.960 --> 01:49:28.499
prose, right?

1818
01:49:28.680 --> 01:49:33.000
Alex Volkov: I have to, I have to pause. I'm so sorry, but I have a question that it's related

1819
01:49:33.040 --> 01:49:35.520
to what you're just saying. You're saying LLMs versus JEV.

1820
01:49:35.760 --> 01:49:39.639
Does JEV have a category now? Like, do we need, do we have a name for JEV and other

1821
01:49:39.920 --> 01:49:43.619
models like JEV that would come out or have come out before, or there's no-

1822
01:49:44.720 --> 01:49:46.340
Allie Laabs: System 1 models. That's what we-

1823
01:49:46.360 --> 01:49:46.619
Alex Volkov: System 1.

1824
01:49:46.639 --> 01:49:52.259
Allie Laabs: We use System 1 models like we use large language models to describe a different kind

1825
01:49:52.480 --> 01:49:52.620
of

1826
01:49:52.800 --> 01:49:52.939
Alex Volkov: Yeah.

1827
01:49:53.120 --> 01:49:57.679
Allie Laabs: a different kind of AI, right? It- AI has means a whole lot of things if we go back

1828
01:49:57.760 --> 01:50:02.319
long before LLMs, right? AI has meant a lot of things, but now we are System 1 model.

1829
01:50:02.440 --> 01:50:08.600
And while we are the first, we expect in the future System 1 models will be the catalyst

1830
01:50:08.639 --> 01:50:10.759
for an economic revolution. We believe that this

1831
01:50:10.800 --> 01:50:10.920
Alex Volkov: Wow.

1832
01:50:10.940 --> 01:50:14.999
Allie Laabs: will be driving a massive amount of automation in our world's future.

1833
01:50:15.200 --> 01:50:20.080
Now, will it be JEV that is doing all of that, or will other people catch up and create

1834
01:50:20.120 --> 01:50:26.039
their own System 1 models? What we have 100% confidence in here is that System 1 models

1835
01:50:26.060 --> 01:50:29.239
will be doing it. We hope it'll be JEV doing all of it, right?

1836
01:50:29.320 --> 01:50:29.580
Obviously,

1837
01:50:29.600 --> 01:50:29.740
Alex Volkov: 100%.

1838
01:50:29.760 --> 01:50:34.120
Allie Laabs: he would love that, but, and and we have a big head start, we've invented it, et cetera,

1839
01:50:34.160 --> 01:50:39.319
et cetera. But we expect other people to make, you know, system, other System 1 schema

1840
01:50:39.440 --> 01:50:43.599
compatible models as time goes on, like what happened with, you know, like what happened

1841
01:50:43.620 --> 01:50:46.800
with ChatGPT, and now all of a sudden, you know, there's, there's a whole bunch of

1842
01:50:46.840 --> 01:50:51.800
these. But we, and so that's why, that's why we decided to coin it as System 1 models.

1843
01:50:51.840 --> 01:50:56.759
We, we want to, we're giving the name System 1 models to the world, like that, this

1844
01:50:56.840 --> 01:50:57.020
is it.

1845
01:50:57.040 --> 01:51:00.439
Alex Volkov: I'm down, I'm down with System 1. Like, we're gonna call this System 1 on the show

1846
01:51:00.560 --> 01:51:04.420
until, like, um, uh, another came, uh, another name comes about.

1847
01:51:04.480 --> 01:51:11.280
I'm showing a, a, a screenshot from TypeSafe.ai, where, uh, you guys are citing 133x

1848
01:51:11.680 --> 01:51:17.579
faster and 444x cheaper than other, like, like competitive level

1849
01:51:17.760 --> 01:51:19.720
models for some of the tasks that you're testing.

1850
01:51:20.200 --> 01:51:23.960
And I think that maybe, maybe, please correct me if I'm wrong, maybe you guys were

1851
01:51:24.040 --> 01:51:27.759
surprised with the amount of stuff that people are shoving this, like, puzzle piece

1852
01:51:27.800 --> 01:51:29.960
that they've been missing and, and using use cases.

1853
01:51:30.360 --> 01:51:34.559
Um, I think I saw Diogo, the, the founder, actually react to some stuff and say, we

1854
01:51:34.600 --> 01:51:37.959
should use, we should have used this as an example versus the other one, uh, which

1855
01:51:38.000 --> 01:51:42.119
is in several. So, so if you don't mind, I'll, I'll bring back Francesco, who is a,

1856
01:51:42.560 --> 01:51:45.279
a co-founder of Kua, which does computer use driving as well.

1857
01:51:45.320 --> 01:51:49.800
And Francesco, I would love for you to tell us, uh, from, from the brief experience

1858
01:51:49.820 --> 01:51:56.040
that you guys had, what effect do you foresee this for computer use and web use,

1859
01:51:56.400 --> 01:52:00.999
uh, uh, interfaces and drivers? Are we talking about like much faster, much more,

1860
01:52:01.720 --> 01:52:03.359
like much better as well?

1861
01:52:04.160 --> 01:52:07.879
Allie Laabs: Yeah, I would say we were really thinking about real time computer use.

1862
01:52:08.080 --> 01:52:12.320
It was like a couple of months ago, some interesting demo with GPT Live, but still

1863
01:52:12.520 --> 01:52:17.800
when, when you're just fitting that in into long horizon trajectory, you know, it's,

1864
01:52:18.320 --> 01:52:23.100
it's hard to have a trajectory of all computer use agents and translating it for a

1865
01:52:23.160 --> 01:52:28.519
real time computer using agents. I think that the, that is actually the missing piece

1866
01:52:28.560 --> 01:52:34.519
here. Um, being able to take this fast decision and being able to live, to work with

1867
01:52:34.560 --> 01:52:40.399
the same space, basically the binary space where really a computer is at.

1868
01:52:40.839 --> 01:52:42.920
I think, honestly, that's, that's the way to go forward.

1869
01:52:43.480 --> 01:52:45.259
And I can tell you about my experience at Microsoft.

1870
01:52:45.320 --> 01:52:45.839
It was like,

1871
01:52:46.839 --> 01:52:50.379
there had to be a better way. We were just like working with a YOLO icon detection

1872
01:52:50.480 --> 01:52:51.819
model at the time, just like throwing

1873
01:52:51.880 --> 01:52:53.540
Alex Volkov: I think it feels like that better way.

1874
01:52:53.640 --> 01:52:58.160
Like, like I, I definitely don't want to overstate, but based on everything that I

1875
01:52:58.240 --> 01:53:02.279
see, and again, as we saw, the algorithm was overfocusing on the stuff that I like.

1876
01:53:02.320 --> 01:53:03.779
I talked about Muse, now everything is Muse.

1877
01:53:04.000 --> 01:53:07.619
Uh, this feels like many of the engineer friends of mine that used LLM for something

1878
01:53:07.640 --> 01:53:12.199
else and are considering costs are now completely just removing costs from the equation

1879
01:53:12.240 --> 01:53:16.359
now because of the, you know, the speed and the fact that you guys are not charging

1880
01:53:16.400 --> 01:53:18.199
anymore for output, which is incredible.

1881
01:53:18.640 --> 01:53:22.379
Uh, Ali, I, you are the devil there, and this is tool for developers as well.

1882
01:53:22.480 --> 01:53:24.660
This is not like for my mom who's gonna like chat with this.

1883
01:53:24.920 --> 01:53:25.459
It's not ChatGPT.

1884
01:53:25.480 --> 01:53:25.899
Allie Laabs: That's right, yeah.

1885
01:53:25.940 --> 01:53:27.359
Alex Volkov: It's not like a user consumer technology.

1886
01:53:27.680 --> 01:53:28.879
Tell us about some of the primitives.

1887
01:53:29.080 --> 01:53:32.919
I wanna hear you kind of explain what, what null is, for example.

1888
01:53:32.960 --> 01:53:36.640
Like, you have like a whole new category, but also like new primitives to build with

1889
01:53:36.680 --> 01:53:40.540
this. Could you please spend the next few minutes, if you don't mind, telling folks

1890
01:53:40.560 --> 01:53:42.519
how to best use this technology? I would love to hear that.

1891
01:53:42.640 --> 01:53:43.439
And Francesca, thank you.

1892
01:53:43.520 --> 01:53:49.359
Allie Laabs: Yeah. So that's the primitives is the way that we, that we've decided to like,

1893
01:53:49.800 --> 01:53:53.899
you know, we want calibrated probabilities from the latent, you know, intelligence

1894
01:53:54.400 --> 01:53:58.199
inside AI. Like, how do then we make that available?

1895
01:53:58.320 --> 01:54:01.240
And the primitives system is what we came up with for that.

1896
01:54:01.400 --> 01:54:07.479
So what you actually get from the Gev API, from the TypeSafe API, is, uh, is you,

1897
01:54:07.600 --> 01:54:11.019
you define the primitives that you are requesting.

1898
01:54:11.080 --> 01:54:12.600
We they, we also call them questions.

1899
01:54:12.680 --> 01:54:14.679
They're called questions actually in the schema, right?

1900
01:54:15.040 --> 01:54:20.279
Um, so think of them as like questions technically, but primitives in like conceptually.

1901
01:54:20.680 --> 01:54:23.680
Uh, and you see them right there. It's choice, score, and newl.

1902
01:54:24.040 --> 01:54:29.480
So each of these is a specific shape of response that you're going to get.

1903
01:54:29.520 --> 01:54:31.679
They are each their own type of response.

1904
01:54:31.800 --> 01:54:35.820
They each return probabilities, but in different sort of ways, and they're appropriate

1905
01:54:35.920 --> 01:54:39.560
for different kinds of questions, different kinds of information you want to get.

1906
01:54:39.800 --> 01:54:42.279
So a choice, uh, a choice is really simple.

1907
01:54:42.360 --> 01:54:45.200
It's choose from a list of discrete options.

1908
01:54:45.360 --> 01:54:50.980
So you provide it a discrete option space, you describe them each semantically, and

1909
01:54:51.280 --> 01:54:55.899
the model will pick, we, if we're explaining it like really high level, we'll say

1910
01:54:55.960 --> 01:54:59.739
the model will pick the answer from that list, but what it actually is doing is it's

1911
01:54:59.800 --> 01:55:04.879
producing the probability of everyone on that list, uh, and you get all of that back,

1912
01:55:05.160 --> 01:55:09.480
right? So in some use cases there, you're always looking for a dominant answer, like

1913
01:55:09.520 --> 01:55:12.520
there is, they they don't have any overlap, you're going to get one, so you're just

1914
01:55:12.560 --> 01:55:17.100
going to go with the choice. But sometimes your your use case actually might be really

1915
01:55:17.200 --> 01:55:21.600
interested if two of the choices are competing for the highest prob- uh pri- uh eh

1916
01:55:22.080 --> 01:55:26.939
probability, in which case you can actually look into the probability array to, uh,

1917
01:55:27.000 --> 01:55:31.580
or the probability map to get, uh, all of them individually and do some more advanced,

1918
01:55:31.640 --> 01:55:33.319
uh, you know, advanced math on that.

1919
01:55:33.680 --> 01:55:35.139
Um, the score

1920
01:55:36.240 --> 01:55:39.299
is the other kind. Let's see, do you have it on your screen there?

1921
01:55:39.480 --> 01:55:39.579
Alex Volkov: Yeah.

1922
01:55:39.640 --> 01:55:40.959
Allie Laabs: Now I started to read off your screen.

1923
01:55:41.960 --> 01:55:42.220
The score.

1924
01:55:42.240 --> 01:55:44.300
Alex Volkov: I just wanted to show that, that you talk about probability.

1925
01:55:44.320 --> 01:55:44.660
Allie Laabs: Oh yeah, we do.

1926
01:55:44.680 --> 01:55:48.179
Alex Volkov: In the case of tweets, for example, I have 98, like, assurance that this tweet is

1927
01:55:48.200 --> 01:55:51.679
about Jeff because Dex here talks about Jeff, but there's other tweets that I'm showing

1928
01:55:51.720 --> 01:55:55.059
here. There's like a 3% probability that this tweet is talking about this category,

1929
01:55:55.120 --> 01:56:00.119
2% about this category, and uh, when choice is basically you guys are telling me based

1930
01:56:00.160 --> 01:56:03.959
on those probability, you still calculate which one is, if I want the specific one,

1931
01:56:04.040 --> 01:56:07.379
does this tweet talk about X, you give like a choice as well, in addition to probability.

1932
01:56:07.420 --> 01:56:08.120
It's very interesting.

1933
01:56:09.120 --> 01:56:15.240
Allie Laabs: Exactly. And so score is about getting back a, uh, uh, a continuous

1934
01:56:15.800 --> 01:56:21.799
value between n 0 and n up to 9, uh, and it is about rating

1935
01:56:22.000 --> 01:56:27.419
something on an axis. So in this case, it's like a, it's like a semantically described

1936
01:56:27.679 --> 01:56:32.700
rubric that you are creating, uh, and asking where, what the answer is along that.

1937
01:56:32.840 --> 01:56:38.760
So you might be asking, like, how, um, you know, how relevant is this to

1938
01:56:39.000 --> 01:56:40.820
something else? That's actually a bad example, but

1939
01:56:41.360 --> 01:56:42.799
Alex Volkov: I can give you an example. Tell me if

1940
01:56:43.020 --> 01:56:44.559
Allie Laabs: Please, please give me an example. Save me.

1941
01:56:44.640 --> 01:56:48.979
Alex Volkov: Uh, previously, in the beginning of the show, we talked about the pacing the frontier

1942
01:56:49.120 --> 01:56:53.539
and where, you know, Dariamo is this, there's like, you know, all the way to Jensen

1943
01:56:53.720 --> 01:56:57.799
and saying, like, let's go. And we also talked about our own kind of like statements.

1944
01:56:58.760 --> 01:57:02.739
Could we take, uh, transcriptions from this podcast and then rate them on this like

1945
01:57:02.860 --> 01:57:06.759
scale of like pace all the way to pause AI on to all the way to go, right?

1946
01:57:06.800 --> 01:57:10.779
This, this will put us as a scale and put every statement that we have categorized

1947
01:57:10.800 --> 01:57:13.040
this as belonging on that scale, correct?

1948
01:57:14.080 --> 01:57:19.339
Allie Laabs: Yes. Yeah, exactly. And, uh, a chat moderation is actually a really great, um, example

1949
01:57:19.360 --> 01:57:24.160
of using where score really, really fits well, because this is a case where you're

1950
01:57:24.180 --> 01:57:28.600
not looking at discrete points, you are looking at, like, a spectrum of answer.

1951
01:57:28.880 --> 01:57:33.120
But the huge power with the score primitive that I think makes really cool, and it

1952
01:57:33.159 --> 01:57:36.599
can be a little tricky for people to wrap their heads around at first, is that the

1953
01:57:37.200 --> 01:57:42.319
levels of the score that, uh, you define them semantically, right?

1954
01:57:42.400 --> 01:57:46.739
So you need a minimum of 2, where it's like between 0 and 1, and maybe you're saying,

1955
01:57:46.760 --> 01:57:49.180
you know, not at all. Let's say it's chat moderation, right?

1956
01:57:49.400 --> 01:57:49.639
Alex Volkov: Yes.

1957
01:57:49.760 --> 01:57:55.819
Allie Laabs: How, how, like, productive and helpful is this comment to the context of the channel,

1958
01:57:55.840 --> 01:57:57.520
right? That'd be a great use of a score.

1959
01:57:57.800 --> 01:58:01.859
And a naive way to do that would be 0, 0 you would define as not relevant, not helpful

1960
01:58:01.880 --> 01:58:03.720
at all, and 1 would be super, super helpful.

1961
01:58:04.240 --> 01:58:09.719
But by using more levels in there, you can semantically describe exactly what a 1

1962
01:58:09.880 --> 01:58:15.920
and a 2 and a 3 and a 4 mean to you, and you should always use as many score levels

1963
01:58:16.319 --> 01:58:18.779
as you can reasonably semantically describe.

1964
01:58:18.840 --> 01:58:22.299
You also don't want so many levels where you're stretching yourself to try to, like,

1965
01:58:22.480 --> 01:58:24.599
figure out what would be the difference between a 3 and a 4.

1966
01:58:24.800 --> 01:58:26.220
If you can't describe it, then don't do that.

1967
01:58:26.319 --> 01:58:28.159
Just do it the levels that you can describe.

1968
01:58:28.200 --> 01:58:33.719
And what this does is, it means that if it's just between 1 and 0, and you're asking,

1969
01:58:33.880 --> 01:58:36.519
we'll get to a newl in a moment, you're just asking, like, is this productive?

1970
01:58:36.920 --> 01:58:42.799
Um, you kind of wouldn't be certain what a 0.6 means from message to message,

1971
01:58:43.280 --> 01:58:47.499
but when you semantically describe what a 1 and a 2 and a 3 means, you know that if

1972
01:58:47.560 --> 01:58:52.319
something scores right at 2, that it's going to fit the criteria that you described

1973
01:58:52.400 --> 01:58:58.340
for what a 2 is. So it allows you to score things, and then when you score other states,

1974
01:58:58.400 --> 01:59:02.359
when you score other things with the same defined criteria, they are directly comparable.

1975
01:59:02.440 --> 01:59:05.280
You can sort, you can threshold, you can create rules around it.

1976
01:59:05.520 --> 01:59:10.779
They can be very, very reliable, um, and, you know, and damn near deterministic, uh,

1977
01:59:10.919 --> 01:59:12.100
when you run these evaluations.

1978
01:59:12.400 --> 01:59:15.779
Alex Volkov: I, I, uh, let's get to null and then determinism before we end this.

1979
01:59:15.880 --> 01:59:17.580
I really want to talk about those two things.

1980
01:59:17.680 --> 01:59:18.559
It's very interesting.

1981
01:59:20.160 --> 01:59:22.999
Allie Laabs: Absolutely. So, null is our third primitive.

1982
01:59:23.240 --> 01:59:26.099
Um, we put it last on the list. In some ways, it's the easiest to understand.

1983
01:59:26.240 --> 01:59:28.719
We're actually going back and forth internally, like, should we describe null first

1984
01:59:28.800 --> 01:59:34.120
or null last? Null is simply a yes or no question or a state, like a true false statement.

1985
01:59:34.200 --> 01:59:39.240
So newl gives you the probability that the answer to your question is yes, or that

1986
01:59:39.280 --> 01:59:41.520
the, or that the statement is true.

1987
01:59:42.040 --> 01:59:47.279
Um, newls are, are, they're, they're a little tricky because on the surface they sound

1988
01:59:47.320 --> 01:59:50.639
like it's a score. It's just like a score with two levels, a true and false, and kinda

1989
01:59:50.760 --> 01:59:54.479
they are. They perform a little bit differently than scores, so it's a kind of thing

1990
01:59:54.520 --> 01:59:59.139
where, like, you should test. Um, but generally my guidance is like, if it, if it

1991
01:59:59.320 --> 02:00:04.620
really is a yes or no question, if your expected answer space really is binary, like

1992
02:00:04.680 --> 02:00:09.680
there should be a yes or a no, and something in between would mean, like, uncertain,

1993
02:00:09.880 --> 02:00:13.800
not enough evidence, something like that, then that's what you would use a newl for.

1994
02:00:13.880 --> 02:00:19.699
And newls are incredibly powerful. A lot of the demos I've made end up just being

1995
02:00:19.760 --> 02:00:24.919
a whole bunch of newls all at the same time to, in, in fact, your tweet classifier,

1996
02:00:25.320 --> 02:00:27.559
um, could be done with a whole bunch of newels, right?

1997
02:00:27.600 --> 02:00:33.680
Like, is this tweet related to JEV, the latest and greatest System 1 model from TypeSafe?

1998
02:00:34.040 --> 02:00:36.900
Um, and that would actually be a way that you would probably want to write it all

1999
02:00:37.000 --> 02:00:39.339
out, because what if they don't say JEV, what if they say TypeSafe?

2000
02:00:39.360 --> 02:00:43.479
So you do want to, like, include semantically, like, what the criteria actually is.

2001
02:00:43.760 --> 02:00:47.740
But you could do that instead of, that would be a case where I would actually want

2002
02:00:47.780 --> 02:00:53.479
to classify tweets by a whole bunch of newels because a tweet, like many things, can

2003
02:00:53.560 --> 02:00:57.780
be multiple things. It could be a comment about Jev and what someone ate for breakfast.

2004
02:00:57.920 --> 02:01:01.420
It might be a 1, 100% that there's talking about.

2005
02:01:01.440 --> 02:01:04.780
Alex Volkov: It could be a tweet from me saying we're gonna talk about Jev and some other things,

2006
02:01:04.800 --> 02:01:06.520
so it's not like specifically Jev, right?

2007
02:01:06.600 --> 02:01:08.479
Like, it's not like I'm talking about Jev only.

2008
02:01:08.520 --> 02:01:11.240
I'm just announcing that this is a list of other things that we're gonna discuss,

2009
02:01:11.280 --> 02:01:11.580
so that's

2010
02:01:11.720 --> 02:01:14.860
Allie Laabs: Precisely. So that would be a case where someone were trying to build like a tweet

2011
02:01:15.000 --> 02:01:17.519
classifier, and they were like, oh, is it, I'm going to do it as a choice.

2012
02:01:17.640 --> 02:01:22.199
Is it JEV? Is it breakfast? I'd be like, no, because those aren't mutually exclusive,

2013
02:01:22.320 --> 02:01:25.600
and if they aren't mutually exclusive, you should be breaking that further into more

2014
02:01:25.760 --> 02:01:28.439
primitives, and this is why we call them primitives.

2015
02:01:28.800 --> 02:01:32.319
Basically, one of the first things we do when we help people, when they're like, oh,

2016
02:01:32.360 --> 02:01:35.639
my, this isn't performing as well as I expected, we're like, let us see the questions,

2017
02:01:35.760 --> 02:01:41.539
and almost always, like the majority of the time, the answer is, ah, they've,

2018
02:01:42.000 --> 02:01:44.040
the question is actually like a compound question.

2019
02:01:44.120 --> 02:01:46.639
You're asking about multiple things in the same question.

2020
02:01:47.000 --> 02:01:50.560
Break that into multiple questions and compose it together in your code.

2021
02:01:50.960 --> 02:01:53.999
Alex Volkov: So, so questions is, my question is about questions.

2022
02:01:54.440 --> 02:01:59.119
Those are things that I, as a human programmer that uses your protocol and model,

2023
02:01:59.800 --> 02:02:04.519
define about what I want to know that your model will tell me, yes or no, or probability

2024
02:02:04.560 --> 02:02:07.100
of yes or no, probability of like spread probability, right?

2025
02:02:07.120 --> 02:02:11.880
Like, I just literally, in natural language, tell you, hey, is this tweet about this

2026
02:02:11.960 --> 02:02:17.600
thing? And then we used to not be able to do this with code, so we wrote about if

2027
02:02:18.080 --> 02:02:21.360
statements, and then we went to BERT models, whatever, those are not super performant.

2028
02:02:21.560 --> 02:02:24.839
Then LLM came by, and everybody's like, oh shit, I can use the general intelligence

2029
02:02:24.880 --> 02:02:27.040
to ask this question in natural language to give an answer.

2030
02:02:27.360 --> 02:02:32.679
Now, basically, what you're giving in JEV is, in System 1 models, is, uh, the ability

2031
02:02:32.720 --> 02:02:38.540
for me in natural language, in semantic space, define what I want, but get, uh,

2032
02:02:38.920 --> 02:02:40.899
predictable outcomes back. Is that

2033
02:02:40.920 --> 02:02:41.259
Allie Laabs: That's right.

2034
02:02:41.400 --> 02:02:41.640
Alex Volkov: Correct?

2035
02:02:41.720 --> 02:02:43.720
Allie Laabs: Exactly. Yes. Yeah, you described it.

2036
02:02:43.800 --> 02:02:44.940
That's exactly what it is, right?

2037
02:02:45.000 --> 02:02:45.160
Alex Volkov: Yeah.

2038
02:02:45.440 --> 02:02:50.940
Allie Laabs: You as the, as the programmer, as the system designer, are defining what the questions

2039
02:02:51.000 --> 02:02:55.000
you need for your domain space is, and you're defining what all of the, you know,

2040
02:02:55.080 --> 02:02:56.439
what all of the criteria is.

2041
02:02:56.640 --> 02:03:01.899
Alex Volkov: So like, yeah, there is a compounding, um, effort here because Gev is not in

2042
02:03:02.760 --> 02:03:08.239
silo. Ge- Gev is also in the era where Astra exists and Fable exists, and being able

2043
02:03:08.360 --> 02:03:13.600
to combine those two, System 1 and System 2, where System 2 significantly advanced

2044
02:03:13.640 --> 02:03:17.019
for the past three years, like it's ab- absolutely incredible what you can ask these

2045
02:03:17.080 --> 02:03:21.620
models to do with the speed and the price performance of something like System1 that's

2046
02:03:21.680 --> 02:03:25.660
built specifically for decision making is, that's why I think many people are getting,

2047
02:03:25.720 --> 02:03:28.799
getting super excited. And thank you so much for, like, telling us about the primitives,

2048
02:03:28.839 --> 02:03:32.019
how to build this, uh, but also about the general, the general category.

2049
02:03:32.040 --> 02:03:35.779
I really appreciate it. I, I do expect that this is not the last time that we're talking

2050
02:03:35.800 --> 02:03:38.799
about JEV or folks from TypeSafe. So thank you so much for coming up.

2051
02:03:39.080 --> 02:03:41.740
Uh, really appreciate you being here, uh, jumping on in the last second.

2052
02:03:41.760 --> 02:03:44.299
So yesterday, in the middle of the crazy lunch that you guys had.

2053
02:03:44.560 --> 02:03:44.779
Allie Laabs: Yes.

2054
02:03:45.160 --> 02:03:49.119
Alex Volkov: So I really appreciate it. Uh, and also I do expect us to tap on you guys more as

2055
02:03:49.160 --> 02:03:53.319
this develops or as new use cases come up or as folks are discovering this is now

2056
02:03:53.720 --> 02:03:56.439
the, the next big thing. Thank you so much for, uh, coming on the show.

2057
02:03:56.720 --> 02:03:57.839
Ali Labs from TypeSafe.

2058
02:03:59.000 --> 02:04:00.179
Allie Laabs: Absolutely. Thank you for having me.

2059
02:04:00.560 --> 02:04:04.220
Alex Volkov: Thank you. All right, folks, we we're nearly the end of our show.

2060
02:04:04.260 --> 02:04:08.060
I'm gonna bring up back the host to kind of talk, uh, about this.

2061
02:04:08.160 --> 02:04:10.879
Nisten, at the beginning of the show, I asked you about if you saw about JEV, you

2062
02:04:10.920 --> 02:04:12.519
were super busy, you went like, no.

2063
02:04:12.960 --> 02:04:13.399
Uh,

2064
02:04:14.720 --> 02:04:17.500
quick comments from you on like what, what, what we're seeing?

2065
02:04:17.560 --> 02:04:22.360
Is this a new paradigm in decision making, in computer use, in whatever else?

2066
02:04:23.000 --> 02:04:27.939
Nisten Tahiraj: Yeah, yeah, I think it bridges the gap between people that are doing like buffer lib,

2067
02:04:28.040 --> 02:04:28.720
uh, RL

2068
02:04:29.680 --> 02:04:33.740
stuff and, uh, actual agentic computer use.

2069
02:04:33.880 --> 02:04:39.439
You do need that level of speed, uh, and it's pretty,

2070
02:04:40.080 --> 02:04:45.240
it's pretty fun that we can get something like 10x faster or 100x faster, and then

2071
02:04:45.280 --> 02:04:49.480
we're just using it. We just, just let it play games or let it handle stuff in real

2072
02:04:49.560 --> 02:04:49.759
time.

2073
02:04:49.960 --> 02:04:51.980
Alex Volkov: Wulfam, uh, thoughts on this comments?

2074
02:04:52.160 --> 02:04:53.059
You've played with this as well?

2075
02:04:53.480 --> 02:04:55.120
Wolfram Ravenwolf: Yeah, I was super excited when I saw it.

2076
02:04:55.160 --> 02:04:59.399
And the thing is, if we are comparing to the cost or to the speed, we should also

2077
02:04:59.480 --> 02:05:02.480
consider that we are comparing to LLMs, so it's much faster than that.

2078
02:05:02.880 --> 02:05:05.839
But, uh, people have been using classifiers for this for a long time.

2079
02:05:06.160 --> 02:05:10.959
So the classifier is not a new technology, but you had to train it, and if you wanted

2080
02:05:10.980 --> 02:05:12.640
to change something, you had to retrain it.

2081
02:05:12.920 --> 02:05:17.399
And here it's all done in natural language, which makes it so, so easy to modify,

2082
02:05:17.560 --> 02:05:22.279
and you just define the questions, and if the question provides bad results, you just

2083
02:05:22.440 --> 02:05:23.199
change the question.

2084
02:05:23.440 --> 02:05:23.579
Alex Volkov: Yeah.

2085
02:05:23.880 --> 02:05:26.660
Wolfram Ravenwolf: And that makes it so, uh, universally applicable.

2086
02:05:26.760 --> 02:05:31.519
So, so you can put it to Doom, you could put it to news, uh, allocation for ThursdAI,

2087
02:05:31.560 --> 02:05:36.059
you can do tweets, uh, analysis, you can do all of these things, and you don't have

2088
02:05:36.080 --> 02:05:36.760
to train a model.

2089
02:05:36.800 --> 02:05:42.360
Alex Volkov: I've been very upset with Gmail not being able to tell tweets from X,

2090
02:05:42.880 --> 02:05:46.659
like phishing, uh, emails from X that somebody accessed my account, and it's like

2091
02:05:46.700 --> 02:05:48.740
the domain is not X. I've been very upset with Gmail.

2092
02:05:48.800 --> 02:05:53.519
Gmail came up with the spam blocking, like, technology back in the 2000s, right?

2093
02:05:54.120 --> 02:05:58.200
And then I've been very easily seeing my assistants that we talked about casually

2094
02:05:58.220 --> 02:05:59.960
going to my inbox and saying, no, this is probably spam.

2095
02:06:00.000 --> 02:06:02.640
Like, they know, they have the intelligence to say this is spam, like, whatever.

2096
02:06:03.120 --> 02:06:09.380
Uh, I saw somebody using JEV on like 100,000 of their emails

2097
02:06:10.320 --> 02:06:15.019
and multi-timing them, multiplexing via their API, and just getting a score on like

2098
02:06:15.080 --> 02:06:19.340
a 100,000 emails with, again, like a cent and a half or something like crazy.

2099
02:06:19.400 --> 02:06:23.320
The, the, the numbers don't make sense, uh, within like a minute or so.

2100
02:06:24.240 --> 02:06:28.619
So this is nearly real time, less than 100 milliseconds decision based on every context

2101
02:06:28.679 --> 02:06:30.500
of whatever. I don't know about context length.

2102
02:06:30.559 --> 02:06:34.879
I don't know about like stuff like that, uh, but it does feel like, hey, the stuff

2103
02:06:34.960 --> 02:06:40.979
that LLMs became good for is now possible without writing or creating

2104
02:06:41.080 --> 02:06:44.879
and running a classifier model at scale, which, you know, many people don't, never

2105
02:06:44.920 --> 02:06:49.640
will do for themselves. Uh, and that, I think, is a great place to end the show for

2106
02:06:49.720 --> 02:06:53.320
today. I have a bunch of demos as well, like Fable cooked up, maybe I should show

2107
02:06:53.360 --> 02:06:57.120
you the demos. Fable cooked up like a bunch of demos specifically for ThursdAI, uh,

2108
02:06:57.520 --> 02:06:57.700
and

2109
02:06:58.800 --> 02:06:59.999
with Jeff. Let me see if I can pull.

2110
02:07:00.080 --> 02:07:02.379
them up, um, because

2111
02:07:03.720 --> 02:07:07.219
I just like, okay, folks, we'll hear about this from the devrel, but like, what are

2112
02:07:07.240 --> 02:07:11.400
the use cases? Somebody's asking in the comments like, hey, uh, can we,

2113
02:07:12.320 --> 02:07:17.840
where's this question? Can we use this for, uh, statements from project chats?

2114
02:07:17.960 --> 02:07:19.759
Yes, you can use this on any type of text.

2115
02:07:19.800 --> 02:07:22.960
You can feed it any type of, again, we don't know the context length.

2116
02:07:23.120 --> 02:07:26.180
You can feed any type of text and get a classification.

2117
02:07:26.440 --> 02:07:28.959
It's not multimodal, so you cannot, like, judge images.

2118
02:07:29.280 --> 02:07:32.299
Uh, this is not the hot dog or not hot dog from Silicon Valley.

2119
02:07:32.480 --> 02:07:37.379
This is not, like, yet multimodal, uh, but it can be with the pairing of an LLM, right?

2120
02:07:37.500 --> 02:07:41.420
We can pair this with Muse, take a description of a picture, and then have it, have

2121
02:07:41.480 --> 02:07:42.280
that text answer.

2122
02:07:43.360 --> 02:07:43.719
Wolfram?

2123
02:07:44.440 --> 02:07:48.659
Wolfram Ravenwolf: Just wanted to, there's another question about specifically for hallucinations, and

2124
02:07:48.680 --> 02:07:51.440
they have this claim that this model can't hallucinate.

2125
02:07:51.800 --> 02:07:55.319
And it's important to realize that what it means is the output it generates, which

2126
02:07:55.360 --> 02:07:59.680
is the JSON, um, that is not, not different like an LLM.

2127
02:07:59.700 --> 02:08:01.640
If you say make JSON, it may not do so.

2128
02:08:02.000 --> 02:08:04.760
But, uh, the questions or the answers can be false, of course.

2129
02:08:04.840 --> 02:08:09.060
So the model has its intelligence, and they benchmark this as well, where it's not

2130
02:08:09.100 --> 02:08:11.920
as good as Astra, for instance, but it's so much cheaper, so much faster.

2131
02:08:11.960 --> 02:08:17.079
So as always, you have to decide uh, what level of intelligence you want and how much

2132
02:08:17.120 --> 02:08:19.159
you want to wait or pay for it, or both.

2133
02:08:19.960 --> 02:08:23.039
Alex Volkov: Yeah, 100%. The hallucination and the probabilistic nature.

2134
02:08:23.160 --> 02:08:25.219
LLMs are not probabilistic even in temperature 0.

2135
02:08:25.319 --> 02:08:27.960
Uh, they're not deterministic at level temperature 0.

2136
02:08:28.319 --> 02:08:31.599
Here is a demo that Fable hooked up with Jeff, just to give you guys an idea of like

2137
02:08:31.640 --> 02:08:33.560
what what's possible. This is like just on the fly.

2138
02:08:34.080 --> 02:08:37.559
All the news that we talk about, I have like buckets in my head on TLDR, right?

2139
02:08:37.600 --> 02:08:41.000
We have voice and vision, we have big companies in LMs, we have AI coding agents,

2140
02:08:41.040 --> 02:08:44.020
we have tools. Sometimes the stuff we talk about co-intermix.

2141
02:08:44.040 --> 02:08:45.919
What is a computer use? Like, where where does it land?

2142
02:08:46.240 --> 02:08:51.280
Uh, this is a categorizer or classification of all the news, uh, based on based on

2143
02:08:51.360 --> 02:08:56.780
Jeff. Google launches Gemini 3 and 8, and I classify this, and it immediately classifies

2144
02:08:56.960 --> 02:09:00.399
voice and vision. This is like a very simple thing that I would use an LM for, but,

2145
02:09:00.480 --> 02:09:03.300
but this is, the number of stuff is very well defined by me.

2146
02:09:03.360 --> 02:09:04.759
I use this in my head all the time.

2147
02:09:05.280 --> 02:09:11.240
Uh, I have another version with live chat, basically, uh, talking about what

2148
02:09:11.600 --> 02:09:16.800
of the statements that we had are related to which of the topics that we talked about.

2149
02:09:17.080 --> 02:09:19.580
So classification, like on the fly, and we can rerun live.

2150
02:09:19.600 --> 02:09:23.080
You can see this happening within 1.3 seconds and 42 calls.

2151
02:09:23.560 --> 02:09:28.080
Uh, live producer, I didn't yet implement this, but you can see rebuilding this.

2152
02:09:28.320 --> 02:09:32.240
So you can see here the chapters that we talked about, the TLDR intros, and then this

2153
02:09:32.280 --> 02:09:36.120
is last week's show. This is the TLDR, this is uh DeepSeek that we talked about for

2154
02:09:36.160 --> 02:09:40.879
the longest time. Here underneath, you guys can see how Jeff scored this.

2155
02:09:41.639 --> 02:09:44.860
Okay, so like instead of me saying, hey, from here to here, we talked about this,

2156
02:09:44.919 --> 02:09:49.280
and then here to here is TLDR, you can see that Jev matches nearly perfectly the kind

2157
02:09:49.320 --> 02:09:52.900
of based on the statements that we made, which chapter are we in, but we can see that

2158
02:09:52.919 --> 02:09:57.099
here, like, it added some stuff about doomerism before the chapter that we talked

2159
02:09:57.139 --> 02:10:02.940
about, uh, Cognition 3.2. And I will absolutely now, I didn't have time, will implement

2160
02:10:03.000 --> 02:10:04.439
this in the live reasoning thing.

2161
02:10:05.480 --> 02:10:09.179
This all costs cents, less than cents.

2162
02:10:09.480 --> 02:10:10.240
Wolfram Ravenwolf: Basically free.

2163
02:10:10.639 --> 02:10:12.819
Alex Volkov: It's ba- no, the outputs are literally free.

2164
02:10:12.880 --> 02:10:16.560
They're not charging for output tokens, like, but inputs are just like basically free.

2165
02:10:17.120 --> 02:10:19.559
42 dollars per 1 billion tokens.

2166
02:10:21.280 --> 02:10:22.000
It's insane.

2167
02:10:22.520 --> 02:10:22.959
Wolfram Ravenwolf: Bro.

2168
02:10:25.520 --> 02:10:29.480
Alex Volkov: I don't know if I've written 42 billion tokens in my life as a human, like, literally

2169
02:10:29.520 --> 02:10:30.759
typing. I don't know if I got there.

2170
02:10:30.800 --> 02:10:36.899
Like, I think that, let me say this, I think that Jev can categorize all of

2171
02:10:37.000 --> 02:10:41.719
my life's work in terms of the text that I outputted as a human with less than 40

2172
02:10:41.800 --> 02:10:42.119
dollars.

2173
02:10:43.280 --> 02:10:45.500
Yam Peleg: What what's the context length if we're already at it?

2174
02:10:45.520 --> 02:10:49.059
Alex Volkov: I don't know. We should have asked, and we will ask Ali if somebody wants to pull

2175
02:10:49.080 --> 02:10:50.800
up an assistant, context length is very important.

2176
02:10:51.720 --> 02:10:54.319
Uh, new paradigm, exciting to play with, uh, folks.

2177
02:10:54.360 --> 02:10:56.120
There's a bunch of stuff that that happened as well.

2178
02:10:56.200 --> 02:11:00.219
Uh, next week is going to be pretty crazy because OpenAI's release for this week was

2179
02:11:00.260 --> 02:11:04.500
pushed to next week. Uh, Grok 4.7 is also coming, they said, uh, they there was, you

2180
02:11:04.520 --> 02:11:08.240
know, Elon said it's gonna come last year, la- last week, but now it's probably pushed,

2181
02:11:08.320 --> 02:11:10.000
uh, so next week is gonna be very, very crazy.

2182
02:11:10.440 --> 02:11:14.359
Uh, hopefully folks got excited about JEV as much as we got.

2183
02:11:14.640 --> 02:11:16.379
Uh, thank you so much for joining, everyone.

2184
02:11:16.440 --> 02:11:18.800
We're way past our 2 hour allotted timeline.

2185
02:11:19.120 --> 02:11:23.759
Uh, this is the world of Fable and Astra and all of the tools that we're living in.

2186
02:11:23.920 --> 02:11:28.899
I I rebuilt completely the Descript app that we use to edit the show with Thursday

2187
02:11:28.920 --> 02:11:34.919
AI specific work. So what you're seeing here is a, like, 0 to 1, not 0

2188
02:11:35.080 --> 02:11:38.399
shot, because I worked for this, and I think I spent like 3 billion tokens of LLMs,

2189
02:11:38.440 --> 02:11:43.399
by the way. I can't imagine how ma- how, how, how much faster Jeff will make this

2190
02:11:43.480 --> 02:11:47.600
in accordance with LLMs, that we have the whole episode with a timeline and every

2191
02:11:48.200 --> 02:11:54.240
an LLM or multiple LLMs can suggest things to cut because these are like live only.

2192
02:11:54.800 --> 02:11:58.639
Can suggest that where Yam is speaking or Nisten said, hey, this is Chris Alexiuk.

2193
02:11:58.720 --> 02:12:00.360
This is, this comes directly from our things.

2194
02:12:00.720 --> 02:12:06.200
This is like a, um, an editor. We also have clips for the clips factory that can play.

2195
02:12:06.520 --> 02:12:09.059
Uh, there's really a bunch of stuff there that I cooked.

2196
02:12:09.360 --> 02:12:13.600
The reason for this is that my agents want to help me with cutting, but I want to

2197
02:12:13.640 --> 02:12:17.420
give my co-hosts, their agents, access to tell me, hey, what is that one clip that

2198
02:12:17.460 --> 02:12:18.720
I said? I want to pull this for me.

2199
02:12:18.760 --> 02:12:23.299
So, like, I'm building that for this ThursdAI editor, and it's nearly ready, uh, and

2200
02:12:23.400 --> 02:12:28.919
uh, Jev, Jev is going to get there and decide, you know, for every um and ah, for

2201
02:12:28.940 --> 02:12:32.220
example. Yeah, we're gonna build Jev into this, and it's gonna be amazing.

2202
02:12:33.160 --> 02:12:34.519
Yam Peleg: Bro, you can sell this.

2203
02:12:34.920 --> 02:12:35.199
Alex Volkov: I-

2204
02:12:35.560 --> 02:12:38.519
Yam Peleg: You, you can sell this. It's insanely good.

2205
02:12:38.880 --> 02:12:42.459
Alex Volkov: Thank you. Uh, we'll see, we'll see if it works when it works, but yeah, it's, it's,

2206
02:12:42.520 --> 02:12:44.100
it's pretty dope. Uh, all right, folks.

2207
02:12:44.160 --> 02:12:45.860
Wolfram Ravenwolf: The best way to show what.

2208
02:12:48.440 --> 02:12:51.600
Alex Volkov: Yep. Wolfram, we lost you there for just a second, it looks like.

2209
02:12:51.720 --> 02:12:56.720
Um, but I think it's a great way to end the show, so I'd be able to go and use this.

2210
02:12:56.920 --> 02:12:58.779
Uh, Wolfram, what did you say? You, we lost you, now you're back.

2211
02:12:59.120 --> 02:13:03.339
Wolfram Ravenwolf: Oh, I wanted to say this is the perfect example what we are preaching here, that you

2212
02:13:03.400 --> 02:13:08.319
are showing the proof that is not just a demo you do and something, but you use this

2213
02:13:08.480 --> 02:13:12.399
technology we are talking about all the time to build something really useful and

2214
02:13:12.600 --> 02:13:16.579
valuable to yourself and everybody who has the same problem.

2215
02:13:17.440 --> 02:13:22.600
Alex Volkov: And I cannot tell you how much JEV makes me excited because I literally can come through

2216
02:13:22.680 --> 02:13:26.899
every sentence and score it based on multiple parameters, whether or not this makes

2217
02:13:27.000 --> 02:13:31.579
sense to the topic that we're discussing, whether or not this is a somebody went on

2218
02:13:31.680 --> 02:13:35.279
a, you know, on a debate and came back, whether or not it's gonna be viral.

2219
02:13:35.360 --> 02:13:39.779
There's like so many things that Jeff can like just judge and process this in a second

2220
02:13:39.839 --> 02:13:43.860
that I cannot wait to get off the street and actually go and and code this, but uh,

2221
02:13:43.920 --> 02:13:46.759
but if you have missed any part of the show, we had an incredible show.

2222
02:13:46.800 --> 02:13:50.560
We had David Paulon in the Assistant Benchmark talking about different assistants

2223
02:13:50.600 --> 02:13:55.020
and how he's scoring them. We had Ali Lab from TypeSafe talking about this, like,

2224
02:13:55.120 --> 02:14:00.640
new and exciting world of AI intelligence that's not LLM based.

2225
02:14:00.920 --> 02:14:05.279
They call them System 1 thinking, and they're incredibly cheap and super fast, and

2226
02:14:05.360 --> 02:14:08.839
many of the things that we started using LLMs for, but LLMs are not fitted for.

2227
02:14:09.120 --> 02:14:14.100
We also had, uh, Francesco from Kua talk about how computer use is like a very important

2228
02:14:14.120 --> 02:14:19.540
thing, both in in both these worlds, and and the TypeSafe JEV, uh, very much helps

2229
02:14:19.800 --> 02:14:23.799
the agents to get better at the tasks that you want the agents to get better at, at

2230
02:14:23.880 --> 02:14:27.440
proactivity, at different things, at prioritizing a a bunch of other stuff.

2231
02:14:27.520 --> 02:14:30.679
So, very exciting show. This feels like a monumental week.

2232
02:14:31.160 --> 02:14:34.620
So, releases from Astra to whatever, they're big and incredible.

2233
02:14:34.720 --> 02:14:38.040
This feels like a monumental week in terms of like what we can do, in terms of how

2234
02:14:38.120 --> 02:14:42.300
big AI's assistants are getting, how fast JEV is getting, and how big computer use

2235
02:14:42.320 --> 02:14:44.380
is getting as well. Folks, thank you so much for joining us.

2236
02:14:44.440 --> 02:14:47.359
Again, if you missed any part of the show, Thursday AI is live.

2237
02:14:47.720 --> 02:14:50.440
Uh, you can rewatch everything on thursdAI.live, by the way.

2238
02:14:50.520 --> 02:14:53.760
Like, once we get off, you can rewatch the whole live stream right here, uh, and see

2239
02:14:53.800 --> 02:14:57.519
the transcripts and everything, but also you are able to, uh, get it as a podcast,

2240
02:14:57.600 --> 02:15:00.839
as a newsletter. Please subscribe, uh, and give us five stars wherever you're listening.

2241
02:15:00.880 --> 02:15:04.060
Please tell your friends that if this is informative for you, maybe they can enjoy

2242
02:15:04.120 --> 02:15:08.180
this as well, um, because we really try really hard to bring you the best possible

2243
02:15:08.280 --> 02:15:11.319
show, uh, and the earliest news, so you'll stay ahead of the game.

2244
02:15:11.400 --> 02:15:12.999
I think we're doing a great job at that.

2245
02:15:13.320 --> 02:15:17.860
Uh, so thank you, co-hosts Wolfram Ravenwolf, Nisten, LDJ, and Yam Peleg for joining

2246
02:15:17.880 --> 02:15:20.960
us. Peter Gosta was here before, and our 3 guests from today as well.

2247
02:15:21.280 --> 02:15:24.500
And, uh, most of all, thank you for tuning in, listening, and if you're still here,

2248
02:15:24.800 --> 02:15:26.779
uh, uh, give us a shout out in the chat.

2249
02:15:26.840 --> 02:15:27.440
Thank you, folks.

2250
02:15:29.520 --> 02:15:29.919
Bye-bye.
