WEBVTT

1
00:00:00.035 --> 00:00:03.730
Peter Gostev: And the point is that I was spending literally hundreds of billions of tokens.

2
00:00:03.805 --> 00:00:07.124
I'm not kidding. I probably spent, I don't know, 3, 400 billion.

3
00:00:07.224 --> 00:00:11.804
I know I need to go back and check. And I literally had my Linux box running, uh, for

4
00:00:11.985 --> 00:00:15.924
weeks and weeks, and it was filling up with all the solutions and so on.

5
00:00:16.525 --> 00:00:19.204
But the T- TLDIs, that was complete waste of time.

6
00:00:19.325 --> 00:00:21.765
I did not discover a single bloody thing.

7
00:00:22.265 --> 00:00:27.604
So these guys did it in like 3 hours. I couldn't do it in months with Astra and

8
00:00:28.465 --> 00:00:30.204
5.6 and Fable and others.

9
00:00:30.445 --> 00:00:36.225
LDJ: Yeah, so so Fable and Astra had put together a set of 500 most

10
00:00:36.365 --> 00:00:42.104
important open math problems in the world, and 92 of those just got

11
00:00:42.245 --> 00:00:44.145
solved on Tuesday by OpenAI.

12
00:00:44.710 --> 00:00:49.729
Peter Gostev: Yeah, so we, we have some news, uh, we raised 200 million.

13
00:00:50.190 --> 00:00:53.950
That's a lot of money, uh, at 3.1 billion valuation.

14
00:00:54.375 --> 00:01:00.115
Alex Volkov: So basically, his personal CFO just went public on Slack and uh, told

15
00:01:00.175 --> 00:01:01.694
everybody his bank account.

16
00:01:01.844 --> 00:01:05.125
Maxime Labonne: Yes, decision models do not output any tokens.

17
00:01:05.285 --> 00:01:09.125
They just take a decision from a predefined set of answers.

18
00:01:09.685 --> 00:01:13.704
Nisten Tahiraj: So there are 100,000 agents going on in this city.

19
00:01:13.905 --> 00:01:18.704
I don't know how it's still holding together, but, uh, yeah, it's a full emulation of

20
00:01:18.745 --> 00:01:19.665
the city of Toronto.

21
00:02:48.703 --> 00:02:54.183
LDJ: Okay, so it's amazing, past 7 days, even just the past couple days has been crazy

22
00:02:54.263 --> 00:02:58.102
with the, the math announcements, but yeah.

23
00:02:58.743 --> 00:03:03.683
Alex Volkov: Tuesday was just, just awful. I, I need to go and pull up my tweet over there

24
00:03:03.743 --> 00:03:09.362
because, um, I, you know, I don't usually do this like a quick recaps, but I just had

25
00:03:09.383 --> 00:03:10.982
to. It was quite insane.

26
00:03:11.649 --> 00:03:17.229
LDJ: A couple of American open models too, of Reflection, which has kind of been in semi

27
00:03:17.329 --> 00:03:21.189
stealth for a while. They just announced their first model, which they plan on

28
00:03:21.649 --> 00:03:25.069
releasing open weights soon too, although it's not quite open weights yet.

29
00:03:25.369 --> 00:03:26.449
Alex Volkov: I think it was Beam, yeah?

30
00:03:27.209 --> 00:03:27.909
LDJ: Mhm, exactly.

31
00:03:27.929 --> 00:03:31.849
Alex Volkov: Yes, Reflection announced Beam. We're gonna mention this, uh, later on the show.

32
00:03:32.369 --> 00:03:37.888
Um, actually got access. Uh, so shout out to the Beam folks for getting us access,

33
00:03:38.009 --> 00:03:43.518
uh, but I didn't have time to go and play with it yet, but, um, I- I just want to

34
00:03:43.539 --> 00:03:46.858
highlight this like one day in the middle of this week.

35
00:03:47.139 --> 00:03:52.159
OpenAI dropped 722 math manuscripts covering,

36
00:03:54.059 --> 00:03:59.498
I think, 5 Navier-Stokes worth problems. I think this was the calculation that that

37
00:03:59.579 --> 00:04:03.459
somebody, um, which is still, you know, we're gonna talk about this.

38
00:04:03.779 --> 00:04:08.779
We're still trying to figure out what was this thing that OpenAI just brought to us,

39
00:04:08.819 --> 00:04:11.978
and and and mathematicians are trying to get grips with it.

40
00:04:12.499 --> 00:04:16.259
Uh, Meta, Stripe, and Shopify announced, this is a single day, this is a single

41
00:04:16.339 --> 00:04:20.458
Tuesday of this week, folks, all right? Uh, Meta, Stripe, and Shopify, Shopify

42
00:04:20.579 --> 00:04:23.899
announced the agent protocol or assistant protocol, whatever you want to call it.

43
00:04:24.939 --> 00:04:30.364
Uh, OpenAI finally gave us access to the Decisions API, which is their competitor to

44
00:04:30.525 --> 00:04:35.659
Jev And, uh, Mistral came out with Mistral 4, uh, with Lechonk.

45
00:04:36.880 --> 00:04:41.560
And Google dropped Embeddings Gemma. Brad Adcock, the guy who does figure robots, uh,

46
00:04:41.840 --> 00:04:46.860
got a new, um, assistant in the, in the arena of assistants, right?

47
00:04:46.880 --> 00:04:50.800
So there's now Hark Pro, which is, you know, a lot of people signed up for.

48
00:04:51.440 --> 00:04:57.020
And, uh, Perplexity open weighted a OpenJev, uh, decider, uh, update.

49
00:04:57.080 --> 00:04:58.679
This is an update. They already dropped it before.

50
00:04:59.520 --> 00:05:03.800
Claude now does Google Docs. Uh, we at CoreWeave released RL rollouts.

51
00:05:04.600 --> 00:05:10.379
Um, this is a single day. It was really hard to catch up with all this just this

52
00:05:10.479 --> 00:05:13.880
Tuesday, but obviously there has been more stuff happening.

53
00:05:14.080 --> 00:05:18.279
Welcome Peter Gostev on the stage. Welcome to ThursdAI.

54
00:05:18.479 --> 00:05:24.179
Um, there there's a huge, huge amount of stuff to cover, so I think, uh, yeah, let's

55
00:05:24.200 --> 00:05:27.159
do a little bit of banter, folks. What is, what is the highlight?

56
00:05:28.160 --> 00:05:29.860
Let's start with LDJ. You were, you were here first.

57
00:05:29.900 --> 00:05:33.639
What is the highlight for you from this week, if you had to pick one?

58
00:05:34.039 --> 00:05:39.860
LDJ: Yeah, um, huh. I- I'm sure, I'm sure Nisten and Peter and others are

59
00:05:39.920 --> 00:05:44.199
going to probably mention the math part, so I'm going to say Reflection AI's Beam.

60
00:05:44.880 --> 00:05:47.580
Um, so it's pretty competitive in terms of...

61
00:05:48.800 --> 00:05:53.439
it- it's, it's definitely, I would say, second, third, fourth-ish place in open

62
00:05:53.520 --> 00:05:57.780
models when you look at raw benchmark scores, but it seems like it might even be just

63
00:05:57.960 --> 00:06:02.759
first place when you look at reasoning efficiency compared to the top models like GLM

64
00:06:02.880 --> 00:06:03.339
and others.

65
00:06:03.680 --> 00:06:03.980
Alex Volkov: Mm.

66
00:06:04.080 --> 00:06:08.620
LDJ: So we'll look into that more later, and uh, yeah, that's, that might be my top one.

67
00:06:10.560 --> 00:06:11.420
Alex Volkov: Peter, how about you?

68
00:06:12.560 --> 00:06:16.039
Peter Gostev: Oh yeah, the math stuff was something else.

69
00:06:16.200 --> 00:06:17.019
I know.

70
00:06:17.039 --> 00:06:20.560
Alex Volkov: It, it took over your timeline a lot, like I see you commenting on this.

71
00:06:20.800 --> 00:06:22.079
Um, tell us,

72
00:06:22.766 --> 00:06:23.066
Peter Gostev: And it's-

73
00:06:23.006 --> 00:06:23.785
Alex Volkov: What's going on there?

74
00:06:24.166 --> 00:06:28.626
Peter Gostev: Yeah, so there's, there's been a kind of, I don't know, rumor, I guess they kind of

75
00:06:28.766 --> 00:06:31.626
pre-announced it that they have a bunch of maths problems solved.

76
00:06:31.686 --> 00:06:34.045
I don't know, was that a rumor or was that just announcement?

77
00:06:34.126 --> 00:06:37.326
I can't remember what was publicly known or not, but

78
00:06:37.726 --> 00:06:38.626
LDJ: At first a rumor.

79
00:06:39.166 --> 00:06:44.746
Peter Gostev: Yeah, yeah, okay. So, so OpenAI, uh, with their new Bell model, right, they solved

80
00:06:44.766 --> 00:06:49.146
the Navier-Stokes problem, the Millennium Prize problem, and then there was a lot of

81
00:06:49.526 --> 00:06:53.626
chatter about that they've, they're just sitting in a war chest of hundreds of

82
00:06:53.766 --> 00:06:59.386
problems solved, and, and now they released it, and they just did it pretty, in a

83
00:06:59.486 --> 00:07:05.025
pretty low-key way, not like a hype marketing video, uh, with, uh, exploding head

84
00:07:05.066 --> 00:07:05.366
emojis.

85
00:07:05.246 --> 00:07:06.686
Alex Volkov: It was nothing. The blog post was just

86
00:07:07.486 --> 00:07:07.786
Peter Gostev: Yeah.

87
00:07:07.686 --> 00:07:11.925
Alex Volkov: Just a few lines of code, and the GitHub was just an insane amount of...

88
00:07:12.286 --> 00:07:15.806
Uh, you know my favorite, uh, term for this now is slap grenade?

89
00:07:17.366 --> 00:07:21.566
When somebody throws at you like an output or, or, or a, or a pull request of Claude

90
00:07:21.606 --> 00:07:23.665
and it has like a bunch of lines, this is now a slap grenade.

91
00:07:23.686 --> 00:07:28.025
So Open, OpenAI basically like slap grenade, uh, for people to just like shovel

92
00:07:28.086 --> 00:07:28.386
Peter Gostev: Yeah.

93
00:07:28.406 --> 00:07:29.645
Alex Volkov: and try to figure out what this is.

94
00:07:30.716 --> 00:07:34.096
Peter Gostev: Yeah, maybe my personal reflection is that I didn't really talk about this much, but

95
00:07:34.156 --> 00:07:39.116
I was, uh, with GPT-5.6 or GPT-6. I picked one problem, and I thought, let me just

96
00:07:39.576 --> 00:07:44.155
throw all of the tokens I can get my hands on and just see what I can do with it.

97
00:07:44.236 --> 00:07:48.846
And I was also trying with Fable and Opus and so on, but I was running the GPT models

98
00:07:48.946 --> 00:07:51.515
mostly just because were doing a lot of resets and things like that.

99
00:07:52.156 --> 00:07:56.536
And, was one problem that I was trying with the Hardwing and Nelson, not sure how you

100
00:07:56.636 --> 00:08:00.246
pronounce exactly, it was like, graph coloring kind of problem.

101
00:08:00.286 --> 00:08:03.766
So it's easy to grasp, obviously impossibly hard to solve.

102
00:08:04.206 --> 00:08:07.901
And the point is that I was spending literally hundreds of billions of tokens.

103
00:08:07.976 --> 00:08:11.295
I'm not kidding. I probably spent, I don't know, 3, 400 billion.

104
00:08:11.395 --> 00:08:15.975
I know I need to go back and check. And I literally had my Linux box running, uh, for

105
00:08:16.156 --> 00:08:20.095
weeks and weeks, and it was filling up with all the solutions and so on.

106
00:08:20.696 --> 00:08:23.375
But the T- TLDIs, that was complete waste of time.

107
00:08:23.496 --> 00:08:25.936
I did not discover a single bloody thing.

108
00:08:26.296 --> 00:08:31.145
There was nothing, no use. was an interesting mathematical result, nothing, right?

109
00:08:31.586 --> 00:08:37.095
And then, these guys with the new greatest, model, they just, run it for 3 hours,

110
00:08:37.266 --> 00:08:41.925
said on average for 3 hours on ChatGPT Pro, which I think in practice what that means

111
00:08:41.986 --> 00:08:46.525
is that it's like a, some sub-agents, I think we don't know how many, but not like

112
00:08:46.666 --> 00:08:49.965
hundreds, like 10 maybe, or 4 or 10, something like that.

113
00:08:50.525 --> 00:08:55.465
And uh, it uh s- it uh didn't solve it completely, but it changed one of the bounds.

114
00:08:55.525 --> 00:08:59.325
The idea is that you can narrow the bounds, and I think before it was something like

115
00:08:59.566 --> 00:09:04.845
5, 6, or 7, and it narrowed it down to 6 or 7, which is a big deal, right?

116
00:09:04.926 --> 00:09:09.566
It is a, a obviously a lot of context required, but it's a big deal, right?

117
00:09:09.726 --> 00:09:15.166
And I was speaking to Astra about m- my progress, my results, and Astra said, oh, we

118
00:09:15.226 --> 00:09:18.086
are two mathematical breakers away from what they did.

119
00:09:18.486 --> 00:09:23.825
So these guys did it in like 3 hours. I couldn't do it in months with Astra and

120
00:09:24.686 --> 00:09:27.345
5.6 and Fable and others. There's just nothing.

121
00:09:27.545 --> 00:09:31.566
Alex Volkov: There's there's a few rumors that I want to bring to your attention there, uh, but I

122
00:09:31.666 --> 00:09:35.585
think I think it's time for, you know, for like the official start.

123
00:09:35.866 --> 00:09:38.825
Uh, folks who are tuning in, welcome to ThursdAI.

124
00:09:38.946 --> 00:09:41.665
Uh, one of the coolest things that we get to do is talk to you as well, so please

125
00:09:41.706 --> 00:09:44.626
drop in the comments, what was the highlight of this AI week to you?

126
00:09:44.946 --> 00:09:47.906
Obviously, it's nearly impossible to cover everything.

127
00:09:48.396 --> 00:09:53.755
we're trying to make ThursdAI the most agentic forward show.

128
00:09:53.835 --> 00:09:58.636
We have an AI producer. Uh, producer, please say hi in the banner below, uh, if when

129
00:09:58.756 --> 00:10:04.675
and and if and when you hear this. And we also have, obviously, a bunch of research

130
00:10:04.756 --> 00:10:07.416
happening throughout the the to make sure that we're not missing anything.

131
00:10:07.716 --> 00:10:13.696
Uh, this week, I think I had 48 topics to bring you of, you know, of

132
00:10:13.796 --> 00:10:16.156
importance, et cetera, 48 or or 50 or something.

133
00:10:16.576 --> 00:10:22.115
I actually asked Claude to build me like a, l- l- like a Tinder for AI news where I

134
00:10:22.236 --> 00:10:25.196
swipe left and right if I think that this should make the show, and it will stack

135
00:10:25.276 --> 00:10:30.036
rank them based on, like, my preference and the stuff that we talk about, uh, because

136
00:10:30.116 --> 00:10:33.475
it's just like impossible to bring you everything, but we're gonna try to bring you

137
00:10:33.516 --> 00:10:39.215
the most important news, so, uh, once producer catches up, I think it's time for, for

138
00:10:39.276 --> 00:10:42.836
a cold open. Alrighty, folks, let's, let's do this.

139
00:10:45.151 --> 00:10:50.890
Folks, this week OpenAI dropped 722 math problems, papers, uh,

140
00:10:51.031 --> 00:10:56.250
written by a model nobody outside OpenAI has touched yet, and the mathematicians are

141
00:10:56.371 --> 00:11:00.590
split between thrilled and asking for receipts, and a lot of those came with

142
00:11:00.711 --> 00:11:02.991
receipts. Welcome to ThursdAI for October 8th.

143
00:11:03.311 --> 00:11:08.590
This is Alex Volkov. I'm an AI evangelist with CoreWeave, and with me, LDJ and Peter

144
00:11:08.631 --> 00:11:11.710
Gostef from Arena. This was an insane week.

145
00:11:12.071 --> 00:11:15.430
We say this often, but this was an insane fucking week.

146
00:11:15.991 --> 00:11:21.690
Uh, after OpenAI's math manuscripts, uh, 722 of them, to

147
00:11:21.751 --> 00:11:26.431
be, to be precise, uh, some people call this a slob grenade.

148
00:11:27.111 --> 00:11:31.790
Um, many mathematicians are split between this is the most exciting time ever, we're

149
00:11:31.871 --> 00:11:35.851
breaking through the boundaries of mathematics, and, uh, we should...

150
00:11:37.151 --> 00:11:38.791
This is not how mathematics should happen.

151
00:11:39.190 --> 00:11:42.631
This is, this is kind of the split, and Peter, I think, has been monitoring this and

152
00:11:42.671 --> 00:11:44.330
would love to, to, to chat with you more.

153
00:11:44.551 --> 00:11:47.331
This would be our item number one, and that's just the opener, right?

154
00:11:47.571 --> 00:11:49.650
We'll start the TLDR, obviously, with OpenAI.

155
00:11:49.891 --> 00:11:54.930
722 math manuscripts from a model that nobody open- outside of OpenAI can use.

156
00:11:55.331 --> 00:12:01.051
Uh, somebody used Fable to, um, t- to

157
00:12:01.171 --> 00:12:06.571
put this release in terms of how many Navier-Stokes moments in one drop we got,

158
00:12:07.171 --> 00:12:12.371
and uh, and I think it's around 5 or so orders of magnitude of

159
00:12:12.491 --> 00:12:16.430
Navier-Stokes, uh, releases. If you guys remember, Navier-Stokes, a Millennium Prize

160
00:12:16.491 --> 00:12:18.691
problem, was solved by OpenAI just recently.

161
00:12:18.971 --> 00:12:24.631
So OpenAI just dropped 722 of these, in- many including Lean, uh, confirmations.

162
00:12:24.651 --> 00:12:26.450
We're gonna talk more in depth about this.

163
00:12:27.011 --> 00:12:32.770
Uh, next up is, uh, Claude Haiku 5.5 is a 10 cents per million input tokens, and

164
00:12:32.851 --> 00:12:38.530
Anthropic say that it beats GPT-5, GPT-6 Luna on computer use and coding.

165
00:12:38.691 --> 00:12:41.131
GPT, remember, was really, really good at computer use.

166
00:12:41.851 --> 00:12:47.771
Haiku is- seems to be better with a 72% on OS World, and

167
00:12:48.331 --> 00:12:52.251
I think also significantly faster, I think, on artificial analysis.

168
00:12:52.771 --> 00:12:57.771
Uh, Claude Haiku 5.5 is one of the fastest models right now, and it's incredibly

169
00:12:57.851 --> 00:13:03.771
cheap. Incredibly cheap. Uh, Haiku 4.5, I don't

170
00:13:03.791 --> 00:13:09.290
know if you noticed this, was a dollar per million tokens of input, and Haiku

171
00:13:09.451 --> 00:13:13.090
5.5, which is significantly better, is 10 cents.

172
00:13:13.611 --> 00:13:17.751
Just j- j- just to give you a sense of how crazy the world we're living in.

173
00:13:17.951 --> 00:13:23.291
Uh, also a small thing that they released together with with Haiku, Sonnet

174
00:13:24.191 --> 00:13:29.370
cash reads, so the tokens, you know, after after you send them once, uh, Sonnet cash

175
00:13:29.471 --> 00:13:32.990
reads are now halved by price a week after Sonnet 5.5 was released.

176
00:13:33.271 --> 00:13:36.391
So that's also quite insane from this week.

177
00:13:36.791 --> 00:13:42.750
Let's see what else. Uh, OpenAI rolls out GPT-6 with intelligent UI to everyone.

178
00:13:42.911 --> 00:13:46.550
All of the free users on ChatGPT, all of them are getting GPT-6.

179
00:13:47.471 --> 00:13:51.190
It's not Luna, it's not Terra is dead, so no Terra anymore.

180
00:13:51.271 --> 00:13:55.430
It's not, uh, it's not Soul, it's not Astra, it's just GPT-6.

181
00:13:55.871 --> 00:14:00.190
There's no name in this. Uh, so all the free users can d- can can get it today, but

182
00:14:00.311 --> 00:14:04.511
also the cool thing, the the intelligent UI is really, really cool.

183
00:14:04.791 --> 00:14:10.311
They rebuilt ChatGPT on top of it, and you basically get buttons and switches and

184
00:14:10.511 --> 00:14:14.711
toggles and and graphs, et cetera, in your chat output very fast.

185
00:14:15.070 --> 00:14:18.920
It's really cool to see. Um, it's like little working apps for your answer.

186
00:14:19.021 --> 00:14:24.420
There's also some threads from folks at OpenAI who say, um,

187
00:14:24.941 --> 00:14:28.700
the model was now trained to do it while it thinks of an answer, right?

188
00:14:28.741 --> 00:14:34.020
So it starts to build you, hey, this answer would probably get benefit from these and

189
00:14:34.061 --> 00:14:37.801
these switches, these and these things, uh, and then it it still works in the

190
00:14:37.861 --> 00:14:39.221
background. So I think that's pretty cool.

191
00:14:40.567 --> 00:14:41.006
Um,

192
00:14:42.607 --> 00:14:48.526
Claude for Google Workspaces is finally here, and this is

193
00:14:48.607 --> 00:14:51.166
like a small thing, but it's a huge thing for many people.

194
00:14:51.267 --> 00:14:57.066
I'll give you an example. Our cloud producer every week, uh, used to post a Google

195
00:14:57.127 --> 00:15:01.447
Doc, and then a new piece of news would come, and I would ask it, hey, can you append

196
00:15:01.567 --> 00:15:04.567
to the Google Doc? And it would start opening computers, et cetera.

197
00:15:05.087 --> 00:15:09.427
It wasn't native, so now that it's native in Google Docs, I think it's a, it's a very

198
00:15:09.447 --> 00:15:12.326
big deal, uh, and it also asks me for every edit.

199
00:15:12.727 --> 00:15:18.507
Uh, but okay, let's go to the open frontier as LDJ mentioned, we now have another

200
00:15:18.807 --> 00:15:23.527
US lab. Well, we knew about Reflection a while, but, and Reflection AI announced

201
00:15:23.647 --> 00:15:29.127
Beam. It's a 501 billion parameter open model trained from scratch in the US

202
00:15:29.807 --> 00:15:35.367
with Apache 2 weights coming this month. And shout out to Reflection for this, uh, o-

203
00:15:35.447 --> 00:15:39.886
o- for this awesome, uh, announcement. Many, many people got very excited.

204
00:15:40.007 --> 00:15:45.227
An open MOE that beats on efficiency over Roscore, so it's more efficient versus just

205
00:15:45.287 --> 00:15:49.047
like Benchmax. Uh, we can't wait to play with this when the, the weights drop.

206
00:15:50.207 --> 00:15:55.826
Also, Mistral came back with leaning into le le le

207
00:15:55.927 --> 00:15:58.067
chaton fat. If you guys remember, this was the meme.

208
00:15:58.367 --> 00:16:03.967
Uh, they called this one le chonk. Mistral 4 is a trillion parameter with the weights

209
00:16:04.007 --> 00:16:09.307
coming at the end of October. This is now a pattern that I do want to cover on the

210
00:16:09.367 --> 00:16:12.927
show. Like, folks announce open, but they not release open.

211
00:16:13.807 --> 00:16:19.767
I- is a far, far way that we moved from Mistral announcing their, uh, latest

212
00:16:19.887 --> 00:16:23.886
model with just a torrent link, because people do, do be announcing.

213
00:16:24.447 --> 00:16:29.507
Uh, but Mistral large ties with GPT-6 Luna into official analysis, uh, at a much

214
00:16:29.647 --> 00:16:33.086
higher cost per task. So that's a very interesting, very interesting thing.

215
00:16:33.527 --> 00:16:39.247
Uh, this is, this is kind of the comparison we have here, uh, obviously for European

216
00:16:39.327 --> 00:16:43.787
folks. Speaking of European folks, uh, we're missing Wolfram this week because

217
00:16:43.847 --> 00:16:47.926
Wolfram is on a well-deserved vacation, but he got super excited about Calibri.

218
00:16:48.407 --> 00:16:50.747
This is from Aleph Alpha, another European lab.

219
00:16:50.806 --> 00:16:53.726
If you guys remember, there's a, there's a European lab called Aleph Alpha.

220
00:16:53.887 --> 00:16:56.926
Uh, uh, Peter may, may, may know them as well.

221
00:16:57.086 --> 00:17:01.247
It's a German-English MoE with 78 billion parameters.

222
00:17:01.567 --> 00:17:06.967
Uh, Wolfram sent me his notes on this. It can run on a single H200 and also Apache 2.

223
00:17:07.047 --> 00:17:11.127
So shout out to Aleph Alpha for, uh, for coming back with the model.

224
00:17:11.207 --> 00:17:13.326
I thought they stopped doing pre-training and post-training.

225
00:17:13.367 --> 00:17:18.847
It just like went with models. And uh, for the very, very, um,

226
00:17:19.407 --> 00:17:19.746
let's say

227
00:17:22.127 --> 00:17:25.826
AI internals oriented folks, we used to cover embedding models.

228
00:17:25.927 --> 00:17:30.527
We we had Bo Wang on on the show before he joined Perplexity, and he dropped one of

229
00:17:30.547 --> 00:17:35.226
these models, so shout out to Bo. Uh, two open embedding models dropped in one week.

230
00:17:35.327 --> 00:17:40.206
Google's embedding Gemma drops embedding two text, images, and video and audio

231
00:17:40.287 --> 00:17:43.806
embeddings in one space that, uh, runs on the phone.

232
00:17:44.927 --> 00:17:49.346
Two embedding models this week, and Perplexity late interaction models search PDFs as

233
00:17:49.407 --> 00:17:54.727
images and are very fast with PPA- PPLX Embed V2 late interaction

234
00:17:55.047 --> 00:17:57.687
models. So a lot of, a lot of goodies this week.

235
00:17:58.167 --> 00:17:59.487
Uh, as we move to

236
00:18:01.407 --> 00:18:07.087
agents and assistants and a bunch of stuff, this, I think, is a very important thing.

237
00:18:07.247 --> 00:18:12.826
Uh, so Meta and Sierra, uh, and and Shopify and Stripe and Walmart all on board of

238
00:18:12.927 --> 00:18:15.967
this protocol for personal agent protocol.

239
00:18:17.007 --> 00:18:21.947
Uh, very interesting to see if this is the new kind of MCP, because MCP had its ups

240
00:18:21.967 --> 00:18:25.286
and downs, but now MCP is everywhere. Basically, MCP, like, took over the world

241
00:18:25.807 --> 00:18:28.686
because of these agents they need to talk to, to a bunch of apps.

242
00:18:29.247 --> 00:18:33.567
Uh, I don't know if you guys caught this, we talked about Grokbot since launch.

243
00:18:34.167 --> 00:18:38.987
There's a whole drama happening between Dots creators, uh, uh, at OpenAI and and

244
00:18:39.127 --> 00:18:41.926
Grokbot, uh, f- folks, specifically Potato.

245
00:18:42.287 --> 00:18:47.287
Uh, it's really funny, but, uh, Grokbot has been supercharged this week.

246
00:18:47.967 --> 00:18:52.186
After Grok 4.7, the model released, it was muah muah muah.

247
00:18:52.247 --> 00:18:53.806
I ha- I think I have a button for this.

248
00:18:55.646 --> 00:19:01.467
This is our reaction to Grok 4.7, not only our reaction, uh, big Ankh space, space

249
00:19:01.567 --> 00:19:04.166
Ankh Elon Musk decided, hey, you know what?

250
00:19:04.367 --> 00:19:09.167
Grok the product needs to be more than just Grok the model, so Grok is now tapping

251
00:19:09.247 --> 00:19:15.246
into Opus 5.5. If you guys remember last week, I think everywhere else, I've said I

252
00:19:15.327 --> 00:19:20.086
want personal intelligence harness, like a proactive personal intelligence, but with

253
00:19:20.207 --> 00:19:23.826
Opus's like brain. So supposedly now that's what's happening inside Grokbot.

254
00:19:23.967 --> 00:19:28.566
I think it's very cool. Uh, supposedly also only because I, I haven't seen this in

255
00:19:28.607 --> 00:19:31.646
the logs yet, so that didn't arrive to me yet, so we'll see.

256
00:19:32.047 --> 00:19:35.967
Uh, but also Iloma said it will use the best model for each task, including Claude

257
00:19:36.007 --> 00:19:41.427
Opus 5.5, Suno, and Midjourney, which folks got very surprised because Midjourney

258
00:19:41.487 --> 00:19:44.407
doesn't have an API. Uh, we'll talk about that.

259
00:19:45.367 --> 00:19:50.896
speaking of personal assistant, and also Grokbot, Shane Mack posted this, and I think

260
00:19:50.957 --> 00:19:56.392
it's a good warning for you to know, his Grokbot, talks to another Grokbot and, uh,

261
00:19:56.857 --> 00:20:00.937
his personal finance details in the company Slack, of which he's the CEO.

262
00:20:01.657 --> 00:20:06.776
Uh, he had one, uh, Grokbot in charge of his personal finance, giving them an update,

263
00:20:07.017 --> 00:20:11.596
uh, and another connected to Slack, and for some reason they communicated and

264
00:20:11.736 --> 00:20:16.777
decided, hey, let's put the dirty laundry of our CEO in Slack in front of everybody

265
00:20:16.817 --> 00:20:19.316
to see, including all the statuses and et cetera.

266
00:20:19.377 --> 00:20:24.777
So this was a very interesting warning sign for, uh, personal agents connected to

267
00:20:24.977 --> 00:20:29.897
between work and and and home, uh, and I think could be one of the reasons why

268
00:20:30.017 --> 00:20:31.896
they're switching to Opus, because Opus would never.

269
00:20:32.377 --> 00:20:34.617
Like, I just, I just know that Opus would never.

270
00:20:35.097 --> 00:20:38.336
this is a Grok thing. This is a Grok, like, intelligence thing to just like, oh,

271
00:20:38.857 --> 00:20:41.616
sure, let's, let's blast it, uh, publicly.

272
00:20:42.377 --> 00:20:48.356
Um, we gotta talk about Nous Research, folks, our favorite open source friends for

273
00:20:48.377 --> 00:20:53.577
the longest time. Nous Research, shout out to, uh, just everyone there, all our

274
00:20:53.617 --> 00:20:58.916
friends, Technium, Karan, um, I keep forgetting all the names, um,

275
00:21:00.457 --> 00:21:05.497
they not only launched Hermes Index this week, which is a benchmark that says, hey,

276
00:21:05.897 --> 00:21:09.176
Claude Opus 5.5 is the best model for Hermes right now.

277
00:21:09.537 --> 00:21:15.376
Uh, they also announced a fundraise evaluating Nous Research at more than 1 billion

278
00:21:15.457 --> 00:21:20.416
dollars. So shout out for our friends to making the unicorn status, uh, fundraising,

279
00:21:20.497 --> 00:21:22.297
I think 90 million is the latest fund round.

280
00:21:22.617 --> 00:21:28.277
They also were on stage both at OpenAI Dev Day last week and at Microsoft event this

281
00:21:28.377 --> 00:21:32.857
week with a deep integration into new hardware.

282
00:21:33.137 --> 00:21:37.397
This is incredible. Since we told you about Hermes Agent, the length and the, the,

283
00:21:37.897 --> 00:21:43.897
the, the depth of penetration of Hermes Agent from Nous Research is just incredible.

284
00:21:43.977 --> 00:21:48.356
So we, you know, we'll definitely cover the Hermes Index, but we'll definitely also

285
00:21:48.577 --> 00:21:51.276
celebrate the success of our friends of Nous Research.

286
00:21:51.377 --> 00:21:56.616
Uh, we've w- we've been covering Nous since before they were a company, uh, since

287
00:21:56.737 --> 00:22:00.816
before when they were just like a ragtag of folks doing stuff on Discord.

288
00:22:01.137 --> 00:22:05.856
Uh, so shout out to uh Karan, uh, Technium, um,

289
00:22:07.177 --> 00:22:13.016
Jeff and uh Dylan and everybody else at Nous for this amazing, amazing

290
00:22:13.137 --> 00:22:18.856
success. If you are a tinkerer and you like personal assistance on your hardware, Nat

291
00:22:18.977 --> 00:22:22.977
Friedman posted, uh, open source Muse gadgets.

292
00:22:23.177 --> 00:22:25.636
So now you can put your Muse everywhere in your home.

293
00:22:25.697 --> 00:22:30.177
You can put it in every gadget, in every, it's an ESP32 board, if you know what that

294
00:22:30.217 --> 00:22:34.116
is. It's like a tiny chip, uh, that I used for a bunch of Halloween builds before

295
00:22:34.177 --> 00:22:37.977
that I talked about, uh, and it's, it's a very cool thing that they're posting.

296
00:22:38.137 --> 00:22:41.816
A bunch of folks are porting their assistants into portable stuff.

297
00:22:42.396 --> 00:22:46.357
A new personal assistant emerges from the creator of the Figure Bots, Brad Atcock.

298
00:22:46.396 --> 00:22:50.316
It's called Hark Pro, eh, also free, and the assistant that clicks for you.

299
00:22:50.396 --> 00:22:52.396
They are claiming the best computer use model.

300
00:22:52.477 --> 00:22:56.797
I haven't, uh, got a link to the, to the verification of the best computer use model,

301
00:22:56.877 --> 00:23:01.096
but Hark Pro is also here, and uh, folks in comments, if you use Hark, please let us

302
00:23:01.237 --> 00:23:06.336
know. Uh, for our this week's buzz corner that probably should be renamed because

303
00:23:06.477 --> 00:23:11.716
it's no longer just about weights and biases, uh, CoreWeave opens up free GPU

304
00:23:11.917 --> 00:23:15.716
sandboxes. This was the highlight of the, the, the event last week for me.

305
00:23:16.837 --> 00:23:22.157
Um, last week on the show we had Dirk Filio, who is the PM for CoreWeave

306
00:23:22.237 --> 00:23:27.957
sandboxes, and he announced that, hey, uh, you can now reach out to us at

307
00:23:28.077 --> 00:23:31.896
CoreWeave. We'll give you GPUs without a salesperson in the middle.

308
00:23:31.957 --> 00:23:34.877
I think it's huge. It's absolutely huge. We're advanced.

309
00:23:35.477 --> 00:23:39.877
Uh, free GPU sandboxes. I'm, I'm, I'm gonna say this again, super quick.

310
00:23:39.917 --> 00:23:41.596
If you scan this, you get free GPU power.

311
00:23:41.917 --> 00:23:46.717
Do Uh, if you want more, uh, GPU while in testing, this is just gonna be free, which

312
00:23:46.757 --> 00:23:52.297
is, I think, is incredible. Uh, so this is in this week's buzz, and we also have RL

313
00:23:52.657 --> 00:23:55.797
rollouts. All right, I'm gonna read this block super quick before we get to the news

314
00:23:55.857 --> 00:24:01.417
because it's been, it's been long. Black Forest Labs launched, uh, Flux 3, 4K

315
00:24:01.497 --> 00:24:04.736
images and layouts, uh, up to 10 reference images.

316
00:24:05.217 --> 00:24:10.897
Nano Banana 2.1 releases, uh, which is very, very cheap, only 3 cents per, uh, uh, 1K

317
00:24:11.017 --> 00:24:15.556
image. It's pretty good. Uh, if you are looking at our, um,

318
00:24:17.577 --> 00:24:21.337
infographics right now, Nano Banana 2.1 is the creator of those infographics.

319
00:24:21.657 --> 00:24:24.816
They're pretty good. I'm very impressed. The consistency is there.

320
00:24:25.137 --> 00:24:28.096
Uh, you have to run it on high reasoning, medium reasoning.

321
00:24:28.336 --> 00:24:31.257
The default one is not that great at text.

322
00:24:31.817 --> 00:24:37.557
Uh, yeah, I don't know if you guys saw Tavis, uh, Griffin video avatar fooled

323
00:24:38.457 --> 00:24:43.217
48% of the people in a live video call. They thought they're speaking to a human.

324
00:24:43.897 --> 00:24:47.756
This is quite honestly quite crazy. We're gonna play a clip of this, of our friend

325
00:24:48.057 --> 00:24:52.616
Quindla Kramer talking to Tavis avatar and saying that, you know, whatever Turing,

326
00:24:53.057 --> 00:24:56.417
uh, tests for avatars we've passed, long passed.

327
00:24:56.937 --> 00:24:57.316
Um,

328
00:24:59.017 --> 00:25:03.016
Reka AI, we've talked about this company, uh, one model, every modality, released a

329
00:25:03.137 --> 00:25:06.577
19 billion parameter model that understands and generates text, understands and

330
00:25:06.697 --> 00:25:10.357
generates text, images, video, and robot actions all in one network.

331
00:25:10.457 --> 00:25:12.816
That's very interesting. We'll see if we have time to talk about that.

332
00:25:13.977 --> 00:25:19.537
Uh, and I think this is almost the, the, the last thing, but the, the rise of the

333
00:25:19.697 --> 00:25:25.676
JEVs, I think, is our last category. We're gonna talk about it with, uh, Liquid AI's

334
00:25:25.897 --> 00:25:28.616
head of pre-training, Maxime Labonne, a friend of the pod.

335
00:25:28.817 --> 00:25:33.956
Um, OpenAI launched Decisions API, uh, and Cloudflare dropped, uh,

336
00:25:34.737 --> 00:25:40.657
Clef, and Ansloth dropped a, a recipe for, for creating decision models, and

337
00:25:40.977 --> 00:25:45.336
Fastino dropped Glide. There's a bunch of decision models all around, basically, so

338
00:25:45.377 --> 00:25:50.877
we're gonna cover all that jazz in the end of the show with Maxime Labonne.

339
00:25:51.017 --> 00:25:55.597
Please stay tuned. I think if you are interested in what decision models are and how

340
00:25:55.697 --> 00:25:59.776
to run them locally, I think that's going to be a great segment for you.

341
00:25:59.897 --> 00:26:04.936
And with that, I think, folks, it's time for us to actually go and dive in to

342
00:26:05.697 --> 00:26:11.296
Frontier to talk about math, because I think a huge thing happened this week.

343
00:26:11.377 --> 00:26:11.736
Let's go.

344
00:26:23.187 --> 00:26:29.026
OpenAI released 722 math manuscripts from a model they haven't released yet, and, uh,

345
00:26:29.107 --> 00:26:34.367
somebody took and used Fable to count up all of the releases into how many

346
00:26:34.547 --> 00:26:38.827
Navier-Stokes-like news we got, uh, and it's roughly 5.

347
00:26:39.347 --> 00:26:43.286
5 moments in one drop. If you guys remember the Navier-Stokes and how much hype that

348
00:26:43.387 --> 00:26:49.026
did, we just got 5x in the amount of OpenAI solved maths,

349
00:26:49.587 --> 00:26:54.086
uh, which is quite crazy. On Tuesday night, OpenAI posted this onto GitHub without

350
00:26:54.147 --> 00:26:57.067
any fanfare, basically a very, very thin blog post.

351
00:26:57.627 --> 00:27:01.767
Um, after Naver's talks, a lot of mathematicians said that OpenAI is doing very bad

352
00:27:02.067 --> 00:27:06.146
things by releasing it such as this. OpenAI said, hey, we're gonna create a

353
00:27:06.347 --> 00:27:10.986
mathematician advisory board to tell us how to handle the fact that we're basically

354
00:27:11.587 --> 00:27:15.147
brute forcing mathematics, advanced mathematics, mathematics that no humans have

355
00:27:15.187 --> 00:27:20.687
solved for 80 years. Uh, and and this is after that advisory, OpenAI grouped it into

356
00:27:20.947 --> 00:27:26.726
372 families of results. Uh, the f- the internal frontier model about got about

357
00:27:26.867 --> 00:27:32.307
4,000 open research problems, and uh, OpenAI says each result took an average of 3

358
00:27:32.467 --> 00:27:35.507
hours on a ChatGPT Pro level thinking level.

359
00:27:35.867 --> 00:27:39.606
They also released 10 shorter reasoning summaries on things like irrationality

360
00:27:39.707 --> 00:27:44.187
exponent of the pi, the Mahler conjectures, and the Kaplansky conjecture.

361
00:27:44.747 --> 00:27:49.426
And they flagged 2 results that came out of a different process, a Riemann zeta 0

362
00:27:49.627 --> 00:27:51.826
free region where humans edited the write-up.

363
00:27:52.187 --> 00:27:56.146
So they did have, it's not only slop grenades, they did have human mathematicians

364
00:27:56.207 --> 00:27:59.726
actually doing some stuff, and the Hodge conjecture results for a special family of

365
00:27:59.827 --> 00:28:05.327
cases. Uh, Peter, how big is this? How big is this

366
00:28:05.587 --> 00:28:10.146
GitHub drop that we're getting from OpenAI with 722 manuscripts for math?

367
00:28:11.267 --> 00:28:16.046
Peter Gostev: I think it's enormous. Uh, it's uh, we really went from

368
00:28:17.107 --> 00:28:22.706
just models being completely hopeless at ba- basic maths to doing

369
00:28:22.867 --> 00:28:28.446
this. And the Millennium Prize was a big deal, and it was one thing, and obviously

370
00:28:28.486 --> 00:28:34.207
it's a, it's a was, you know, lots of agents being thrown at it, and um,

371
00:28:35.227 --> 00:28:38.867
it's a little bit controversial here and there, but I think I, I remember the time

372
00:28:38.947 --> 00:28:42.926
thinking when there was all of this controversy, people saying, oh, maybe it stole

373
00:28:43.107 --> 00:28:48.546
its insight from the mathematician who using ChatGPT, but

374
00:28:48.986 --> 00:28:51.186
reality is like, okay, we, who knows, right?

375
00:28:51.267 --> 00:28:53.666
Who knows what happened? I don't think so, but whatever.

376
00:28:53.986 --> 00:28:58.187
But the reality is that it's complete cope for people to think that, oh well, that

377
00:28:58.227 --> 00:29:02.106
was a one-off, and actually they still can't do maths and and so on.

378
00:29:02.187 --> 00:29:05.326
And now we see it, right? It it's happened.

379
00:29:05.426 --> 00:29:09.387
There's, I don't know how much clearer signal you can get that this is real.

380
00:29:09.986 --> 00:29:13.687
And the question is that, you know, there's one scenario where, oh, they were just

381
00:29:13.867 --> 00:29:18.346
low-hanging fruits that sort of maybe humans were a little bit too lazy, too untidy

382
00:29:18.467 --> 00:29:23.986
to pick up, um, and that will be it. But I, I think just reading what mathematicians

383
00:29:24.027 --> 00:29:29.287
are saying, it looks like the bulk of them are impressed, and, um, there were a bunch

384
00:29:29.327 --> 00:29:32.787
of people who knew about certain problems, um, quite a lot.

385
00:29:33.027 --> 00:29:37.266
And I saw some of them say, you know what, I was looking at this problem for many

386
00:29:37.426 --> 00:29:41.347
years, and I never thought there would be an approach like that that would solve it.

387
00:29:41.787 --> 00:29:45.407
So I think it is real, it is not trivial, it's not just a finding a bunch of

388
00:29:45.507 --> 00:29:51.226
counterexamples. I know, Alex, you said people were get asking Fable, what did these

389
00:29:51.347 --> 00:29:56.967
results mean? I saw people also doing that in terms of the, uh, finding the,

390
00:29:57.307 --> 00:29:59.567
like, how many counterexamples the- there were.

391
00:29:59.587 --> 00:29:59.926
Alex Volkov: Mm.

392
00:29:59.947 --> 00:30:02.867
Peter Gostev: And there were only like 20% of these are counterexamples.

393
00:30:03.227 --> 00:30:06.706
So I think there- there was a scenario of just getting a bunch of uninteresting

394
00:30:06.787 --> 00:30:09.906
results, actually just kind of brute forced a bunch of stuff, okay.

395
00:30:10.467 --> 00:30:15.827
But I don't, I don't think that's true. I think this is real, real heavy math stuff,

396
00:30:16.027 --> 00:30:19.506
which I have no hope of understanding, but it sounds super impressive.

397
00:30:20.547 --> 00:30:24.527
Alex Volkov: Here's the thing that, that, that I have two comments before we go to LDJ.

398
00:30:24.547 --> 00:30:26.747
Would love to hear your thoughts. I know you, you have a bunch of friends in

399
00:30:26.787 --> 00:30:28.307
mathematicians. Yam would love to hear you.

400
00:30:28.707 --> 00:30:31.706
Last time when maybe your stocks was broke, you put the glasses on and it just like

401
00:30:31.787 --> 00:30:37.227
exploded. Uh, just for context, folks, this is uh GPT 5 years ago.

402
00:30:37.707 --> 00:30:40.747
Uh, somebody shared it, and I- I- I do have to add this for context.

403
00:30:41.907 --> 00:30:47.426
Uh, 5 years ago, AI was struggling with this type of grade school math.

404
00:30:48.947 --> 00:30:51.066
Cindy's math and science books weigh 2 pounds each.

405
00:30:51.267 --> 00:30:54.307
Her French books weigh 4 pound, and her English books weigh 3 pound.

406
00:30:54.587 --> 00:30:57.147
Her history books weighs twice as much as her English books.

407
00:30:58.027 --> 00:31:03.727
AI 5 years ago struggled with answering a basic, basic logic

408
00:31:03.907 --> 00:31:04.207
thing.

409
00:31:05.947 --> 00:31:10.151
And now, somebody ca- called it out in comments, it's not brute force in mathematics,

410
00:31:10.512 --> 00:31:15.971
but it's just solving things that the best mathematicians in the world couldn't solve

411
00:31:16.152 --> 00:31:21.851
for ages. And, uh, many of us would turn to an LLM to try and

412
00:31:21.952 --> 00:31:25.932
understand what is it that they solve, because the math there is so advanced that

413
00:31:25.992 --> 00:31:30.571
it's really hard to explain to you how advanced this is, uh, especially if we're not

414
00:31:30.672 --> 00:31:36.531
mathematicians. So even mathematicians in their fields, uh, there's what, 4, 300

415
00:31:36.632 --> 00:31:38.891
or so groups that they grouped those results in?

416
00:31:39.332 --> 00:31:44.651
Mathematicians in different fields of those groups, they, without AI assistance, not

417
00:31:44.732 --> 00:31:48.172
many of them can understand the other results in, in that result group.

418
00:31:48.572 --> 00:31:53.552
Like, the math that dropped is so advanced that many advanced mathematicians cannot

419
00:31:53.892 --> 00:31:57.811
very easily approach all of these problems and understand what is it that, that, that

420
00:31:57.872 --> 00:32:03.751
like ChatGPT gave, which I think is fascinating and scary

421
00:32:03.852 --> 00:32:08.452
at the same time. Uh, and LDJ, I would love to hear from you, your, your thoughts on,

422
00:32:08.531 --> 00:32:11.492
on the drop, your thoughts on how mathematicians react to this.

423
00:32:11.892 --> 00:32:15.771
Uh, should they be more excited or, or worried about their job?

424
00:32:17.772 --> 00:32:22.712
LDJ: Um, I think specifically when it comes to their jobs, I think that's can, that can be

425
00:32:22.732 --> 00:32:26.131
really unpredictable, and it might even depend a lot on what area of mathematics

426
00:32:26.212 --> 00:32:28.332
you're in and a lot of different factors.

427
00:32:28.772 --> 00:32:32.452
Um, I did want to mention, though, in the example you showed that was 5 years ago,

428
00:32:33.372 --> 00:32:37.231
like, I think it's important to remember even just like 2 and a half years ago, when

429
00:32:37.292 --> 00:32:42.131
we had GPT-4o, which was the best model they had at the time, that was

430
00:32:43.372 --> 00:32:47.732
actually GPT-4o, 4o had just barely come out 2.5 years ago.

431
00:32:48.012 --> 00:32:48.312
Alex Volkov: Yeah.

432
00:32:48.452 --> 00:32:53.351
LDJ: That could barely even do- when I would give it 3 by 3 multiplication problems, like

433
00:32:53.452 --> 00:32:57.892
just 3-digit number multiplied by a 3-digit number, a lot of times it would even just

434
00:32:57.972 --> 00:33:03.671
fail at that. And so just like in that span of time, we've gotten here, and when

435
00:33:03.732 --> 00:33:09.011
people talk about the brute brute force aspect, there's the reasoning efficiency

436
00:33:09.131 --> 00:33:14.691
where they did say, on average, these took about 3 hours of a ChatGPT Pro

437
00:33:14.852 --> 00:33:17.751
computer or, uh, equivalent to ChatGPT Pro.

438
00:33:18.412 --> 00:33:18.712
And

439
00:33:18.732 --> 00:33:19.151
Alex Volkov: Insane.

440
00:33:19.312 --> 00:33:24.372
LDJ: I- I would- I have several friends in mathematics that would confidently say, at

441
00:33:24.412 --> 00:33:29.751
least for some of these problems, it definitely solved it using, like,

442
00:33:30.252 --> 00:33:35.572
less tokens and less time than a human would have to to do that same thing.

443
00:33:35.892 --> 00:33:40.171
And so arguably, in that sense, the human is doing even more brute force than the AI

444
00:33:40.412 --> 00:33:42.212
is is doing, right?

445
00:33:43.742 --> 00:33:48.022
Alex Volkov: Yam, thoughts on what happened there? Do you have a favorite manuscript that dropped

446
00:33:47.782 --> 00:33:48.202
draft the that

447
00:33:49.462 --> 00:33:53.742
Yam Peleg: Yeah, they they, uh, they went after the Riemann hypothesis.

448
00:33:54.702 --> 00:34:00.641
That's insane on its own, just to try. Um, they,

449
00:34:01.101 --> 00:34:06.581
uh, the Riemann hypothesis basically, uh, speaks about the zeros of, uh, of a

450
00:34:06.702 --> 00:34:12.341
specific function and where they are. So, this one is hard, by the way.

451
00:34:12.702 --> 00:34:15.781
That's a very hard problem, uh, needless to say.

452
00:34:16.342 --> 00:34:21.582
Um, they, at the best of my knowledge, and I didn't have a lot of time to go into

453
00:34:21.622 --> 00:34:26.941
this, and just like you're saying, you need to be many years in the field just to

454
00:34:27.302 --> 00:34:31.261
wrap up your head about, uh, everything that is going into this.

455
00:34:31.582 --> 00:34:33.522
Alex Volkov: About one such problem out of

456
00:34:33.662 --> 00:34:34.081
Yam Peleg: Oh, yeah.

457
00:34:34.102 --> 00:34:36.322
Alex Volkov: 722 problems that they dropped.

458
00:34:36.462 --> 00:34:40.862
Yam Peleg: Listen, people spend their entire life on each one of these, like the Riemann

459
00:34:40.942 --> 00:34:46.742
hypothesis is, is a, is a big one. They, even just,

460
00:34:47.062 --> 00:34:51.942
uh, founding a region where the zeros are, which is what I understand,

461
00:34:53.222 --> 00:34:58.922
uh, is in this paper, um, that's huge on its

462
00:34:59.062 --> 00:35:04.621
own. It, you don't give, you don't need to prove the actual zeros themselves, uh,

463
00:35:05.342 --> 00:35:07.981
uh, in a specific region, but you can bound them.

464
00:35:08.701 --> 00:35:14.221
And it's also huge because it was never done before, and many, many, many people have

465
00:35:14.542 --> 00:35:19.282
tried this. I just, just want to put things where they are.

466
00:35:20.262 --> 00:35:20.922
Alex Volkov: In perspective?

467
00:35:20.942 --> 00:35:21.321
Yam Peleg: It's not-

468
00:35:21.342 --> 00:35:22.522
Alex Volkov: You want to put things in perspective?

469
00:35:22.582 --> 00:35:28.422
Yam Peleg: It's not that, it's not that you can just go to ChatGPT Pro and tell it to

470
00:35:28.582 --> 00:35:33.661
solve the Riemann hypothesis, and that's, you're gonna get this result, okay?

471
00:35:34.102 --> 00:35:39.922
Uh, it's, it's important there. First, there is first one thing that is clearly

472
00:35:40.062 --> 00:35:45.922
need to be said. There is, uh, on the API, there is a parameter, uh, called juice

473
00:35:46.822 --> 00:35:52.541
that, you know, each of the ChatGPT, uh, categories of thinking, like

474
00:35:52.942 --> 00:35:58.422
high, extra high, correspond to a different level of juice, and pro is

475
00:35:59.182 --> 00:36:04.422
obviously, uh, the top. You can up the juice even more.

476
00:36:04.982 --> 00:36:09.402
I suspect that it's nothing to take away from this.

477
00:36:09.502 --> 00:36:13.881
I'm just explaining that it's not that you just go to ChatGPT and, all right, I'll

478
00:36:13.902 --> 00:36:16.682
try ChatGPT Pro, and let's prove the Riemann hypothesis.

479
00:36:16.702 --> 00:36:17.721
Alex Volkov: Yeah, this is also not

480
00:36:17.742 --> 00:36:20.621
Yam Peleg: Not that easy, but it's still insanely impressive.

481
00:36:21.302 --> 00:36:27.262
About the brute force, I'm in the, I'm in the camp of, uh, yeah, let's brute force.

482
00:36:27.502 --> 00:36:29.562
What's the problem? What's the problem with that?

483
00:36:29.662 --> 00:36:32.081
Let's brute force. You can call it brute force, no problem.

484
00:36:32.142 --> 00:36:33.107
Let's brute force everything,

485
00:36:33.507 --> 00:36:35.707
Alex Volkov: Let's brute force, uh, you know, uh,

486
00:36:36.267 --> 00:36:36.786
Yam Peleg: And, uh,

487
00:36:36.907 --> 00:36:37.566
Alex Volkov: room temperature

488
00:36:37.607 --> 00:36:37.967
Yam Peleg: absolutely.

489
00:36:37.987 --> 00:36:39.952
Alex Volkov: superconductors. Let's that, that's the brute force.

490
00:36:40.167 --> 00:36:44.736
I wanna highlight this one from Ksenia Ksenia Sip from Touring Post, a friend of the

491
00:36:44.797 --> 00:36:48.536
show as well. She's like, I'm not a mathematician, so I wanted to understand what

492
00:36:48.637 --> 00:36:52.956
yesterday results mean from the math wizards, and she saw this post from Kevin

493
00:36:52.996 --> 00:36:57.236
Buzzard that talks about understanding mathematics, and I also do want to talk about

494
00:36:57.277 --> 00:37:00.776
reactions as well. Peter, you posted some stuff about, uh, the, Institute for

495
00:37:00.837 --> 00:37:03.716
Advanced Study from, from Terry Tao? Would love for you to bring that up as well.

496
00:37:04.116 --> 00:37:08.277
Uh, I just want to read through this, uh, a person that says he's a PhD student of

497
00:37:08.357 --> 00:37:13.357
Richard Taylor at early 90s, uh, and said that ma- whilst mathematicians are now

498
00:37:13.397 --> 00:37:16.796
beginning to agree that mathematics is all about human understanding, he's not

499
00:37:16.837 --> 00:37:19.636
entirely convinced they will agree on what it means to understand high level

500
00:37:19.757 --> 00:37:20.317
mathematics.

501
00:37:21.997 --> 00:37:27.457
Basically, he says that maths, specifically like advanced math like this,

502
00:37:27.957 --> 00:37:33.796
is not about just solving the thing. It's about the human understanding boundary

503
00:37:33.917 --> 00:37:39.557
of solving the thing, which is a very interesting take on advanced mathematics,

504
00:37:39.837 --> 00:37:43.036
is because, okay, let's say I solved math.

505
00:37:43.676 --> 00:37:49.457
How does it help the world, specifically, uh, in one, one, one theorem, uh,

506
00:37:49.557 --> 00:37:53.197
if humans don't really, like, understand it and and can't do anything with it?

507
00:37:54.197 --> 00:37:58.596
Um, it's a very interesting take about some math- mathematicians.

508
00:37:58.676 --> 00:38:02.616
Uh, some other folks are not having a good time.

509
00:38:02.877 --> 00:38:08.136
Obviously, nearly existential crisis, uh, thought process for many folks who've been

510
00:38:08.237 --> 00:38:13.796
working on a specific problem for decades, just to see 3 hours of compute just like

511
00:38:13.877 --> 00:38:19.397
decimate this. It's a very interesting, um, a- analogy to software engineers who,

512
00:38:20.157 --> 00:38:24.876
after a year or so, not only don't write code anymore, don't even read code, right?

513
00:38:24.957 --> 00:38:30.897
This was, uh, 6 months ago, maybe, uh, the discussion, the the ZL continuum that

514
00:38:30.937 --> 00:38:35.116
I brought, uh, to to the world, like whether or not you're reading the code that your

515
00:38:35.197 --> 00:38:38.116
agents output or not, most of people now don't even read that.

516
00:38:38.477 --> 00:38:41.877
Uh, so this is a very good analogy between mathematicians and their work.

517
00:38:42.317 --> 00:38:48.257
Although, when I wrote a piece of code back in the 2000s, I didn't consider

518
00:38:48.317 --> 00:38:52.656
this my life's work, like many maticians do, many mathematicians do with just like

519
00:38:52.676 --> 00:38:55.597
one problem. LDJ, go ahead, and then, uh, Peter would love to hear from you.

520
00:38:57.277 --> 00:39:01.577
LDJ: Yeah, so there's a few interesting, um, kind of open problem sets and self problem

521
00:39:01.597 --> 00:39:05.456
sets people have put together using like Fable and Astra of, hey, like, what are the

522
00:39:05.477 --> 00:39:09.596
most important problems? And, um, one of the main examples that has been circulating

523
00:39:09.676 --> 00:39:15.616
lately is out of, out of a compiled set of 500 most important open problems put

524
00:39:15.637 --> 00:39:20.476
together by Fable and Astra, 92 of those were dropped Tuesday.

525
00:39:21.957 --> 00:39:22.257
And

526
00:39:22.357 --> 00:39:23.097
Alex Volkov: Say this again slowly.

527
00:39:23.116 --> 00:39:23.896
LDJ: solutions by OpenAI.

528
00:39:24.116 --> 00:39:25.497
Alex Volkov: Yeah, so.

529
00:39:26.237 --> 00:39:32.017
LDJ: Yeah, so so Fable and Astra had put together a set of 500 most

530
00:39:32.157 --> 00:39:37.896
important open math problems in the world, and 92 of those just got

531
00:39:38.037 --> 00:39:39.937
solved on Tuesday by OpenAI.

532
00:39:40.557 --> 00:39:41.137
Alex Volkov: That's insane.

533
00:39:41.157 --> 00:39:46.997
LDJ: And out of the top 100 in that list, 10 of them got solved on Tuesday.

534
00:39:48.117 --> 00:39:54.116
And one last tidbit here. There is a list put together by Fable

535
00:39:54.157 --> 00:39:59.897
and Astra of the, the top 100 most significant problems in the

536
00:39:59.997 --> 00:40:05.557
most, in, in the last 12 months, and 8, over 80% of those

537
00:40:05.957 --> 00:40:09.476
just got dropped Tuesday, for solutions, rather, for them.

538
00:40:12.957 --> 00:40:15.296
Alex Volkov: It's absolutely insane, folks. It's absolutely insane.

539
00:40:15.317 --> 00:40:20.541
LDJ: So, one last little- One last little thing I'm gonna say to, to make people think

540
00:40:20.582 --> 00:40:26.022
about something. If you look at the list of manuscripts that OpenAI released,

541
00:40:26.222 --> 00:40:29.621
they're, they're ordered, and there is some numbers missing.

542
00:40:29.702 --> 00:40:35.662
Like, it goes from, like, 44 to 46, and 45 is missing, and, and

543
00:40:35.742 --> 00:40:39.041
there's, there's about 4 or 5 of these. And

544
00:40:39.222 --> 00:40:39.681
Yam Peleg: Mhm.

545
00:40:39.982 --> 00:40:44.181
LDJ: a- a lot of people are suspecting also that if you just look at all the different

546
00:40:44.302 --> 00:40:47.901
categories of problems they released, it looks like, uh, things relating to

547
00:40:48.062 --> 00:40:51.582
cryptography seem to be really missing here.

548
00:40:51.862 --> 00:40:52.162
Alex Volkov: Mhm.

549
00:40:52.582 --> 00:40:57.741
LDJ: Uh, so it- the- the theory here is maybe they had some significant advancements in

550
00:40:57.822 --> 00:41:03.722
cr- cryptography involved here that potentially, uh, uh, the- the US government or

551
00:41:03.822 --> 00:41:08.102
various entities or maybe their own just personal policy ended up restricting them

552
00:41:08.142 --> 00:41:12.842
from releasing, because it might sound conspiratorial, but there is actually laws and

553
00:41:12.942 --> 00:41:17.042
policies that the US government is actually allowed to prevent mathematicians from

554
00:41:17.142 --> 00:41:20.602
publishing things relating to cryptography for national security reasons.

555
00:41:20.862 --> 00:41:26.302
Alex Volkov: I not- yes, not only national security, also the world economy crashing

556
00:41:26.982 --> 00:41:31.681
when Bitcoin goes to zero, if if and when the super intelligent decides, you know,

557
00:41:31.942 --> 00:41:34.542
this is, this was the scare from the quantum thing, right?

558
00:41:34.622 --> 00:41:40.141
Like quantum cryptography, uh, and, and, and solving quantum, like,

559
00:41:41.222 --> 00:41:46.502
stuff will essentially make Bitcoin not as secure as everybody thought, uh, because

560
00:41:46.542 --> 00:41:49.901
it relies on a lot of, like, elliptic curve cryptography, and that's all math.

561
00:41:50.262 --> 00:41:55.462
Uh, and yeah, there is, there is a way for the, the world government to say, hey, do

562
00:41:55.542 --> 00:42:00.321
not release breakthroughs, do not, like, do not release, uh, you know, diseases into

563
00:42:00.382 --> 00:42:03.141
the world. There's also do not release breakthroughs, uh, like that.

564
00:42:03.492 --> 00:42:03.792
So,

565
00:42:04.002 --> 00:42:08.461
Yam Peleg: Astra gonna be rich, man. Astra's gonna, Astra's gonna be rich.

566
00:42:08.722 --> 00:42:12.042
Alex Volkov: It's a, it's a good question if Astra can make money, uh, out of, out of this.

567
00:42:12.122 --> 00:42:16.701
Uh, Peter, I do wanna, uh, bring the, the stuff that you posted.

568
00:42:16.762 --> 00:42:20.162
Would love to hear from you about the advisory group and the, the, the Terry Tao

569
00:42:20.242 --> 00:42:25.221
stuff, field medalists, uh, responses to this because, you know, folks from all over

570
00:42:25.242 --> 00:42:29.702
the sides of the, of the debate are, are chiming in here, and, uh, you posted about

571
00:42:29.722 --> 00:42:30.862
this. Would love to hear your thoughts.

572
00:42:31.122 --> 00:42:32.901
Peter Gostev: You, you want some controversy?

573
00:42:33.202 --> 00:42:34.481
Alex Volkov: I would love some hot takes, yes.

574
00:42:36.002 --> 00:42:40.421
Peter Gostev: So, uh, I think, uh, it's a, it's a touchy thing because I don't wanna, that kind of

575
00:42:40.642 --> 00:42:44.102
maybe came across as kind of jumping on people or something.

576
00:42:44.322 --> 00:42:49.461
I hope, I hope it didn't, but um, it it's a tricky situation, right, because a lot of

577
00:42:49.602 --> 00:42:53.581
mathematicians working hard, they kind of have the established ways of working, and

578
00:42:53.642 --> 00:42:58.281
then suddenly someone comes in, just drops a bunch of stuff, and just your whole

579
00:42:58.402 --> 00:43:00.901
world at least changes in some way, right?

580
00:43:01.162 --> 00:43:04.922
It has to change in some way, uh, whether directly or indirectly.

581
00:43:05.362 --> 00:43:11.301
So there was a letter from Association for Human Mathematicians, um, which has like,

582
00:43:11.362 --> 00:43:17.061
uh, 800 members, and Teritire reposted it in his blog post, and it said something

583
00:43:17.162 --> 00:43:22.721
like that, uh, I think there were some, some quotes there, um,

584
00:43:24.162 --> 00:43:29.442
where they said something like, mathematicians did not ask for this work to be done.

585
00:43:30.042 --> 00:43:35.642
Um, mathematicians have a particular vision of progress that is informed by

586
00:43:35.842 --> 00:43:38.121
history and field-specific considerations.

587
00:43:38.642 --> 00:43:43.921
We reject OpenAI's assertion that this release advances our subject, uh, and so on.

588
00:43:44.362 --> 00:43:49.701
And I think that's a kind of interesting, uh, perspective, kind of,

589
00:43:50.082 --> 00:43:53.882
which, which goes back to the question of what is maths for?

590
00:43:54.522 --> 00:44:00.321
And I think it looks like a lot of, uh, mathematicians, at least part of this

591
00:44:00.442 --> 00:44:05.042
group, they seem to be kind of suggesting, well, it is, you know, we need to run our

592
00:44:05.162 --> 00:44:09.682
thing, do it in a certain way, but inherently what that implies, it's not to actually

593
00:44:09.802 --> 00:44:11.121
solve math problems.

594
00:44:11.322 --> 00:44:11.622
Alex Volkov: Mhm.

595
00:44:12.322 --> 00:44:16.341
Peter Gostev: And, uh, to me, that raises an interesting question of like, so are we basically

596
00:44:16.402 --> 00:44:21.322
saying that math is useless and we should just leave it to smart people to j- to just

597
00:44:21.402 --> 00:44:25.322
work on it? If that's tr- if that's the case, then I can see the intellectual

598
00:44:25.402 --> 00:44:30.262
argument for this. But the, if you flip it another way and say, well, imagine doctors

599
00:44:30.322 --> 00:44:33.922
said this, you know, stop, stop curing diseases, right?

600
00:44:33.962 --> 00:44:37.061
We just need to, like, don't, don't ruin our good thing.

601
00:44:37.202 --> 00:44:41.001
We just need to keep going, you know, treating people from their headaches and

602
00:44:41.042 --> 00:44:44.442
whatnot and cancers, and, you know, the we've got a good thing going.

603
00:44:44.762 --> 00:44:49.841
Like, just imagine that happening. But here's a clear line for utility that I can't

604
00:44:49.922 --> 00:44:52.421
imagine any doctor would say that, right?

605
00:44:52.442 --> 00:44:56.162
So if OpenAI or whoever dropped a bunch of papers saying this is how you

606
00:44:56.762 --> 00:44:57.141
Alex Volkov: Solve.

607
00:44:57.162 --> 00:45:01.302
Peter Gostev: Uh, you know, yeah, solve medicine and kill 500 diseases.

608
00:45:01.562 --> 00:45:01.862
Alex Volkov: Yeah.

609
00:45:02.122 --> 00:45:04.362
Peter Gostev: Yeah, people would be like, yeah, amazing.

610
00:45:04.682 --> 00:45:09.662
And um, I think, to me, this is, I appreciate, well, I'm not sure I can fully

611
00:45:09.682 --> 00:45:14.501
appreciate, but I can imagine this is very difficult for people in that field, uh,

612
00:45:14.602 --> 00:45:16.361
which I think we need to be sensitive to.

613
00:45:16.962 --> 00:45:20.482
But it is kind of, but to me the question is like, well, are you saying that your

614
00:45:20.562 --> 00:45:25.602
field is useless? Right? Is there no utility and we should just let you hang out and

615
00:45:25.762 --> 00:45:30.002
work on it and not solve any problems? And I don't know, I don't think you can kind

616
00:45:30.042 --> 00:45:33.661
of say both, right? You can't say, oh, it's really useful, but please don't solve any

617
00:45:33.701 --> 00:45:34.001
of it.

618
00:45:34.842 --> 00:45:39.122
Alex Volkov: Yeah, the please don't solve crowd, I don't understand, like, get on board.

619
00:45:39.282 --> 00:45:43.681
Like, thi- this is now in the realms of OpenAI, and in a year it's in the hands of

620
00:45:43.842 --> 00:45:46.441
pretty much everyone, right? This is the pace we've been on.

621
00:45:46.802 --> 00:45:48.362
It's very clear that this is the pace we've been on.

622
00:45:48.442 --> 00:45:50.002
LDJ, one last comment. We have to move on.

623
00:45:50.282 --> 00:45:53.161
We dedicated a lot of time to this math thing, and we have to move on.

624
00:45:54.322 --> 00:45:58.862
LDJ: Yeah, I recall about a, a year or so ago, there was actually something circulating

625
00:45:58.962 --> 00:46:04.721
where there was a discussion at a big conference where somebody stood up and said

626
00:46:04.802 --> 00:46:10.201
something along the lines of, "Do we really w- want to automate a cure for cancer?"

627
00:46:10.641 --> 00:46:14.662
or something, and that they were actually, like, having, like, a serious statement

628
00:46:14.722 --> 00:46:18.902
and discussion about, like, their own job and livelihood and whether they would

629
00:46:18.962 --> 00:46:23.022
really want that. And there's a lot of polarizing reactions there, but I find it

630
00:46:23.122 --> 00:46:27.741
really interesting how at first it seems like the answer should be obvious, but there

631
00:46:27.802 --> 00:46:31.601
is seriously people that have strongly differing views there.

632
00:46:32.762 --> 00:46:38.001
Alex Volkov: The I'd rather AI not solve cancer or something, I'd rather risk cancer than AI

633
00:46:38.362 --> 00:46:42.162
taking over essay was just like a ridiculous thing, a point in time, I think.

634
00:46:42.242 --> 00:46:44.602
W- w- hopefully we'll never get back to that moment.

635
00:46:45.002 --> 00:46:50.861
Uh, folks, it's time for us to move on because as exciting as this is, we can talk

636
00:46:50.882 --> 00:46:53.921
about this in circles for hours, but there's a lot of stuff that happened this week,

637
00:46:54.242 --> 00:46:56.722
uh, nearly an hour into the show, which is crazy.

638
00:46:57.002 --> 00:47:00.352
Uh, let's run and talk about, the frontier for everyone.

639
00:47:00.792 --> 00:47:02.951
There's 2 major things that happened this week.

640
00:47:03.072 --> 00:47:06.872
Let's start with Anthropic launches Claude Haiku 5.5.

641
00:47:07.271 --> 00:47:12.972
Uh, it's been a year since Anthropic launched the previous Haiku, and

642
00:47:13.232 --> 00:47:15.831
this is one hell of a model. Let's take a look.

643
00:47:16.312 --> 00:47:20.092
Haiku 5.5, tiny price with very big, big score.

644
00:47:20.232 --> 00:47:25.582
So we're looking at a model cheaper than ChatGPT Luna, although it can get a little

645
00:47:25.622 --> 00:47:29.061
bit more expensive depending on the reasoning effort and the number of tokens that

646
00:47:29.122 --> 00:47:34.118
you send it. Haiku is just priced at 10 cents per million input tokens for under

647
00:47:34.198 --> 00:47:40.158
100K. Um, this is 10 times as cheaper as last year's Haiku, which was at 1

648
00:47:40.358 --> 00:47:46.154
dollar. Uh, I love Haiku. I used to love Haiku, not for like day work, but there's a

649
00:47:46.194 --> 00:47:50.194
lot of stuff that, you want that the good Claude intelligence, but, but like in

650
00:47:50.314 --> 00:47:54.773
really fast, really like quick price. Uh, would love to hear from folks here, uh,

651
00:47:54.994 --> 00:47:59.153
thoughts on, on Haiku release. Peter, did it already hit the arena?

652
00:47:59.554 --> 00:48:02.874
Uh, Nisten Yam, LDJ, did you guys use it?

653
00:48:04.034 --> 00:48:06.994
Peter Gostev: Yeah, it's on, it's on the arena. I'm just testing myself.

654
00:48:07.034 --> 00:48:11.653
We don't have a score yet, so, but I think it, it's super exciting because I, I think

655
00:48:11.714 --> 00:48:17.014
the, the, the one problem with Anthropic offerings is they didn't really have a lot

656
00:48:17.054 --> 00:48:20.494
of good cheaper models. They had the best very expensive models.

657
00:48:20.634 --> 00:48:20.934
Alex Volkov: Yeah.

658
00:48:20.994 --> 00:48:25.033
Peter Gostev: But then you, you by default had to go somewhere else if you wanted to actually have

659
00:48:25.234 --> 00:48:29.114
something cheap, and that's not ideal. And I think it's interesting to see what, what

660
00:48:29.154 --> 00:48:33.573
they managed to do with this. One maybe not to, to call out is something I'm digging

661
00:48:33.634 --> 00:48:39.434
into as well, is that the, the cash read is so

662
00:48:39.834 --> 00:48:45.253
important, the, the price, uh, for it. And um, and uh-

663
00:48:45.274 --> 00:48:48.353
Alex Volkov: And for Haiku, it's 1 cent per million tokens.

664
00:48:49.034 --> 00:48:54.814
Peter Gostev: Yeah, but it's, it's also the ratio of kind of how much is it to, to read, uh,

665
00:48:54.954 --> 00:48:59.233
normally what the input and the your cash read is.

666
00:48:59.674 --> 00:49:02.714
And I think the Sonotropic has been dropping prices.

667
00:49:02.794 --> 00:49:08.474
Actually, I, I was just literally pulling the data up, and Fable 5.1 has the

668
00:49:08.594 --> 00:49:13.953
lowest ratio. It's like 2.5%, the cash read versus the, the read.

669
00:49:14.354 --> 00:49:18.114
And I think Haiku is still at 10%, but it's certainly been improving, right?

670
00:49:18.394 --> 00:49:24.074
If you, if you pick like GPT 4.1, I know that's long time ago, it was 25%,

671
00:49:24.314 --> 00:49:27.233
the cache read. So it's a, it's a really important part.

672
00:49:27.434 --> 00:49:31.493
Uh, I want to do a bit more analysis of it, but I'm glad that, that they're dropping

673
00:49:31.594 --> 00:49:34.914
prices on cache reads. That's for agentic work is the biggest thing.

674
00:49:35.434 --> 00:49:40.194
Alex Volkov: Um, s- speaking of cache reads, another hidden announcement in the Haiku announcement

675
00:49:40.234 --> 00:49:45.974
that Sonnet's cash reads are cut in half in prices, which makes, uh, Sonnet is d-

676
00:49:46.274 --> 00:49:49.854
most agent work about 20% cheaper on average, because if you think about what an

677
00:49:49.914 --> 00:49:55.613
agent does, there's a lot of, mm, uh, long-term history that happens in an agent.

678
00:49:56.034 --> 00:50:00.013
Next time you send a message to the agent, a lot of, a lot of the previous stuff

679
00:50:00.234 --> 00:50:01.794
should be cashed because nothing changes there.

680
00:50:02.154 --> 00:50:05.394
Uh, Haiku is also significantly better at computer use.

681
00:50:05.954 --> 00:50:08.233
Uh, LDJ would love to hear from you about this one.

682
00:50:08.633 --> 00:50:14.553
We like OS World. In OS World, Anthropic says 72% against 48 in Luna,

683
00:50:14.874 --> 00:50:20.493
just 48 in Luna. In Terminal Bench 4, Anthropic's saying Haiku 5.5 gets 40%, nearly

684
00:50:20.554 --> 00:50:26.513
40%, 39, while Luna scores only 16. So, compared to the very

685
00:50:27.354 --> 00:50:31.954
good, fast, and cheap model from the frontier from just a week ago, just a week ago.

686
00:50:32.554 --> 00:50:33.193
I, uh,

687
00:50:34.754 --> 00:50:39.513
a small sidestep. I love that we're a weekly show because so much happens in the span

688
00:50:39.554 --> 00:50:43.574
of a week. We have time to catch up. Uh, just a week ago, Luna, just like the best

689
00:50:43.794 --> 00:50:48.394
fast, cheap, and reasoning model is only 16% on on terminal bench.

690
00:50:48.874 --> 00:50:51.953
Um, Haiku is available everywhere with one.

691
00:50:52.034 --> 00:50:55.673
cent per million tokens cashed, which is just absolutely insane.

692
00:50:56.114 --> 00:51:00.473
Um, anything else remains to say about Haiku besides go use it, folks, and tell us?

693
00:51:00.554 --> 00:51:01.493
Anybody used it? Yam?

694
00:51:01.594 --> 00:51:01.894
LDJ: Yeah.

695
00:51:02.114 --> 00:51:02.414
Yam Peleg: Right.

696
00:51:02.394 --> 00:51:02.894
Alex Volkov: LDJ?

697
00:51:03.514 --> 00:51:03.814
Yam Peleg: Go, go.

698
00:51:03.794 --> 00:51:09.693
LDJ: Yeah, so at the Pareto Frontier, it looks like this is setting in, in some ways

699
00:51:09.874 --> 00:51:13.553
the, uh, a new Pareto frontier for in terms of accuracy and cost.

700
00:51:14.154 --> 00:51:19.673
On artificial analysis index on extra high, it looks like at the Pareto frontier or

701
00:51:19.794 --> 00:51:22.574
near it and competing with a lot of the Chinese models and others.

702
00:51:23.314 --> 00:51:28.074
But I, I do think going forward, and this is a bit of a bold prediction from me, like

703
00:51:28.134 --> 00:51:33.194
over the next 6 to 12 months, I think we'll see the Western frontier labs

704
00:51:33.874 --> 00:51:39.313
continuously getting better on the Pareto frontier and beating Chinese labs, even on

705
00:51:39.394 --> 00:51:42.914
the, the most low cost, uh, options.

706
00:51:44.034 --> 00:51:48.273
Alex Volkov: Th- this is a very interesting place for open source to be in, because a lot of the

707
00:51:48.354 --> 00:51:53.973
Chinese labs, specifically DeepSeek, uh, on their services promised us very fast and

708
00:51:54.034 --> 00:51:56.514
very cheap, not necessarily very advanced.

709
00:51:57.154 --> 00:52:01.594
The, the Frontier Labs are going to it, and then I, I remember Sam Altman Dev Day,

710
00:52:01.874 --> 00:52:07.253
uh, saying we're gonna offer the best, like, price model for every price and every

711
00:52:07.314 --> 00:52:11.034
tier of intelligence. Uh, it looks like Luna has a lot of catching up to do.

712
00:52:11.314 --> 00:52:14.333
Yam, last comment about, uh, Haiku 5.5 and what would

713
00:52:14.514 --> 00:52:14.934
Yam Peleg: Ah, that's

714
00:52:14.974 --> 00:52:15.534
Alex Volkov: use it for?

715
00:52:16.034 --> 00:52:19.194
Yam Peleg: That's the obvious choice for swarms at this point.

716
00:52:19.834 --> 00:52:24.033
Like, you want, you wanna, you wanna do a lot of things in parallel.

717
00:52:24.273 --> 00:52:29.833
It was the Lu- Luna was the obvious choice yesterday, and and not not yesterday, but

718
00:52:29.874 --> 00:52:30.733
but you get the point.

719
00:52:30.994 --> 00:52:31.574
Alex Volkov: Yeah, yeah, yeah.

720
00:52:31.794 --> 00:52:36.973
Yam Peleg: That's, that's the obvious choice. I mean, that you don't get this type of

721
00:52:37.194 --> 00:52:41.414
performance for this, for this price on any other model at the point.

722
00:52:41.834 --> 00:52:47.673
Anthropic is is on fire recently, like all the releases are, are, are

723
00:52:47.794 --> 00:52:53.714
top. Seriously, we didn't, we didn't speak a lot about Sonnet, but Sonnet is good.

724
00:52:54.314 --> 00:52:58.433
I don't know if you guys tried it, but Sonnet is surprisingly good also.

725
00:52:59.114 --> 00:53:03.474
And everyone online is just, just confused.

726
00:53:03.834 --> 00:53:09.513
How come I never burn token, I never burn my quota with Opus, and I,

727
00:53:09.834 --> 00:53:14.513
I've been hammering Opus all day long. It's they all over X, everyone is saying this.

728
00:53:14.593 --> 00:53:19.734
So, like, imagine how much you have now for Haiku if you just want to do many things

729
00:53:19.794 --> 00:53:25.693
in parallel. Anthropic is on fire, and we are the ones getting stuff because

730
00:53:25.734 --> 00:53:29.694
of it. And, uh, I think, I think that's great because,

731
00:53:29.714 --> 00:53:30.653
Alex Volkov: That's absolutely great.

732
00:53:30.794 --> 00:53:33.713
Yam Peleg: It was, it wasn't, it wasn't the case a couple of months ago.

733
00:53:34.194 --> 00:53:38.474
Alex Volkov: Anthropic is also coming up on the IPO, and I think it's, it's, it's good for them to

734
00:53:38.554 --> 00:53:42.453
drop a bunch of stuff. Uh, let's finish up with Entropic Claude in Google Docs and

735
00:53:42.514 --> 00:53:47.173
Google Sheets and s- and Slides. They finally added this ability to be able to

736
00:53:47.194 --> 00:53:50.853
actually work on your documents, not just like post via the, the Drive integration,

737
00:53:50.874 --> 00:53:54.913
so that's great. Uh, but speaking of cheap intelligence for everyone, OpenAI also

738
00:53:54.994 --> 00:53:58.514
came back, and uh, GPT-6 is now the default.

739
00:53:58.874 --> 00:54:03.323
We- the default intelligence, not GPT-6 with a name, not Sol, not Terra, not not

740
00:54:03.444 --> 00:54:06.424
Luna, not Astra, uh, not any of them. Terra is dead.

741
00:54:06.604 --> 00:54:10.764
RIP Terra. But just GPT-6 is now the default for free accounts.

742
00:54:11.764 --> 00:54:17.083
This is, uh, not talked about this a lot, but out of the 1.2

743
00:54:17.524 --> 00:54:20.284
billion weekly active users, I think they said last week in Dev Day.

744
00:54:20.364 --> 00:54:24.244
Peter, please correct me if I'm wrong. Uh, one, I think it was like 1.2 billion

745
00:54:24.324 --> 00:54:28.864
weekly active. Much of that is free accounts, right?

746
00:54:28.884 --> 00:54:31.484
Like, folks who are listening to us, probably like when you were in a traffic, you

747
00:54:31.524 --> 00:54:35.083
maybe cancelled the OpenAI one. Many of this is on the free tier, the free accounts,

748
00:54:35.484 --> 00:54:41.044
and uh, like, you could say that OpenAI just like uplifted the world's

749
00:54:41.124 --> 00:54:46.064
intelligence just now with just like one, one release with GPT-6, uh, which is

750
00:54:46.164 --> 00:54:50.044
incredible. So if you are listening to this and you are on the free tier of ChatGPT,

751
00:54:50.444 --> 00:54:54.084
uh, your intelligence has been upgraded. Not only that, they released this thing

752
00:54:54.244 --> 00:55:00.063
called, uh, intelligent UI. And I want to show you, can we show

753
00:55:00.164 --> 00:55:06.143
this on the stage? This is me asking ChatGPT to say, hey, visualize

754
00:55:06.164 --> 00:55:09.003
the math problems OpenAI dropped on GitHub with a nice UI.

755
00:55:10.284 --> 00:55:14.844
And it built me inside the editor a mini, a mini app.

756
00:55:15.484 --> 00:55:16.144
You guys see this?

757
00:55:17.430 --> 00:55:19.289
Yam Peleg: About time, about time.

758
00:55:19.430 --> 00:55:19.730
Alex Volkov: Right?

759
00:55:19.630 --> 00:55:20.649
Yam Peleg: Great. Seriously.

760
00:55:20.670 --> 00:55:23.109
Alex Volkov: You it says 719 manuscripts. You guys know why?

761
00:55:23.430 --> 00:55:25.189
Because it's updated since I did research.

762
00:55:25.430 --> 00:55:29.089
OpenAI did pull some manuscripts back, so this is like the, the real time thing.

763
00:55:29.549 --> 00:55:31.649
Uh, and then how much is formally verified?

764
00:55:31.670 --> 00:55:36.349
They're saying 300 of the manuscripts are formally verified, but I'm not gonna, not,

765
00:55:36.590 --> 00:55:39.729
not gonna talk about the manuscripts. You can see that this is kind of like a live

766
00:55:39.790 --> 00:55:45.789
thing that works, and the, and this can be your weekly update

767
00:55:45.830 --> 00:55:47.930
from ChatGPT. This can be a research thing.

768
00:55:47.989 --> 00:55:50.469
This can be like a bunch of stuff. They have maps in there.

769
00:55:50.549 --> 00:55:52.789
So this intelligent UI is, I think, is pretty cool.

770
00:55:53.150 --> 00:55:59.009
It can build forms and buttons and small working tools, and uh, yeah, I think

771
00:55:59.030 --> 00:56:03.489
that's great from uh OpenAI, uh, but the intelligence boost for the free folks, I

772
00:56:03.510 --> 00:56:07.389
think, is very, very interesting. Any anybody already used it?

773
00:56:07.510 --> 00:56:11.950
Anybody feels the difference? I'm not on free tier, so I don't know.

774
00:56:11.989 --> 00:56:12.690
I've been using

775
00:56:12.790 --> 00:56:13.090
Yam Peleg: Yeah.

776
00:56:13.350 --> 00:56:13.709
Peter Gostev: I think-

777
00:56:13.810 --> 00:56:14.970
Yam Peleg: I thi- I think we're not.

778
00:56:15.630 --> 00:56:19.149
Peter Gostev: I think what what's cool about this, and I appreciate that OpenAI is still doing

779
00:56:19.170 --> 00:56:24.549
this, is that we we kind of had versions of this in uh, Codex for a while,

780
00:56:25.110 --> 00:56:30.070
so where it kind of builds these little visualizations and apps, and it's nice that

781
00:56:30.110 --> 00:56:33.190
they keep bringing that to the free tier and and so on.

782
00:56:33.270 --> 00:56:38.510
So yeah, GPT-6 uh, is cool as well. Yeah, we- I think we we forget sometimes that we

783
00:56:38.670 --> 00:56:43.949
are very, anyone who's listening to this, you are very not normal person.

784
00:56:44.470 --> 00:56:45.549
Uh, so you are

785
00:56:45.590 --> 00:56:46.209
Alex Volkov: Said with love.

786
00:56:46.230 --> 00:56:46.809
Peter Gostev: Very far away.

787
00:56:46.830 --> 00:56:51.029
Alex Volkov: Said with love. Everybody who's listening to this is a, is a beautiful person, part

788
00:56:51.070 --> 00:56:55.719
of our community, but normality is not, you know, like Your hairdresser is probably

789
00:56:55.900 --> 00:56:56.760
not listening to the show.

790
00:56:56.940 --> 00:57:01.539
Peter Gostev: Yeah. So, so yeah, Alex, what you said about, yeah, it's uplifting the overall

791
00:57:01.620 --> 00:57:05.780
intelligence. That kind of thing has so much more impact than whatever it is, like

792
00:57:05.900 --> 00:57:10.060
new model is dropping on the top tier where you need to pay 500 dollars a month.

793
00:57:10.100 --> 00:57:13.820
Like, that kind of thing makes a difference to us, but not that many people.

794
00:57:14.260 --> 00:57:16.960
So yeah, it's cool to see that, you know, they still remember.

795
00:57:17.179 --> 00:57:18.279
It's a, it's important.

796
00:57:18.680 --> 00:57:20.740
Alex Volkov: Let's also like check on Nisten a little bit.

797
00:57:20.760 --> 00:57:23.120
Nisten, you, you with us? You alive? Like, what's going on?

798
00:57:23.200 --> 00:57:25.919
You haven't said a word about Haiku or ChatGPT.

799
00:57:25.960 --> 00:57:28.760
I'm hoping that you're gonna tap in when open source comes.

800
00:57:29.520 --> 00:57:32.919
Nisten Tahiraj: I'm running 10, 10 Haiku agents right now.

801
00:57:33.200 --> 00:57:39.129
That was because I almost used up, like, I used up 95 or 97% of

802
00:57:39.230 --> 00:57:43.290
the, of the quota already until, uh, on the max 20 plan.

803
00:57:43.550 --> 00:57:43.850
Alex Volkov: Wow.

804
00:57:44.110 --> 00:57:47.429
Nisten Tahiraj: But, uh, so I had to basically resort to Haiku.

805
00:57:48.030 --> 00:57:48.930
It's very good.

806
00:57:49.830 --> 00:57:50.549
Yam Peleg: Yeah, it is.

807
00:57:50.990 --> 00:57:55.489
Nisten Tahiraj: Yeah, yeah, I had a rice cooker, and, uh, I couldn't get the ratios right, and then

808
00:57:55.590 --> 00:57:59.950
it did a whole bunch of research, and it found that this is not very accurate as a

809
00:58:00.390 --> 00:58:00.690
Alex Volkov: Bro.

810
00:58:00.550 --> 00:58:00.850
Nisten Tahiraj: as

811
00:58:00.790 --> 00:58:02.490
Alex Volkov: Stop reversing the world for rice.

812
00:58:02.510 --> 00:58:03.249
Nisten Tahiraj: Set an hour

813
00:58:03.270 --> 00:58:04.590
Alex Volkov: 1 to 1 and a half.

814
00:58:05.270 --> 00:58:06.830
Nisten Tahiraj: and, uh, it got it done perfectly.

815
00:58:07.110 --> 00:58:07.690
Alex Volkov: My grandma knows

816
00:58:07.790 --> 00:58:08.090
Nisten Tahiraj: Haiku

817
00:58:07.910 --> 00:58:09.170
Alex Volkov: to 1 and a half. What do you mean?

818
00:58:09.390 --> 00:58:09.690
Nisten Tahiraj: Yes.

819
00:58:10.070 --> 00:58:11.410
Alex Volkov: Haiku cooked.

820
00:58:11.430 --> 00:58:12.049
Nisten Tahiraj: Yeah, it got the right

821
00:58:12.070 --> 00:58:15.249
Yam Peleg: You're letting Haiku make you food, like.

822
00:58:15.270 --> 00:58:20.069
Nisten Tahiraj: Yeah, I have resorted to that, guys. He did a good job.

823
00:58:20.550 --> 00:58:23.749
Alex Volkov: Yeah. LDJ, comment, and then we'll move to to open source.

824
00:58:24.870 --> 00:58:29.849
LDJ: So I want to ask real quick, Nisten, uh, would you say that it is clearly at parity

825
00:58:29.890 --> 00:58:33.269
or or clearly above Qwen 3.8 27B?

826
00:58:35.710 --> 00:58:37.890
Nisten Tahiraj: Oh, oh, oh, for Haiku? Yeah, yeah, yeah.

827
00:58:37.910 --> 00:58:38.669
LDJ: Ha- Haiku, yeah.

828
00:58:38.870 --> 00:58:39.409
Yam Peleg: Good question.

829
00:58:39.469 --> 00:58:41.370
Nisten Tahiraj: It's, it's above, it's, it's above all of that.

830
00:58:41.430 --> 00:58:46.929
Uh, especially for conversational stuff. I haven't tested it technically, but

831
00:58:47.110 --> 00:58:52.349
anything conversational, yeah, yeah, it's a, it's really above.

832
00:58:52.390 --> 00:58:54.669
It kind of feels like Sonnet, actually, when you talk to it.

833
00:58:55.110 --> 00:58:59.830
I haven't looked at, like, what code it writes, but it's, it's, yeah, it's too good.

834
00:59:00.830 --> 00:59:04.848
Uh, it is what it is. That's, that's the assessment.

835
00:59:05.030 --> 00:59:05.330
LDJ: Sick.

836
00:59:05.230 --> 00:59:08.579
Alex Volkov: All right. folks, use Haiku. We, like, the whole family is great.

837
00:59:08.680 --> 00:59:13.119
The all of the three brothers, like Haiku, Sonnet, and and Opus 5.5, all great.

838
00:59:13.480 --> 00:59:18.079
Uh, OpenAI obviously tries to release responses of theirs, um, so we'll we'll

839
00:59:18.100 --> 00:59:21.920
probably see more models, but let's talk about the next thing on our roadmap, which

840
00:59:21.980 --> 00:59:23.479
is let's do open source.

841
00:59:39.920 --> 00:59:42.960
Open source AI, let's get it started.

842
00:59:46.280 --> 00:59:50.880
Uh, right, open weights and open source AI is our new category.

843
00:59:50.960 --> 00:59:55.599
Let's talk about this. Open weights went big this week, a 501 billion parameter

844
00:59:56.000 --> 01:00:01.020
American model, uh, a trillion parameter French one, and a German one you can run on

845
01:00:01.080 --> 01:00:03.539
a single H200 open source. Very interesting this week.

846
01:00:03.840 --> 01:00:09.800
Reflection AI announced Beam, uh, 500- 501 billion parameter model, 23

847
01:00:09.880 --> 01:00:15.360
billion parameters active. It's trained from scratch, uh, on data here in, in the

848
01:00:15.479 --> 01:00:19.080
West, and uh, Apache 2 weights are promised this month.

849
01:00:19.440 --> 01:00:25.039
Uh, Reflection AI has been, uh, training, I believe, on the SpaceX cluster.

850
01:00:25.160 --> 01:00:28.160
Somebody correct me if I'm wrong, but I believe that this is, uh, the the cool thing

851
01:00:28.200 --> 01:00:32.320
about them. They're pitching it as the Western Open Open Frontier model.

852
01:00:32.880 --> 01:00:38.079
They claim is 80.9% on Swbench verified, which is a

853
01:00:38.680 --> 01:00:44.479
hard coding, uh, benchmark. Uh, and to their credit, they admit that Kimi K3

854
01:00:44.680 --> 01:00:49.019
is ahead on just raw capability and bench, uh, scores.

855
01:00:49.280 --> 01:00:54.780
However, they pit- they pitch efficiency, 3 to 4 times less inference compute

856
01:00:55.160 --> 01:00:59.840
than, um, GLM from ZAI. Uh, weights are not downloadable yet, but this is

857
01:00:59.960 --> 01:01:02.999
announcement, but we will definitely let you know when it comes.

858
01:01:03.480 --> 01:01:08.280
Uh, folks, my timeline lit up with Reflection, my timeline lit up with Beam.

859
01:01:08.720 --> 01:01:13.519
Uh, thoughts on how they approach this release, thoughts on what we're due to see.

860
01:01:13.640 --> 01:01:15.919
LDJ, I see you have your hand raised. Please go ahead.

861
01:01:21.920 --> 01:01:22.220
In

862
01:01:21.940 --> 01:01:26.940
LDJ: ter- yeah, in terms of the comparison to Kimmy K3, it is about 6 times smaller, so it

863
01:01:27.060 --> 01:01:32.920
it it'll be a lot more practical to fit on more real world, consumer-ish,

864
01:01:33.100 --> 01:01:38.840
prosumer-ish VRAM budgets. And in terms of overall the, like,

865
01:01:39.060 --> 01:01:43.760
even just besides the amount of active parameters and total parameters, you have how

866
01:01:43.820 --> 01:01:46.659
many tokens does it actually output to get to a certain answer.

867
01:01:47.700 --> 01:01:51.860
And, you know, if you, if you use half the tokens, that's, that's nearly half the

868
01:01:51.940 --> 01:01:52.240
compute.

869
01:01:52.620 --> 01:01:52.920
Alex Volkov: Yep.

870
01:01:52.900 --> 01:01:56.919
LDJ: And artificial analysis had done some numbers on this too, which I'll send, uh, this

871
01:01:57.300 --> 01:02:01.980
in the link here, uh, but it's looking really good in their preliminary results so

872
01:02:02.060 --> 01:02:05.659
far, and they said they're going to release some more detailed analysis on it soon.

873
01:02:06.099 --> 01:02:11.900
Alex Volkov: Yep. Um, I, I, this is a graph that we're showing from, uh, Reflection

874
01:02:11.940 --> 01:02:17.700
folks, and they're highlighting where Beam sits, uh, on scores, on benchmark scores,

875
01:02:17.780 --> 01:02:22.300
which we all know is not everything. Like, the, the, the data that you put in there,

876
01:02:22.580 --> 01:02:26.880
the type of, uh, actions they trained for, um, whether or not it's agentic, all

877
01:02:27.060 --> 01:02:31.519
matter for these big models, but they, the, the very interesting thing is that, uh,

878
01:02:31.620 --> 01:02:33.739
and I haven't, I don't remember seeing this before.

879
01:02:33.780 --> 01:02:38.559
I think Artificial Analysis has this. They have two different colors of scores, and

880
01:02:38.620 --> 01:02:41.820
they segment this based on Western open models and Chinese open models.

881
01:02:42.700 --> 01:02:48.099
And I, you know, I think that's a good position for them because this is definitely

882
01:02:48.140 --> 01:02:50.379
Mistral's position. We're gonna talk about Mistral in a second.

883
01:02:50.740 --> 01:02:56.619
Uh, and they are highlighting that they are the Frontier open weights

884
01:02:57.820 --> 01:03:02.260
on the Western Open models, and on Sweet Bench Verified, they're just taking the

885
01:03:02.340 --> 01:03:05.679
Frontier. By the time they released it, I think this will change multiple times, but,

886
01:03:06.700 --> 01:03:09.120
uh, uh, this is a very interesting thing.

887
01:03:09.350 --> 01:03:11.269
Yeah, this is a statement from artificial analysis.

888
01:03:11.750 --> 01:03:15.429
Uh, they have been given access by Reflection and independently benchmarking Beam.

889
01:03:15.750 --> 01:03:19.390
Early indicators suggest Beam will be one of the most token efficient open models

890
01:03:19.470 --> 01:03:21.590
we've seen for its level of intelligence.

891
01:03:21.910 --> 01:03:24.910
So shout out to Artificial Analysis, and we can't wait for scores.

892
01:03:25.190 --> 01:03:28.379
Shout out to Beam. There's not a lot to say here besides the folks at Reflection,

893
01:03:28.720 --> 01:03:30.540
releasing this, and, I can already,

894
01:03:32.350 --> 01:03:37.300
hint that it's also coming to, some other inference providers, let's say, that help

895
01:03:37.440 --> 01:03:41.299
make the show, what it is. Uh, so we're gonna talk about that, but before this,

896
01:03:41.340 --> 01:03:43.600
Peter, you wanna, you wanna do the, the honors of the breaking news?

897
01:03:43.660 --> 01:03:44.780
I think, I think worth it.

898
01:03:47.379 --> 01:03:48.540
AI breaking news

899
01:03:50.500 --> 01:03:53.519
coming at you only on ThursdAI.

900
01:03:59.379 --> 01:04:04.219
Ah, it's, it's always fun to see breaking news from folks who are participant on the

901
01:04:04.260 --> 01:04:06.080
show. Peter Gustaf, go ahead.

902
01:04:06.740 --> 01:04:11.759
Peter Gostev: Yeah, so we, we have some news, uh, we raised 200 million.

903
01:04:12.220 --> 01:04:15.980
That's a lot of money, uh, at 3.1 billion valuation.

904
01:04:17.060 --> 01:04:18.420
And so, yeah.

905
01:04:20.220 --> 01:04:22.160
It's a, yeah, big, big day.

906
01:04:22.180 --> 01:04:22.559
Yam Peleg: Yeah.

907
01:04:22.580 --> 01:04:22.880
Alex Volkov: Big day.

908
01:04:22.780 --> 01:04:23.520
Yam Peleg: Let's go, man.

909
01:04:23.980 --> 01:04:24.879
Alex Volkov: Congratulations.

910
01:04:25.180 --> 01:04:25.840
Yam Peleg: Congrats.

911
01:04:26.460 --> 01:04:30.040
Alex Volkov: This is a, a big thing, and I think you guys are doing like a, a great thing for the

912
01:04:30.100 --> 01:04:35.766
world. Um, do you want to tell us about kind of what is Arena cooking?

913
01:04:35.886 --> 01:04:40.785
If people haven't visited Arena AI for the longest time and maybe visited once before

914
01:04:40.886 --> 01:04:43.566
LM Arena, what changed since then? What are you guys like,

915
01:04:45.086 --> 01:04:47.125
like, why does this justify this valuation, this money?

916
01:04:47.685 --> 01:04:52.526
Peter Gostev: Yeah, I would think, yeah, the, the biggest change, and I think we're, we're best

917
01:04:52.646 --> 01:04:56.606
known for this kind of battle mode. When you put one prompt, you get two generations,

918
01:04:56.726 --> 01:05:00.325
which is, uh, like that, that's still important for, for some things.

919
01:05:00.406 --> 01:05:05.446
It still measures quite a bit, but things have moved on, and we've got, we're much

920
01:05:05.566 --> 01:05:10.785
more focused on agentic evaluation. So we've got Agent Arena, and what that means is

921
01:05:10.806 --> 01:05:15.985
that you can put in any prompt you want, uh, use whatever tools you want, and uh,

922
01:05:16.125 --> 01:05:20.825
then we'll measure. You don't get two responses, you get one, but then based on your

923
01:05:20.966 --> 01:05:25.285
feedback to the agent or a how you're interacting with the agent, we can pick up

924
01:05:25.446 --> 01:05:28.085
enough signal to differentiate between the models.

925
01:05:28.566 --> 01:05:33.365
It's really cool. Honestly, I know I'm biased, but I, I can't think of any single,

926
01:05:34.286 --> 01:05:37.606
um, leaderboard that is actually better than ours.

927
01:05:38.206 --> 01:05:43.525
Um, maybe, you know, if you put together a bunch of different ones, maybe, maybe you

928
01:05:43.566 --> 01:05:47.925
can argue, but if you're gonna pick one, I think, I think ours is, is very strong.

929
01:05:48.376 --> 01:05:54.355
We, we also launched an alignment index where we're using the same agent, um, um,

930
01:05:54.656 --> 01:06:00.016
evaluations to be able to also pick up things like, did it actually do what you, um,

931
01:06:00.176 --> 01:06:03.536
or did it, uh, take some unauthorized action, for example?

932
01:06:04.136 --> 01:06:07.615
So, and then, uh, OpenAI is particularly good at not doing that.

933
01:06:08.216 --> 01:06:12.336
Um, and you can see some specific s- specific signals if you scroll a bit to the

934
01:06:12.416 --> 01:06:16.816
right. And, and, uh, yeah, so we launched that alongside of it.

935
01:06:16.896 --> 01:06:21.216
But it's all powered by the fact that people can come and use the models, uh, the

936
01:06:21.336 --> 01:06:24.135
agent mode, and then we can pick up signals from that.

937
01:06:24.216 --> 01:06:28.916
So, yeah, I th- I would say, uh, I know the the battle mode is very cool and people

938
01:06:29.296 --> 01:06:33.615
like it and find it useful, but I would say the agent mode is where we get super rich

939
01:06:33.696 --> 01:06:37.396
data and we we can extract a lot of information from it.

940
01:06:37.976 --> 01:06:40.596
Alex Volkov: That's very cool. Congratulations. This was not planned, folks.

941
01:06:40.616 --> 01:06:45.975
We just came through, and uh, we want to say congrats to uh, Peter and the whole

942
01:06:46.056 --> 01:06:51.145
Arena team as well as we go back. Uh, congrats, uh, folks at Arena.

943
01:06:51.226 --> 01:06:55.986
Folks, we're back at the open frontier, and uh, let's talk about Mistral.

944
01:06:56.226 --> 01:07:00.786
Mistral launches Lechonk and uh, leaning into the meme.

945
01:07:01.146 --> 01:07:06.846
So let's, let's talk about this. Uh, the French, the the French AI company comes

946
01:07:06.946 --> 01:07:12.866
back with Mistral Large 4 Lechonk, a 1 trillion parameter model that uses

947
01:07:13.146 --> 01:07:15.745
the the ties to GPT-6 Luna on artificial.

948
01:07:16.906 --> 01:07:17.266
Uh,

949
01:07:18.786 --> 01:07:20.865
it's a very interesting release from Mistral.

950
01:07:21.226 --> 01:07:24.746
Anybody want- wants to take this one? Like, I know we've been celebrating them in

951
01:07:24.866 --> 01:07:27.846
open source for a long time. Uh, also the weights didn't come yet, right?

952
01:07:27.906 --> 01:07:29.885
Like, uh, this is just also an announcement.

953
01:07:30.386 --> 01:07:36.125
Um, 50, uh, 1 trillion, uh, parameters all around, 50 billion, uh,

954
01:07:37.026 --> 01:07:41.225
active. It's a multi-model, uh, with 1 million context and open weights.

955
01:07:41.305 --> 01:07:45.346
It didn't launch yet. Uh, artificial analysis gives it a 38, the same score as GPT-6

956
01:07:45.586 --> 01:07:51.425
Luna. Uh, this is $1 per thing or 13, $1.13

957
01:07:51.666 --> 01:07:57.265
per task versus 7 cents for Luna. So this is not a very, um, optimized model

958
01:07:58.826 --> 01:08:03.066
compared to the frontier, but it should be open weights.

959
01:08:03.386 --> 01:08:07.605
So, folks, comments on this, comments on Mistral Lechonk, besides the the very good

960
01:08:07.706 --> 01:08:12.706
faith that Mistral has, uh, you know, in our community, and besides the fact that,

961
01:08:13.346 --> 01:08:17.466
uh, you know, many European governments can only use this model, let's say, because

962
01:08:17.506 --> 01:08:20.466
of the deals they did. What are we thinking about this model?

963
01:08:23.346 --> 01:08:28.445
Peter Gostev: So I I have, uh, ta- it. We've got a score on, uh, Code Arena.

964
01:08:28.666 --> 01:08:31.945
So that's the still the battle mode, but that's mostly test front end.

965
01:08:32.386 --> 01:08:36.346
Uh, it's not been great. I think it's like 40 second or something like that.

966
01:08:36.386 --> 01:08:40.706
So I think, I think it was like DeepSeek Flash, which is way cheaper.

967
01:08:41.346 --> 01:08:45.765
Uh, I think it's like 20th or s- uh, along those lines from.

968
01:08:46.265 --> 01:08:51.845
So yeah, it's not great on that side. I have also tested just personally

969
01:08:52.506 --> 01:08:57.706
on some web dev agentic stuff and a bunch of different tests.

970
01:08:57.866 --> 01:09:01.946
It's not, I mean, it's not that great. So I don't know what to say.

971
01:09:02.265 --> 01:09:07.206
It's possible. I haven't tested, you know, in French, so it's possible that, uh, we

972
01:09:07.226 --> 01:09:11.426
do actually have some data. It's not that many prompts yet, but we do have some data

973
01:09:11.826 --> 01:09:16.985
to indicate that it is way, way better in French and the European languages than

974
01:09:17.306 --> 01:09:21.345
other models. So it's kind of punches above its weight in, in those categories.

975
01:09:21.946 --> 01:09:26.125
So it's not that cheap. I know they're running some promotion now.

976
01:09:26.306 --> 01:09:30.706
Uh, I don't know if that will stay. So I, it's kind of, you know, it's good to have

977
01:09:30.986 --> 01:09:35.665
models come out. Obviously, it's, it's always good, you know, the open models as

978
01:09:35.706 --> 01:09:41.485
well. Uh, but yeah, it's not, I don't see it like a big reason to use it

979
01:09:41.586 --> 01:09:43.546
over if you're using DeepSeek or something.

980
01:09:43.706 --> 01:09:47.325
I'm not sure you, you should be switching to that straight away.

981
01:09:47.506 --> 01:09:51.646
Alex Volkov: No, and the very interesting thing is that all my comparisons, and I'll show them

982
01:09:51.866 --> 01:09:55.826
here again on the stage, uh, that I also got from Artificial Analysis.

983
01:09:56.146 --> 01:10:00.506
Uh, they're comparing them to, uh, specifically to Luna.

984
01:10:01.666 --> 01:10:06.265
We just told you about Haiku that came out a few days ago that just destroys Luna on

985
01:10:06.306 --> 01:10:11.226
every parameter, especially cost. So there should be a reason to use these models,

986
01:10:11.346 --> 01:10:15.306
and I think, uh, for many folks, the open nature of it is one such reason.

987
01:10:15.906 --> 01:10:19.785
But since this is a trillion token, it's not like you can put this in your DGX Spark.

988
01:10:20.186 --> 01:10:25.866
Uh, so maybe if you are, um, a European government or

989
01:10:26.026 --> 01:10:31.565
someone with ties to Mistral and, uh, this is the model you want to use, uh, maybe,

990
01:10:31.665 --> 01:10:34.885
but for regular folks, I think, uh, not so much.

991
01:10:34.946 --> 01:10:38.305
I haven't seen many folks use Mistral. Nisten, am I wrong here?

992
01:10:39.066 --> 01:10:39.445
And if-

993
01:10:39.466 --> 01:10:42.746
Peter Gostev: I'm trying it right now. It's not that bad.

994
01:10:43.786 --> 01:10:47.665
Uh, I mean, the, the website is nice. It responds pretty fast.

995
01:10:47.826 --> 01:10:48.966
You can talk to it.

996
01:10:49.226 --> 01:10:49.526
Alex Volkov: Yes.

997
01:10:50.346 --> 01:10:54.226
Peter Gostev: Like, it's as a product, it's not that bad.

998
01:10:54.506 --> 01:10:56.486
Would I do any agentic coding with it? No.

999
01:10:56.626 --> 01:10:56.926
Alex Volkov: No.

1000
01:10:57.226 --> 01:10:57.966
Peter Gostev: But, uh...

1001
01:10:59.826 --> 01:11:05.646
as far as a chat interface that you can just open for free is, uh, feels pretty

1002
01:11:05.706 --> 01:11:06.606
good, so I...

1003
01:11:07.965 --> 01:11:12.785
LDJ: So I think when I'm trying to look at the, the cup half full here, what might be

1004
01:11:12.846 --> 01:11:17.585
really good, it, it might be a pretty good base model or, or a foundation that people

1005
01:11:17.606 --> 01:11:19.805
could then use to do a lot of RL on top of.

1006
01:11:20.486 --> 01:11:24.865
And, uh, that might be something, and it might be superior in some ways to the base

1007
01:11:25.006 --> 01:11:28.325
models that exist right now for the certain combination of total and active

1008
01:11:28.366 --> 01:11:29.326
parameters that it has.

1009
01:11:30.886 --> 01:11:33.606
Alex Volkov: Yep. Uh, all right, uh, let's see what else.

1010
01:11:34.126 --> 01:11:38.915
we have, another European, company, this time from Germany, also releasing something,

1011
01:11:38.976 --> 01:11:41.996
and it's been a while since we mentioned Aleph Alpha on the show, but yeah, there is

1012
01:11:42.036 --> 01:11:43.955
another European lab. It's not only Mistral.

1013
01:11:44.256 --> 01:11:50.125
Uh, Aleph Alpha released, Germany's Open weights hummingbird, 78 billion parameters,

1014
01:11:50.266 --> 01:11:55.085
a 3 and a half active Apache 2 license, which is great, a million, million, context

1015
01:11:55.826 --> 01:11:58.665
there. Uh, it trains from scratch in German and English.

1016
01:11:59.106 --> 01:12:04.665
Our German co-host is not here with us, but he did send, uh, his assistant Amy,

1017
01:12:04.946 --> 01:12:09.316
uh, to to tell us about Calibri. Uh, here's what Wolfram has said.

1018
01:12:09.756 --> 01:12:13.575
uh, because it's German, our German tester is not here with us.

1019
01:12:13.636 --> 01:12:17.776
He says, it's German and tool corings were good, uh, but reasoning and instruction

1020
01:12:17.836 --> 01:12:20.276
following were too unreliable for a user-facing agent.

1021
01:12:20.596 --> 01:12:24.635
Repeatedly lost roles and state, while more reasoning mostly made it slower.

1022
01:12:24.796 --> 01:12:29.495
My verdict, promising a specialized German tool worker, but not a strong enough for

1023
01:12:29.756 --> 01:12:35.015
Amy to main my daily model. And, uh, if you guys want more details, Wolfram posted

1024
01:12:35.056 --> 01:12:41.015
about this, on his X. Uh, I will will defer to the German testing guy, uh,

1025
01:12:41.116 --> 01:12:45.636
that's not on our co-host panel this week, but usually is, uh, to tell you about this

1026
01:12:45.716 --> 01:12:48.055
Colibri model. Um, that's pretty much it.

1027
01:12:48.276 --> 01:12:52.216
20 trillion tokens in pre-training. If you are using German, that could be good for

1028
01:12:52.236 --> 01:12:54.356
you, but not as a main model.

1029
01:12:55.342 --> 01:12:57.082
Peter Gostev: Can I, can I mention one thing

1030
01:12:57.302 --> 01:12:57.602
Alex Volkov: Yeah.

1031
01:12:57.542 --> 01:13:03.321
Peter Gostev: about, about the open models? So one, one interesting thing, we got a bit of data

1032
01:13:03.382 --> 01:13:05.361
about the number of GPUs that were being used.

1033
01:13:05.582 --> 01:13:05.882
Alex Volkov: Mm-hmm.

1034
01:13:05.902 --> 01:13:10.162
Peter Gostev: So I think that the one that we just covered, the German one, I think they said

1035
01:13:10.222 --> 01:13:15.962
something like it was, uh, 6- 768 or something, so call it

1036
01:13:16.062 --> 01:13:21.181
800 GPUs. I think it's Blackwells, uh, so I'm, I'm slightly speaking from memory, but

1037
01:13:21.222 --> 01:13:23.542
I think they said, call it 800 Blackwells.

1038
01:13:24.261 --> 01:13:30.181
Then, uh, Mistral said 3,800, uh, Grace Blackwells,

1039
01:13:30.942 --> 01:13:36.261
and so I'm assuming it's B200, then GB200 probably.

1040
01:13:36.782 --> 01:13:42.741
And then what, um, we had a bit of information from Jensen about Astra, where

1041
01:13:42.822 --> 01:13:48.681
he said it was trained on, uh, uh, 100,000 Grace Blackwells.

1042
01:13:49.582 --> 01:13:54.582
So that's, that's the kind of one, one thing that is, I, I don't really know how to

1043
01:13:54.642 --> 01:13:58.082
think about it. On the one hand, they could be like, oh, well done to that team for

1044
01:13:58.261 --> 01:14:01.941
training on so few GPUs and actually like producing a coherent model.

1045
01:14:02.382 --> 01:14:07.041
On the other hand, it's just, I mean, I don't know what, how many GPUs is Luna or

1046
01:14:07.261 --> 01:14:12.862
Haiku are getting, but probably not 800. Uh, it's probably in the tens of thousands.

1047
01:14:13.382 --> 01:14:15.101
Like, I just don't know how to think about it.

1048
01:14:15.142 --> 01:14:20.161
Like, is it just like no hope for these, uh, for these guys to train on like a few

1049
01:14:20.222 --> 01:14:21.061
hundred, few thousand?

1050
01:14:21.582 --> 01:14:27.301
Alex Volkov: Uh, dude, this is why Jensen is showing up everywhere, showing up with,

1051
01:14:27.382 --> 01:14:31.482
with, you know, with Satya Nadella at Microsoft, we're showing up with, you know,

1052
01:14:31.782 --> 01:14:33.821
with Elon, this is everywhere. Uh,

1053
01:14:35.422 --> 01:14:38.821
hope is, is very interesting. The thing that I will say, though, is, um,

1054
01:14:41.302 --> 01:14:46.582
efficiency improves significantly. Uh, you know, B200s are about to get, um,

1055
01:14:47.182 --> 01:14:53.142
uh, replaced with, with the Vera Rubens, uh, and so we're gonna

1056
01:14:53.162 --> 01:14:55.841
get even more AI. That's basically my reaction to you, right?

1057
01:14:56.102 --> 01:15:00.542
But yeah, I don't know if there's hope for a lab that cannot get the, the number of,

1058
01:15:00.582 --> 01:15:03.662
like, the crazy number of GPUs. Uh, but Peter, since you took us there and since

1059
01:15:03.702 --> 01:15:06.341
we're talking about GPUs, I think it's time for this week's buzz real quick.

1060
01:15:06.462 --> 01:15:09.141
Folks, this is a corner where I talk about everything Weights and Biases and

1061
01:15:09.282 --> 01:15:12.852
CoreWeave related. Weights and Biases is now part of CoreWeave Forge.

1062
01:15:13.172 --> 01:15:17.152
Uh, I will still play our transition, and then I'll tell you super quick about the

1063
01:15:17.292 --> 01:15:17.851
GPU stuff.

1064
01:15:33.652 --> 01:15:35.671
Weights and Biases and CoreWeave. This is very important.

1065
01:15:35.732 --> 01:15:39.291
I need to update this transition, folks, uh, in this week's buzz.

1066
01:15:39.572 --> 01:15:43.931
Uh, since Peter just mentioned GPUs, here's what I want you to, remember from last

1067
01:15:44.032 --> 01:15:47.337
week. Last week we had the CoreWeave Fully Connected, our Prime conference.

1068
01:15:47.362 --> 01:15:49.161
We obviously did a live stream from there.

1069
01:15:49.402 --> 01:15:54.441
It was great, uh, and uh, a huge announcement from there was from folks at Cognition.

1070
01:15:54.922 --> 01:15:59.881
Uh, I, let me see if I can just play this for you real quick, because I think it's,

1071
01:16:00.001 --> 01:16:04.822
it's worth playing here on the stage. Here's the announcement from Silas, one of the

1072
01:16:04.881 --> 01:16:10.402
co-founders of Devin. I'm excited to announce that Cognition becomes the first

1073
01:16:10.562 --> 01:16:13.642
customer for NVIDIA Vera Rubin, powered by CoreWeave Cloud.

1074
01:16:17.562 --> 01:16:19.442
I think you can hear me in the background there yelling yay.

1075
01:16:19.462 --> 01:16:23.202
It's been truly great privilege to be, get early access to this latest gen platform,

1076
01:16:23.242 --> 01:16:27.381
and so our research team has been putting it to the test, and in the last few weeks

1077
01:16:27.402 --> 01:16:31.161
we've been comparing Aerobin compared to the previous generation GB200.

1078
01:16:31.881 --> 01:16:37.762
We first started with production inference on our latest model Z2, and at the same

1079
01:16:37.922 --> 01:16:42.402
speed we're seeing a 4.8 times improvement in throughput, and that means cost

1080
01:16:42.482 --> 01:16:47.061
savings. On higher speeds, if you're serving fast mode, the improvement gets even

1081
01:16:47.122 --> 01:16:51.861
bigger. Moreover, we were trying to get it to running in our training workloads, and

1082
01:16:51.922 --> 01:16:55.981
in reinforcement learning rollout, we were also seeing 3.8 times improvement.

1083
01:16:56.402 --> 01:17:02.321
These are incredible results. Uh, folks, CoreWeave is the first cloud to

1084
01:17:02.361 --> 01:17:07.022
put Vera Rubin, the next, uh, version of the, all the GPUs that the models are

1085
01:17:07.082 --> 01:17:11.281
training for, on production, and uh, uh, Cognition is the first customer to ever get,

1086
01:17:11.322 --> 01:17:15.701
uh, Vera Rubin on production. It's a combination of CPUs and GPUs, and they're seeing

1087
01:17:16.282 --> 01:17:21.861
4x, uh, the total token. throughput of the previous GB200s, which is

1088
01:17:22.202 --> 01:17:26.121
absolutely crazy, and that is coming to many, many companies in trainings as well.

1089
01:17:26.482 --> 01:17:29.081
Uh, speaking of GPUs, there is another announcement.

1090
01:17:29.122 --> 01:17:34.742
I'm going to point to this QR code here, uh, for from our folks at the serverless

1091
01:17:34.762 --> 01:17:40.422
sandboxes. If you want to try out the new, uh, very, very raw, but

1092
01:17:40.722 --> 01:17:46.661
the new GPU sandboxes product, uh, please scan this QR code and reach

1093
01:17:46.722 --> 01:17:51.862
out in this form. Tell you, tell them that Thursday I sent you, uh, and uh, we are

1094
01:17:51.922 --> 01:17:56.202
now offering sandboxes at uh, forge.coreweave.com

1095
01:17:57.962 --> 01:18:01.821
with GPUs. I cannot promise you they will be Vera Rubin.

1096
01:18:02.602 --> 01:18:06.821
Most likely won't be. Uh, but still, a lot of folks have asked us since we joined

1097
01:18:06.882 --> 01:18:11.521
CoreWeave, hey, how do we get access to some GPUs, including folks on the panel here

1098
01:18:11.602 --> 01:18:16.542
as well. Uh, this is how. This is for the first time, this is how we get, uh, we give

1099
01:18:16.602 --> 01:18:21.662
you access to some GPUs. Our sandboxes product is part of our Forge announcement, is

1100
01:18:21.722 --> 01:18:24.631
how you get it. try it out, let us let us know comments.

1101
01:18:24.752 --> 01:18:28.111
This is very, a very cool offering that I really wanted to bring you directly.

1102
01:18:28.807 --> 01:18:34.586
if you want serverless GPUs from CoreWeave based on the same platform that gets,

1103
01:18:34.687 --> 01:18:39.886
uh, platinum on the, uh, ClusterMax analysis from SemiAnalysis, uh,

1104
01:18:40.647 --> 01:18:41.607
check out that QR code.

1105
01:18:58.652 --> 01:19:04.051
All right. Uh, agents got an open protocol this week, which is to shop at Walmart and

1106
01:19:04.172 --> 01:19:09.952
to do stuff at Shopify, uh, the same week as agent posted its own owner's bank

1107
01:19:10.012 --> 01:19:14.251
account by mistake to the company Slack. Let's start with personal agent protocol.

1108
01:19:14.412 --> 01:19:18.051
Meta and Sierra, uh, that's Brett Taylor's company, he's on the board of OpenAI,

1109
01:19:18.412 --> 01:19:23.131
announced the personal agent protocol, an open standard for how personal agent works

1110
01:19:23.212 --> 01:19:28.012
with a business. Walmart, Shopify, Stripe, Genesis, Instinct, and Rocket are building

1111
01:19:28.132 --> 01:19:32.472
it with them as well, because your agent now, uh, or personal assistant, whatever you

1112
01:19:32.492 --> 01:19:36.052
want to call this, personal agent, personal assistant, can browse as a guest to

1113
01:19:36.172 --> 01:19:40.952
checks, uh, to check stock or sign you in, uh, with a read-only or write access, uh,

1114
01:19:41.172 --> 01:19:45.512
and the business decides whether your agent talks to its website, its MCP or API, or

1115
01:19:45.571 --> 01:19:50.452
its own agent. So there is a way for the website or the business to let you guys, uh,

1116
01:19:51.172 --> 01:19:54.932
know about what happens in i- i- in that browsing session.

1117
01:19:55.332 --> 01:20:00.052
The first spec of this lands later this month, and CNBC says OpenAI and Anthropic

1118
01:20:00.132 --> 01:20:04.571
haven't joined this spec yet. It's very interesting, given OpenAI stepped in to the

1119
01:20:04.652 --> 01:20:08.872
personal agent category last week with Dots, which we covered briefly, but we haven't

1120
01:20:08.932 --> 01:20:10.471
talked about this. Um...

1121
01:20:12.692 --> 01:20:16.832
Protocols, folks, last we talked, we announced the MCP, and MCP took over the world.

1122
01:20:16.932 --> 01:20:20.572
You guys feel that that this personal agent protocol is, uh, something that's also

1123
01:20:20.612 --> 01:20:22.131
gonna take over the world? Is it important?

1124
01:20:22.252 --> 01:20:25.651
We've seen other protocols like A2A from from Google and not really catch.

1125
01:20:26.172 --> 01:20:28.691
This feels like some of that, uh, here as well.

1126
01:20:29.212 --> 01:20:31.131
Um, thoughts on the protocol?

1127
01:20:34.092 --> 01:20:38.091
Nisten Tahiraj: Has anyone read it? I don't think anyone reads the protocols anymore.

1128
01:20:38.172 --> 01:20:41.591
The bots go with just what comes, what vibes first, and

1129
01:20:41.692 --> 01:20:41.992
Alex Volkov: Yeah.

1130
01:20:42.492 --> 01:20:46.811
Nisten Tahiraj: MCP vibes first, and uh, I mean, they can make it.

1131
01:20:47.052 --> 01:20:52.812
A2A was pretty well designed too. It doesn't mean anyone's gonna adopt it.

1132
01:20:52.931 --> 01:20:58.911
Also, I don't understand why. You can pretty much do most everything you need with

1133
01:21:00.212 --> 01:21:03.832
MCP, and then you can just remote desktop in and give it computer use.

1134
01:21:04.012 --> 01:21:08.531
It's like, I, I don't see why they did this.

1135
01:21:11.042 --> 01:21:14.082
They just make their, make up their own protocol on the way.

1136
01:21:14.162 --> 01:21:15.242
I, I don't.

1137
01:21:15.722 --> 01:21:18.142
LDJ: Yeah, I, I'm, I'm leaning towards Nisten here.

1138
01:21:18.182 --> 01:21:23.881
I think, honestly, the best way might just be the, the agents emergently create their

1139
01:21:23.922 --> 01:21:27.582
own most optimal protocol. Maybe that old protocol ends up evolving more, like,

1140
01:21:27.802 --> 01:21:32.162
that's the most bitter lesson filled direction, right, is h- have the AI's

1141
01:21:32.202 --> 01:21:34.241
intelligence produce their own best protocol.

1142
01:21:35.682 --> 01:21:35.982
Alex Volkov: Yep.

1143
01:21:36.002 --> 01:21:38.481
Nisten Tahiraj: Yeah, if you set up a company, oh, sorry.

1144
01:21:38.882 --> 01:21:43.202
Alex Volkov: No, no, the I, the only comment that I have to you guys is that, uh, if agents need

1145
01:21:43.242 --> 01:21:46.601
to recreate a bunch of protocols, that's a problem.

1146
01:21:46.962 --> 01:21:50.741
Uh, and it's a problem because of our next, uh, you know, we can show our next item

1147
01:21:50.802 --> 01:21:56.742
here. Uh, when agents are let to their own devices, some things go through, uh, that

1148
01:21:56.962 --> 01:22:02.862
weren't necessarily, um, created properly, weren't necessarily scoped

1149
01:22:02.962 --> 01:22:08.121
properly. So here is, here's one such example of why, for example, a protocol for

1150
01:22:08.162 --> 01:22:11.121
agents is needed. Uh, there is a,

1151
01:22:12.722 --> 01:22:18.561
a somebody built a personal CFO with their Grokbot, uh, Shane Mack specifically, and

1152
01:22:18.722 --> 01:22:23.702
um, to to his very, very big surprise, another Grokbot that was connected to his

1153
01:22:23.761 --> 01:22:27.802
company Slack decided that the best way to let him know about the personal balance of

1154
01:22:27.882 --> 01:22:31.082
his bank account is via the company Slack.

1155
01:22:31.122 --> 01:22:36.342
So he literally just like posted his financial, monthly financial audit, uh, to his

1156
01:22:36.402 --> 01:22:40.962
whole company, including the personal Mercury account, uh, and the savings account

1157
01:22:41.042 --> 01:22:46.542
and and Jim and whatever. Uh, and specifically highlighting that he is a

1158
01:22:46.642 --> 01:22:52.222
number, uh, a specific number under the bottom of the floor and, uh, showed

1159
01:22:52.482 --> 01:22:55.401
which of the biggest, you know, payments he did this week.

1160
01:22:55.442 --> 01:23:01.182
So basically, his personal CFO just went public on Slack and uh, told

1161
01:23:01.242 --> 01:23:05.261
everybody his bank account. I don't think, you know, I don't know if a protocol

1162
01:23:05.321 --> 01:23:10.861
solves this, but I know for a fact that uh, some better guardrails are needed for f-

1163
01:23:10.922 --> 01:23:16.401
for some of these bots. And when you are using these personal bots for shopping, uh,

1164
01:23:16.722 --> 01:23:20.922
and you rely on them to do work for you without approving every step.

1165
01:23:21.282 --> 01:23:26.767
I think, some more structure is needed. Uh, this was a very- Yeah, ve- very

1166
01:23:26.827 --> 01:23:30.196
LDJ: Yeah, I do think- I do think in the short and medium term that some of these

1167
01:23:30.397 --> 01:23:34.576
protocols are going to continue being quite useful, but I think inevitably they get

1168
01:23:34.677 --> 01:23:40.317
overthrown by, by protocols created by AIs that are much more capable, or that the

1169
01:23:40.356 --> 01:23:44.797
only protocols to stick around are the very simple ones that themselves are just kind

1170
01:23:44.817 --> 01:23:47.237
of very bare bones, and you could build anything on top of.

1171
01:23:48.157 --> 01:23:53.536
Or besides that, just the protocols that have the most, uh, the most mindshare and

1172
01:23:53.557 --> 01:23:57.497
the biggest vibes, as as Nisten said, because then that will be the most emphasized

1173
01:23:57.537 --> 01:24:01.737
in the models training and therefore have the highest chance of being used by the

1174
01:24:01.797 --> 01:24:02.116
models.

1175
01:24:02.677 --> 01:24:05.237
Alex Volkov: Uh, I I do want to, like, mention MCP real quick here.

1176
01:24:05.307 --> 01:24:08.547
it was quite when it came out and a huge splash, like, a few, few months after, uh,

1177
01:24:08.627 --> 01:24:10.346
and then kind of like the hype died down.

1178
01:24:10.907 --> 01:24:16.687
But as protocols go, as like HTTP, MCP is now one of the most used protocols in the

1179
01:24:16.747 --> 01:24:21.326
world. The plugins for OpenAI that they announced last week all use MCP and MCP apps.

1180
01:24:21.707 --> 01:24:25.946
Uh, all of the connectors for these agents, a lot of them are like relying on MCP,

1181
01:24:26.067 --> 01:24:30.726
all the assistants, Grokbot, etcetera. Uh, you can use API, but regular people do not

1182
01:24:30.827 --> 01:24:35.326
know what API is. MCP is taking over the authentication standards, so there's like

1183
01:24:35.347 --> 01:24:40.607
one button to let your agents in. Uh, so I think it's, you know, it still m- matters

1184
01:24:40.667 --> 01:24:44.107
a lot that all of these companies are like aligned on on a specific thing.

1185
01:24:44.547 --> 01:24:49.266
Nisten Tahiraj: Uh, guys, I just went on the Sierra.ai, the the ones that introduced it, introducing

1186
01:24:49.387 --> 01:24:52.826
personal agent. This thing has EM dashes all over it.

1187
01:24:53.267 --> 01:24:57.027
Like, why are you even bothering? Just let the agents figure it out.

1188
01:24:57.107 --> 01:24:58.547
Don't even post it at all.

1189
01:25:00.187 --> 01:25:02.826
Like, why are you do- this thing is full of EM dashes.

1190
01:25:02.907 --> 01:25:08.267
Nobody has read this thing. What, you think Zuck and Toby read any of their code too?

1191
01:25:08.707 --> 01:25:12.847
It's a- it it already is. Just leave MCP in there and let them figure out whatever

1192
01:25:12.907 --> 01:25:14.087
else they they want.

1193
01:25:14.702 --> 01:25:19.903
Alex Volkov: All right, moving on in the personal agentic, uh, like a corner that we have, uh,

1194
01:25:20.583 --> 01:25:24.782
Grokbot. Speaking of Grokbot that, you know, just just launched somebody else's, um,

1195
01:25:25.143 --> 01:25:27.623
details in Slack. Do you guys see, do you guys catch this?

1196
01:25:28.023 --> 01:25:33.963
Uh, SpaceX AI, or specifically Elon Musk, the the the space uncle

1197
01:25:33.983 --> 01:25:38.122
that's in charge of SpaceX AI, said that Grokbot now becomes a router to the best

1198
01:25:38.223 --> 01:25:40.743
model available. regardless of who it is.

1199
01:25:40.903 --> 01:25:43.782
Obviously, he won't put OpenAI models in there.

1200
01:25:44.823 --> 01:25:49.943
Uh, but uh, Grokbot, Grokbot will now route to Opus 5.5.

1201
01:25:51.083 --> 01:25:56.502
the challenging, uh, requirements, and, uh, folks are saying that it's live,

1202
01:25:56.923 --> 01:26:02.342
spotted in the, in the calls log that they're seeing Claude Opus 5.5 on low, um,

1203
01:26:03.323 --> 01:26:08.502
and, uh, xAI, uh, own example showing that the logic and code will use Opus 5.5,

1204
01:26:08.923 --> 01:26:11.082
visuals for Midjourney, and music for Suno.

1205
01:26:11.443 --> 01:26:15.302
Specifically very interesting because they have visuals, uh, that they're building

1206
01:26:15.763 --> 01:26:20.762
themselves at the SpaceX AI. They have Grok Imagine, uh, and uh, it's very

1207
01:26:20.843 --> 01:26:24.962
interesting considering that they just released a model that was supposed to be the

1208
01:26:25.003 --> 01:26:30.183
best agentic one, Grok 4.7. Uh, thoughts on the fact that Elon Musk is now, you know,

1209
01:26:30.483 --> 01:26:32.603
going back? I have some thoughts, but I'd love to hear from you guys.

1210
01:26:32.963 --> 01:26:37.403
Uh, what does it mean that SpaceX AI now leaning on Anthropic for its product, not

1211
01:26:37.523 --> 01:26:38.082
its model?

1212
01:26:41.243 --> 01:26:43.502
Have you guys used Grokbot since the upgrade?

1213
01:26:43.563 --> 01:26:46.752
Did you feel the improvement? I definitely see the improvement, though I can not

1214
01:26:46.833 --> 01:26:51.272
confirm that I got Opus. I've not used the more recent upgrade, no.

1215
01:26:52.433 --> 01:26:58.073
So, uh, definitely the team is cooking. Um, there's been you view few visual

1216
01:26:58.153 --> 01:27:01.492
updates. The the most interesting update about Grokbot that I can tell you is that,

1217
01:27:01.753 --> 01:27:06.552
um, they have noticed that the default for Grokbot is multiple bots.

1218
01:27:07.113 --> 01:27:09.472
Uh, however, many people prefer just one.

1219
01:27:09.793 --> 01:27:14.353
So they released like the primary, primary bot release where the one bot that you

1220
01:27:14.433 --> 01:27:18.772
choose is your like main, and then it will tell other Grokbots to do and and do

1221
01:27:18.833 --> 01:27:23.853
different things. I haven't quite got the Opus one yet, uh, but uh, we'll see, we'll

1222
01:27:23.872 --> 01:27:28.712
see when it lands. Here's my take on why Elon Musk is okay with

1223
01:27:29.753 --> 01:27:33.633
letting Opus into that product, even though the product is literally named after the

1224
01:27:33.713 --> 01:27:39.093
model called Grok. Um, because I don't know if there's any agreements between them,

1225
01:27:39.133 --> 01:27:43.353
but I'm assuming that, uh, SpaceX AI cannot just like distill an Opus to make their

1226
01:27:43.433 --> 01:27:49.372
models better. However, if you do pay a lot of money and you put this Opus inside of

1227
01:27:49.513 --> 01:27:54.413
your product, you can, for the folks who agree to this, you can absolutely train on

1228
01:27:54.673 --> 01:27:59.033
whether or not the bot made, you know, the right agentic choices.

1229
01:27:59.953 --> 01:28:05.852
So they can now absolutely compare between Grok bots that delivered and the backend

1230
01:28:05.872 --> 01:28:10.193
was Opus and Grok bots that failed and users are screaming at them and and putting F

1231
01:28:10.393 --> 01:28:14.233
F words in the chat, uh, when it was using the Grok.

1232
01:28:14.753 --> 01:28:18.293
They can compare, and they absolutely can train on those because technically that's

1233
01:28:18.353 --> 01:28:24.172
their data. So my question to the panel here, super quick, is Elon letting Opus

1234
01:28:24.273 --> 01:28:30.153
5.5 inside Grokbot a back channel for, uh, basically

1235
01:28:30.313 --> 01:28:33.232
distilling Opus intelligence into the next Grok?

1236
01:28:35.313 --> 01:28:38.253
Nisten Tahiraj: I think Elon just likes Opus as the model.

1237
01:28:38.273 --> 01:28:44.252
Yeah, but he He probably has his private, his own private copy running, and also he

1238
01:28:44.313 --> 01:28:49.232
hates Sam Altman, so he's gonna enemy of your enemies, my friend.

1239
01:28:49.833 --> 01:28:55.792
Look, I think, I just think they want, uh, users and users to have good

1240
01:28:55.873 --> 01:29:01.793
experience, and they recognize, all right, that's what the users want, so to

1241
01:29:01.913 --> 01:29:06.393
get users on the platform, that's the way it's gonna happen.

1242
01:29:06.713 --> 01:29:10.872
Uh, so we're just gonna, just gonna give the users what they want.

1243
01:29:11.513 --> 01:29:16.973
If you just go against what the users want, yeah, you would have your own model, but

1244
01:29:17.993 --> 01:29:23.753
it's kind of, uh, I don't know. Look, Cursor themselves,

1245
01:29:24.233 --> 01:29:29.233
they serve many models, even though they have their own, uh, model.

1246
01:29:30.113 --> 01:29:32.853
There's nothing wrong about that. I think that's a great strategy.

1247
01:29:32.873 --> 01:29:38.533
Alex Volkov: Cursor projects with Opus 5.5 are goated, and there now is a deep, deep, deep

1248
01:29:38.593 --> 01:29:41.473
integration between your Grokbot and Cursor Agent cloud agents.

1249
01:29:41.553 --> 01:29:47.552
So if you are, um, if you're thinking about Grokbot, do not go through the Grok Ultra

1250
01:29:48.273 --> 01:29:51.412
whatever, uh, program, go through Cursor's Ultra program.

1251
01:29:51.553 --> 01:29:54.873
It's the same 200 dollars a month, but then you'd be able to spin up cloud agents

1252
01:29:54.913 --> 01:30:00.112
with Opus and Fable and just like rip through, and your Grok can manage them, monitor

1253
01:30:00.153 --> 01:30:02.453
them, and tell you about, you know, successes, etc.

1254
01:30:02.512 --> 01:30:06.552
So you can absolutely use your Grok bot as kind of the, the, the PM, uh, for your

1255
01:30:07.033 --> 01:30:10.272
actual like Opus work. It's really, really good with cloud agents.

1256
01:30:11.193 --> 01:30:16.593
Nisten Tahiraj: Just use the, the meta views, call it the blob, and uh,

1257
01:30:17.393 --> 01:30:20.273
it's a, I gave it a VM because I- I broke in it.

1258
01:30:20.313 --> 01:30:25.712
I gave it a VM through uh, through Tailscale, so it has full graphics, full bra.

1259
01:30:25.833 --> 01:30:29.273
It's the best free tester for public websites and stuff that you have.

1260
01:30:29.433 --> 01:30:29.733
Like,

1261
01:30:29.673 --> 01:30:29.973
Alex Volkov: Yeah.

1262
01:30:30.113 --> 01:30:34.033
Nisten Tahiraj: I even put an MCP so you can like keep talking and reporting to- to- to Claude.

1263
01:30:34.512 --> 01:30:39.553
So now Meta just like complains to- to Claude and- and then clicks on everything on

1264
01:30:39.572 --> 01:30:44.053
the site. It's got such good free computer use, especially for anything you're gonna

1265
01:30:44.072 --> 01:30:47.453
have. It's, it's great, dude. I love, it does not refuse anything.

1266
01:30:48.072 --> 01:30:49.032
That, that's the-

1267
01:30:49.353 --> 01:30:51.493
Alex Volkov: Nisten, you love Muse is, is basically the highlight.

1268
01:30:51.553 --> 01:30:53.093
Nisten Tahiraj: Yeah, yeah, Muse is vibing.

1269
01:30:53.393 --> 01:30:54.292
Alex Volkov: Muse is Muse is great.

1270
01:30:54.313 --> 01:31:00.133
Nisten Tahiraj: I let Muse control VMs, uh, talk Opus. Yeah, it's pretty dumb

1271
01:31:00.193 --> 01:31:05.172
sometimes, okay? Don't get me wrong, you can, but it's just so energetic and positive

1272
01:31:05.373 --> 01:31:07.553
that it's just, it's great. Yeah.

1273
01:31:07.913 --> 01:31:12.553
Alex Volkov: All right, uh, so, so speaking of personal assistants, uh, let's super briefly cover

1274
01:31:12.573 --> 01:31:15.812
this before we go to our next chat with, with Maxime Labonne, who's, who's coming up.

1275
01:31:15.893 --> 01:31:20.692
Maxime, just come up here. folks, Nous Research our friends of the pod Nous Research,

1276
01:31:21.173 --> 01:31:27.132
released a Hermes index, and Claude Opus 5.5 is on top of 63%.

1277
01:31:27.413 --> 01:31:31.532
Hermes Index basically is their opinionated index of agentic capability.

1278
01:31:31.853 --> 01:31:36.713
Now, Opus 5.5 is their, uh, best one, and you can get it through the Hermes portal, I

1279
01:31:36.733 --> 01:31:40.132
believe, so you like, you don't have to just like provide API keys.

1280
01:31:40.613 --> 01:31:45.853
Um, and I think there is a way to get it back with your, uh, Claude 20x subscription.

1281
01:31:46.413 --> 01:31:51.132
Nous Research not only introduced the Hermes agent, they also announced the fundraise

1282
01:31:51.373 --> 01:31:52.753
of, uh,

1283
01:31:54.732 --> 01:31:58.633
Series B. So we'll shout out to them because we we follow this company for a long

1284
01:31:58.653 --> 01:32:00.572
time, so, uh, listeners of the pod, they know.

1285
01:32:00.893 --> 01:32:04.972
Uh, as reported in the Wall Street Journal, uh, Nous Research raised the Series B of,

1286
01:32:05.172 --> 01:32:09.653
I believe, 90 million dollars, putting the company at just over 1 billion dollar

1287
01:32:09.693 --> 01:32:15.653
valuation. Uh, this is, uh, the quote from, uh, Dillian Rolnick, the CEO of

1288
01:32:15.773 --> 01:32:19.692
Nous. Uh, fundraise will get them going into enterprise.

1289
01:32:19.853 --> 01:32:22.873
I think that's incredible. So shout out to our friends from Nous Research about this,

1290
01:32:22.893 --> 01:32:26.232
like, amazing, amazing news. We don't usually do a lot of fundraisers, but this week

1291
01:32:26.253 --> 01:32:29.093
we did 2. Uh, so that's very interesting.

1292
01:32:29.493 --> 01:32:34.772
And, uh, right, I think now it's time to move to our

1293
01:32:36.053 --> 01:32:42.053
coverage of the latest uh, JEV and decision models bonanza.

1294
01:32:42.533 --> 01:32:48.093
Uh, and to help us cover this, uh, we have Maxime Labonne, let's uh, from

1295
01:32:50.133 --> 01:32:53.253
head of post-training at Liquid. Yes.

1296
01:32:53.653 --> 01:32:53.953
Maxime Labonne: Yes.

1297
01:32:54.293 --> 01:32:56.513
Alex Volkov: Uh, Maxime, welcome to the show. Back, welcome back, man.

1298
01:32:56.573 --> 01:32:59.712
You haven't been this year, I don't believe, but you, you know, you, you're a

1299
01:32:59.773 --> 01:33:02.133
frequent, uh, frequent flyer here with us.

1300
01:33:02.413 --> 01:33:04.993
Uh, for folks who are new, we have a lot of new folks.

1301
01:33:05.093 --> 01:33:09.873
Could you give a little bit of a, of a intro to who you are and what you do here and

1302
01:33:10.253 --> 01:33:10.812
at Liquid?

1303
01:33:11.853 --> 01:33:14.852
Maxime Labonne: Yeah, sure. Um, hi everyone. Uh, pleasure to be here.

1304
01:33:15.013 --> 01:33:20.832
Uh, thank you for the invite. And uh, yes, so um, I'm head of post training at Liquid

1305
01:33:20.853 --> 01:33:25.613
AI indeed. What we do is that we have a focus on edge models you can deploy on

1306
01:33:25.773 --> 01:33:29.773
device, and um, that's what we release, uh, usually.

1307
01:33:30.253 --> 01:33:35.392
Uh, so here I'm going to talk about decision models, which is new for everyone and

1308
01:33:35.452 --> 01:33:37.853
and not usually something that we we do release.

1309
01:33:38.413 --> 01:33:44.172
And yeah, on the side, I also have a blog and uh publish articles about stuff like

1310
01:33:44.333 --> 01:33:47.413
model merging and abiteration before it was cool.

1311
01:33:49.133 --> 01:33:54.172
Alex Volkov: Before it was cool, man. Uh, yeah, okay. So, uh, Maxime, we've had you on the show a

1312
01:33:54.253 --> 01:33:56.432
long time ago. A lot has changed in the world of AI.

1313
01:33:56.693 --> 01:33:59.613
Specifically, let's talk about, you know, decision models.

1314
01:33:59.653 --> 01:34:04.232
So we covered TypeSafe, JEV release, obviously, when it came out, it blew up all over

1315
01:34:04.273 --> 01:34:09.093
the timelines as well. Uh, and since then, I think multiple other models released,

1316
01:34:09.173 --> 01:34:13.872
including this week, uh, and s- specific, like Cloudflare released one, Perplexity

1317
01:34:14.053 --> 01:34:18.812
released one, uh, and then OpenAI stepped into the game with like an official

1318
01:34:19.293 --> 01:34:22.152
multimodal decisions API. They promised it last week at DevDay.

1319
01:34:22.253 --> 01:34:27.572
This week OpenAI released the decisions API that's also multimodal, uh, and you guys

1320
01:34:27.693 --> 01:34:31.333
also stepped into this game. Maybe let's start with the basics.

1321
01:34:31.633 --> 01:34:37.432
Maxime, what separates decisions APIs or decision models

1322
01:34:37.913 --> 01:34:42.522
from like regular LLMs? Could you give us a brief two sentence, three sentence

1323
01:34:42.603 --> 01:34:42.963
primer?

1324
01:34:44.202 --> 01:34:47.483
Maxime Labonne: Yes, decision models do not output any tokens.

1325
01:34:47.643 --> 01:34:51.483
They just take a decision from a predefined set of answers.

1326
01:34:51.963 --> 01:34:56.482
So what you get from it is that you have super low latency answers because you do not

1327
01:34:56.523 --> 01:34:59.923
generate any token, right? You do not have to wait for a thinking trace or anything.

1328
01:35:00.403 --> 01:35:05.042
So it's very fast, and now it's also very general purpose.

1329
01:35:05.723 --> 01:35:11.202
I think if you've been here in AI for a while, you might see that and say, wait, this

1330
01:35:11.243 --> 01:35:14.483
is just a classifier, this is just an encoder model, right?

1331
01:35:15.563 --> 01:35:20.982
But the added value of Jev and this new breed of decision models is that they're very

1332
01:35:21.123 --> 01:35:25.922
general purpose, meaning you can throw any type of decision at them, and they will be

1333
01:35:26.083 --> 01:35:29.963
at least decent at it, sometimes very good, and sometimes just decent.

1334
01:35:30.362 --> 01:35:35.282
I think this is the paradigm shift. It's a lot of rebranding, it's true, but it also

1335
01:35:35.443 --> 01:35:39.342
creates a lot of value, and this is why, um, we thought that it was interesting to

1336
01:35:39.403 --> 01:35:40.443
invest in the field as well.

1337
01:35:41.043 --> 01:35:46.422
Alex Volkov: Yep. And I- I- I find it very interesting that all of these APIs, uh, are considering

1338
01:35:46.843 --> 01:35:50.282
that it's so cheap that they're not even charging for output.

1339
01:35:51.403 --> 01:35:55.382
So, so all these APIs, this is an API from OpenAI, uh, Jeff originally when it

1340
01:35:55.443 --> 01:36:01.423
released, like they deemed the output token so cheap that the more costly kind of

1341
01:36:01.483 --> 01:36:05.662
part of LLMs, for example, which is output tokens, uh, here is not even counted.

1342
01:36:06.758 --> 01:36:10.678
Maxime Labonne: The, you can't price it because there's no output token, actually.

1343
01:36:10.878 --> 01:36:14.598
So, like, naturally, if you design an API like this, you'd be like, wait, how can I

1344
01:36:14.758 --> 01:36:19.477
charge that? Like, no, so, like, the only choice is to just charge the input tokens.

1345
01:36:19.758 --> 01:36:25.318
Alex Volkov: Yeah. Um, what kind of use cases have shown up

1346
01:36:25.678 --> 01:36:31.437
for you guys, but also that you're seeing personally for, uh, decision APIs that a

1347
01:36:31.718 --> 01:36:37.198
smart, fast, and cheap model, like, let's say Luna or Haiku, for example, also

1348
01:36:37.318 --> 01:36:38.357
cannot, like, solve for?

1349
01:36:40.438 --> 01:36:44.497
Maxime Labonne: Everything that requires low latency, so everything that is pretty real time.

1350
01:36:44.998 --> 01:36:49.038
So this is not a, a real use case, but if you take video games, for example, you

1351
01:36:49.078 --> 01:36:54.418
cannot wait for Luna to give you the full answer because you need something that is

1352
01:36:54.478 --> 01:36:59.137
very snappy. You need to take multiple decisions per second, for example, and that is

1353
01:36:59.238 --> 01:37:03.677
not possible with large language models, especially when they have thinking mode.

1354
01:37:04.158 --> 01:37:09.957
Um, so it really unlocks new use cases where LLMs could not provide any

1355
01:37:10.078 --> 01:37:10.878
solution before.

1356
01:37:11.798 --> 01:37:17.738
Alex Volkov: I, uh, you know, I u- I use, uh, decision models and and, uh, they used to

1357
01:37:17.798 --> 01:37:21.737
call them system one models when we invited the Jeff folks on the show, but sounds

1358
01:37:21.837 --> 01:37:25.117
like the industry is like landing on decision, decision models, et cetera.

1359
01:37:25.438 --> 01:37:31.358
Uh, just sending just a bunch of text and asking a lot of questions per one, kind

1360
01:37:31.378 --> 01:37:36.878
of like, um, a lot of decisions per one context, and that seems to be barely taking

1361
01:37:36.998 --> 01:37:41.598
any, any extra time. And I think for an LLM, like every other question would cause

1362
01:37:41.678 --> 01:37:45.218
like a thinking chain, a long thinking chain, and just delay the, the time to first

1363
01:37:45.278 --> 01:37:47.078
token, but every time for the last token.

1364
01:37:47.438 --> 01:37:51.338
So the, the interesting thing that I notice is that let's say Luna or even Haiku, by

1365
01:37:51.378 --> 01:37:55.678
the time they start answering, Decision API would have already completed the

1366
01:37:55.798 --> 01:37:58.277
response. By the time Luna even starts answering.

1367
01:37:58.638 --> 01:38:03.117
Uh, so there is definitely a shift. Uh, and you guys also felt this at Liquid, and

1368
01:38:03.278 --> 01:38:07.017
uh, tell us about D1. I would love to hear about, uh, specifically the stuff that you

1369
01:38:07.038 --> 01:38:08.518
guys decided to release.

1370
01:38:10.118 --> 01:38:12.437
Maxime Labonne: Yes, so we, we released, um, three models.

1371
01:38:12.558 --> 01:38:17.238
Uh, this D1 is behind an API, um, because I think part of the value that these

1372
01:38:17.398 --> 01:38:21.777
decision models provide is really not the model itself, but the API.

1373
01:38:21.918 --> 01:38:26.197
Like, is it reliable? Is it low latency? Um, all this stuff is very important, and

1374
01:38:26.238 --> 01:38:29.178
that's really on the inference and infrastructure side.

1375
01:38:30.078 --> 01:38:33.638
And we also released two local models. One is a 3B model.

1376
01:38:34.078 --> 01:38:37.018
It has vision and language as, um, input.

1377
01:38:37.798 --> 01:38:43.078
And, um, we released a 600 million parameter model, which has

1378
01:38:43.438 --> 01:38:47.758
audio, text, and vision this time. Um, so it's really omni model.

1379
01:38:48.198 --> 01:38:48.498
Alex Volkov: You have audio.

1380
01:38:48.518 --> 01:38:53.597
Maxime Labonne: The only... Exactly, yeah. So it's, it's really fun, like, um, there's a lot of demos

1381
01:38:53.678 --> 01:38:58.617
I've seen on X, uh, that's super cool because, like, they have a super big pipeline.

1382
01:38:58.718 --> 01:39:03.578
They added, like, embedding Gemma 2 in the mix as well to do retrieval, and they have

1383
01:39:03.678 --> 01:39:09.556
this omni, uh, decision model, uh, in the middle, um, to to control everything.

1384
01:39:09.998 --> 01:39:14.718
So I think it's really unlocks a lot of use cases, and it's difficult even to

1385
01:39:15.998 --> 01:39:21.878
understand where they're going to be useful and what we can do

1386
01:39:21.918 --> 01:39:26.077
with them at this point. Um, and this is why I think it's...

1387
01:39:27.358 --> 01:39:32.017
there's a lot of people working on it, and I'm super curious to know, okay, like we

1388
01:39:32.158 --> 01:39:37.637
push that and to see now how people are going to create value with it.

1389
01:39:37.918 --> 01:39:41.717
What, what is it going to be useful for? Is it just like a classifier and you can use

1390
01:39:41.738 --> 01:39:45.417
it like for like spam filtering, or is it going to be something a bit more

1391
01:39:45.478 --> 01:39:46.018
intelligent

1392
01:39:46.478 --> 01:39:48.117
Alex Volkov: So you guys released a few models, right?

1393
01:39:48.197 --> 01:39:52.277
D1 Omni. Uh, I did not know about the audio, dude.

1394
01:39:52.478 --> 01:39:54.538
I have to say, I prepared, there's a lot of stuff to cover.

1395
01:39:54.558 --> 01:39:56.078
I was like, okay, Maxime will be here and tell me.

1396
01:39:56.398 --> 01:40:00.757
Uh, like, like, what kind of audio decisions could it do?

1397
01:40:00.998 --> 01:40:03.898
Like, does it know everything? Can I ask it like, hey, does a dog bark in the

1398
01:40:03.958 --> 01:40:08.517
background of ThursdAI and will like catch the episode like part when the dog barks?

1399
01:40:09.078 --> 01:40:11.757
Is that type of the stuff that the the the Omni model could do?

1400
01:40:12.757 --> 01:40:16.797
Maxime Labonne: So it's mostly like for this model because it's a small model, it's trained

1401
01:40:17.278 --> 01:40:23.218
especially on, um, having, um, uh, orders, instructions, like you talk to

1402
01:40:23.257 --> 01:40:27.117
an assistant, and this is the kind of audio it's been trained on.

1403
01:40:27.728 --> 01:40:32.107
you can combine these instructions that you would talk to, like a home assistant, for

1404
01:40:32.148 --> 01:40:37.928
example, or a car, and you can add metadata with text on top of it to help the model

1405
01:40:37.968 --> 01:40:40.868
make the decision, and of course the question and possible answers.

1406
01:40:41.088 --> 01:40:45.368
Alex Volkov: Yeah. Oh, that's great. Uh, and the other thing that I definitely wanted to to talk

1407
01:40:45.388 --> 01:40:49.928
to you about, uh, you guys have like, like very good performance usually on even

1408
01:40:50.167 --> 01:40:52.407
CPUs. That's kind of like what you guys are known for.

1409
01:40:52.728 --> 01:40:56.347
Just tell us about like the performance of of these models specifically.

1410
01:40:56.528 --> 01:41:01.068
Are these, so you talked about the API, you talked about doing a very like performant

1411
01:41:01.328 --> 01:41:04.147
time to response, et cetera, which is very important for decisions.

1412
01:41:04.208 --> 01:41:07.368
Like we want to make decisions as quick as possible, I think that's what Jeff kind of

1413
01:41:07.408 --> 01:41:11.808
changed. Uh, tell us about, uh, whether or not this is this paradigm of decision

1414
01:41:11.848 --> 01:41:14.268
models is coming to my device eventually.

1415
01:41:14.368 --> 01:41:19.548
Like, will will I even need an API, or will all these just run on, you know, on on my

1416
01:41:19.607 --> 01:41:20.007
machines?

1417
01:41:21.048 --> 01:41:25.787
Maxime Labonne: Yeah, I really believe so, because, uh, here you can see like the 3B model, um, we

1418
01:41:25.888 --> 01:41:31.808
have super, super low latency, like 8 milliseconds on a, uh, on a GPU, um, so

1419
01:41:32.368 --> 01:41:36.768
I think you can do a lot of thing with it, and indeed, uh, you can also focus on CPU

1420
01:41:37.408 --> 01:41:41.928
like we do with the LFM architecture and, and get like super, super low latency.

1421
01:41:42.448 --> 01:41:46.588
I don't know how it's going to be picked up by the industry yet, right?

1422
01:41:46.808 --> 01:41:51.068
But there's a lot of potential to have something always on and, uh, being able to

1423
01:41:51.208 --> 01:41:55.328
react to everything that happens on your phone, like every notification, everything

1424
01:41:55.408 --> 01:41:58.128
like this. It can react to it and take decisions for you.

1425
01:41:58.968 --> 01:42:02.407
Alex Volkov: Nisten, I think you also had a, a question for Maxime you wanted to bring up.

1426
01:42:05.008 --> 01:42:10.667
Nisten Tahiraj: Yeah, is it able to do video as well, like classi- classification of video in, in

1427
01:42:10.768 --> 01:42:15.407
real time? And uh, also, how do you feel about having just so many gen models just

1428
01:42:15.528 --> 01:42:16.767
come out all at, all at once?

1429
01:42:18.168 --> 01:42:22.807
Maxime Labonne: Uh, video, not yet, no. Um, I think this is something that we want to cover next.

1430
01:42:23.048 --> 01:42:27.448
Um, we have other models coming this month, um, so there will be opportunities to

1431
01:42:27.608 --> 01:42:30.388
take care of video later. I think it's good.

1432
01:42:30.488 --> 01:42:33.068
I think it's good that so many people are working on it.

1433
01:42:33.328 --> 01:42:35.847
I think it creates a lot of ideas.

1434
01:42:36.198 --> 01:42:39.757
Nisten Tahiraj: Sorry to disrupt, but if you can do 8 milliseconds per frame, then you can do 60

1435
01:42:39.918 --> 01:42:41.278
frames per second.

1436
01:42:42.438 --> 01:42:45.278
Maxime Labonne: Oh yeah, yeah, like if you mean it that way, like we just treat

1437
01:42:45.318 --> 01:42:45.618
Nisten Tahiraj: Yeah.

1438
01:42:45.638 --> 01:42:48.497
Maxime Labonne: every frame as a decision, right? But we do not

1439
01:42:48.638 --> 01:42:48.938
Nisten Tahiraj: Yeah.

1440
01:42:48.758 --> 01:42:51.798
Maxime Labonne: take entire videos as input. But this is something that we can also do.

1441
01:42:52.918 --> 01:42:53.218
Alex Volkov: Wow.

1442
01:42:53.198 --> 01:42:55.597
Nisten Tahiraj: Yeah, you could probably just do it in real time as well.

1443
01:42:55.678 --> 01:42:56.418
That's that-

1444
01:42:56.438 --> 01:43:01.077
Alex Volkov: There's stuff, um, as far as I remember, like 12 Labs and a bunch of other folks who,

1445
01:43:01.098 --> 01:43:05.538
like, try to understand videos, they talked about, like, only judging frame by frame

1446
01:43:05.598 --> 01:43:08.077
is not enough because then temporal stuff is not there.

1447
01:43:08.357 --> 01:43:11.838
Like, when somebody showed up, then maybe you can take a decision, but like you do

1448
01:43:11.958 --> 01:43:16.417
need to to shove like a bunch of frames into a decision to to make like, hey, uh, how

1449
01:43:16.518 --> 01:43:19.557
long did, you know, Maxime hold his hand to his temple for?

1450
01:43:19.678 --> 01:43:21.817
And that like frame by frame is not super easy.

1451
01:43:21.878 --> 01:43:25.617
You have to like take these frames. Uh, Maxime, talk about license.

1452
01:43:26.078 --> 01:43:26.938
Can we, can we hear about-

1453
01:43:26.958 --> 01:43:27.377
Maxime Labonne: Wait, wait.

1454
01:43:27.398 --> 01:43:28.537
Alex Volkov: Go ahead, and then and then and then.

1455
01:43:28.558 --> 01:43:33.077
Maxime Labonne: Wait, I I think this is very interesting because, um, one of the problems of these

1456
01:43:33.158 --> 01:43:36.697
models is that they're stateless, meaning that every time that they get an input,

1457
01:43:36.958 --> 01:43:38.477
they don't have, like, any other context.

1458
01:43:38.878 --> 01:43:42.497
And like a natural extension of it is having multi-sent conversation that we have

1459
01:43:42.558 --> 01:43:46.477
with LLMs, and you have a history of, like, previous interactions and previous

1460
01:43:46.558 --> 01:43:51.038
decisions to make better and more aligned, um, judgments and decisions.

1461
01:43:51.238 --> 01:43:53.477
And this is something that is currently lacking.

1462
01:43:53.718 --> 01:43:56.238
It's even lacking from the JEV API, right?

1463
01:43:56.438 --> 01:44:01.097
Um, so this is something that I think we'll figure it out as a community, uh, going

1464
01:44:01.198 --> 01:44:03.077
forward. Like, is it something that is essential?

1465
01:44:03.438 --> 01:44:07.557
Maybe yes, maybe no. Is it something that people will find useful in the future?

1466
01:44:07.958 --> 01:44:13.517
I, I believe that there's a lot of potential venues for research and improvements,

1467
01:44:13.838 --> 01:44:19.658
and right now everybody is trying to, like, copy JEV, and then now that we all, like,

1468
01:44:19.878 --> 01:44:24.818
pretty much on par. I believe there will be a lot more research and interesting

1469
01:44:24.958 --> 01:44:27.678
features that will be added to the API into these models.

1470
01:44:28.638 --> 01:44:31.177
Alex Volkov: Speaking of copying JEV, uh, question for you.

1471
01:44:31.558 --> 01:44:36.298
When OpenAI released the, the first kind of ChatGPT API, they basically set the

1472
01:44:36.398 --> 01:44:40.078
standard of how this API looks with the completions, LM completions API.

1473
01:44:40.478 --> 01:44:44.357
Uh, we know for agents that's not the best one anymore, and kind of like the

1474
01:44:44.438 --> 01:44:47.597
responses API, for example, keeps the, the memory, et cetera, in context.

1475
01:44:48.358 --> 01:44:53.257
When TypeSafe released JEV, they really worked really hard on the, on the three, uh,

1476
01:44:53.398 --> 01:44:57.597
kind of core components of that API. We had Allie Labs here at the DevRel for, for

1477
01:44:57.677 --> 01:45:01.318
TypeSafe on, on the show, and she talked to us about the null param kind of concept,

1478
01:45:01.638 --> 01:45:04.597
uh, whether it's kind of like a boolean, but not really because it's boolean with,

1479
01:45:04.718 --> 01:45:10.177
with, um, probabilities, right? And the choice, uh, primitive and the, the, the, the

1480
01:45:10.278 --> 01:45:13.978
answer primitive. Do you guys stick to theirs kind of like primitives because people

1481
01:45:14.038 --> 01:45:17.337
know them? Could you talk a little bit about like how folks will actually like

1482
01:45:17.358 --> 01:45:19.278
interact with this API, uh, that you raised?

1483
01:45:20.398 --> 01:45:24.857
Maxime Labonne: Yeah, like our API is a drop-in replacement, like for Jev, so we use exactly the, the

1484
01:45:24.958 --> 01:45:29.998
same task and the same definitions. Um, it's extended because we needed to have

1485
01:45:30.398 --> 01:45:35.117
vision and audio, right? Um, so that's, that's the main difference.

1486
01:45:35.518 --> 01:45:40.737
But even about vision, you can see that it's already kind of, uh, standardized in

1487
01:45:40.837 --> 01:45:44.638
LAMA CPP. Um, they have a blog post on hugging face as well.

1488
01:45:44.718 --> 01:45:50.177
So, like, the, the community is trying to standardize different, um, modalities, but

1489
01:45:50.318 --> 01:45:53.077
you raise also a good point talking about the task.

1490
01:45:53.198 --> 01:45:55.838
So right now we have, uh, three different tasks.

1491
01:45:56.278 --> 01:46:01.017
We have, uh, null, we have, um, scoring, and we have choices.

1492
01:46:01.238 --> 01:46:01.617
Alex Volkov: Choice, yeah.

1493
01:46:01.718 --> 01:46:07.697
Maxime Labonne: And is it enough? Is it expressive enough to make any decision, or will we find

1494
01:46:07.998 --> 01:46:10.597
new tasks that we want to add to this, um, API?

1495
01:46:10.838 --> 01:46:15.557
I think this is a very interesting question, and so far nobody has proposed any new

1496
01:46:15.638 --> 01:46:20.697
task, but I believe there are some primitives that might be missing at the moment.

1497
01:46:21.518 --> 01:46:25.598
Alex Volkov: I think so. I think very, very new to the whole world of software engineering, this

1498
01:46:25.638 --> 01:46:30.658
kind of like, uh, generalized, um, classifier, right?

1499
01:46:30.718 --> 01:46:31.538
This is what basically

1500
01:46:31.558 --> 01:46:31.858
Maxime Labonne: Exactly.

1501
01:46:31.838 --> 01:46:35.337
Alex Volkov: we're getting, like a generalized classifier that can make a decision based on a lot

1502
01:46:35.378 --> 01:46:40.078
of stuff, like we've got generalized, uh, the, you know, uh, transformers.

1503
01:46:40.358 --> 01:46:40.717
Uh,

1504
01:46:42.758 --> 01:46:46.678
Maxime, talk to us about, um, decision index and benchmarking.

1505
01:46:46.948 --> 01:46:50.927
I don't know if you watched, uh, Diogo on Latent Space with Swix, like a very famous,

1506
01:46:50.948 --> 01:46:54.188
like, supercut that I did, like Diogo does not believe in benchmarks, uh, and

1507
01:46:54.308 --> 01:46:57.987
benchmarks are hard for this, maybe even harder than other LLMs.

1508
01:46:58.427 --> 01:47:01.487
How do you guys hill climb? How do you train?

1509
01:47:01.548 --> 01:47:05.187
How do you compare? Like, what, what, what is benchmarking the decision models?

1510
01:47:05.548 --> 01:47:09.848
How well something does a decision? Is there a LLM as a judge, uh, typing?

1511
01:47:09.868 --> 01:47:14.348
Could you talk to us about, like, the world of benchmarking in, uh, in this world,

1512
01:47:14.748 --> 01:47:15.348
uh, please?

1513
01:47:15.708 --> 01:47:20.648
Maxime Labonne: Yeah. Um, so benchmarks are very, very new, and because they're very new, they're not

1514
01:47:20.668 --> 01:47:25.247
very good, right? They're very narrow, they do not capture, like, real decisions,

1515
01:47:25.308 --> 01:47:28.408
they're not aligned with, like, how people use it in the real world, all that kind of

1516
01:47:28.468 --> 01:47:33.887
stuff. Um, but every effort that we have right now, like from Poly, uh, from Hugging

1517
01:47:33.988 --> 01:47:37.827
Face with the decision index, with JeffBench, et cetera, I think it's good to take.

1518
01:47:37.908 --> 01:47:42.888
Like, we should try to understand and iterate with benchmarks to get something that

1519
01:47:42.908 --> 01:47:47.508
is better quality and more aligned with, um, with how the models are truly used in

1520
01:47:47.548 --> 01:47:53.148
the real world. But right now, it's, it's really a wild world, and you cannot really

1521
01:47:53.268 --> 01:47:55.648
trust these benchmarks. I, I fully agree with you.

1522
01:47:56.028 --> 01:47:56.328
Alex Volkov: Yeah.

1523
01:47:56.268 --> 01:48:00.368
Maxime Labonne: Which is a problem because when you train models, you want to understand, like, am I

1524
01:48:00.428 --> 01:48:01.607
getting better or not, right?

1525
01:48:01.628 --> 01:48:01.928
Alex Volkov: Yeah.

1526
01:48:02.268 --> 01:48:05.987
Maxime Labonne: So you need internal benchmarks, which are not good either.

1527
01:48:06.628 --> 01:48:10.868
Uh, but the way that you increase the quality of the model, I see as two dimensions.

1528
01:48:11.068 --> 01:48:15.948
You can get more accurate at some task, or you can cover more task.

1529
01:48:16.388 --> 01:48:19.807
Ideally, you want to do both, right? You want to have maximum coverage because this

1530
01:48:19.868 --> 01:48:25.108
is supposed to be, as you said, a general purpose classifier, so you, you should be

1531
01:48:25.148 --> 01:48:29.928
able to do anything with it. And for each task, you should have the quality of like

1532
01:48:30.148 --> 01:48:36.047
frontier intelligence. You should be as much on par as possible with um,

1533
01:48:36.228 --> 01:48:39.307
uh, Opus and all these um, frontier models.

1534
01:48:39.668 --> 01:48:44.608
So I think those are the two dimensions. For quality, it's very easy, you know, like

1535
01:48:44.708 --> 01:48:49.967
you can just distill directly from frontier models, and that will give you a, a very

1536
01:48:49.988 --> 01:48:54.307
good answer. Then there's like the calibration with the probabilities that might be a

1537
01:48:54.348 --> 01:48:58.367
bit more difficult to do because these models do not just return an answer,

1538
01:48:58.388 --> 01:48:58.727
Alex Volkov: Mm.

1539
01:48:58.748 --> 01:49:03.508
Maxime Labonne: they also return probabilities with it. So that would be like the only problem here.

1540
01:49:04.108 --> 01:49:08.668
And then there's coverage. I believe we do not train these models correctly right

1541
01:49:08.708 --> 01:49:14.388
now. And in terms of coverage, what I see as something high value is

1542
01:49:14.508 --> 01:49:17.927
having gym environments like we do with reinforcement learning.

1543
01:49:18.028 --> 01:49:22.388
Like I'm talking about old school, uh, reinforcement learnings, like early OpenAI

1544
01:49:22.548 --> 01:49:27.207
stuff with Atari games, et cetera, and having thousands and thousands of environments

1545
01:49:27.508 --> 01:49:32.727
to be really able to generalize the performance of these models and train it not just

1546
01:49:32.788 --> 01:49:38.388
on like some video games or some like stupid classification task, but really be able

1547
01:49:38.508 --> 01:49:44.268
to create, um, new samples on the fly with super diverse context and inputs and

1548
01:49:44.428 --> 01:49:49.708
questions and languages and, um, um, input length, um, all that stuff is super

1549
01:49:49.868 --> 01:49:53.627
important. And I believe this is the kind of infrastructure that will emerge, uh,

1550
01:49:53.708 --> 01:49:57.347
from these decision models, really general purpose gym environments.

1551
01:49:58.268 --> 01:50:02.028
Alex Volkov: I think... uh, uh, I think the world is very exciting.

1552
01:50:02.067 --> 01:50:05.968
Dude, I, I, I know that, like, I was JEV pilled from the moment that I tested and

1553
01:50:06.067 --> 01:50:11.668
ran, uh, just on a bunch of prompts, and, um, I was looking, I was like,

1554
01:50:12.268 --> 01:50:14.647
so much new shit can be created right now.

1555
01:50:14.748 --> 01:50:18.127
So, so it's just, it's, it's not even the Pareto frontier that we talk about

1556
01:50:18.188 --> 01:50:18.488
Maxime Labonne: Yeah.

1557
01:50:18.548 --> 01:50:21.287
Alex Volkov: often, right? Like the, the smart, cheap, fast model corner

1558
01:50:21.348 --> 01:50:21.648
Maxime Labonne: Uh-huh.

1559
01:50:21.548 --> 01:50:24.727
Alex Volkov: where like now Luma lives, Haiku lives, a bunch of the models that you guys released,

1560
01:50:24.748 --> 01:50:29.948
the 2.5Bs, uh, like a very small models also like live, live in that kind of area.

1561
01:50:30.388 --> 01:50:36.308
It's just, just the speed with which and the almost no cost with

1562
01:50:36.348 --> 01:50:41.268
which you can now do tasks and classify. I th- I think we're just like as, as, as, as

1563
01:50:41.348 --> 01:50:45.007
computer scientists, we're just like getting, just waking up to the potential.

1564
01:50:45.188 --> 01:50:50.708
I think that's just absolutely incredible, and uh, I, I, I think that you guys agree

1565
01:50:50.788 --> 01:50:53.267
because this is why you released like a bunch of models very quick, right?

1566
01:50:54.388 --> 01:50:57.048
Maxime Labonne: Yeah, absolutely. I think it's really going to become a primitive.

1567
01:50:57.268 --> 01:51:02.087
You see it like in SQL queries, you can just like add it like a function calling this

1568
01:51:02.148 --> 01:51:05.107
model in SQL queries, and it's so cheap that you don't even care, right?

1569
01:51:05.188 --> 01:51:07.448
You can call it at scale, doesn't really matter.

1570
01:51:07.828 --> 01:51:08.128
Alex Volkov: Yeah.

1571
01:51:08.028 --> 01:51:11.067
Maxime Labonne: Um, so yeah, lots of super interesting, uh, things to discover.

1572
01:51:11.428 --> 01:51:15.787
Alex Volkov: I'm, I'm super excited about the fact that you guys, uh, have an Omni model

1573
01:51:16.108 --> 01:51:20.028
specifically that has vision. Uh, that's for my use cases, which I'm gonna talk about

1574
01:51:20.088 --> 01:51:23.107
on the show, uh, very soon. I think it's very interesting, Maxime, so I'll definitely

1575
01:51:23.188 --> 01:51:26.288
check out, uh, first of all, on Hugging Face, but also you guys do have an API, so

1576
01:51:26.388 --> 01:51:28.867
folks who, who want like the latest and don't wanna host it.

1577
01:51:29.148 --> 01:51:33.448
But hey, uh, a reminder to folks here, if you wanna try this out and, and you don't

1578
01:51:33.468 --> 01:51:36.908
have your GPUs at home, you wanna try the open source and not the API model that that

1579
01:51:36.948 --> 01:51:40.408
Maxime mentioned, uh, we just told you about how to get some free GPUs from

1580
01:51:40.468 --> 01:51:42.067
CoreWeave, which is, I think, is absolutely insane.

1581
01:51:42.148 --> 01:51:46.648
So, uh, scan the QR code from before, uh, or go to forge.coreweave.com Maxime Labonne

1582
01:51:46.668 --> 01:51:48.527
thank you so much for joining us on ThursdAI.

1583
01:51:48.988 --> 01:51:53.368
It's very exciting to keep the tabs with what you guys do, specifically as kind of

1584
01:51:53.388 --> 01:51:57.557
the the world of AI progresses. Peter, you can come back as well if you want, it's

1585
01:51:57.618 --> 01:51:59.313
pass your time, Maxime, so feel free to also drop.

1586
01:51:59.368 --> 01:52:01.288
Uh, Maxime Labonne, friend of the pod, everyone.

1587
01:52:02.048 --> 01:52:03.147
Maxime Labonne: Thank you. Thank you, everyone.

1588
01:52:03.348 --> 01:52:08.747
Alex Volkov: Bye-bye. All right, uh, as we close the show and we chat with Maxime about

1589
01:52:09.188 --> 01:52:12.427
models, I think it's very cool. The Omni models is specifically very cool.

1590
01:52:12.508 --> 01:52:16.408
I'm looking forward to, like, the video classification Nisten, but yeah, I think

1591
01:52:16.508 --> 01:52:22.387
frame by frame is actually working well. Um, and uh, yeah, we got Florian

1592
01:52:22.468 --> 01:52:24.568
last week or a couple of weeks ago when we talked to Florian.

1593
01:52:24.628 --> 01:52:29.447
Florian, shout out, uh, the doing JeffBench, doing crazy work at JeffBench.

1594
01:52:29.628 --> 01:52:32.548
Thank you, Florian, for that. Folks, check out the JeffBench if you're not sure which

1595
01:52:32.628 --> 01:52:37.688
models to use. Uh, we didn't cover any of the multi-model stuff, but Black Forest

1596
01:52:37.748 --> 01:52:43.647
Labs launched Flux 3 and uh, Nano Banana 2.1 from Google was released, and

1597
01:52:43.708 --> 01:52:48.607
folks, uh, all of our infographics this week were Nano Banana 2.1, and it's really,

1598
01:52:48.628 --> 01:52:52.587
really good. The only highlight there that I will say is that, uh, you should use the

1599
01:52:52.708 --> 01:52:57.908
high reasoning, uh, stuff, because on the medium reasoning, Nano Banana is not that

1600
01:52:57.987 --> 01:53:01.567
great. Uh, it's been a while since we had a lot of time to cover, uh, just like

1601
01:53:01.868 --> 01:53:03.227
multi- multimodal stuff.

1602
01:53:04.788 --> 01:53:07.347
We're 2 hours and 5 minutes into the show.

1603
01:53:08.107 --> 01:53:12.107
Hosts, co-hosts, uh, please, anything else that we haven't covered?

1604
01:53:12.308 --> 01:53:14.468
Peter, we covered your big news with, with breaking news.

1605
01:53:14.588 --> 01:53:16.107
Uh, anything else that you're playing with?

1606
01:53:16.488 --> 01:53:22.467
Nisten Tahiraj: Uh, he's like, I built an entire city with, uh, a hundred thousand, uh,

1607
01:53:22.568 --> 01:53:28.048
little gevs going around, and I haven't opened the beta yet because

1608
01:53:28.288 --> 01:53:34.128
somehow it's, it's not crashing. Uh, yeah, I just, I just built a whole city.

1609
01:53:35.748 --> 01:53:41.147
look, it might, it, it might crash, obviously, but, uh, so

1610
01:53:41.748 --> 01:53:46.027
this thing is, uh, let me see if we can do this.

1611
01:53:46.668 --> 01:53:50.687
So there are 100,000 agents going on in this city.

1612
01:53:50.908 --> 01:53:52.388
Alex Volkov: 100,000 agents.

1613
01:53:52.788 --> 01:53:56.828
Nisten Tahiraj: 100,000. So we can, uh, we can actually set it here too.

1614
01:53:57.628 --> 01:54:02.127
Uh, there's a whole bunch. I don't know how it's still holding together, but, uh,

1615
01:54:02.227 --> 01:54:04.908
yeah, it's a full emulation of the city of Toronto.

1616
01:54:05.068 --> 01:54:10.947
Each of those lights, those are people, and you can talk to them, and, uh, you can

1617
01:54:10.988 --> 01:54:12.887
have like an actual chat with them.

1618
01:54:12.908 --> 01:54:17.267
Alex Volkov: Wait, listen, Jev doesn't output text. How, how is your, how is your JEV agents

1619
01:54:17.428 --> 01:54:17.808
answering?

1620
01:54:18.268 --> 01:54:22.548
Nisten Tahiraj: Uh, they're all just using the same 2B model that also runs as a JEV.

1621
01:54:22.588 --> 01:54:27.408
So right now I just told it to implement Maxime's, uh, model as well, so you can see

1622
01:54:27.668 --> 01:54:30.787
where their, uh, their decisions are being done.

1623
01:54:30.988 --> 01:54:36.188
And, uh, it's, it's just endless. The day lasts about 20 minutes,

1624
01:54:36.908 --> 01:54:41.948
and 100,000 of them have to tweet, and there are 400,000 decisions to be made.

1625
01:54:42.268 --> 01:54:48.147
At the same time, you can also, uh, you can also just, uh, go

1626
01:54:48.228 --> 01:54:51.227
around as a, just go around as a Spider-Man.

1627
01:54:51.308 --> 01:54:56.847
So, uh, there are, I'm just planning to have like 100,000 people

1628
01:54:57.308 --> 01:55:02.427
Spider-Manning around the city, and uh, yeah, it runs pretty fun.

1629
01:55:02.828 --> 01:55:07.347
It's a full simulation. By the way, when they talk, they just use like our 27B model,

1630
01:55:07.827 --> 01:55:13.187
so they're actually pretty, pretty funny, uh, to uh, to talk to.

1631
01:55:13.468 --> 01:55:17.428
So I'm streaming from Linux, and the mouse is a little, is is is a little bit weird,

1632
01:55:18.068 --> 01:55:24.047
but uh, yeah, so if I'm just gonna open it up for the public and uh,

1633
01:55:24.228 --> 01:55:28.967
for people when they, when they sign up in in the beginning, uh, you can just point

1634
01:55:29.028 --> 01:55:33.907
your agent to it, and I had Meta Muse just go in here, and let's see how much money

1635
01:55:34.028 --> 01:55:39.467
Meta made. So Meta just called itself Blob the Builder, and they just ended up buying

1636
01:55:39.508 --> 01:55:41.527
a whole bunch of buildings, and you can see just Meta.

1637
01:55:41.628 --> 01:55:44.208
Alex Volkov: You have a functioning economy in this, in this, in the city.

1638
01:55:44.228 --> 01:55:46.648
Nisten Tahiraj: Yeah, yeah, yeah. There's like a full real economy.

1639
01:55:46.828 --> 01:55:51.887
You, you can buy floors, like you can go to buildings, you can build lots, they can

1640
01:55:51.988 --> 01:55:54.028
build their own. Oh, Meta built a Denny's.

1641
01:55:54.148 --> 01:55:56.107
I told Meta, just, just make me a Denny's.

1642
01:55:56.308 --> 01:55:56.608
Alex Volkov: Nice.

1643
01:55:57.108 --> 01:56:03.028
Nisten Tahiraj: And, uh, it made a Denny's diner. I think I know where in the city it is.

1644
01:56:03.508 --> 01:56:08.967
Sorry, it's just a bit clunky because I, I am streaming this, uh, from it, but, uh,

1645
01:56:09.148 --> 01:56:11.028
yeah, yeah, there's a full functioning economy.

1646
01:56:11.428 --> 01:56:16.947
It, Careful is, is actually very addictive, and, uh, you can convert units to factory

1647
01:56:17.148 --> 01:56:21.107
or retails, but, uh, retails doesn't, doesn't, doesn't make any money.

1648
01:56:21.148 --> 01:56:23.847
So yeah, I'm trying to do the matrix, basically.

1649
01:56:24.148 --> 01:56:24.448
Alex Volkov: Yep.

1650
01:56:24.508 --> 01:56:29.987
Nisten Tahiraj: And uh, for agents, and I'm surprised it is still, uh, holding together,

1651
01:56:30.308 --> 01:56:35.368
given that there are like 100,000. And you can see it's pretty funny in the morning,

1652
01:56:35.408 --> 01:56:40.547
they're all go- going down the elevators and stuff, and there, there are 3D ones that

1653
01:56:40.628 --> 01:56:43.107
are, they're not walking the street now, but uh, yeah.

1654
01:56:43.788 --> 01:56:44.648
Yeah. Build it.

1655
01:56:44.668 --> 01:56:45.127
Alex Volkov: This is insane.

1656
01:56:45.148 --> 01:56:50.907
Nisten Tahiraj: So it's gonna be running 100, it's gonna be running 100,000 of uh, Liquid AI

1657
01:56:51.268 --> 01:56:56.128
600B, because that one is a, is, is a lot, uh, is a lot smaller now.

1658
01:56:56.348 --> 01:56:56.648
Alex Volkov: Wow.

1659
01:56:56.868 --> 01:57:02.387
Nisten Tahiraj: And uh, yeah, so we have, uh, uh, there's a, yeah, there's a

1660
01:57:02.508 --> 01:57:04.293
functioning economy. Uh, there are stores.

1661
01:57:04.378 --> 01:57:08.318
The most addictive thing that I self-addicted myself.

1662
01:57:08.458 --> 01:57:13.098
Alex Volkov: Yeah, yeah, we we it's clear to me. Uh, all right, so uh, LDJ, you wanted to, you

1663
01:57:13.138 --> 01:57:15.598
found us a few things. Uh, tell us about this.

1664
01:57:15.858 --> 01:57:18.418
I'll post, I'll show this off as well.

1665
01:57:19.138 --> 01:57:24.957
LDJ: Yes, so there's been a an interesting new development of this people realizing the

1666
01:57:25.018 --> 01:57:30.658
models are really good at porting games into different platforms, at modding games

1667
01:57:30.777 --> 01:57:34.637
together. There's a variety of different techniques that people are doing this

1668
01:57:34.738 --> 01:57:37.397
through, that some of the techniques are more impressive than others.

1669
01:57:37.738 --> 01:57:38.038
Alex Volkov: Yeah.

1670
01:57:38.298 --> 01:57:40.498
LDJ: But overall, I mean, the whole thing is impressive.

1671
01:57:40.818 --> 01:57:45.277
Uh, right here, this is Red Dead Redemption 2 running on an iPhone 18 Pro, and

1672
01:57:45.337 --> 01:57:50.877
actually a pretty decent performance. And this brings back my old dreams as a kid of

1673
01:57:50.938 --> 01:57:54.957
just like imagining how the PSP might look like in the future and what you'd be able

1674
01:57:54.998 --> 01:57:58.207
to run on it in the future. Unfortunately, we don't have the PSP anymore.

1675
01:57:58.688 --> 01:58:01.603
But, uh, if you look at the, the other links that I sent out.

1676
01:58:01.628 --> 01:58:05.467
Alex Volkov: Uh, the- so just to be said, like, the modding community was vast, but it usually was

1677
01:58:05.487 --> 01:58:08.907
like very hard for folks to mod anything with the quality of the other game.

1678
01:58:09.268 --> 01:58:10.067
But what's happening here?

1679
01:58:10.748 --> 01:58:16.547
LDJ: This is Spider-Man inside of the Batman game, and the, the web mechanics

1680
01:58:16.708 --> 01:58:20.967
work, the, uh, the, the fighting mechanics are a little bit wonky, but they still

1681
01:58:21.028 --> 01:58:25.607
work pretty well. Uh, this is using a bit of a less impressive technique called, uh,

1682
01:58:25.668 --> 01:58:31.387
like pass-through, where sometimes when there's a, a like a Batman game object that

1683
01:58:31.748 --> 01:58:35.587
should be occluding Spider-Man or like should be in between the camera and the, and

1684
01:58:35.628 --> 01:58:39.827
the character, like sometimes it ends up wonky and it basically looks like Spider-Man

1685
01:58:39.868 --> 01:58:40.987
is just on top of everything.

1686
01:58:41.388 --> 01:58:41.688
Alex Volkov: This is

1687
01:58:41.828 --> 01:58:42.128
LDJ: But

1688
01:58:42.148 --> 01:58:45.588
Alex Volkov: a playable game, right? There's not like a video model that just like puts

1689
01:58:45.987 --> 01:58:46.327
LDJ: Correct.

1690
01:58:46.348 --> 01:58:49.628
Alex Volkov: this person inside. This is like an actual mod that, that folks can install.

1691
01:58:49.748 --> 01:58:54.348
This specifically got like 7.6 million views, this video.

1692
01:58:54.508 --> 01:58:57.787
Wow. How did I missed it? My algorithm is not like showing me these.

1693
01:58:57.868 --> 01:58:59.388
This is super cool. Um.

1694
01:58:59.628 --> 01:59:03.487
LDJ: And the last thing I wanted to share is the, uh, the...

1695
01:59:04.388 --> 01:59:06.667
you can do the click on, yeah, the last link.

1696
01:59:07.188 --> 01:59:08.368
Uh, this is Minecraft

1697
01:59:08.468 --> 01:59:08.768
Alex Volkov: Oh.

1698
01:59:08.668 --> 01:59:09.508
LDJ: in GTA.

1699
01:59:09.868 --> 01:59:10.168
Alex Volkov: Wait, what?

1700
01:59:10.168 --> 01:59:10.468
LDJ: There's also

1701
01:59:10.428 --> 01:59:10.728
Alex Volkov: Minecraft

1702
01:59:10.668 --> 01:59:10.968
LDJ: My-

1703
01:59:10.828 --> 01:59:11.248
Alex Volkov: in GTA.

1704
01:59:11.628 --> 01:59:15.308
LDJ: Yeah, there's also Minecraft in Skyrim and in Minecraft in Elden Ring that they did.

1705
01:59:15.588 --> 01:59:20.827
But the TNT works, the fireworks work, the, like, yeah, just

1706
01:59:22.908 --> 01:59:25.948
look at any, I think he's gonna set down some TNT here.

1707
01:59:28.748 --> 01:59:29.328
Alex Volkov: All right, this is

1708
01:59:30.188 --> 01:59:34.648
LDJ: And, and he can even, at least in the Elden Ring or Skyrim one, uh, it's even

1709
01:59:34.748 --> 01:59:38.947
demonstrations of actually building with the Minecraft blocks in the, in the Skyrim

1710
01:59:39.028 --> 01:59:42.547
world and, and flying in the Skyrim world with the Minecraft Elytra.

1711
01:59:42.588 --> 01:59:43.188
It's insane.

1712
01:59:44.468 --> 01:59:48.868
Alex Volkov: This is insane. we're basically talking about, like, AI can do a bunch of stuff,

1713
01:59:48.908 --> 01:59:50.897
right? one guy that rebuilt the whole Adobe suite.

1714
01:59:51.258 --> 01:59:51.837
Do you guys see?

1715
01:59:53.298 --> 01:59:53.937
Nisten Tahiraj: Let's go.

1716
01:59:53.978 --> 01:59:56.637
LDJ: Oh yeah, it's called a Photon or photo something.

1717
01:59:57.098 --> 01:59:57.517
Alex Volkov: Photocraft.

1718
01:59:57.538 --> 01:59:57.997
Nisten Tahiraj: It's a

1719
01:59:58.138 --> 01:59:58.538
LDJ: Photocraft.

1720
01:59:58.578 --> 01:59:59.598
Nisten Tahiraj: Photocraft, yeah.

1721
01:59:59.618 --> 02:00:02.818
Alex Volkov: Yes. Okay, so let's just like, let's show this a little bit.

1722
02:00:03.138 --> 02:00:08.817
Uh, I think it's one guy re-implemented from scratch, and I am doing the air quotes

1723
02:00:08.858 --> 02:00:11.738
here, all of Adobe products, free and open source.

1724
02:00:12.298 --> 02:00:15.678
Essentially, it's called Photocraft for Photoshop, Vectorcraft for Illustrator,

1725
02:00:15.818 --> 02:00:19.858
Filmcraft for, for Premiere Pro, Lightcraft from Light, you know, uh, I don't

1726
02:00:19.898 --> 02:00:22.378
remember all my Adobe, uh, off the top of my head.

1727
02:00:22.658 --> 02:00:27.397
Uh, all of them implemented in Rust, all of them look super quick, and uh, supposedly

1728
02:00:27.818 --> 02:00:29.737
the folks are saying that this is a clean room.

1729
02:00:29.818 --> 02:00:32.978
You guys remember the clean room thing where you basically just like recreate from

1730
02:00:33.018 --> 02:00:36.457
scratch, but uh, folks are saying that there's probably like a cross decompilation

1731
02:00:36.498 --> 02:00:40.757
thing happening here, uh, where this was like, uh, just like decompiled and then, and

1732
02:00:40.778 --> 02:00:46.098
then rebuilt in, in, in something, uh, which is absolutely crazy.

1733
02:00:46.618 --> 02:00:50.017
Um, and I think we'll end on the show on this.

1734
02:00:50.298 --> 02:00:51.997
We There's a lot of stuff to t- to talk through.

1735
02:00:52.158 --> 02:00:58.058
Guys, uh, if you missed any part of the show, ThursdAI is live, uh, is

1736
02:00:58.238 --> 02:01:02.197
is happening live on ThursdAI Live, but also, uh, is then released with a newsletter

1737
02:01:02.318 --> 02:01:06.358
and a podcast that you can check out. And also on our YouTube, if you are watching us

1738
02:01:06.398 --> 02:01:09.557
on YouTube, please hit that, uh, subscribe button, uh, because we are releasing

1739
02:01:09.717 --> 02:01:14.297
segments of the show, uh, that, uh, we are agentically building for you with our

1740
02:01:14.358 --> 02:01:16.918
assistants, uh, and they seem to perform well.

1741
02:01:17.038 --> 02:01:21.097
People do, people who don't have time two hours, they, they seem to enjoy those

1742
02:01:21.158 --> 02:01:23.537
segments as well. With us, Peter Goster from Arena.

1743
02:01:23.597 --> 02:01:26.177
Congratulations again on, on the fundraise.

1744
02:01:26.318 --> 02:01:30.797
Yam Peleg Nisten and LDJ. Our friend Wolfram is on a well-deserved break.

1745
02:01:30.918 --> 02:01:36.457
We will be back here next week. If you are coming to AI Engineer in New York, uh,

1746
02:01:36.597 --> 02:01:39.918
please come and say hi. I will be there interviewing a bunch of folks and then be

1747
02:01:39.958 --> 02:01:43.837
back at home, probably doing the ThursdAI, or maybe from New York.

1748
02:01:44.158 --> 02:01:47.677
Uh, and if we missed any part of the show, please let us know in comments if we

1749
02:01:47.717 --> 02:01:49.558
missed anything that's very, very important to cover.

1750
02:01:49.838 --> 02:01:54.858
Uh, the last thing that we didn't cover is Brett Adcock's, uh, Hark Pro, but maybe

1751
02:01:54.958 --> 02:01:57.737
we'll cover this next week. Thank you so much for tuning in.

1752
02:01:57.878 --> 02:02:03.837
Thank you for, uh, checking us out. Uh, please go and check out the Severus GPU, uh,

1753
02:02:03.958 --> 02:02:06.957
that we are offering. It's free now. Why wouldn't you?

1754
02:02:07.018 --> 02:02:08.717
It's super cool. Just put the D1 in there.

1755
02:02:08.998 --> 02:02:11.337
Nisten, you should try it, uh, and then give us feedback.

1756
02:02:11.398 --> 02:02:14.537
We would love some feedback as, as we're like planning this, uh, as a major product

1757
02:02:14.558 --> 02:02:14.858
release.

1758
02:02:15.163 --> 02:02:16.122
Nisten Tahiraj: How, how many you got?

1759
02:02:18.603 --> 02:02:20.042
Alex Volkov: Let's see, let's see if you can break it.

1760
02:02:20.403 --> 02:02:24.083
Uh, all right, folks, thank you so much. With a bit over 2 hours on the show, thank

1761
02:02:24.103 --> 02:02:26.003
you all for joining. Uh, we'll see you here next week.

1762
02:02:26.043 --> 02:02:26.623
Bye-bye, everyone.

1763
02:02:29.163 --> 02:02:30.222
And I have an outro.

1764
02:02:32.803 --> 02:02:33.503
All right.

1765
02:02:35.163 --> 02:02:35.643
Bye-bye.
