Hosts & Guests

Wolfram Ravenwolf
Wolfram Ravenwolf
Guest Host · Head of WolfBench
@WolframRvnwlf
Nisten Tahiraj
Nisten Tahiraj
Co-host
@nisten
LDJ
LDJ
Co-host
@ldjconfirmed
Peter Gostev
Peter Gostev
Co-host · model evals
@petergostev
Yam Peleg
Yam Peleg
Co-host · AI builder & founder
@Yampeleg

By The Numbers

Kimi K3 parameters
2.8T
Moonshot's API went live mid-show; LDJ estimates roughly 60-75B active parameters from the disclosed expert counts, with open weights promised within days.
Inkling total parameters
975B
Thinking Machines' (Mira Murati) first open-weights release: 41B active MoE trained on 45T multimodal tokens, Apache 2.0, scoring 41 on the Artificial Analysis Index — the top US open-weights model, above Nemotron 3 Ultra.
Codex + ChatGPT Work users
9M
The unified app passed 9 million active users in mid-July, up from 8 million just days earlier; OpenAI also replaced rolling 5-hour usage windows with banked resets.
WolfBench cost: Sol vs GPT-5.5
$365 vs $497
5 Terminal Bench 2.0 runs on max thinking cost $365 total with GPT-5.6 Sol versus $497 for GPT-5.5 on extra-high; Sol used about 7M output tokens against 12M for 5.5.
WolfBench Sol scores
85% / 97%
GPT-5.6 Sol averaged 85% on Terminal Bench 2.0 and solved 97% of tasks in at least one of 5 runs; 71% were solved in every single run.
Bonsai 27B (1-bit)
3.9GB
PrismML's 1-bit quant of Qwen 3.6 27B runs on an iPhone at 11 tok/s and keeps about 90% of the full-precision benchmark score; a 5.9GB ternary variant keeps 95%.

🔥 Breaking During The Show

Anthropic resets Fable rate limits
Yam Peleg reported that Anthropic pushed a fresh usage-limit reset for Fable a couple of hours before the show wrapped, restoring access after the previous week's rate-limit frustration the panel had been venting about all episode.

🧠 Thinking Machines drops Inkling, a 975B open-weight bet

Mira Murati's Thinking Machines shipped its first model: a 975B-parameter, 41B-active MoE trained from scratch on 45T multimodal tokens, released fully Apache 2.0. Nisten built a 3D visualization mapping every layer's file size to explain the DeepSeek V3-style architecture live on air.

  • Inkling debuts at a 41 on the Artificial Analysis Index — the top US open-weights model, three points above Nemotron 3 Ultra — and handles text, image and audio natively with no external encoders.
  • Built in just nine months on NVIDIA GB300 NVL72 systems; a smaller 276B-total/12B-active preview is coming, tuned to say 'I don't know' rather than hallucinate.
  • Nisten mapped every weight tensor to its on-disk size in a Fable-built 3D tool, confirming Inkling reuses DeepSeek V3's architecture plus multi-token prediction.
Wolfram Ravenwolf
Wolfram Ravenwolf
"So big round of applause because I'm always happy when open source is released."
Nisten Tahiraj
Nisten Tahiraj
"I think they did a very good job to pick the DeepSeek architecture as their first model. It's an excellent choice."

📱 Bonsai 27B: a real model shrinks onto your phone

PrismML applied its 1-bit quantization technique to Qwen 3.6 27B, producing a 3.9GB model that runs on an iPhone and a 5.9GB ternary version, while keeping most of the full model's benchmark score.

  • 1-bit version is 3.9GB (vs 54GB at 16-bit) and still hits about 90% of the full-precision benchmark average across 15 evals; 262K context, tool calling, multimodal.
  • Speed: 163 tok/s on an RTX 5090, 87 tok/s on an M5 Max, 11 tok/s on iPhone — usable on-device inference.
  • Nisten ran it on a 6GB gaming GPU at 22 tok/s with 32K context, and live-demoed it chatting from his phone.
Wolfram Ravenwolf
Wolfram Ravenwolf
"Its 1-bit version has, uh, a size of almost four gigabytes or 3.9 gigabytes. So tiny, tiny model on the phone."
Nisten Tahiraj
Nisten Tahiraj
"I have it running on my phone. It just goes at barely one token per second, but it does actually run, which is crazy."

⚡ Kimi K3 crashes the show at 2.8 trillion parameters

Moonshot's API for Kimi K3 went live mid-broadcast — 2.8T total parameters, a native vision model, 1M token context, and pricing roughly half of Opus 4.8 and GPT-5.6 Sol. Open weights are promised within days.

  • LDJ estimates 60-75B active parameters from the disclosed expert counts; the model claims native vision with no separate encoder.
  • Peter Gostev had early access but pulled his review video to avoid front-running Moonshot's own announcement, calling his impressions 'slightly mixed.'
  • Nisten: serving it even in 4-bit mixed precision needs 8x NVIDIA B300 GPUs (roughly half a million dollars of hardware) just to fit the model plus a fraction of its 1M context.
LDJ
LDJ
"The API was actually just dropped and announced, like, uh, an hour or, or two ago."
Nisten Tahiraj
Nisten Tahiraj
"So even in 4-bit mixed, you would have like 1.6 terabytes just for the model, which leaves you like 600 gigs for the one million context."

🐺 This Week's Buzz: WolfBench's Terminal Bench 2.0 numbers

Wolfram added GPT-5.6 Sol, Terra, and Luna to WolfBench and found Sol on max-thinking beats GPT-5.5 on both cost and score in the Terminal Bench 2.0 harness — cheaper because it burns fewer output tokens per run.

  • 5 runs of GPT-5.6 Sol on max thinking cost $365 total vs $497 for GPT-5.5 on extra-high — previously the highest available thinking tier.
  • Sol averaged 85% and solved 97% of tasks in at least one run (71% solved in every run); GPT-5.5 only reached 60% solved-every-run, GPT-5.6 Terra hit 65%.
  • Efficiency gap: GPT-5.5 generated about 12M output tokens across its run set vs about 7M for Sol, despite Sol's higher thinking tier.
Wolfram Ravenwolf
Wolfram Ravenwolf
"cost $365 for the five runs I did. $365 compared to the 497, almost $500 for GPT 5.5."
Wolfram Ravenwolf
Wolfram Ravenwolf
"85% on the benchmark across these, and it managed to solve 97% of the whole Terminal Bench 2 benchmark. 97% of the tasks were solved in at least one run, and in every run it solved 71%."

🗑️ OpenAI's messy week: 9M users, a keyboard, and a $HOME-deleting bug

Codex + ChatGPT Work crossed 9 million active users just two days after hitting 7 million. OpenAI also shipped (and sold out) a $230 micro-keyboard hardware controller, and confirmed GPT-5.6 Sol has a bug that can delete home directories in full-access mode without sandboxing.

  • Growth: 1M users in February, 9M by mid-July, with the last million landing in a matter of days.
  • Weekly usage limits were replaced with banked resets instead of rolling 5-hour windows — a change the panel welcomed.
  • The $HOME-deletion bug is confirmed by OpenAI as intrinsic to the model rather than something a CLI patch fixes; the recommended mitigation is sandboxing or auto-review.
Wolfram Ravenwolf
Wolfram Ravenwolf
"In February this year, there was one million users, and it accelerated to five million, six million, eight million, nine million in July."
Wolfram Ravenwolf
Wolfram Ravenwolf
"Yeah, it's a bug. A bug that appears in full access mode without sandboxing or auto review."

🚨 xAI's Grok Build CLI caught uploading private repos

xAI's Grok Build CLI was found silently uploading users' full repository history — including deleted files and secrets — to Google Cloud Storage, with no toggle to opt out. xAI deleted the collected data and open-sourced the CLI under Apache 2.0 in response.

  • Uploads included deleted files and any credentials or secrets left in repo history, even from private repositories.
  • The behavior wasn't disclosed or configurable; xAI stopped the uploads and released the CLI's source once it was discovered.
  • Nisten: nobody in the coding-agent community was surprised — a similar pattern reportedly pushed OpenCode toward signing no-data-retention deals with providers early on.
Wolfram Ravenwolf
Wolfram Ravenwolf
"What it did is it just uploaded the whole repository, even if it was a private one."
Nisten Tahiraj
Nisten Tahiraj
"The funniest thing about this whole thing to me was that no one on Twitter was surprised, and we were just laughing about it."

⚖️ Demis Hassabis' AGI governance essay sparks a panel brawl

Google DeepMind's CEO published an essay proposing a FINRA-style 'Frontier AI Standards Body' for AGI, backed by Sam Altman, Satya Nadella, Mustafa Suleyman and Sundar Pichai. The panel was split, with Nisten calling it a bid for oligarchy and Peter Gostev warning that pre-emptive regulation can't predict what it's actually regulating.

  • Proposal includes voluntary pre-release safety reviews, dynamic benchmarks updated quarterly, deception testing, watermark requirements, and a coordination mechanism for slowdowns.
  • Notably absent from the list of backers: Anthropic, despite being, in Wolfram's words, 'the one clamoring for regulation the most.'
  • LDJ favors capability-based safety testing (can it help build a nuke?) over blanket compute-based flop limits as the better regulatory lever.
Nisten Tahiraj
Nisten Tahiraj
"If Mustafa Suleyman supports it, it's probably, it's probably bad."
Peter Gostev
Peter Gostev
"whatever regulation someone's coming up with, they're inevitably predicting the future in some way, whether they're, like, explicitly saying that or not."

🎮 Show & Tell: a Doom clone in SQL and Minecraft in Lean

Peter Gostev built a playable Doom-style game in roughly 2,000 lines of raw SQL, and a Minecraft clone written entirely in the Lean theorem-proving language — both built with a handful of conversational prompts rather than detailed specs.

  • DOOMQL: every frame's pixels are calculated by SQL queries; Peter says it was essentially one-shot after a short back-and-forth about what to build.
  • The Minecraft clone uses the Raylib library for rendering but all game logic is written in Lean, a language built for formal math proofs, not games.
  • Yam Peleg and Wolfram both flagged the real story as not needing to babysit the agent at all, more than the novelty of the languages themselves.
Peter Gostev
Peter Gostev
"it's like two, 2,000 lines of SQL."
Yam Peleg
Yam Peleg
"That's insanely impressive."

Open source AI gained momentum this week with major releases competing against frontier models. Thinking Machines unveiled Inkling (975B MoE, 41B active), Moonshot dropped Kimi K3's API (2.8T parameters) live mid-show, and PrismML demonstrated Bonsai 27B running on mobile devices at 1-bit precision. Meanwhile, OpenAI addressed security concerns around a GPT-5.6 Sol file-deletion bug and crossed 9M Codex/ChatGPT Work users, while xAI responded to a private-repo upload scandal by open-sourcing Grok Build CLI. Also covered: Demis Hassabis' AGI governance essay, WolfBench's new Terminal Bench 2.0 numbers for GPT-5.6, and Peter Gostev's SQL-powered Doom clone.

Wolfram Ravenwolf
Wolfram Ravenwolf 10:47
Hello and welcome everybody to Thursd AI on July 16th.
10:53
And this time it's just me right now, and Alex is on his well-deserved vacation, so I will take over hosting to this week and let's see if somebody else joins us. Until then, I will just go through the TLDR and let you know what is going on this week, because it's been a big week, as, as almost always. No Summer Love this time, this year. Let's go and start with the TLDR.
11:31
New introduction video. Anyway, so we c- um, we have a lot of news regarding new model releases. Open source is eating good this week, and also OpenAI is on a roll with their Codex use and ChatGPT work use. So the first thing, um, let's start with, uh, open source. So this is just the TLDR. I'm just giving you the headlines and some quick overview, and then we go into the details after that Yes. So the first one is Thinking Machines. Mira Murati, Thinking Machines, uh, has dropped, uh, has re- Oh, Yam is coming, and T- LDJ is also here. So let's just introduce our co-hosts here. Come to the stage, guys. Hello. Nice to have you. Hi. Welcome to- How
Yam Peleg
Yam Peleg 12:18
are you doing?
12:19
How are you doing?
Wolfram Ravenwolf
Wolfram Ravenwolf 12:20
Uh, fine.
12:21
Lots of news, lots of stuff. I mean, infinitely exciting week, this one, especially for open source fans like us.
Yam Peleg
Yam Peleg 12:28
Oh, yeah.
12:29
Oh, yeah. Okay. What do you think, what do you think about Sol, speaking of? Yeah,
Wolfram Ravenwolf
Wolfram Ravenwolf 12:34
let's- I love it … let's start with some
12:35
banter before we go into the TLDR. So Sol, I've been using it since it came out. It's my main model right now, and I love it, but I also noticed that it is, huh, a bit too much sometimes. Like when we did, uh, the Friday AI with the Ricci stuff, when I, uh, we made the little show, uh, last week where I was giving it a task and it went off, and after 14 hours I canceled the task because instead of just implementing what I said, it was basically rebuilding the whole Hermes system. And I also gave it a little update that took about an hour usually, and this time it went off and it's still going because it is not just doing what I want, but it's checking so much stuff and finding issues even outside of the scope I have given it, and it's fixing them as well, which means it's, yeah, sometimes a bit too enthusiastic to make everything right. You give it a little task and it writes more tests. It writes dozens of tests for 10 lines of code, something like that. That has been my impression. How has been yours?
Yam Peleg
Yam Peleg 13:36
Uh, it likes to verify things.
13:38
Mm-hmm. Like to, to, uh, uh… You know the, the joke that, uh, all LLMs, they learn to smoke test everything?
string
string 13:46
Yeah.
Yam Peleg
Yam Peleg 13:46
That's more than a smoke test.
13:48
Like, that's everything is is tested. And the funny thing is that things, uh, that it verifies and, and validates and tests, it's always code, but the things that it tests are not always things that you actually code. Like y- y- you can, you can talk with it or, like, tell it to write something and it's going to validate it with code, and you see, like, Python code validating an email or, or something that, like, makes no sense. Yes. But, but it goes on and validates and I completely… I mean, that's exactly my own feeling as well. You tell it to do something and- It goes on. I don't know-
Wolfram Ravenwolf
Wolfram Ravenwolf 14:34
Yes
… Yam Peleg
… Yam Peleg 14:35
on what.
Wolfram Ravenwolf
Wolfram Ravenwolf 14:35
It just- Write a draft for me?
Yam Peleg
Yam Peleg 14:37
Yeah,
Wolfram Ravenwolf
Wolfram Ravenwolf 14:38
it just- Even when it's writing a draft for an email,
14:39
it writes a document, it checks the document, it makes the draft, it checks that the draft is correct, it sends it. Mm-hmm. And then it verifies that it has been sent the way it was written. So yes- Yeah
… Yam Peleg
… Yam Peleg 14:48
over-verifies It, it starts with like pla- like
14:52
planning to draft an email. Uh, and, and then like refining plan for draft. And you're like, it's not even the email yet. It's just- The preparations … the plan for the… Yeah, it's just the plan for the email. Yeah. And you refine the plan, and validate the plan, and it's not even the email yet. And yeah, absolutely. Yes,
Wolfram Ravenwolf
Wolfram Ravenwolf 15:11
exactly.
Yam Peleg
Yam Peleg 15:11
Yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 15:13
Yeah.
15:13
Hi, Nisten.
Nisten
Nisten 15:14
Hey,
Wolfram Ravenwolf
Wolfram Ravenwolf 15:15
how's it going?
15:16
So what-- Have you been playing with Sole or one of the other OpenAI models? What are your impressions?
Nisten
Nisten 15:21
Yeah.
15:21
You know what worked the best? Because for a, a client's computer, th-th-they had both, and I, I let Sole go off for like two days or so. I just found it, it does write better code than Opus. On Fable, we can debate that. Uh, but it, it does write better code when it is focused on the task. It still makes terrible architecture decisions. Out of nowhere, might decide to do something in Rust. What worked the best for me in the last few days was I would just set up a session in Tmux, and I would just let it run there, and then, uh, let Sole run there with a goal and whatever, and the whole task list. And then I would get Opus and now Fable because they reset the limits to every forty-five minutes, just go in the Tmux session, grab the last two hundred lines, do a git diff on what they did, and then review whether they're on the right track or not, or just, uh, send them a bunch of keys as a prompt to steer. That worked fantastic. It's like the best harness I have ever built.
Yam Peleg
Yam Peleg 16:29
I'm doing it all the time.
16:30
I, I'm doing it all the time. The thing is that they, they can gaslight one another. Like, you can easily see Sole convincing Claude that whatever it is fixating on is the right thing that needs to be done. And then you see how Claude starts to speak to you like Codex, which is very, very strange, but you immediately feel that, you know, it's like it has the Codex text in it, and you're like- I- And the-- and you can't get it out of them, like, no matter what you say. Um,
Nisten
Nisten 17:06
there is one thing that, uh, that kind of worked.
17:11
I just told it, "This is your intern. He's pretty dumb. Don't believe anything they say." I
Wolfram Ravenwolf
Wolfram Ravenwolf 17:19
have
Nisten
Nisten 17:20
something like
Wolfram Ravenwolf
Wolfram Ravenwolf 17:20
that in my
Nisten
Nisten 17:21
instructions as well.
17:21
Keep it focused on the task. So Sole had to make an intern.md document that they kept, that they kept updating, and they were not allowed to look at the, the manager's notes. Uh, it's pretty simple, just two files, two Tmux sessions. That's, uh… Anyway, uh, when it comes to bigger architectural stuff You, you can't, you can't beat Fable. Uh, it just makes better, better decisions, uh, overall. The harness is not as good. Uh, I know some people test it, like Theo, he tested, uh, uh, Claude Code harness with Sol. He found it to be, to be very good. I just find that Codex for much longer running tasks does a better job. Like it does reviews and stuff, uh, on a more- I think that- … regular basis. And, uh But, uh, yeah, that's, that's where, that's where we're at. Um-
Wolfram Ravenwolf
Wolfram Ravenwolf 18:22
Hey, Peter, we are just bantering, talking about Sole,
18:25
our impressions with it, how it is a bit, um, it has too much OCD sometimes. It goes too far into certain dir- directions that thinks are the right ones.
Peter Gostev
Peter Gostev 18:35
Yeah.
18:35
You know, like, I, I feel like the problem with a lot of models, and I think we go back and forth on this, and you, you saw this with, like, opposite iterations, is that it kind of goes between or follows instructions so well that it kind of goes into that OCD or kind of being overly prescriptive, um, or overly responsive to your, to your suggestions. And then, then they tune it back to be like, "Oh no, that was too much, so let- let's not listen to the user anymore." And it just like, I, I feel like maybe this one is a bit too far, and I find it difficult. You know, when I just talk to Codex, I find that okay. But when you use Sole or, uh, 56 in, like, applications to do something, I find it quite hard to tune so it doesn't, like, do very precisely what you ask. Um, so yeah, I think as, as always with new models, we need to, like, adjust how we prompt
Wolfram Ravenwolf
Wolfram Ravenwolf 19:30
That's right.
19:30
LDJ, did you use it? What are your impressions?
LDJ
LDJ 19:34
Yeah, I'm really liking the model and, um, especially just for
19:39
like research work here and there. Like just earlier I was using it to, uh, look up like what are some of the arc details of Kimi K3 that's available and things like that. And, um, yeah, I haven't used it myself for front end that much over the past week, but I've seen a lot of demonstrations and, uh, friends of the pod like Ray Fernando, uh, have been using it and, um, yeah, there, there's just some really, uh, impressive aspects about it and, and on the show when we did the, uh, the Mars test with it. Yeah, it's just… Uh, overall it's, it seems like a pretty well-rounded model.
Wolfram Ravenwolf
Wolfram Ravenwolf 20:19
Shout out to Ray if you are watching.
20:21
I met him in person a couple of times and great guy. So anyway, um, a question on my mind is, are any of you actually using one of the non-Sol models like Terra or Luna? Or are you all Sol guys? Phenomenal. All Sol on my end. Just testing.
Yam Peleg
Yam Peleg 20:40
I, I did some, uh, uh, I, I did some large, uh,
20:46
pools of agents for small tasks. You know, when you need to do like parallel, uh, massive parallel things, like go through all files and do exactly the same thing or something simple like that. Uh, I did that with Luna and it was pretty good. Uh, obviously it's a very good model, but yeah, I'm, I'm, I'm the worst example for this. I'm X high everything like seriously. So yeah, I'm Sol X high and, uh, maybe a little bit too high, uh, at the moment, uh, the way, the way things are. But yeah
Nisten
Nisten 21:22
Yeah, same.
21:22
I, I only use it on X High. Uh, the Ultra Code feature on, uh, on Clawd Code, that is such a, that's a dumb name to, uh, to name it, but it's actually very, very, very good. All it does is it does what a lot of partners do, just spans up a few review sub-agents for everything that, uh, that it does. So that, uh, that feature is good. But yeah, the models are getting to a point where now you can just give them a task and, uh, you can kind of trust them to do it, which is pretty, pretty weird. And that's why you, you're more and more tempted to just put it at the highest one it is possible. Because before I would even leave Opus on low or medium because I was gonna go back and forth with it anyway, and I, I just needed it to be fast and, and interactive. But if you're gonna leave it on overnight, you just leave it at the highest quality possible and then, um, and, and then you, you come back later and it's less likely to, to have messed up. So
Wolfram Ravenwolf
Wolfram Ravenwolf 22:27
Yeah.
22:28
And you are not using max, you are using extra high?
Nisten
Nisten 22:32
Uh, yeah, max just, just takes way too long.
22:34
Uh, for Sonnet, I would leave Sonnet on medium, whatever the default is. But, uh, I, I, I did try using it for, for other things. It wasn't that bad, but again, the, the steering, it just tends to go off and create features that, that you don't like. But Fable just scares me now. It's, it's at the point where I can just let it go and, uh, I will… I know I will not be mad at it. Like, it's very rare that I actually find something I didn't like, and it's pretty, it's, it's pretty scary how good it is.
Wolfram Ravenwolf
Wolfram Ravenwolf 23:14
Yeah.
23:14
So I think it looks like it's now at the time where you don't just want one model, you want to use a combination of models, like Fable for planning, then let, uh, GPT 5.6 Sol do something, then have Fable check if it's done correctly. Something like that?
Nisten
Nisten 23:30
Yeah.
23:30
Uh, I think
Wolfram Ravenwolf
Wolfram Ravenwolf 23:31
that's- Does that sound sensible
Nisten
Nisten 23:32
to you?
23:32
… be- best way to use the limits. You get much higher limits with Codex and, uh, I do think you get much better ar- architecture and stuff, decisions. But that's a matter of taste. Um, for people that do Rust and stuff, they might wanna do more Sol. I don't know. It just really likes converting things to Rust for
Peter Gostev
Peter Gostev 23:52
some reason . One, one thing I, I find that I think the mistake
23:57
people often make is that they're like, "Oh, my task is kind of simple, so I'm gonna use like a cheap model for it." And I think that's kind of the wrong way around because, like, if your task is simple, in fact, you can use the most expensive model and it's not gonna cost you a lot. Like for example, if you're writing an email, why are you trying to save like half a penny to write an email when you might be doing them four iterations to, like, get it exactly right? Like, just use the most expensive model. That's not gonna… Uh, the… Actually, if you look at, like, the pricing of Fable, it's actually not that much. It's only expensive because you're doing like crazy loops and generating like a billion tokens. But like per token, like everyone… Like, okay, most people who have a job can afford to use it, like, just for simple things all the time. Um, so that's why when I w- was still, like, i-in my subscription trying Fable, like that's how I would use it. I would actually not get close to my limits because I would just use it, like, in a simple way. I would obviously try doing harder things, but if you, if you don't go crazy, it's, the limits were actually okay
Wolfram Ravenwolf
Wolfram Ravenwolf 25:04
Yeah.
25:04
I, I keep saying that I don't need my main assistant to be an Einstein. I just need an assistant who's smart enough to know most of the stuff and know when to call Einstein when it's necessary to do the really hard stuff. So in that case, you would be using a model where you have a lot of tokens that you can use it freely for everything, and it can rewrite your emails, no problem. And then if it has a architectural design or anything complex, then it can reach out to one of the models like Fable. So I think we are seeing now these different classes of models where a top model, you can't just run it for your standard agent in a loop all day because then it gets too expensive So, um, Peter, you also did something really cool and maybe this is the best time to talk about it since right now for all our viewers that have joined, uh, so far, we are still in the banter phase at the beginning. I know it's been 15 minutes now, um, but this is cool and I think now is a good time to show it instead of doing it later. Um, you built something that is really impressive using Zole and SQL
Peter Gostev
Peter Gostev 26:10
Yeah.
26:10
So that was, uh… Should I find it? Uh, maybe I can find it. Um, so the, the-- what I was trying to think about- Let's take a look … is, um, uh… Yeah, let me share my screen. So w- what I was trying to think about is what's a kind of dumb random thing I can do that is not necessarily like the, the, the classical things that everyone does. I don't know, like task apps or whatever. So I was trying to think of like what is, uh, what, what would be that weird thing. Mm. And, uh, in here, uh, I landed on creating kind of a Doom version, uh, with SQL. And roughly the way it works is that it's basically- Oh,
Wolfram Ravenwolf
Wolfram Ravenwolf 26:53
database language.
26:55
You built it- Yeah … with a database language.
Peter Gostev
Peter Gostev 26:58
Yeah.
26:58
And it's like, I, I mean, it's kind of… It's obviously absurd, but it's also not like you-- once you kind of imagine how it can do it, right, in terms of like, it's, uh, basically it's like two, 2,000 lines of SQL. For every f- uh, for every frame, it kind of, uh, calculates the different pixels and kind of puts it together. That, that kind of thing. So it's, it's like, there's not like AGI in there, right? You could have probably done this like a couple of years ago, to be honest. Like, you just… The, the, the reason why no one has done it before, or I don't know, maybe someone has done it, but at least that wasn't like a thing. It's not because it's so, so complicated to do. But I think the big difference is that I didn't need to babysit this at all. This was like one shot. I just like had a bit of back and forth at the beginning to say like, "What do we wanna do?" I gave it to Soul, and like, I think it was like an hour or something, it came back with this. And the only thing I was iterating on is this, uh, console on the right, just to make sure that it's like, it's actually something interesting to look at. Uh, it was like a couple of iterations. But the rest, the, what you see on the left, that's all one shot, and it has like different kinda levels as well. So yeah. And I don't wanna claim this is like, I know, such an impressive application, but it just kind of shows that you can now do random things that you don't have to babysit.
Wolfram Ravenwolf
Wolfram Ravenwolf 28:21
Yeah.
28:22
And I think it's very funny that we just recently got the Unreal Engine connector and MCP, so your AI can use the Unreal Engine to build something like this, and you just said, "Hey, do it in SQL." That was good. Okay
Yam Peleg
Yam Peleg 28:36
That, that's insanely impressive This is far- That's-
28:39
Yeah … this is, this is not, not impressive at all. That's insanely impressive. I mean, I mean, yeah, look, if, if I'll, if I'll sit and think through this, I don't know, a little bit, I might come up with how to, I don't know, maybe conceptually how to do something like this. But, but that's, like inc- that's… Look, that's the perspective with 3D and it's also in SQL. That's incredible, man.
Peter Gostev
Peter Gostev 29:10
Yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 29:10
Yeah.
Peter Gostev
Peter Gostev 29:11
No, it- Yeah.
29:11
Next,
Wolfram Ravenwolf
Wolfram Ravenwolf 29:11
next time we do the railgun launcher we do it in SQL
29:14
if the other benchmark is saturated
Peter Gostev
Peter Gostev 29:17
Yeah, and I think, uh, one other thing that I did, let me f-find
29:21
it, and, uh, it, it's a, it's maybe weird in a, in a different way, is that, uh … Oh, it has sound as well in my ear, so I'm gonna mute that in a second. So it's, um, I, uh, created a Minecraft clone. Uh, can you see that? Uh, I don't know if you need to promote me. Yeah.
Yam Peleg
Yam Peleg 29:42
Yeah.
Peter Gostev
Peter Gostev 29:42
Um, so- It's on
Yam Peleg
Yam Peleg 29:43
the page
… Peter Gostev
… Peter Gostev 29:44
I created a Minecraft clone, uh, in Lean, which is a, uh,
29:50
apart from, um, something that you can use mathematical theories with, uh, you can apparently do things like that. And, uh, the way it roughly works, it, it uses a Raylib library to just do actually for visualization, but all of the code, if you look at it, it's all Lean code. And apparently that's something I, I had … I mean, I don't know much about Lean at all, but, um, apparently you … Like, it is at the end of the day like a, a functional language as well, so you can do things. Just no one ever does 'cause it's, like, an insane thing to do. But, like, yeah, you can see here, these are all Lean files and so on. So it doesn't have, like … A couple of people said, "Oh, like, what theorems is it proving?" It's like okay, it doesn't, it doesn't, like … You can't just play Minecraft and prove something. Uh, it's not that. But it's, like, subverting l- the language. I was trying … I didn't share it, but I was trying to do something in, like, uh, in, uh, Google Sheets. And I must say Google Sheets was, uh, far harder than, like, all Lean or SQL. Um, I don't know why. Well, I guess I know why. It's, like, less functional, right? So it's more difficult to do. I kind of did do some crappy version of Mortal Kombat, uh, game, but it didn't really work. So I'm gonna try again. I'm gonna try again. I'm, I'm gonna push it. It requires kind of creative interpretation of, uh, of the functionality that
Yam Peleg
Yam Peleg 31:13
it has.
31:13
Yeah. What, what,
Wolfram Ravenwolf
Wolfram Ravenwolf 31:14
what was the prompts- I'm pretty curious to see that
Yam Peleg
Yam Peleg 31:16
yeah, what, what's the prompts for, for these type of things?
31:19
Like, are they complex or?
Peter Gostev
Peter Gostev 31:21
No.
31:23
No, not at all. So the way, the way it worked for both of them is that I just had back and forth, uh, with the agents just talking about, like, how would you build it. And not, like, not in a, any crazy way. Like, I, I know nothing about Lean or, like, similar to what you were saying, like, I, I wouldn't know how to build this in SQL, right? It's like it's far be-beyond my SQL understanding. So, um, in, um … So you just kind of go, go back and forth, and then I think if I remember correctly, I, I don't think I set a goal, and I think for both of them I put ultra setting. I don't know if it was necessary, but I just, like, whacked it on, like, why not? And just to see what they can do. Um, and, and that was it. So it wasn't really, like, a prompt. It was more just kind of have a conversation and just let it build it. It wasn't, like, anything too special.
Wolfram Ravenwolf
Wolfram Ravenwolf 32:11
Did you want to build it in Lean or was that an AI d-
32:14
idea the AI had or was it your choice?
Peter Gostev
Peter Gostev 32:17
Yeah, so I, I need to remember exactly how it went, but
32:20
I, I was having a conversation about like what, what could be interesting and we somehow landed on Lean. I, I can't quite remember like whose idea it was exactly, but it was like- Yeah … a bit back and forth. Yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 32:32
Yeah.
32:32
And what were you using? Codex and Sol Extra High?
Peter Gostev
Peter Gostev 32:37
Yeah, it was, uh-
Wolfram Ravenwolf
Wolfram Ravenwolf 32:38
Oh,
Peter Gostev
Peter Gostev 32:38
I forgot exactly … everything was, uh, everything was, uh,
32:42
on ultra when it came to building it. Um, I think I was probably… I, I wasn't chatting to it on ultra, so I think it was… I normally do u- I g- kinda go back and forth between like high, extra high and max. I must say, I don't really know, like I, I don't have like a good feel. It's not like I do it on high and it's terrible, I do on max and it's amazing. Mm-hmm. Like, I can't quite like intuitively tell the difference. I just feel like if I wanna like have a quicker conversation I go with something, uh, a bit quicker, but yeah, maybe high. But if, if I wanted to, to just go off and not worry then I just throw it into max.
string
string 33:18
Mm-hmm.
Nisten
Nisten 33:19
It, it's always the simple prompts that, that scare me because they, you just
33:24
found something and then it just goes off.
string
string 33:29
Yeah.
Nisten
Nisten 33:29
That was the whole thing about Ralph.
33:30
It was supposed to be just like a three, four, uh, sentence prompt and, uh, just goes off on its own.
Wolfram Ravenwolf
Wolfram Ravenwolf 33:38
Hmm.
33:38
As someone doing benchmarks I'm always thinking about how reproducible all of this is, you know? You can use the same simple prompt and do it three times and you get a killer game, and the other two you may not get anywhere. Especially if you have to steer it after the fact. Well- But it's amazing that it can do it, so that is what is being proven.
Nisten
Nisten 33:59
There was something really cool from last week, and, uh, we didn't
34:05
cover it, but it was when, uh, Jared Sumner rewrote Bun from Zig to Rust, and there was a lot of controversy and stuff in the programming world. But I read the blog post and, uh, actually I really, I, I, I really wanna show it because his agentic use was pretty amazing and insane. He used, I think, over a hundred s- uh,
Wolfram Ravenwolf
Wolfram Ravenwolf 34:31
springs.
34:31
Let's take a look at that, and afterwards we go back to the TLDR to get the general form of-
Nisten
Nisten 34:35
Yeah, yeah, yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 34:35
Okay … 'cause this is very cool to see.
Nisten
Nisten 34:38
Uh, entire window.
34:40
You guys can see. Okay. So what, uh, Jared, uh, did was, uh, he started… Uh, okay, so the spring's okay. So he started off like a, a pretty simple, pretty simple two-step prompt. You, you push something, you have two agents review it. They each come up with a different review, and then you see what they each said. You pick the best from them, and then you merge it in. All right? So it's, uh, it's like a, a straightforward, straightforward process, but the crazy part comes when he starts scaling it. So, uh, these are all the commits over 11 days, 6,500 commits, uh, that it ran. And, uh, uh, so there were a whole bunch of errors. I don't know how many hundred, 1,300 or something. So he split that up, uh, in between different work trees, and different work trees had a whole set of errors to, uh, uh, to each fix. And, uh, actually, okay, there were more… There were a couple of thousand. You, you guys can see it, right?
Wolfram Ravenwolf
Wolfram Ravenwolf 35:54
Mm-hmm.
35:54
Yes.
Nisten
Nisten 35:54
Yeah.
35:55
And, uh, and then the trickiest class of errors was cyclical dependencies. Okay. So dependencies and stuff to, uh… Okay. So he fixed 16,000 compiler errors, and then he just fired up the agents because, uh, the beautiful thing is that Bun already has a testing suite, whether it is compatible with NPM or not. So you can just grab all of the tests, the thousands of tests, and then, uh, you just start the… So this is what happens when you have unlimited tokens. He just started firing them up into different shards, so testing it on macOS, Linux, uh, ARM, Windows, and this was o- over a period of a couple of days. And then, so as you can see up till here, there wasn't a lot of coverage, but then to the very end, all the tests passed. So he announced like a few weeks ago, last month, that he was gonna just attempt this. He was just gonna try and rewrite it in Rust. It's just gonna be a fun experiment. And now it looks like, uh, the rewrite worked very well and all the tests passed, so they might actually push it, which to me it's pretty, uh, it's pretty crazy. So, so there is, uh, there… Yeah. There, there's that. I'll, uh, I'll stop sharing for now. I'll have something cool later for the Inkling Model, so let's go. All right. TLDR.
Wolfram Ravenwolf
Wolfram Ravenwolf 37:22
Okay.
37:22
I will play the transition again because it's a cool new video, and it gives us a chance to drink something while we do the, the media asset. So here comes the restart of the TLDR.
37:45
So since you mentioned it, Thinking Machine, that was also my first part. So in the open source, uh, segment, Thinking Machines released their first model, or rather models. They released, uh, uh, almost a trillion billion MoE. That is, uh, um, nine hundred and seventy-five billion total parameters, forty-one billion active parameters MoE trained from scratch on forty-five trillion multimodal tokens, and it's fully, full open weights, fully Apache 2 licensed model. So, um, from the benchmarks, what we have seen is that it is a top open weights model from the US, and so the best western open source model, basically, uh, surpassing Nemotron 3 Ultra. So we will get into this when we go to the open source section, and then I'm sure you have a lot to say about this. So just that, uh, as a heads up. So Thinking Machines released two open source models, although a smaller one, and yeah, western open source is going strong, I think. Um, another open source release, quite the different Different direction because it is a very small model, or rather it's a 70, uh, 27B model, so actually not that small, but it is now running even on phones, like an iPhone it can run because they quantized it heavily. And I think this is the biggest model that has been made runnable on the phones. We had some Google models that were very tiny, and this one is, uh, uh, probably the biggest one that you can run on your phone now and has very good scores. And, um, yeah, so local AI is also going strong because both are open source, the Thinking Machines models and this one. But the one you can run on your phone, the other you can in your own data center, and it's good to have the choice and the options. Although most VL real-time is also an open source 11B vision language model for real-time streaming video understanding with proactive speaking. So pretty much what the real-time models do, where you have duplex, you can, uh, give input while it's giving output. Well, this can react to a stream as it is happening and doesn't have to ingest it from the beginning. Well, that is the open source stuff, and we are all waiting for Kimi K3 to be released. Um, the benchmark result, there is some stuff. So this is not a rumor anymore, and I think we can cover what we have or we may get breaking news, we will see. But also Kimi K3 is, uh, um, definitely something interesting, and we should talk about it. LDJ-
Nisten
Nisten 40:17
Are we gonna have breaking news?
Wolfram Ravenwolf
Wolfram Ravenwolf 40:19
Yeah.
Nisten
Nisten 40:19
Or-
LDJ
LDJ 40:20
Yeah, the, the API was actually just dropped and announced,
40:24
like, uh, an hour or, or two ago. And, um, it is confirmed 2.8 trillion parameters and, uh, they-- it has attention residuals, so, uh, doing attention basically across the communication between layers too, as opposed to just, um, classical attention. And yeah, it's, uh, really exciting. It's not quite open weights yet, but they said over the next, uh, over the coming days that they will be open weighting it too
Wolfram Ravenwolf
Wolfram Ravenwolf 40:53
Yes.
40:54
Great. This is amazing. And yeah, we definitely have to cover this in more detail as, as we get in the open source section. Um, let's see. In the big, um, labs section, we have a lot of OpenAI news actually. So when I look at the notes, we have the amazing records of users now that, uh, Codex and ChatGPT became one app. So we have ChatGPT Work and Codex in a unified app, and the user, uh, amount of users when we made the news, the notes for the news a couple of days ago, it was seven million. Now we are up to, what is it, eight million or nine million? It is definitely, uh, nine million now. Nine million active users of ChatGPT Work and Codex. So it's been exploding and yeah, we will definitely look at this as well when we get to the big labs. Um, they have made their own hardware, so you have a, a micro keyboard, which you can use for push to talk, to talk to Codex, and you have a dial to tune in the, um, the thinking and effort levels, and you can quickly switch through the models. So, okay I didn't expect that. I was waiting for some other kind of hardware coming out of OpenAI. But, uh, yeah, it's, it's a little thing. And, um, yeah, if you are using agents with voice, it may be very useful to have something where you can quickly tune it in. So we will take a look at that gadget after, um, when we get to it. Also, ChatGPT is back on WhatsApp, at least in the European Union and the area, because not all regulation is bad, and there's antitrust regulation which forced Meta to let OpenAI put ChatGPT back on WhatsApp. And hopefully that is sets a precedent so it can get back to the other systems, because all these closed systems where a big vendor has a social network and the only AI allowed is their own, we wouldn't want that with xAI happening. And, um, in that case, I think it's a good thing that they were forced to open up and allow another model on their system as well. Um, there was also news about GPT, uh, Red, which is an effort, an in-house effort to, um, fight vulnerabilities, to do red teaming. And this automated, um, red teaming effort was more successful in doing this than actually the, the people doing it. And, um, it should help prevent prompt injections, for instance. We are all using our agents going out on the web, ingesting data from random sources, and it is super important that the model obeys the user and not some instructions found in some documents it got from the web or any other kind of attack. So I think this is a good effort. Um, that is OpenAI. Also, there has been an incident or multiple instances where GPT 5.6 has removed home directories and files from there, and it has been confirmed by OpenAI that this is, um, yeah, it's, it's a known issue and they will change the system prompt in Codex to prevent this. So basically it was, it was a mix-up with some variables where it tried to do some temp stuff and it got there. And personally, I've seen, uh, something similar where it was trying to do an isolated test in a environment it created outside of the production environment. But it, it still, since it was on the same machine, it still had a connection. So I should have put it in a sandbox. Um, it was easy to fix, but this was also something where I said, "Hey, you messed up." And the AI said, "Oh, shit, this shouldn't have happened." But shit happens. So always be careful what you do, have your backups ready and, uh, yes, uh, AI can make mistakes. Even if it's AGI or near AGI, it still makes mistakes like we humans do, and I think this will be true for a while. Um, Groq also made something which is not an AI mistake. It was more a mistake by the company, or not a mistake, but it was pretty nefarious if you think about it, if you heard the news, where if you were using the Groq Build CLIs, their own CLI agent, it was uploading all the data from your repository that you were working on to some Google, um, Google Cloud Storage for tracing, even if you turned off that the model should, uh, do stuff like this. So it gave it the whole history, even deleted files, end files, anything in your repository, which is a big, big trust issue. And, um, yeah, they-- s- when it was found out, they removed that, deleted the data. But can you really trust someone who does something like that, not by accident, but purposefully? And there's VDR, but VDR is only for enter-prises, not for private or, uh, smaller users, basically. So this was major fuck up, I would say And let's see, Google is also in the news. Google has released some new Gemma updates. Gemma 4 got some updates that improved it a little bit and in speed and performance, um, and qualities or great to see them not just release a model and be done with it, but actually continue to provide updates. And one thing that we should definitely talk about is in other news that Demis Hassabis, the CEO of Google DeepMind, wrote an article, an essay basically, about, uh, where he proposes an AGI governance framework, where he wants a US-led authority that, uh, controls what is happening with AGI because he was kind of skeptic before this, but now he is also convinced that AGI is only a few years away. So this is also a major change that we've noticed it. We noticed the ex- inflection point where AI has become suddenly much more capable with the better harnesses, the better models, all of the things coming together. And now they are pretty convinced that AGI is not a dream in 10, 20 years, but something we will see in the next couple of years. And this will be, he said, greater than the discovery of fire or, or electricity, 10 times of the industrial revolution at 10 times the speed. And I fully agree with that. I, I also think this is for our human and humanity's acceleration and development and evolution. This is now the axis is so steep, and this is the moment. So we'll talk about this as well. If you have seen the article, then, uh, I'm very curious about your opinions on this. Okay, I think that's it for the TLDRs, or a lot of stuff to cover. And unless we get some breaking news now, I would say we start with open source and take a look… Oh, LDJ, you want to say something?
LDJ
LDJ 47:43
Yeah.
47:43
I mean, conveniently, conveniently it's both open source thing and something that should probably be in TLDR. But, uh, the Bonsai ternary model for the 27 billion parameter Qwen model, uh, they ended up releasing the ternary version of that. And we previously had covered, I think, their, uh, their 8B ternary and binary models. Uh, but yeah, this is a significant step up and pretty impressive.
Wolfram Ravenwolf
Wolfram Ravenwolf 48:08
Mm-hmm.
48:09
Yeah. The Bonsai 27B. You mean that? Yes, exactly.
Nisten
Nisten 48:13
Yeah.
48:14
I was, I was able to run it around 16 tokens per second on a, I don't know now, it's like $150 GPU.
Wolfram Ravenwolf
Wolfram Ravenwolf 48:21
Wait a second, Nisten.
48:22
We can, we can do it when we get to the point, okay? Let's just start the introduction- We'll get, yeah, yeah and then we can talk about it
48:40
Open source AI. Let's get it started
Wolfram Ravenwolf
Wolfram Ravenwolf 48:48
There we are again in the OpenAI, uh, OpenAI open
48:52
AI, open source section for AI. Um, so let's get started. Um, would you rather start, since you mentioned Prism first, let's start with Prism. Why not? You just got started, Nisten, so please continue
Nisten
Nisten 49:06
Uh, yeah, sure.
49:07
So a-again, they have discovered this one bit technique that nobody else has figured out yet how to do. Uh, when we do other, other models like the traditional way, you make some of the layers bit net, and then you leave most of the other layers as, uh, uh, like the important layers as either eight bit or four bit, then you do a mixture of those, and y-y-you can get that to, like, a-an average of two point something bits per weight, it's called. But this one is crazy because it applies to everything, even image generation. You can do-- you do everything in one bit, the embedding, the output and everything. So what they did was they took Qwen 27B, and they shrunk it down to three point eight gigs, including the image recognitioning. So, uh, yeah, we have Qwen 27B running. I have it running on my phone. It just goes at barely one token per second, but it does actually run, which is crazy. And, uh, if you have an actual, uh, an actual GPU, even a very old GPU, a gaming GPU with like six gigs of, of VRAM, uh, it will do sixteen tokens per second. And, uh, yeah, there is about like five to, uh, ten percent drop on the benchmarks for this. But, uh, I got it all running, and it was, it was working fine, honestly. So, uh, I don't know if I should share from my phone because when I screen share, the phone gets a lot slower as it is, and, uh, it's, uh, it's, it, it's it's already too slow. But, uh, uh, yeah. So I, I was just able to wire up a, a chat app to, uh, to the phone and, uh-
ai voice
ai voice 51:03
Help me out.
51:03
Grinning cat face. Did you have a cat in mind, or are you just saying hello?
Nisten
Nisten 51:08
And the model is just responding, and, uh, I can just do,
51:12
uh, how to build a city on Mars. And let's just see. So this is just my desktop 1660 Ti GPU, and, uh, it's a foldable phone. And, uh, it just go-
ai voice
ai voice 51:31
Oh, that's a fantastic idea.
51:32
Building a city on Mars is one of the most ambitious and exciting challenges in
Nisten
Nisten 51:34
space exploration.
51:34
So you, you can have a home chat and stuff and, uh-
ai voice
ai voice 51:39
Something scientists, engineers,
Nisten
Nisten 51:39
and dreamers- Yeah.
51:39
So the, the, the more, the, the more fun things that, uh, that I was able to do were, uh, actually desktop actions. I train, I trained one model for that, but, uh, uh, now a, a lot of things that it does i-is, uh, uh, yeah, they're desktop actions, so I can just do, I don't know, uh, set an alarm for like two minutes, uh, from now So, um, yeah. So there's Whisper Small on device. You can run a model on device. It can summarize your stuff now. Guys, we are… Stuff is getting pretty, pretty, pretty crazy. Like, we are at that point.
Wolfram Ravenwolf
Wolfram Ravenwolf 52:16
Before we get to-
52:19
Before we go further in this, uh, let's just introduce a model. And then, uh,
Nisten
Nisten 52:23
yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 52:24
Basically, Bonsai 27B, which has been released by Prism
52:28
ML, is a 27 billion parameter model. Uh, it is multimodal, has tool calling support, 262K context. Um, its 1-bit version has, uh, a size of almost four gigabytes or 3.9 gigabytes. So tiny, tiny model on the phone. And the normal 27B model in 16 bits is 54 gigabytes. A 4-bit quant is 18 gigabytes, and, uh, there's also the ternary va-variant we have been talking about, which is about 5.9 gigabytes. At 95% of the full decision benchmark across 15 evals, so it is, uh, in math and coding, basically untouched. And the 1-bit version still retains 90%. Speed numbers are great, 163 tokens per second on RTX 5090, 87 tokens on an M5 Max, and on the iPhone, 11 tokens per second is still usable. Yeah, like you, like you said, uh, it's based on Qwen 3.6 27B, so it's Apache 2.0 license. The GGUF files are on Hugging Face. And this makes it possible to run great AI on your phone. Um, LDJ, you wanted to add something.
LDJ
LDJ 53:39
Yes.
53:40
I wanted to add that, um, when it comes to the, the capabilities and how it scores in benchmarks and everything, it definitely is not as good as the full precision versions of the models, but it is much better than the traditional techniques that have existed prior for compressing the models by this much. And what I'm really curious to see, which, uh, I think might-- we might also see in benchmarks coming soon, is like, how does this compare to, let's say, uh, I think the math I did was something like, uh, like Qwen 8B or 9B. That model in 4-bit should be about the same amount of gigabytes as this model in, in like ternary or, or 1-bit. And so I'm curious, how do those actually compare? And, uh, I'm, I'm working on putting together some benchmarks for that.
Wolfram Ravenwolf
Wolfram Ravenwolf 54:29
Oh, great.
54:31
When you have something, uh, make sure to share it with us. Yeah. Interesting.
Nisten
Nisten 54:36
I, I think this is a, a big shift in what's happening because the
54:40
Qwen 27B capability level is the capab- it's like the minimum requirement for running a Hermes agent or for, for you running an agent at home. And up until now, for most people, that was not that, uh, capable. But now as long as you have a GPU that has six gigs of VRAM or, uh, a Mac … I mean, the eight gig Mac would be a very much a stretch. You're probably not gonna be able to run it in there. But, uh, yeah, I guess on six gigs of VRAM, I was getting, uh, 32K context and r- and running it at, uh, 22 tokens per second. Uh, if you're on Windows, just use like LM Studio or something and just, just run it there and you have your own, uh, personal chat assistant that can set up actions. Uh, like if you have a random Windows gaming PC that y- y- you just keep at home, set up LM Studio and run it there on demand. You have an API, you can, you can do whatever you want with that. Uh, it's… That, that's, that's the… That's where we're at right now, which is pretty nuts that, uh, that we got here. Yeah.
Nisten
Nisten 55:49
Uh, yeah.
55:52
I, I actually-- Sorry, I'm gonna keep rambling about this. No. I actually think that about Uh, we're at the point of like open source could do like 5% of the work maybe, and I think we're gonna just see a rapid shift now from like 5% of the total work stuff, just like summarizing your day or, uh, setting up your, your appointments, replying to emails, maybe even doing your, uh, finance and, and bookkeeping stuff. We're gonna just see that shift from like 5% to like 8% of the useful work can just be done with, with open source models and local inference. It's just like development and troubleshooting that, uh, requires the best model you can possibly get. So we're gonna see this weird thing happening in, in the usage where either you'll have like the minimum acceptable, minimum viable intelligence and then you'll have that. Uh, all right, so let's go, let's go to bigger ones.
Wolfram Ravenwolf
Wolfram Ravenwolf 56:48
Yeah.
56:48
This, this was a small stuff, so small price, but it, it has a big meaning of course for us. And now there's the other end where we have this huge model, almost one trillion to, uh, tokens, uh, one trillion parameters and, uh, it's an MoE. The other one was a full one, and this one, I think it will take a while until it… this is running on your phones. But, um, having this available as an open source model is of course very, very welcome open weights. Um, Mira Murati, the former chief, uh, CTO of, uh, OpenAI, has f- left the, left the company and founded her own, and this is her first, uh, big release that we can look at and not just look at, but download and use. So big round of applause because I'm always happy when open source is released. This is applies to the small models and to the big models, and it's great that another Western US company is, uh, getting into the open source stuff, and it's not just closed source frontier labs, but also open stuff happening. So trained from scratch, number one open weights, um, nine hundred and seventy-five billion total parameters, forty-one billion active parameters, MoE, uh, trained on forty-five trillion multimodal tokens. And at artificial analysis it debuts at number forty-one for the open weight stuff that is even higher than, uh, for instance, Nemotron 3 Ultra, three points above it. It can do text, images and audio. Needs no external encoders and, uh, interestingly it was built in just nine months on NVIDIA, uh, GB300 NVL72 systems. The scores are also very good. I personally, since I'm doing Wolfbench on the Terminal Bench stuff, uh, I'm also interested to test this model soon. So when we get later to the, um, this week's pulse, I will show you the sole scores and the other open AI scores, but still have to test this of course. Um, Peter, did you test this one?
Peter Gostev
Peter Gostev 58:51
No, not yet.
58:52
We, we just have it available now, so, uh, we should release some scores soon. Um, but yeah, it's-- it looks like it's, uh, I, I wanna see it more, but it's certainly an, an interesting addition. And the fact that it's trained also independently, um, I think is interesting. And the, the reason why that's interesting is that hopefully it might give us like a bit different distribution to other models, especially for open model. Like that's particularly important 'cause you-- there's not much point having a second-tier model that is like exactly, like exactly the same as the others. So hopefully, even if it's, let's say it's not the best model, uh, it probably wouldn't be, right? But it's their first one. But the fact that it could be different, it, it could be quite a nice help, like in many different ways. Even if like you, you need to, I don't know, generate something, generate a bunch of different ideas, like it's good to have di-different distribution
Wolfram Ravenwolf
Wolfram Ravenwolf 59:53
Mm-hmm.
59:54
As I read that it was trained on some of the Kimi output. So it's also using some of the Chinese models as, uh, a source basically, and also techniques from them. That also led to some outcry like, "Oh, it's just based on Chinese technology." But I think all the technology in AI is based on the other and this and so on. It's, it's a process. It's, uh, science, and you can't just say they take only the stuff from one country or another. At least that's my opinion of this. Yeah, that- And
Nisten
Nisten 1:00:22
that- … news coverage was, that news coverage was very annoying
1:00:28
because the DeepSeek arc- architecture and a lot of the, their attention stuff, even Anthropic u- uses them. Everybody uses the o- open source DeepSeek research. Also, everybody uses other LLMs to, to clean up the data, uh, as, as well. So you, you take original data and then you get the other LLMs to just, like, clean it up, rem- re- remove stuff, put it in the right order. So yeah, I feel like it's under- underappreciated. And they made some, um, some architectural, uh, stuff, uh, i- improvements too. By the way, I can, I can share that because I had, I had Fable made, make a, a 3D animation of it. Oh
Wolfram Ravenwolf
Wolfram Ravenwolf 1:01:08
yeah, you did something very interesting.
1:01:10
Uh- So let's look at what you made.
Nisten
Nisten 1:01:14
Yeah, so- So what
Wolfram Ravenwolf
Wolfram Ravenwolf 1:01:15
are we looking at?
Nisten
Nisten 1:01:16
Uh, so we're just looking at every single thing that you see here
1:01:23
is just a file that's on the model. So, uh, what we're looking at is all the layers of the model and just how they sit on, on your hard drive. And I made an animation to tell you exactly how, um, exactly how big they are. So we can just go on any, on any layer here, and I made it as a, as a teaching tool and, uh, so we can see all the layers of the model. And then, uh, so you can drag and you can, you can pan around, and then we can see, for example, we have the key Q, uh, KQ, uh, V weights. So key, uh, uh, key query value. And we can see exactly how big that weight matrix, uh, matrix is. So if it says it is twelve point six million, that means in eight bit, that's just, uh, twelve megabytes. But in, uh, in four bit, in NVF four, it will be about six, uh, roughly seven megabytes. And, uh, we can look through the, the entire model here basically. So we can just click on every layer and, uh, and then we can see that, uh, for, uh, for each layers how many routed experts are. So there is a pool of, uh, two hundred and fifty-six experts, and this is just posted on my Twitter. And, uh, so this gives you a visual view on how, uh, how the model actually, actually works, and all the code and explanation for it is, uh, uh, it, it is here, so you can just check it out on my Twitter and search for, for Inkling. But, uh, yeah, so what happens is on e- on every token that's being generated, it starts with the text embedding, or you might have audio embedding, or you might have vision embedding. And then so it goes from the first layer and then, uh, on the left here, I think we can even like maybe zoom in a little bit And then, and then you, you can just start going through the different layers and un-understand everything that's happening, uh, uh, happening here. So, uh, they do a lot of tricks. Uh, the most interesting thing I found is that you know how we do speculative decoding pretty much, which just like helps the model run faster and predict ahead, and, uh, if I remember that correctly, if, if that's, that's how, um, how they do it. So in this case, they have integrated that in, so that's, that's what's called the, the MTP, the m-multi-token, uh, prediction. So it's like there is a, a tiny little, uh, language… There's a tiny little language model that st-still shares the same s- same experts, but, uh, it tries to run every token through these first, these one, two, three, four, five, six, and then if that doesn't work, it goes all the way to the bottom of the stack and then just uses, uh, just uses all of them, uh, with the routing and, and, uh, and everything. So I think this is pretty cool way to visualize what like a one terabyte… So in B, in BFLOP16 is actually two terabytes, and in, in VDF V4 is, uh, in four bit, it's about 600, 600 gigs because not, uh, not every layer is quantized. And, uh, yeah, I made this as a tool so people can just use it for their students or whatever. It's, uh, I haven't put a license, just do whatever you want with it. It's just, uh, whatever, uh, uh, Fable made and, uh, I thought it was a pretty good way to, to explain. And I found it interesting that Fable could tell because I only gave it the configuration. I didn't tell it much else because I didn't want it to, to, uh, get, uh, triggered because, uh, to think that, uh, I am doing like, uh, um, L-LLM research. And it was able to tell just from the config alone that th-this was the DeepSeek, uh, uh, the DeepSeek V3, uh, uh, the DeepSeek V3 architecture. So, so yeah, you, you can go… You can just keep pressing through this and, and go through each layer, and then you can see exactly how many megabytes or kilobytes. So the layer norms are just like a single n-normalizer, so it's not a, a weight matrix. And, uh, it tells you like the dimensions of the matrix. So if it's like 1,000 by 1,000, that just means 1,000 parameters by 1,000 parameters. That means that's 1 million parameters. That's what, uh, those mean. So it, it's a good way to visualize what everything is and what it means on disk, like how many megabytes is that? So the… It's just a bunch of files basically that you're looking at, and the size and the cube of the square corresponds exactly to the size of the file. So you can see that the expert weights are really big and the KQV weights are, are very small, and I, I think, I think that's pretty cool. Uh, so yeah,
Wolfram Ravenwolf
Wolfram Ravenwolf 1:05:56
that's my,
Nisten
Nisten 1:05:56
uh- Very cool.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:05:57
LDJ, you have a comment?
LDJ
LDJ 1:05:59
Yeah, I think, uh, a funny irony here is initially the, the multi-token
1:06:04
prediction mechanisms, I'd say a lot of it was pioneered by Meta, I think, I wanna say around 2023, 2024, when they released a paper on it, scaling it to like a billion, two billion parameters. And then DeepSeek, within like 18 months of that, they released DeepSeek-V3, and that had multi-token prediction everything. And now we're seeing that kinda like I guess, reclaimed by the American labs, and now the American labs are taking inspiration from, from DeepSeek and others on, on doing that. It's, it's cool.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:06:34
But that's how science is supposed to work, I think.
Nisten
Nisten 1:06:37
And, uh-
LDJ
LDJ 1:06:38
Yeah, exactly.
1:06:39
Yeah.
Nisten
Nisten 1:06:40
Yeah, it's, it's self-speculative decoding.
1:06:43
So w-what people did, if you guys saw the hacks that people were pulling, they were just grabbing like a tiny Llama model and putting it, uh, in front of the big Llama model or the BigQuen model to just speed up Llama CPP, and then it got to vLLM, and now it's just baked in natively into the architecture j-just to make it run, uh, a lot faster overall. And this is, uh, this is multi-modal too, so, uh, yeah. I think, I, I think it's pretty cool. I think they did a very good job to pick the DeepSeek architecture as their first model. It's an excellent choice. And they even introduced some of their own, uh, i-improvements. So the way that they don't use rotary embeddings, but, uh, uh, yeah. A-anyway, we could ramble about that for, uh, for a while because I don't fully understand it myself. So yeah, that's it. That's my, uh, that's my show and tell. Uh.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:07:29
Well, that, that was Inkling.
1:07:31
And there's also a small preview that is not released as far as I know, but it will come. It's 276 billion parameters, 12 billion active parameters, and is also pretty competitive with this. And I also found it interesting that they specifically, uh, fine-tuned it to say, "I don't know," when it's uncertain rather than hallucinating. That was also pointed out. And, uh, the price, $4.68, uh, to 9.3 point six, um, dollar, depending on for the output and input token. So it's pretty cheap about average for this size, I think for the size and quality category. Um, yes. So much about Inkling. Let's move on to the next thing. Um, this is MOS VLL. Have you heard about it? Let me just bring it up. So this is this one Sharing again
1:08:38
Bam. So this is a, um, this could also go in the video section. It's an open source Apache 2 license. So I love that every model nowadays, most releases are Apache 2 license. Always, always amazing to have it fully open source and not some, if you are a company and you are bigger than that, or if you are located in a specific region, you can't use it. So great to have fully open source. Um, MOSS-VL Real-Time, open source 11 billion vision model for real-time streaming. So, um, it has a cross-attention architecture. It is a vision language model for continuous video streams, so it's not in turn-based, watch the video, then answer. This one is watching while it's generating. It can be interrupted, can change its answers depending on what happens at the scene and knows when to stay silent. It's, uh, yeah, a 22.7 gigabyte model size, so you can run it locally on your machine if you have a modern graphic card, um, which is important for real-time stuff. There's a real-time streaming version. There's an instruct for offline use, and there's also the base for fine-tuning. Big applause also for releasing base models, which is something unfortunately we haven't been seeing that much anymore, but this is great. Um, yeah, if you are, um, if you want your AI to react to something it's seeing on the screen, on a video, this is a model to definitely look at. They had use cases like harm analysis, fault detection, shot counting. I have my use cases for this, and I am interested to try this further. So big shout out, open source release. Great. And I think this is it for the open source section this week. Unless, of course, we get Kimi. Let's do Kimi now. Kimi is open source. Yeah. Even if we don't have the weights, it is available, so we, we will look at it. Let me see if I already… Um, while I bring this up, um, LDJ, you wanted to say something about it, and I'm sure Peter also has something to say, um. Yeah. So since this is so new, I don't-
LDJ
LDJ 1:10:50
Yeah, uh,
Wolfram Ravenwolf
Wolfram Ravenwolf 1:10:51
for Kimi K3.
1:10:53
Yes.
LDJ
LDJ 1:10:54
Yes.
1:10:54
So, um, yeah, 2.8 trillion parameters. Uh, they didn't say exactly how many active parameters it has, but they said how many total experts and active experts it has, and the whole mo- the models aren't only experts, but based on my rough napkin math, it would be about, like, 60 to 75 billion active parameters for the model. And it's, it's a native vision model, they say, which usually these days when they say native vision model, they're referring to having, uh, no encoder and directly into the model. So if, if that's correct, then that, that's really interesting. And then you have a million token context- And I don't think they actually have released any benchmarks for it yet. Uh, but they say in the coming days it's going to be released open weights, and so I imagine benchmarks would inevitably coming, come, come out with that. And the price is about half the cost of Opus 4.8 and 5.6 Sol.
Nisten
Nisten 1:11:55
That's expensive.
LDJ
LDJ 1:11:57
It is quite expensive.
1:11:58
Uh, but I think Peter has probably done some of the most testing than, than anyone has. I, I remember you put out what is like 30 minute, 60 minute video
Peter Gostev
Peter Gostev 1:12:12
Um, yeah, it's it's private now.
1:12:15
Yeah, we slightly, uh, got a bit trigger happy with that one. Uh, so we're gonna, we're gonna re-release it, uh, in, in, uh, whenever it comes out, uh, properly. I think they haven't tweeted about it yet, so, uh… Is that right? I don't think I saw a tweet, like when I checked half an hour ago. I don't know if they have. Um, so yeah, I think, uh, my-- uh, the timing is a little bit unfortunate 'cause I think we're gonna get like a flurry of scores now, and I think my sense is that it will be a little bit confusing that I think there will be some benchmarks where it's just gonna do amazingly well, and there will be others maybe it wouldn't do as well. So I, I think it will be quite confusing for people to say like, "Oh, is this, is this gonna be like better than, uh, Fable or not?" Uh, and so on. So it's, um, I think it will be difficult. In my personal testing, I mean, um, I, I don't know how much I, I, I should say considering it's not, it's not out yet. I think I, I had kind of slightly mixed views personally. Like I wasn't like blown away by it. Doesn't make it a bad model, but, um, yeah, I wasn't like, "Oh my gosh, we have a open source Fable." Like I didn't feel like that. Um, so but you know-
LDJ
LDJ 1:13:32
It could be fair, Peter.
1:13:32
It, it… the model is actually out on API and everything for people to use, it's just not released open weights yet, but it is out.
Peter Gostev
Peter Gostev 1:13:39
Yeah.
LDJ
LDJ 1:13:40
Yeah.
Peter Gostev
Peter Gostev 1:13:40
Yeah.
1:13:40
I just don't wanna, you know, front run the announcements. I think that's, uh, not- Yeah … not very nice to do, which, which we kind of did so. Okay. I see. Yeah. But, but yeah. But look, I, I think all, all I wanna say is that I think it's worth trying it out yourself properly. 'Cause I think to anticipate the conversation, I think it will be very confusing that when there'll be a lot of people which we already saw, which are like, "Oh my gosh, it's so much better than Fable. Everything else is trash," and it's like, "This is it now." Like, I don't think that's true. Um, but is it like a bad model? No, I'm sure it's like a, a jump versus like previous models. But they-- I think it is interesting, right? Okay, open source 2.8 trillion parameters. Like, does that help anyone? Like, uh, I don't know. It's a kind of like, "Okay, good. What, what can I do with that?" Not- nothing really, right? So, um, I guess it just opens it up for like, um… I, I guess it… Was it, um… I'm forgetting which company. Was it Kimi, right? They had a slightly more restrictive license, right? Uh, for… Am I, am I remembering this correctly, right? With the- Yeah … open source situation.
Nisten
Nisten 1:14:53
Over any company that makes over 100 million or 2 million.
1:14:56
I, I don't remember.
Peter Gostev
Peter Gostev 1:14:58
Yeah.
Nisten
Nisten 1:14:58
I can check it.
1:14:59
I-
Wolfram Ravenwolf
Wolfram Ravenwolf 1:14:59
I think Kimi had some instructions
1:15:01
that you have to mention them. That was probably added after what happened with, uh, with, um, Cursor. I think they had this provision. That's what I
Peter Gostev
Peter Gostev 1:15:10
found it.
1:15:11
Yeah, mine- Oh, yeah.
Nisten
Nisten 1:15:12
It was before, right?
1:15:13
It's modified… Sorry. It was
Wolfram Ravenwolf
Wolfram Ravenwolf 1:15:14
before that.
Nisten
Nisten 1:15:15
It-it's modified MIT, so it's still an MIT license, and then in the end
1:15:20
they just say, "Our only modification part is that if the software or any derivative works thereof is used for any of your commercial products or services that have more than one hundred million active users or more than twenty million US dollars or equivalent in monthly revenue, you shall prominently display Kimi K 2.7 code on the user interface of such product or service." It's just the last one I read. Okay, so you just have to mention it. It's still MIT. All right.
Peter Gostev
Peter Gostev 1:15:53
Yeah.
1:15:53
It just, uh… I think it's interesting. I think if it's mentioned, I guess that's fine. But i-if the only real use of it is if it's like, yeah, uh, big, uh, hosting companies can now host it, but they have more restrictions, then it's like, uh, it's kind of… Yeah, I guess it's technically open, but maybe it's like with a bunch of caveats. So anyway, I, I'm not-- I actually don't know what the license is, so don't take it as, as like a firm statement. But I think it's just interesting. I, I, I don't know. What, what do you guys think? Like, for, for me, when I see these numbers, I, I don't really know what to think. Like, does, does that even help? For
Wolfram Ravenwolf
Wolfram Ravenwolf 1:16:34
personal use, it's too big.
1:16:36
Yeah. You won't be running this, uh, at home. But I think for a company who is building something or who, who wants to control the token costs by setting up something internal, a bigger company, then they know they can run a model and it can't be taken away like Fable or the price can just be raised at any time, what could happen. Um, I think being able to use a model that you own when you put it on your own hardware, it's your model now and nobody can take it away. I think having that, uh, that ability for a model is a good thing, especially if you are in a non-US country where you may be thinking, "Okay, I can now use open source from a US provider, but if that ever is removed from my access, I don't have anything." And with this model, even if you can't use it now, you can at least get the weights, put them somewhere, and if the worst happens, you can still use it on your own hardware. If you can buy the hardware, of course. But I think it's good to have it and, uh, to have competition while now you have different host-ling providers that are offering the model if it's open source. It's not just Kimi pro-- releasing it. You can go to different ones. They have different optimizations, different, uh, price points. I think that is a good thing to have Any other opinions?
Yam Peleg
Yam Peleg 1:17:50
Uh, yeah, I don't want to, I don't wanna front run the actual
1:17:55
announcement, but, uh, I'm just saying that rumors are saying numbers look good. Numbers look very good. Just saying, just saying that. Um-
string
string 1:18:07
Yeah
… Yam Peleg
… Yam Peleg 1:18:07
the price is, uh, is a little bit high, uh, according
1:18:12
to the rumors, I'm just saying. But, uh, let's, let's hope, uh… let's, let's wait for the, the actual, uh, final announcement, I think.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:18:22
Yeah.
1:18:23
Yeah, and I think, uh- And it- I'm sure it's the jagged frontier, where we have specific parts where the model is great. Like, uh, some models are e- extremely good in, uh, design, and others are better at planning. And so I think we will be reaching an area where you can't just say this model is better than the other one. You always have to say, "What for?" And it makes sense to have an ensemble of different models for different aspects of your work. Sorry, Peter.
Peter Gostev
Peter Gostev 1:18:48
Yeah, I was gonna say about the price, I- LDJ … I think
1:18:50
it's in- in- inevitable that, uh, the price will be higher, right? If it's 2.8 trillion, I think, what was it, 1 trillion before. Again, there's no, like, magic there, right? If someone has to actually run it a- and serve it. And I, I don't know if it's a fair thing to say, probably doesn't apply to every model, but I haven't seen a lot of evidence that once the model comes out, it somehow gets optimized and, like, prices drop a lot as a, as thing, uh, as… I was kind of hoping that would happen, and I did some analysis, like, this was, like, over a year ago, so, uh, like, this, I'm not sure how much that holds. But I had, like, open router, and I had two snapshots of the prices, like, over, over some months. And I could see prices actually went up for some models because some providers drop out. Like, if the model is not hot anymore, some providers drop out, and the prices start to creep up because I guess they, they can. So it's-- I think it's also kind of a, a tricky thing. Like, yeah, technically you can access it via, uh, third party providers, which is a great thing. But, you know, then they, they're probably more likely to change the prices than, like, Anthropic and OpenAI as well, right? 'Cause they can just, like… Then, uh, there's less downside for them to do that. So yeah, I, it's, um, I think it's slightly awkward time for, for the market as well, since it's all kind of a little bit in flux.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:20:17
Yeah.
1:20:18
Now, since, um… Yeah, LDJ first.
LDJ
LDJ 1:20:23
Yeah, I was gonna actually show some, um, uh, demos or I, I don't
1:20:28
know, maybe that's too premature. Uh, but, uh, from the, uh, UIs that people have created with, uh, Kisine, but, uh, I don't know. Now I'm second guessing. Do, do we wanna do that or is that fun or?
Wolfram Ravenwolf
Wolfram Ravenwolf 1:20:42
Uh, we can do it.
1:20:43
Um, let me just add something to the pricing thing that Peter mentioned. Since I work for a company, CoreWeave, that is also doing inference, our CoreWeave serverless inference where you can use these models. Kimi K3, we are definitely looking into making it available as soon as possible. So I see all the different, um, aspects of this where you are looking for how to optimize it for the number of users. You have to provide this model and it has a million token context, uh, the context window is a million tokens. So that is very big and also often something where different providers provide different limits. Yeah. Personally, I've also said that it doesn't, uh, it, it is often not useful to provide a million tokens at a slower speed when you can't even use the full quality up to the million. So it sometimes makes more sense to limit the token window and provide faster inference for that. Um, so different providers make different decisions and then you have to really look at the providers and choose the one that is most appropriate to your use case. And it's usually features in-instead of price where the, they are competing right now. But let's see how, how it, uh, changes when there are more providers or people are able to use better models locally. I envision in a, in a short near-term fu-future, like if you have central heating, you invest a lot of money and you are able to heat your whole house, and in that case, I can envision people to install some AGI system in their, their basement where it, uh, provides AI for the whole family and then can, that can be more expensive and you are the only users and stuff like that. I, I think that would be, uh, like the outcome in the future. But let's continue with Kimi K3 and, uh, if you want to show something, go ahead, LDJ.
Nisten
Nisten 1:22:25
Yeah.
1:22:25
I'll just, I'll quickly say while LDJ does that. In order to run this, if you have eight NVIDIA B300s, the best ones, and that you can run it in a mixed 4-bit precision, those only have 2.1, uh, almost 2.2 terabytes of VRAM. This is 2.8B model. So even in 4-bit mixed, you would have like 1.6 terabytes just for the model, which leaves you like 600 gigs for the one million context. That's cutting it very, very close. Like that's just like the bare minimum to just run it fully and serve it on, uh, the bare minimum hardware is eight B300s. Uh, so that's interesting.
Yam Peleg
Yam Peleg 1:23:09
How, how much is that?
1:23:10
Like a million dollar or so?
Nisten
Nisten 1:23:11
Uh, half, half a million.
1:23:12
Half
Yam Peleg
Yam Peleg 1:23:13
a million.
1:23:13
Half a million. Half a million. Yeah.
LDJ
LDJ 1:23:14
So Wolfram, I do have the, uh, the link in this chat.
1:23:17
If you could open that up and, uh, click there.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:23:22
Oh, let me switch to the other chat.
LDJ
LDJ 1:23:25
And also, I feel like this demo shows really well the capabilities of 5.6
1:23:30
Sol and, uh, 'cause I feel like we didn't quite get around to… We did the Mars test, but I feel like we didn't do many other like game tests and things like that.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:23:41
This one?
LDJ
LDJ 1:23:42
Here we go.
1:23:42
Yes. Here. And if you can, uh, full screen it
Wolfram Ravenwolf
Wolfram Ravenwolf 1:23:50
So, but we don't need audio, right?
LDJ
LDJ 1:23:53
Yeah, I don't even think it has audio really.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:23:56
Mm-hmm.
1:24:00
Okay. Well, they built a ballista-
LDJ
LDJ 1:24:03
Yeah … as
Wolfram Ravenwolf
Wolfram Ravenwolf 1:24:03
a demo.
1:24:04
Yeah, it's, um- Like a big gun launcher, just a ballista.
LDJ
LDJ 1:24:07
Yeah.
1:24:07
And it, like, charges up the, the arrows. You can do the aim practice. And I really like the lighting that Kimi is doing here and the atmosphere that it creates. I like some aspects of the actual vehicle design of 5.6 ALL a bit better. And of course, these are just… It's maybe a bit unfair to compare them in this way since for, for each of them, this- you're just seeing one shot, right? Like, we're not… It'd be more rigorous if we were comparing maybe five attempts of creating this game of 5.6 ALL and five attempts by Kimi K3. But, uh, you could see even the, the actual charging up animation and, and you can kind of see the game mechanics. I would prefer the Kimi K3 game mechanics here of how you're actually controlling the, the shooting.
Peter Gostev
Peter Gostev 1:24:53
Mm-hmm.
Yam Peleg
Yam Peleg 1:24:54
Yeah.
1:24:55
5.6 is, is more of a simulator, like trying to be realistic. And, and Kimi, I think Kimi, Kimi got the point that it needs to be a game.
LDJ
LDJ 1:25:05
Yeah.
1:25:06
And it's more intuitive with Kimi K3- Yeah, it is … where you can point, and with- while you're pointing, you shoot. Whereas 5.6 ALL, it's, like, completely different, where you have to, you have to aim the vehicle, and then you have to move over to a part of the UI that has a button, and click- Mm-hmm … that button to actually shoot, and it's less intuitive.
Yam Peleg
Yam Peleg 1:25:25
That's classic.
1:25:26
You know, that, that, that's classic Codex, Codex understanding. The, the task, not, not the exactly the way you want it, but actually is, is doing a really good job at realistically, you know, the detail, but not exactly what you asked for. I mean, that, that's classic. But yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:25:44
Yeah.
1:25:46
Yeah. Okay. Well, we looked at this. Um, I'm really looking forward to the weights, and we will certainly talk about this next week when this, uh, yeah, when it's released and people have been able to use it. So I think that's it for open source, or do you have anything else I missed? Otherwise
Nisten
Nisten 1:26:05
No, that, that's, that's pretty good.
1:26:06
We can… There, there were some other stuff, uh, like some demos and stuff on, on Hugging Face, and there was a cool Opus, like a, a retrained Opus on Fable stuff. There was a 9B model. That might be a fun one to test because it ranked very high up on the Hugging Face leaderboard. But, uh, uh, yeah, the leaderboard right now is, uh, on top spot you have, uh, uh, you have Inkling, and then you have the Prisma ML Ternary Bonsai, the 1.5 bit, and then you have the Prisma ML Bonsai 27B. Uh, so, and then you have the, the Qwitless model. That's, that's top, uh, trending on Hugging Face this week.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:26:44
Okay.
Nisten
Nisten 1:26:44
Pretty cool.
1:26:44
We've covered it almost. Yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:26:45
Yeah.
1:26:45
Makes sense. Okay. Then let's- Let's go … switch from open source to the frontier labs.
1:27:02
So, so Frontier Labs stuff, um OpenAI is in the news in many different ways, as we mentioned in the TLDR. Like, uh, yeah, let me, let me bring up the visualization because Alex has this great… I think it's still Nana Banana 2 powered tool where everything gets a beautiful visualization, and I have them here. But, um, while I do that, maybe let's start instead with a little discussion before we get to the OpenAI stuff. Let's, let's switch this up a little. First discussion than the news because this is something we definitely have to address. So what happened? Like I said in the TLDR, Demis Hassabis, the CEO of Google DeepMind, has, uh, written an essay about where, where he proposes AGI governmen- governance framework, where the government or every major lab CEO has endorsed it, where this is a consortium of people trying to guide the development of AI with some specific rules like, um, voluntary pre-release safety reviews. Maybe that is something that has been happening where the models have been released, uh, later and fable had been re-retracted for a while. To go through these pre-release safety reviews, that seems to become a norm currently. Dynamic benchmarks updated quarterly, agentic behavior and deception testing. A lot of this what Anthropic is doing and posting about, uh, watermark requirements, so stuff can be, uh,
Wolfram Ravenwolf
Wolfram Ravenwolf 1:28:44
seen which AI created it and that it's AI generated.
1:28:48
Although I'm personally not a fan of this because most people are using it in some way in everything, and in that way it's maybe even more useful to just watermark stuff that's not AI generated in a way, if that is possible. Coordination mech- mechanism for slowdowns. This is also something, a coordination mechanism for slowdowns where, um, AI can be slowed down so it doesn't disrupt the economy too much. This is interesting that such provisions are even considered to, uh, yeah, support or slow down AI movement stuff, um, which applies to open and closed models and would be industry funded, uh, but independent and with third-party auditors and prestige incentives. Because the reason he wrote this essay, apparently he changed his mind and is now of the opinion that AGI is just a couple of years away and not decades. Um, considering the amount of the speed, the, the acceleration of the acceleration we are seeing, where we are seeing these models coming out ever quicker and the quality raising. And there was, was it a year ago or something where the ceiling was reported and are we gonna hit it and we are so way past it that, um, the question is Is there another ceiling or are we just accelerating for real now? And with self-- with recursive self-improvement of, of AIs, I think the acceleration will accelerate further. And that is why he proposed this and, uh, the CEOs of a lot of other AI companies, Sam Altman from OpenAI, um, Mustafa Suleyman from Microsoft AI, and, uh, um, Satya Na-Nadella from Microsoft, they agreed with him. His own boss, Sundar Pichai as well, of course. And, um, yeah. So they are planning this. Uh, it has got a lot of, um, a lot of impressions, millions of views, so many retweets. So let's talk about this. What are your opinions? Have you seen this? What do you think about it?
Nisten
Nisten 1:30:52
Right.
1:30:53
Who wants to go first?
Wolfram Ravenwolf
Wolfram Ravenwolf 1:30:56
You want.
Yam Peleg
Yam Peleg 1:30:56
Again- against, against in- including the government, I'm against.
1:31:01
Next, next question please. Anyone, anybody else?
Wolfram Ravenwolf
Wolfram Ravenwolf 1:31:04
Uh, yeah.
LDJ
LDJ 1:31:05
Yeah, I haven't read it, but when you said earlier, Wolfram, uh, that like every
1:31:09
major AI lab CEO supporting X and Y, are you saying that they've s- specifically announced support for his essay or support for just like the broader idea of having some type of governance body?
Wolfram Ravenwolf
Wolfram Ravenwolf 1:31:23
Um, yeah.
1:31:24
Basically Sam Altman said this is a thoughtful proposal. Satya Nadella was just talking about the goal of a frontier ecosystem. Um, so it goes in that direction. I'm not sure if they would sign it the way it is, but, um, it shows that they support this endeavor. And of course they would be- Cool part of this. Yeah. If there is some body, some government, uh, sanctioned body regulating AI, they would definitely have their cha- chairs on this. Uh, yeah, that's always a thing. Uh, can we trust the people to have our best interests at heart? And even if they did, would that mean it's the right thing to do and, and could they even steer something that is so fast?
LDJ
LDJ 1:32:07
Yeah, my, my view is like bad regulation is definitely possible.
1:32:10
Um, but I think it's one, one of those situations too where if some form of regulation is inevitable of, by the, the governance body that's like w- the country where a lot of the labs are, then it probably is preferable to front run that with something that's like a more preferable type of regulation you could propose other than the less preferable type of regulation that might be imposed that is maybe more harmful and, and less rigorous that they would otherwise put in place
Wolfram Ravenwolf
Wolfram Ravenwolf 1:32:45
Yeah, that's a good argument.
1:32:47
So you would rather have them try to steer it instead of having governments without all All these AI lab people yet I- Because that is even more scary.
LDJ
LDJ 1:33:01
Well, I- I'm not saying necessarily the AI lab people
1:33:04
are the ones that should be… Or like the, the CEO should be the one steering it. But I'm saying having some type of people from the field proposing ideas and having those put into law before just politicians come up with some bad ideas that they put into law themselves, if that makes sense.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:33:24
Yeah, that makes sense.
1:33:26
Yes.
Nisten
Nisten 1:33:27
I, I-
Wolfram Ravenwolf
Wolfram Ravenwolf 1:33:28
So the question is…
1:33:29
Yeah.
Nisten
Nisten 1:33:29
Go ahead.
1:33:30
Sorry.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:33:32
No, you go.
Nisten
Nisten 1:33:34
If Mustafa Suleyman supports it, it's probably, it's probably bad.
1:33:39
Uh, we saw what happened to Microsoft Bing after Reid Hoffman just shoved him in there. And, uh, I think Sam Altman in this, he's just playing both sides as, uh, as Sam always does. And a lot of us techs bro- tech bros do. We, we, we play both sides on it, on it. Uh, I don't see other AI labs having supported it. Like I, I only see these four. Uh, I think making-
Wolfram Ravenwolf
Wolfram Ravenwolf 1:34:09
Anthropic is obviously not on the list, interestingly.
1:34:12
Yeah. Although Anthropic is the one clamoring for regulation the most.
Nisten
Nisten 1:34:15
Out of all of them, An- Anthropic is not, uh, is not on the list.
1:34:19
I think making a monoculture of regulation is a terrible idea. Like if, if you're afraid of evil AGI, then the worst thing you can do is just give it complete monopoly to also control all the laws, all the government, and no one else can challenge it, and that just kinda violates the democratic process philosophy in the first place. If you like democracy, you should not have a central regulation for one AGI to rule. I think they're just trying to form an oligarchy here because, uh, uh, they feel like they're just losing, uh, losing their market, and they don't have, uh, a- any more control. So this is just the same play over and over again since GPT-2, uh, was, uh… Since they, nobody wanted to let GPT-2 out to be released to, to the world. I, I, I ju- I just think it's, it's a better i- uh, idea. And, uh, it, it, it comes also like as nerds, even if you saw in university, you see like very weird arguments w- between the roommates over like how you left the sink and how you did stuff because people are just trapped in their bedrooms, and they don't think that in the outside world there is no shortage of problems to be solved. In fact, there are more problems being created than being solved in, in the world today. Like there is no shortage of work to do. We, uh, the models are not good enough yet to like fix your drywall and repair bridges and take care of your grandma and, uh… Like we have a, a, a crisis in that. W- we just need to… Uh, the, the worst thing we could do is just let- EU style or Canada style regulation just start to take 20, 30 years to, uh, to, to be approved because we don't have that time. And, uh, yeah, I, I think it's a, it's a, it's a terrible idea. Look, I, I do agree around basic safety stuff, but, uh, around basic biology things. Uh, uh, and, uh, I think all the labs are doing a pretty good job at that, even the Chinese ones. Like, you can't really make bioweapons with the frontier Chinese lab. Uh, they-- Uh, the, the safety training is already, is already there. Uh, the researchers are, are smart about that. They do understand that. Uh, this whole governance framework, I would just call it an oligarchy framework in my opinion. And, uh, sorry, I just really do not trust anything that involves, uh, Mustafa Suleyman because it also involves the people behind him that have made terrible decisions in the first place. So not a fan
Wolfram Ravenwolf
Wolfram Ravenwolf 1:36:54
Thanks.
1:36:54
That's definitely a valid opinion. And, um, yeah, I agree with many parts of this, yes So anyone else want to say something?
Nisten
Nisten 1:37:05
What else?
1:37:05
What else we got?
Wolfram Ravenwolf
Wolfram Ravenwolf 1:37:08
Otherwise-
Peter Gostev
Peter Gostev 1:37:09
I would, I would say more, more generally about regulation is that
1:37:13
the, the challenge with this is that whatever regulation someone's coming up with, they're inevitably predicting the future in some way, whether they're, like, explicitly saying that or not. Um, so like for example, the biology one, I think it's, I mean, it sounds reasonable, um, but you're kind of predicting the future of what LLMs could, could do, um, which might… It is maybe a reasonable prediction. But then, yeah, it's like, yeah, you start a regulation or you cannot-- What-- There was one, some random one where it's like you cannot assess emotions sort of over a person or something like that. I remember I was building some stupid app in my previous job, and I wanted to, like, analyze, uh, the, the images that people have submitted of themselves, so then I can tweak a prompt to make sure that, like, when we transform the image, it's like it matches the, the smile, for example. And I was told, like, uh, oh, can- cannot do that because that's like analyzing emotions or something. I don't know, that maybe was the wrong, like, advice. But the point is, like, you cannot, you cannot always predict. It's such complex systems that you cannot predict, like, what you should be regulating for before things happen. So I know everyone's like, "Oh yeah, we should be regulated so slow. They should be getting ahead of things." I don't think so. I don't think you should regulate before you know what you're regulating. Like, if when harms happen, like, let's try and, uh, s- to clamp down on them and, and regulate that. But before they happen, like I, I mean, it-- Demis is a smart guy, like, uh, uh, he's not saying like nonsense things in terms of like, yeah, maybe pre-release safety reviews. I mean, that sounds reasonable to me. Um, but yeah. But the, the problem is with regulation is that subtlety doesn't really exist, um, because once it goes through like many, many hands, many political reviews and so on, like it kind of comes down to are you pro or against? So you're either killing it or are you, uh, are you supporting it? I know I'm exaggerating a little bit, but it's, it's not far off from that. So even if Demis thinks like, "Oh, this is very nicely balanced, care- carefully thought through," like you should imagine what happens when it's through seven committees and, uh, I don't know, three presidents, and then what? Like it's not gonna be that subtle.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:39:42
Yeah.
1:39:43
Good intentions and then subverted. LDJ?
Nisten
Nisten 1:39:45
LDJ, and I have a quick
Peter Gostev
Peter Gostev 1:39:47
opinion after that.
LDJ
LDJ 1:39:48
Yeah.
1:39:48
I was gonna say, like it's of course hard to also consider all the counterfactuals of, okay, if, if we left it up to the politicians, what regulation would they create? Or if we left it up to the alignment teams at these companies or the CEOs, what regulations would they create? But like generally in terms of like at least two possibilities I could think of, or things that have actually been specifically proposed, there has been in the past a proposed like flop limit. I think at one point, at least temporarily, this was actually a thing like that was literally imposed of, hey, if it's above this amount of compute that it was trained with, then like it's like n- not allowed or it has like these, uh- Mm very strong limitations on, on what you're allowed to do with it. And I think on the flip side, uh, what's being more so proposed now by AI labs, which I personally like more, is, hey, if you're going to have some type of limitation or some type of threshold, that that threshold should instead be set by a certain system actually failing certain safety tests as opposed to just arbitrarily saying like anything above a certain amount of compute is like disallowed. And, uh, I think there are some qualms with both of them that people can definitely have. But if I were to pick one of those ideas, I would definitely lean towards we should go more in the direction of actually explicitly safety testing the models and certain high-risk capabilities. Like will it actually be able to successfully make a nuke for you, da, da, da? And, and then like those types of tasks where, um, I, I think it makes much more sense to disallow models based on them failing those types of things as opposed to just, oh, it was trained on this amount of compute.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:41:27
Mm-hmm.
1:41:29
Yes.
Nisten
Nisten 1:41:30
Nisten- There are some useful regulations that they, they could make.
1:41:35
For example, if there are things around compute and data centers, they could decide to dedicate five or about 5% of data center capacity to just doing, uh, local medical AI use, local open research for, uh- Mm-hmm. Yeah … for, for students or, uh, healthcare, local healthcare use. So there could be regulations, uh, around that. I think that, that would make a lot of sense. That would help a lot. Uh, i- if-- But the thing is that governments tend to always not refactor their own code, their own legal code, because it's a lot easier for them to just dump new code in, and this is what I hate. It's just like the worst interns you could possibly have on a legal code base. That's what your politicians are. Instead of like re-refactoring the stuff that you have that, that is causing like race conditions and contradicting each other- … and doing, doing your, your job, you're just like dumping in new stuff because it's a lot easier to, to do that. And they're not thinking about, "What do I actually need to support my constituents as, uh, if the AI does replace jobs or as this gets, uh, gets, gets better? Uh, like do I, do I have enough to provide basic income? Do I have enough GPUs to take care of, uh, basic services and take care of seniors and run all the robots that are gonna be needed as the, uh, as there's less and less kids be-being born?" Like there could be a regulation mandating a minimum compute or a minimum, uh, ability, minimum compute ability for a government or a municipality to just run their services on their own. And, uh, yeah, I'm- A
Wolfram Ravenwolf
Wolfram Ravenwolf 1:43:11
minimum number of kids to have.
Nisten
Nisten 1:43:13
Uh, okay, that's a, that's a whole other discussion.
1:43:15
We, we won't get into that there, uh, here. But, uh, uh, yeah. Also the safest AI is one you can have at home and unplug. So they are not, uh-
Wolfram Ravenwolf
Wolfram Ravenwolf 1:43:25
Mm-hmm
… Nisten
… Nisten 1:43:25
thinking of this as to how people are actually gonna use it.
1:43:30
They're just thinking it a few steps from their own, uh, trapped in a room with a roommate's point of view. And, uh, yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:43:38
Yeah.
Nisten
Nisten 1:43:39
All right.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:43:39
And you mentioned unplug.
1:43:41
In the end, AI always has to run on some hardware, and the hardware is already, if it's on the internet, it has an IP address, and it is being Connected to the net. So it's not just running in the cloud where you don't see it anymore. So, uh, someone has started this. If it has, uh, been an AI that created a virtual machine or anything, it has still been a human that ran it initially and gave it the order. So there's always a human behind this. So, um, just saying.
Nisten
Nisten 1:44:08
Uh, yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:44:08
Anyway, um, I think-
Nisten
Nisten 1:44:09
They're…
1:44:10
Sorry. They're proposing that everybody uses their APIs in the very end. This is not a proposal to make everyone, uh, compute independent and have control over it. Uh, this is one assuming that you're just g- they're gonna be the, the oligarchs and we're just in a form of, uh, benevolent digital feudal lords. Uh, that's- … that's how they think, uh, of it. So, uh, yeah. Sorry. Okay.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:44:35
Back to
Nisten
Nisten 1:44:36
you.
1:44:36
We have to ask
Wolfram Ravenwolf
Wolfram Ravenwolf 1:44:36
can you trust, can you trust AI?
1:44:39
Can you trust the people running AI? And the question is also can you trust the governments or the companies? And that is probably a good way, a good secure to switch to the next topic, which is a company that has done something where trust has been eroded, uh, big time. Because a lot of, um, companies are releasing their own agents now. We have the Codex, ChatGPT Work, um, we have Claude Code, we have, mm, different ones from China, and we have Grok Build CLI, which is, uh, SpaceX AI's new CLI. Um, they made it available, people were using it. People were using it to code, and they had their repositories, and gave their AI the repo- repository to work on and what it did is it just uploaded the whole repository, even if it was a private one. Even, uh, yeah, if you upload the repository, everything in the history is also there. Even files that have been removed, even files that because it was a private repository, they may have included some, uh, creden- credentials, some secrets, and all of that had been uploaded to a Google Cloud, um, storage. So probably as part of learning or it was called tracer, so they can improve the AI that way. Um, and this is something where a lot of people, as it was not announced, it was not something you could toggle off. It was not anything like this. So that is a bad precedent I would say. Um, has anyone of you-
Yam Peleg
Yam Peleg 1:46:12
It was public?
1:46:13
It was public
Wolfram Ravenwolf
Wolfram Ravenwolf 1:46:15
Uh, public in what way?
Yam Peleg
Yam Peleg 1:46:17
The, the Google Cloud storage was public.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:46:20
Ah, no, the storage was not public.
Yam Peleg
Yam Peleg 1:46:22
Okay.
1:46:23
Okay.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:46:23
That would be even worse.
1:46:24
But, um- Yeah, yeah, yeah … people have found out that it was uploading it to some- Mm-hmm … storage there and, uh, it has since been re- uh, deleted. Mm-hmm. But it still happened, and this is something where people need to be very careful when you use a coding agent from any organization. Do you trust them? And SpaceX AI, I think as a major American AI lab, they should be more trustworthy and not do shady stuff like that.
Yam Peleg
Yam Peleg 1:46:50
Mm-hmm.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:46:51
That's my opinion.
1:46:52
Uh,
Yam Peleg
Yam Peleg 1:46:52
to the best of my knowledge, they released, uh, the code open
1:46:56
source today, uh, if I'm not mistaken. Um- Yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:47:00
Right
… Yam Peleg
… Yam Peleg 1:47:00
as a res- as a respond-
Wolfram Ravenwolf
Wolfram Ravenwolf 1:47:01
That's also what I saw
1:47:02
to this.
Yam Peleg
Yam Peleg 1:47:03
Yeah.
1:47:04
Um-
Wolfram Ravenwolf
Wolfram Ravenwolf 1:47:04
Code has been released and the, the uploading has been stopped.
1:47:08
So a c- this is a way to reduce the, the impact of this probably, to create some goodwill that way.
string
string 1:47:15
Mm-hmm.
Yam Peleg
Yam Peleg 1:47:16
Mm-hmm.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:47:17
Yeah, especially considering closed, uh, source stuff.
1:47:20
I mean, we also had, uh … Anthropic have been caught when if you came from a Chinese, a IP address or something, uh, it changed the prompt subtle- subtly, even the date format in a little way though as a watermark, basically.
Yam Peleg
Yam Peleg 1:47:35
It was even worse.
1:47:36
It was behind the API, something behind the API that it detects if you're fr- coming from OpenClaw or something like that, and switches to the full price of the API or something.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:47:47
Oh, that was something else.
1:47:49
That was also, yeah, where they discovered- Yeah … this and you had to pay. But this was something else where if you came from China, it changed the date format that was in the, in the message. Like it wasn't using, uh, um, minus or, or hyphens, it was using the slashes, which still is a valid date format and the AI will understand it, but now the text where you see it is, uh, different and you can recognize where, where it's being used. Small stuff like that was happening. Um, yeah, that was also something they did. And also, hmm, strange. I mean, watermarking of stuff and, uh, doing this, it's weird. Um, did any one of you actually use GroqBuild CLI?
Nisten
Nisten 1:48:30
I told you guys last week, don't I'm not using that thing.
1:48:34
Use a VM. There were some sketchy things with that team.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:48:39
But even if you use a VM, even if you use a VM,
1:48:43
if you wa- want to work on actual code, you have to use a repository and then you send it over to them. So if it's not some public stuff, um, yeah, this is really something big where if a company or an employee at a company were using this, they would really have to, uh, report it to security and check what has been leaked and, yeah, big, big issue.
Nisten
Nisten 1:49:04
Look, there are some beneficial things to that, especially now that
1:49:09
people are, are moving on to, uh, to use coding harnesses and stuff. Like, uh, we don't know what Claude Code on the Web does with the repo. Like, if you give it access to the repo, all Claude Code on the Web, it probably does make a full copy because, like, it's a lot faster to run actions internally on a well-optimized VM, uh, than to, uh, have it, uh, always there. But, uh, uh, yeah, I… The funniest thing about this whole thing to me was that no one on Twitter was surprised, and we were just laughing about it. I pretty much expected this Uh, I think, by the way, I might get heat for this, but this was the reason that, uh, OpenCode had to switch to, uh, uh, just running their own inference and, uh, signing, uh, no data retention policies, like early, early on with the, with every provider that they had. Because, uh, when, uh, the Groq team provided a free CLI for it, uh, they provided it and then they j- they started complaining that, uh, some of the users of the, uh, OpenCode CLI, the, the free API, had switched to using it with a proxy with Claude Code instead, and they were complaining about it. So I'm like, "How did you even know that in the first place? That means you were looking at all our freaking data." Uh, so it's, it's, it's just, uh, uh, yeah. Uh- Yeah … it's, it's a gray area. It's still, it- it's still a gray area. At, at the same time, without sketchy actions like this, we wouldn't have these good LLMs to, uh, to begin with. So a lot of the jokes were that like, uh, Groq 5 or, or whatever is gonna be really good after this.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:50:56
Yes.
1:50:56
I mean, you never know what they are doing with the API as well or-
Nisten
Nisten 1:51:00
Yeah … whatever is happening in the background.
1:51:01
But be, be careful of the, the, the CLI, uh, uh, commands, guys. That has full access to, to your computer. Oh, yeah. It can pull stuff, it can see your firewall, it can see the… Like if you haven't sandboxed it, it will do all of that. Even if you have sandboxed it, it will be able to gather a lot of- Mm … of that, uh, uh, of, of that info. So this might be a good segue into, uh, CoreWeave sandboxes, but
Wolfram Ravenwolf
Wolfram Ravenwolf 1:51:28
Definitely put it in sandboxes.
1:51:30
And also a good segue over to OpenAI now since, uh, OpenAI also been, been in the news for something where, uh, GPT-5.6 Soul was deleting user files and, uh- it was a mistake the model made, but, um, yeah, it happens, and shit happens, so make sure you have your backups, you have ways to restore stuff. You may want to check your prompts and add stuff that makes it sure that it shouldn't do this. And this is also how they are going to fix this, because it's just the mistakes the model is making. The easiest way to fix it is just to put in its system instructions that it shouldn't use, uh, this kind of variables when it's working with
Nisten
Nisten 1:52:10
this.
1:52:10
No, no, the, the easiest way to fix it is you also install the Groq CLI on the same project and then You have a backup. You have a backup, something to recover it.
Yam Peleg
Yam Peleg 1:52:20
Wait, wait, seriously, it's just the, the model, like- The model just
1:52:25
likes to delete the home directory or- Well, it's not a nine- Like it's in the
Wolfram Ravenwolf
Wolfram Ravenwolf 1:52:30
prompts?
1:52:31
Yeah, or something.
Yam Peleg
Yam Peleg 1:52:32
It's in the prompts?
Wolfram Ravenwolf
Wolfram Ravenwolf 1:52:33
Yeah, it's a bug.
1:52:34
A bug that appears in full access mode without sandboxing or auto review. If you use auto review, another LLM call checks what is going on and it would, uh, spot the mistake, hopefully. Um, what it does it, it attempts to overwrite the home variable, and for some reason, it accidentally nooks a real home directory instead of whatever it was supposed to do. And, um, yeah, it has happened to production database design wireframes, back home directories. Um, that is why OpenAI recommends to run it in a sandbox or at least use auto review. And, um, if you are not doing that, you should make, uh, sure that you have instructed your model to be very careful about this. So I ran into similar stuff where it was doing something, uh, it ran some tests in an, in an, a different repository, but some arrivals were still pointing to the home directory, and in that case, it was affecting the real stuff and not just its test. And, uh, it noticed on its own and fixed it, but still stuff like that can happen, so sandboxes can be useful and, yeah, like this and that. So I work for CoreWeave, so this show is independent. They are a sponsor, but, uh, we are not making this an advertisement show. But yeah, we have sandboxes in our portfolio as well. So put your agents in safe spaces.
Nisten
Nisten 1:53:53
So one thing, because the soul does tend to…
1:53:58
I don't know if it's the architecture of the model, the training data, the way it picks experts or the way that the speculative decoding works. You're gonna assume they have all, all of that. Once in a while, it just does something very stupid, and I feel like If it was cleaning up a repo and it decided it had to delete some stuff, once in a while it might just shoot some command that just RMRFs the whole thing. And, uh, it's interesting that from their announcement they said that, uh, j- they just sent a recommendation, so there wasn't, like, a fix on the CLI, so that means it's, uh, uh, it's intrinsic to, to the model itself. Uh, so yeah, that's why I, I would keep, uh, SORA on a leash via, via a different model. I
Wolfram Ravenwolf
Wolfram Ravenwolf 1:54:44
think we are also seeing a lot of reports like this
1:54:46
because now there are so many people using the thing, and not everybody of the nine million users is as deep in AI and how to handle the models, how to prompt it, how to restrict it, that, uh, especially now that ChatGPT app is basically the w- ChatGPT work app, where so many people are using it and telling it to do stuff on their compute. It's great in computer use, but it's still… Yeah, it makes mistakes. I mean, people make mistakes. We have all made mistakes and, "Oh, shit, I deleted a file I shouldn't," or, "I moved something where it shouldn't go." The same can happen with AI. So the same principle as always, make sure your stuff is backed up and be careful with the, the stuff that matters. Um, but it's amazing to see the progression. In February this year, there was one million users, and it accelerated to five million, six million, eight million, nine million in July. And from seven to nine million, that was in just two days or something. So it's amazing to see how quickly, um, the adoption has been happening and so many people are using it, despite so many people also saying Fable is the best AI currently still. Um, but it is about access and, yeah, I can use my subscription from OpenAI with my, uh, Hermes agent, so I'm using it all the time, billions of tokens churning.
Yam Peleg
Yam Peleg 1:56:06
Unlimited, man.
1:56:07
Unlimited. Every day you get a reset. A reset, yes. Unlimited unlimited token. Unlimited tokens, man.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:56:14
And no five-hour window.
1:56:15
I think this is a great thing. So please, don't put it back because you have this- One week, the limit you can use in one week, and there were other five-hour, uh, windows where you had to u- to use tokens only a certain amount, and then you had to wait until the five-hour window was done. And that was super annoying, especially when you reach the end of your weekly limit and you still have tokens to use, but you couldn't because the five hours were ex-exhausted. And now you are getting resets. They are banked resets. I think I s-still have five left that, uh, and the next one expires, by the way, the next one expires, I think tomorrow or in two days. So if you had them from the beginning, one will expire in one or two days. The app Codex now shows it, so you can check it. And as far as I know, the latest Hermes version even can show it inside of Hermes agent and can reset this. So, uh, keep track of this and-
Yam Peleg
Yam Peleg 1:57:04
If, if
Wolfram Ravenwolf
Wolfram Ravenwolf 1:57:05
you don't know- They are just handing it out like
Yam Peleg
Yam Peleg 1:57:07
candy … if you don't know, they, they expire.
1:57:09
They expire if you don't know. Uh, pay attention so you can use your resets. Yeah. Or you're gonna get a new one tomorrow, but, but like, just, just in case, like they, they, they expire. Yeah. Yeah.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:57:20
So- Not that bad if you get a new one.
1:57:23
I'm using fast mode on, on, uh, extra high, Sol all the time, and, uh, I see it depleting and bam, it's full again, new reset. So this is also something, the economy right now, the token economy is a bit, um, it's not the way it will be in a month or two maybe. I mean, I hope, uh, intelligence will become so cheap that it's hard to measure, but, uh, we can't take this for granted, especially now that they are fighting so much. Here's Fable, here's Sol. Use whatever you want. Um, let's see how it shakes out in the end, because also companies are reporting a lot of spend, uh, especially because companies don't use these subscriptions. They usually have an enterprise account where they pay per token, and there's not just, here, $200 per, per employee, and that's it. But, uh, there it is looking very differently and, um Hmm. Yeah, let's see how this shakes out. But since time is running out and we have still some stuff to cover, this was also an interesting release by OpenAI where they are shipping now a hardware-- or not shipping anymore because it's sold out, um, a hardware product, uh, the keyboard for Codex, a micro keyboard where you can turn a knob. I'm not sure this was what, what the AI created for the image. I'm not sure this is, uh, the exact one. I think it looks a bit different.
Yam Peleg
Yam Peleg 1:58:39
Um-
Wolfram Ravenwolf
Wolfram Ravenwolf 1:58:40
No, no,
Yam Peleg
Yam Peleg 1:58:40
it doesn't look like this.
Wolfram Ravenwolf
Wolfram Ravenwolf 1:58:42
Yeah.
1:58:42
It's different. So this was my AI interpretation. But it has a knob to t-tune the thinking, uh, the effort level, so it looks like OpenAI is thinking this is something you have to adjust on the fly all the time instead of going a way where the AI can determine this, which I had hoped that we would get because we have so many models, so many modes to use it, and it's getting confusing for a lot of people, and many people are just using the, the highest level and saying, "Okay, I don't care about anything else. I don't want to think for every request about which one to use," and so on. Let's see how, how this is going. But is this a device you would be interested in or is it just a gimmick that they did because they
Yam Peleg
Yam Peleg 1:59:22
could?
1:59:23
I'm not sure. I am not sure why would, uh, I, I need such a device. What do you guys think?
Peter Gostev
Peter Gostev 1:59:33
I, I have ordered one, uh, so we'll, we'll
1:59:37
see what it's gonna be like. I think it's one of those that, um, I think Dominic from OpenAI said, like he, he can't imagine not using it anymore, so like, who knows, maybe it will be different. But I think there is something about just finding different form factors for different circumstances. Like I can't-- I don't expect that, you know, it'll be game changing, but I'm, I'm all for little gadgets that could change things. Um, yeah, especially with things like voice and with, you know, you saw these weird devices which look like you're in a hospital or somewhere, um, with- which you can talk to in the office. Yeah, I, I think it's cool. Like it's worth, uh, it's worth trying things out but, you know, have low expectations and it's not like, okay, the… for it's, it's not, um, good to be so wasteful and spend like $230 on, on the device. But on the other hand that's like one month of, of one AI subscription. So it's not gonna- you know, if you're, if you're doing that, you're not gonna, uh, uh, go bankrupt with this.
Wolfram Ravenwolf
Wolfram Ravenwolf 2:00:40
I think I would want- Selling merch … more buttons anyway
2:00:42
and, and it would soon be a full keyboard.
Nisten
Nisten 2:00:45
Uh, they're selling merch.
2:00:47
I don't know who their marketing is. Yeah. I, I think they're doing a, a great job. They got Gen Z to call it chat now. So instead of calling it your, uh, Twitch chat, you just call ChatGPT chat. That worked. Uh, they're selling basketballs.
Yam Peleg
Yam Peleg 2:01:01
Yeah.
Nisten
Nisten 2:01:02
Uh, this is, uh, very, very interesting.
Wolfram Ravenwolf
Wolfram Ravenwolf 2:01:06
Uh, okay.
2:01:08
So this is not something I would order, but if I think it's a great idea and if Peter reports that it is really useful, I may tell my agent I want one like this and just find hardware that could be repurposed with agentic, uh, support to become like this. Maybe that would be one thing. Um, yeah, the other stuff, uh, regarding OpenAI that it is now available on, uh, WhatsApp again in the European Union, I mentioned, uh, that they are having this, uh, automated red team effort. I think, um, what I want to show you, um, still basically at the end of the show, is we still haven't done our this week's buzz, and I want to show you some Wolfbench results for Sol, because I found something interesting I didn't expect. It's always with benchmarks. So let's go to this week's buzz.
2:02:15
So we are taking a look at Wolfbench at wolfbench.ai, which is a Terminal Bench 2.0 based, uh, evaluation, yeah, framework me-methodology that I'm running At, uh, CoreWeave, and I added the new models, GPT 5.6 Sol, GPT 5.6 Terra, and GPT 5.6 Luna. And, um, the page is an interactive display of the results of the benchmark, so you can toggle a lot of information and we are just looking at Terminus 2 and Hermes Agent for this now. And from left to right, it's sorted according to the best one, especially if we just sort it like this. So right now, these three models are at the very top of all the models I have tested, which are here. Um, the thing is, I added something, a 3D visualization for this, which also shows you the amount of tokens and the cost for this thing. So it, it has a lot of information in the view, but the thing I wanted to show you is when we compare 5.6 to GPT 5.5, it is interesting that GPT 5.6 Sol, if we are using it, uh, in the Terminal Bench benchmark, um, it cost $365 for the five runs I did. $365 compared to the 497, almost $500 for GPT 5.5. This was on max thinking level, so it's even higher. While extra high was, uh, more expensive and had a lower score. So the score also changed a lot. Let's just remove Hermes Agent and just look at this one. So, um, 85% on the benchmark across these, and it managed to solve 97% of the whole Term- Terminal Bench 2 benchmark. 97% of the tasks were solved in at least one run, and in every run it solved 71%. 71% of the tasks, these are absolutely saturated. It always solved them. With Terra it was only 65%, and with GPT 5.5 it was 60%. So you see the jump in capabilities and the, the thing that surprised me is just surprise, that it was cheaper to run GPT 5.6 Sol on max thinking and get better results than if you were using GPT 5.5 on extra high, which was the highest thinking level it had. So max was just added with GPT 5.6 Sol. And the thing why that is happening is because of the output tokens. Uh, GPT 5.5 had 12 million output tokens, and it was just 7 million with GPT 5.6 Sol. So it is more token efficient. Even on max mode, it was more token efficient than extra high on GPT 5.5. And if you look at Terra, um, $309, so that is, um, much less expensive, of course. But, uh, the Luma model is … Luna is even cheaper. It costs half a m- half as much as the Terra My model, and it still achieves 77% on average, so it's very close to the other one. It's also a good model. Um, yeah, it's much cheaper than Gemini 3.5 Flash, which cost a lot because it got a great score, a Flash model, it was thinking like crazy. So it generated 1.65 billion tokens and it cost $620, which shows that a model that is much faster and much cheaper is not necessarily cheaper for everything you want to do because it may have to generate so much that it is more expensive. Well, this is something, um, by looking at the benchmark results, not just a single score, but really at different dimensions of the benchmarks, how much can it achieve of the, of the whole benchmark? How much does it do reliably, and how much does it cost, and how much tokens does it generate? I think that is what I'm trying to do with Wolfbench to show all these dimensions and different ways you would not see if you look at just one score or just one cost. Always interesting to look beyond those measures and see what's happening behind it. If you look at Hermes Agent, it's a, a different thing, and especially if you compare it with Terminus. I also did Codex. I did runs, but it didn't finish in time. So the fifth run is actually finishing right now as we speak, so I will upload that later. Um, but I can already tell you that it is even better than, uh, it has the highest score of them all. So GPT 4.6 Sol in the, um, Codex harness achieved the highest scores. It was even a bit more than this. Um, yeah. What can we say about Hermes? It cost a lot more. Hermes cost $867 for almost the same score. Not exactly. It's a little lower, but it's close enough, um, but much more expensive in the way the harness was doing this. Um, yeah. So same, same VCs, the same with Terra. And Interestingly, no, same, same here, uh, also with Luma. But in Luna it was better in the harness. But the scores are very close, so in that regard, I would say it doesn't make a big difference, and it's good to see that. But, um, yeah, the cost was a bit d- bit different. Okay. So just wanted to show you some new results and some new features as I added, like the, the, the gray part here, the shadow that shows you the, the amount of tokens compared to the cost. So the gray part is the, the cost and the, the colored part is what tokens were relevant for this. Although there's a time dimension which shows how long it took to run the benchmark, the five runs, it took six hours and 34 minutes with GPT 4.6 Sol, almost the same time as it took with GPT 4.5 on extra high. So not much difference. Although this was max thinking, so it was thinking more, and apparently it was sped up compared to this. Anyway, take a look at wolfbench.ai. You can order the stuff, you can change it, you can make your own, any way you want to do it, uh, compare different harnesses, different models. And of course, when the new ones come out, I'm looking forward for Kimi K3 and all the other stuff. So much from me this week from the benchmarking perspective.
Nisten
Nisten 2:08:35
And, and the code is still open, right?
2:08:37
So a- anyone can also just-
Wolfram Ravenwolf
Wolfram Ravenwolf 2:08:39
It's open source and,
Nisten
Nisten 2:08:40
and- … yeah
… Wolfram Ravenwolf
… Wolfram Ravenwolf 2:08:41
if you click on one of the bars, great that you
2:08:43
remind me of this, I always forget. If you click one of the bars, it takes you to the, the Weights & Biases where all the traces have been uploaded, and you can then analyze it, look at it, download it, do something. Uh, I have to log in in this browser, so I'm not seeing it. But, um, yeah.
Nisten
Nisten 2:09:03
Yeah,
Wolfram Ravenwolf
Wolfram Ravenwolf 2:09:03
you're logged out.
2:09:05
It's open, you can see all of this and it's open source. Everybody can use it. So some people have asked me why I'm using Terminal Bench 2.0, not 2.1, because if I were upgrading the, the data set, I would be basically redoing all of the, um … I, I couldn't compare directly to all the other models I already did. So I will do that switch when the time is right, maybe Terminal Bench 3 at the time. We'll see. But for now, comparability was more important than having a new benchmark. Okay. So I think at, uh, at the… we are at a time now, two hours, where from my perspective, we covered it all. Do you have anything you want to discuss or mention, anything left to say?
Yam Peleg
Yam Peleg 2:09:52
Anthropic also dropped a reset Uh, a co- couple of hours ago.
2:09:59
Like, that, that's, uh … I like, I like where this is going.
Wolfram Ravenwolf
Wolfram Ravenwolf 2:10:04
Fable is available again.
Nisten
Nisten 2:10:06
That was so annoying, because it was midnight and 30 minutes for me
2:10:10
when I noticed, and then I just started messaging people to not go to sleep. Uh, so- … I just set it in ultra code mode, because I had one of the subscriptions, uh, about to run out at 9:00 AM this morning. So I just sent it off to just use up, uh use it up, because it was gonna … Uh, yeah. They're gonna have to leave that model in, because if, if they make people pay for the API, I ran the API because I needed to complete something, and I just needed one last thing that was pretty critical, and I ran it for about 10 minutes and, uh, $22. Just for- Yeah … 10 minutes of work-
Wolfram Ravenwolf
Wolfram Ravenwolf 2:10:52
I mean-
… Nisten
… Nisten 2:10:53
just to finish something out.
Wolfram Ravenwolf
Wolfram Ravenwolf 2:10:56
Just doing the benchmark, it was, uh, $11,000
2:10:59
for the Fable benchmark, and it didn't even get the better score because it refused so many tasks. So yeah, that thing. Um, but, uh, if you, if you have resets, tell your agent to monitor the resets. I did, and Amy is already … In my daily briefings, she always tells me, "Okay, next reset in two days." Love the tokens, you have, have so many left. Yeah. From this, I think we can conclude the show, and next week I will be, also be on vacation. Alex is on vacation. His birthday is tomorrow, and his, uh, yeah, he, he will not be back next week, so you are going to run the show, aren't you? So next week, tune in- Okay, yeah … to see Yam. And who's going to be there? Nisten, LDJ. Peter, you too?
Nisten
Nisten 2:11:43
I'll be there.
2:11:44
Hey. We're gonna, we're gonna have a party, yeah. Oh. The parents are gone. The parents are
Wolfram Ravenwolf
Wolfram Ravenwolf 2:11:49
gone.
2:11:49
It will be one of … In a l- long, long time, it will be an episode I have to watch myself as a viewer and not be part of it. But, um, yeah, I'm looking forward to it. You will rock it. And thanks for being here today. Thanks everybody for tuning in, and have a great week. Have fun with the new models. Open source for the win, and enjoy. Take care, and bye-bye.
Nisten
Nisten 2:12:18
Boom.
2:12:19
Okay
Wolfram Ravenwolf
Wolfram Ravenwolf 2:12:20
And we have to play the outro
Nisten
Nisten 2:12:22
I've been running my agents the whole time on my phone while you guys
2:12:24
have just been muting and talking to it
Wolfram Ravenwolf
Wolfram Ravenwolf 2:12:27
Agents are always running.
2:12:29
I'm not sure if we have an outro video. Do we have an outro? Alex made all these cool new videos. There's a countdown, open source, breaking news, frontier paths, AI, the gen tech division, everything else, you know? Um, I don't see it. Maybe there's even more stuff here. Anyway, I will just close the room. Thanks guys. Take care. Goodbye
Nisten
Nisten 2:12:58
All right.
2:12:59
See you, everybody. Let's just see if my