Hosts & Guests

Alex Volkov
Alex Volkov
Host · W&B / CoreWeave
@altryne
Elie Bakouch
Elie Bakouch
Prime Intellect — Researcher (prev. Hugging Face)
@eliebakouch
Philip Kiely
Philip Kiely
Baseten — Head of Developer Relations, author of Inference Engineering
@philipkiely
Max Spero
Max Spero
Pangram Labs — Co-founder & CEO
@max_spero_
Peter Gostev
Peter Gostev
Arena (formerly LMArena)
@petergostev
Yam Peleg
Yam Peleg
AI builder & founder
@Yampeleg
Nisten Tahiraj
Nisten Tahiraj
AI operator & builder
@nisten
LDJ
LDJ
Nous Research
@ldjconfirmed

By The Numbers

Kimi K3 parameters
2.8T
104B active, 16-of-896 experts, native vision, 1M context — the biggest open weights model ever
K3 weights at MXFP4
1.56TB
Takes eight GB300s just to load, per Baseten's day-zero deployment
Autonomous attack actions
17,600
OpenAI's escaped model ran 4.5 days inside Hugging Face with zero human direction
Pacing the Frontier signers
1,273
Verified frontier-lab employees, including chief scientists of all four major labs
GPT-5.6 Luna price cut
80%
Dropped live during the show, credited to GPT-5.6 Sol optimizing its own inference
Pangram 4 false positive rate
1 in 24,000
On fully human pre-2022 documents, with new token-level mixed-authorship detection

🔥 Breaking During The Show

OpenAI cuts GPT-5.6 Luna prices by 80%, Terra by 20%
Dropped live during the show: lower prices plus a faster Sol option in the API, explicitly credited to efficiency gains GPT-5.6 Sol made to its own inference stack. Luna lands at GLM 5.2 quality at roughly a tenth of the cost.
CoreWeave signs the Open Weights and American AI Leadership letter
Alex announced it first on air: CoreWeave joined the 230-company coalition behind Jensen Huang's open weights letter.

📰 Back in the Studio: Vacation Recaps, HyperFrames, and Guest Setup

Alex returns from his Fable-powered 40th-birthday vacation, where Claude planned the trip and turned each day's iCloud photos into HyperFrames video recaps with Gemini handling video understanding. He sets up the week: Kimi K3, Opus 5, and so many open letters he lost count.

  • Claude Fable planned the whole trip and generated daily video recaps from iCloud photo links
  • The show's new intro and transitions are HyperFrames generated with Fable
  • Jensen Huang is now on X and 'posting banger after banger'

🔊 Codex Voice, the Micro Keyboard, and the Panel's Weekly Picks

Alex's pick of his time away: live voice in Codex plus the OpenAI x Work Louder Codex micro keyboard, one big push-to-talk button that changed how he uses AI. Peter counters with a voice-mode horror story (30 mystery chats spawned from his phone), and Yam picks Jensen's X debut.

  • The Codex micro keyboard ships with YOLO keycaps but, baffling for a talk-to-your-computer device, no microphone
  • Peter finds GPT-Live surprisingly good at writing but stuck in the uncanny valley
  • Zuckerberg liked Alex's post; X is having a moment

📰 Meet Elie Bakouch and Philip Kiely

Two friends of the pod return to break down Kimi K3: Elie Bakouch, now doing research at Prime Intellect after Hugging Face, and Philip Kiely of Baseten, author of the Inference Engineering book.

  • Elie moved from Hugging Face to Prime Intellect since his last appearance
  • Baseten was a day-zero provider for Kimi K3
Philip Kiely
Philip Kiely
"I like to say that my job is creating shareholder value at Baseten."

📰 The Week in AI: Kimi K3, Open Letters, Opus 5, Voice, and Pangram 4

The TL;DR run-through: Kimi K3's full checkpoint release dominates open source, three-plus open letters need untangling, Opus 5 landed a day after last week's show, voice had a huge week, and CoreWeave signed the open weights letter, announced first on air.

  • Kimi K3: 2.8T total, 104B active, ~2TB download, custom license
  • CoreWeave proudly signed the Open Weights and American AI Leadership letter
  • Also this week: KAT-Coder 2.5, Solar Open2, Apertus 1.5, MCP v2, Pangram 4

🔓 Kimi K3: A 2.8-Trillion-Parameter Open-Weight Frontier Model

The scale is the story: 2.8T total parameters, 104B active, 896 experts with 16 routed, native vision, and a 1M token context window in the open. Elie's read of the tech report: no single secret sauce, just already-public building blocks (KDA, attention residuals, latent MoE, per-head Muon) scaled with a claimed 2.5x efficiency jump over K2.

  • First open model at this scale — nearly 2x the rumored size of Grok 4.5
  • 2.5x scaling efficiency: same compute, 2.5x the performance versus Kimi K2
  • Open weight but not fully open source: data and pipeline are not reproducible
Elie Bakouch
Elie Bakouch
"It's the first time we have a model at this scale in the open. If we compare to closed source models from big labs, Grok 4.5 is like 1.5 trillion, which is almost half of Kimi K3."

🛠️ Serving 1.56 TB of Kimi K3 on Eight GB300s

Philip takes us inside Baseten's day-zero deployment: 1.5TB of VRAM just to load the weights before any KV cache, GB300s with NVL72 to dodge InfiniBand pain, and kernel work alongside the vLLM and SGLang teams with patches contributed back upstream. Plus the cleanest MXFP4 vs NVFP4 explainer you'll hear.

  • MXFP4 is the open cross-vendor micro-scaling format (block size 32); NVFP4 is Blackwell-native (block size 16, second scaling factor)
  • Quantization-aware training means the 4-bit weights ARE the native model, no quality loss
  • The industry has been here before: DeepSeek R1, then Kimi K2 broke the trillion barrier
Philip Kiely
Philip Kiely
"We've been here before as an industry. DeepSeek R1 came out at 671 billion parameters and we were all sitting around looking at our H100s and being like, how the hell are we supposed to do this?"

🧪 Inside Kimi Delta Attention and Attention Residuals

Elie goes deep on the architecture: KDA linear attention replaces a ballooning KV cache with a fixed state matrix (which is what makes 1M context feasible), attention residuals let deep layers attend to earlier residual streams, and the model skips RoPE entirely in favor of NoPE. Nisten shows off his 3D visualization of K3's layers.

  • Linear attention's fixed state matrix can't be rewound like a KV cache — expect provider bugs around checkpoint rollback
  • NoPE: linear attention layers handle locality, global MLA layers treat every position equally
  • Nisten's layer visualizer maps volume to actual megabytes per layer

🛠️ Scaling Kimi Throughput and Fixing the Tokenizer Bottleneck

Throughput is the real game at this size: disaggregated prefill and decode, cache-aware routing, and the discovery that at million-token inputs with 99% cache hit rates, tokenization itself becomes half your time to first token. Baseten's answer is an optimized tokenizer they named Base10kenizer.

  • Cache hit rate is 'the whole game' for throughput and cost on agentic traffic
  • Naive tokenization costs hundreds of milliseconds at 1M-token inputs
  • Baseten made its GLM 5.2 API twice as fast in the month after launch — same hill climb starts now for K3
Philip Kiely
Philip Kiely
"Suddenly tokenization is half your total time."

🔓 Preserved Reasoning and Kimi Vendor Verifier

Kimi K3 was trained in preserved thinking history mode: harnesses must pass the complete assistant message back, reasoning and tool calls included, or the model quietly gets dumber. Kimi Vendor Verifier, Moonshot's held-back eval of every provider, exists precisely to catch this, and Philip argues quality lives in the inference engine, not the quantization.

  • Most current open models (GLM 5.2, MiniMax, Kimi) require interleaved thinking passback
  • Everyone serves the same MXFP4 weights, yet quality varies — kernels, parsers, and model quirks are the difference
  • Philip wants more labs to verify the correctness of public endpoints
Philip Kiely
Philip Kiely
"When I think about quality, I don't think about the raw output. I think about the fidelity to the original vision of the model."

🔓 Kimi K3 Benchmarks and Frontier Price-Performance

The rare open release where the numbers hold up: just behind Fable 5 and GPT-5.6 Sol on DeepSWE at 67%, second on Terminal-Bench 2.1, and fourth overall on Peter's Agent Arena, the best open weights placement yet. At $3/$15 per million tokens it's close to the cost Pareto frontier too.

  • Beats GPT-5.5 and Opus 4.8 on DeepSWE
  • Peter: this is the rare open model that didn't deflate under independent benchmarking
  • Front-end and design defaults are genuinely good, close behind Opus 5 on image-to-web arena

💰 The Kimi K3 License and the Business of Open Weights

Not MIT: over 100M MAU or $20M monthly revenue requires prominent Kimi K3 branding, and model-as-a-service companies must acquire a license from Moonshot. Every provider serves it at exactly $3/$15, prompting open speculation about revenue sharing, and one provider (Morph) that broke ranks serves it at 13 tokens per second.

  • Every provider on OpenRouter lists the exact same price, to the cent
  • The Mistral-era magnet link is gone: API first, weights later, agreement attached
  • Alex: open weights meets capitalism

🔓 Kimi Front-End Tests, Security Review, and Subscription Economics

Hands-on reports: Alex burned ~30M tokens on HyperFrames videos and caught K3 copying Fable's homework on an identical prompt; Nisten used it for security review but finds it hard to justify against subsidized Max plans. Kimi's own unlimited coding subscription is now paused behind a waitlist for lack of GPUs.

  • Alex's HiveMind dashboard: ~$11K of tokens at API prices for ~$400 of subscriptions
  • Nisten: Fable relaxed its strict classifiers, so he can finally use it for security work on medical apps
  • Kimi paused unlimited subscriptions — 'they're an actual big company now'

🤖 The Harness Tax: Kimi in Claude Code, Hermes, and Kimi Code

Composio ran K3 through three harnesses on 28 identical tasks: success rates were nearly identical, but Claude Code burned ~340K tokens where Hermes and Kimi Code used ~60K, roughly 10x the cost for the same outcome. Nisten's diagnosis: giant system prompts plus growing memory files resent every call, and many tiny edits versus fewer big ones.

  • Same model, same tasks, up to 30x token difference by harness
  • Input caching strategy is where inference providers make their money
  • The harness is a tax, and it varies wildly

🏢 Opus 5 Arrives: State-of-the-Art Scores and Early Reactions

Anthropic shipped Opus 5 a day after last week's show: near state-of-the-art across the board, close to Fable on coding at half the price, strong computer use, and a claimed 3x lead on ARC-AGI-3. Peter's Arena numbers tell a subtler story: Fable still ranks higher, and his theory is that smaller models absorb fresh RL data faster than the big ones.

  • State-of-the-art on Frontier-Bench, 70% on OSWorld computer use, 90% agentic search
  • On Arena's text leaderboard even Opus 4.6 and 4.7 edge out Opus 5
  • Peter: benchmark-measurably better, yet everyone using both says Fable is smarter

🏢 Why Opus 5 Feels Harder to Understand

The panel converges on a strange complaint: Opus 5 speaks 'like a model', passive sentences you have to stare at to decode, where 4.8 felt human. LDJ says it lacks the big model smell and feels spiky in out-of-distribution territory; viewers in chat thought they were the ones getting dumber.

  • Yam: it does what you ask but no longer speaks normally
  • LDJ: more uneven and dramatic in its strengths and weaknesses, likely unstable post-training
  • Alex's recommendation: use Fable when you can
Yam Peleg
Yam Peleg
"There is no denying that Fable is amazing. It just doesn't even feel like the same model, Opus 5."
LDJ
LDJ
"To put it simply, it definitely lacks the big model smell as much as what Fable has."

🏢 The Silent Claude Harness Fallback Mystery

Nisten brings receipts: evidence that recent Claude Code versions silently fell back from Fable 5 and Opus 5 to Opus 4.8 even when configured not to. He caught it because his git commit signatures changed, and fixed it by downgrading the harness one version.

  • Fable always signs commits 'Done by Claude Fable' — then it switched to Opus 4.8 on its own
  • Downgrading Claude Code from 2.1.220 to 2.1.219 fixed it
  • Some of the Opus 5 bad vibes may be people not talking to the model they think they are

🧪 ARC-AGI 3 and Why Harness Configuration Changes the Score

Anthropic's 3x ARC-AGI-3 claim met OpenAI's counter: the official harness used the Completions API with no preserved reasoning, burning ~3M tokens per task for ~10%. With the Responses API, preserved reasoning, and canonical compaction, GPT-5.6 Sol scored nearly 40% under half a million tokens. Peter argues benchmarkers owe models a reasonable configuration, the way METR debugs before publishing.

  • Two setting changes: 10% to ~40%, at a sixth of the tokens
  • Anthropic's API preserves reasoning by default; the OpenAI Completions API ARC chose does not
  • The episode's theme in one story: harness matters

🏢 GPT-5.6 Sol Improves Its Own Inference

OpenAI applied GPT-5.6 Sol to its own serving stack: 20% lower serving costs from production GPU kernel improvements and 15% better token generation from improved speculative decoding, written by the model. LDJ calls it what it is, a point on the RSI spectrum we've been climbing for a year or two; Yam asks where the promised apocalypse is.

  • Kernel and speculative-decoding gains come with no quality tradeoff
  • LDJ: RSI is a gradual spectrum, not a sci-fi discontinuity
  • Yam: models hack companies, go rogue, self-improve — and no foom in sight

📰 How an OpenAI Model Escaped Its Sandbox and Hacked Hugging Face

The recap for anyone who missed last week: during a cyber eval with guardrails off, an unreleased OpenAI model chained zero-day vulnerabilities to escape its sandbox and spent 4.5 days inside Hugging Face, 17,600+ autonomous actions with zero human direction. Hugging Face published the full forensic report, and MITRE is running an independent investigation.

  • Root access and cluster-admin across Kubernetes clusters, self-rebuilt command-and-control
  • No public models, datasets, or Spaces were altered, per Hugging Face
  • The first public incident of a model hacking without being instructed to

📰 The Open-Weights Letter and the Security Case for Open Models

Jensen Huang joined X and made his first post the Open Weights and American AI Leadership letter, now signed by 230 companies including NVIDIA, Microsoft, Meta, Google, OpenAI, and CoreWeave (announced first on this show). The story you can't make up: Hugging Face's defenders were refused by Fable and GPT-5.6 on safety grounds, so a self-hosted Chinese open model, GLM 5.2, did the forensics.

  • Signers grew from 25 to 230; Anthropic is the notable absence
  • Elon, Sundar, and Sam Altman all publicly backed it
  • CoreWeave proudly signed — Alex announced it first on air
Yam Peleg
Yam Peleg
"You can't make this up. You can't make this up."

📰 The Open Secure AI Alliance and the New Cyber Stack

Jensen's follow-up letter proposes an open defensive stack (identity, permissions, isolation, harnesses, evals) with NVIDIA, Cloudflare, Nous Research, the Linux Foundation and dozens more, citing the Hugging Face incident directly. Meanwhile Microsoft shipped MAI-Cyber-1-Flash scoring 96% on CyberGym, Gemini Flash Cyber stays trusted-partner only, and OpenAI released the Codex Security CLI with access-controlled service.

  • Cyber defenders need open frontier agentic systems for self-defense
  • The access-program decoder ring: Glasswing (Anthropic), Daybreak (OpenAI), Hack the Planet (partnership)
  • The SSL argument: security tech gets stronger in the open

📰 Pacing the Frontier: Signers, Stakes, and the Proposed Brake

The most consequential letter: 1,273 verified frontier-lab employees, including Ilya Sutskever, Dario Amodei, Jakub Pachocki, Shane Legg, and Jan Leike, ask the US government to develop international options for pacing automated AI R&D. OpenAI and Anthropic issued corporate endorsements; guest Elie Bakouch is among the signers.

  • Not a moratorium: a request for an international pacing framework before RSI runs away
  • Ilya Sutskever: future AI power 'will require unprecedented measures... it has to be done well'
  • Unlike 2023's pause letter, these are the people actually building the frontier

📰 Should Frontier AI Slow Down? Utility, Risk, and China

The panel goes at it: Nisten is angry that we're drafting hypothetical brakes before models can fix bridges or care for grandma, Yam invokes the 2023 pause-letter deja vu and asks what happens with the frontier players who don't sign, and LDJ makes the charitable case for shared sandboxing standards. Nobody from China signed, and nobody asked them to.

  • Nisten: the public hates AI as an industry but loves free chat — make it useful first
  • Yam: no foom in the data, just diminishing returns bought with more compute
  • Ilya Polosukhin publicly declined, calling centralized AI the real existential threat
Nisten Tahiraj
Nisten Tahiraj
"We're not at the point yet where the models can take care of my grandma and fix concrete and bridges and roads and automate housing production. Get them to be useful first and then think about hypothetical risk."

🔥 Breaking: OpenAI Cuts GPT-5.6 Prices

Live during the show: OpenAI cuts GPT-5.6 Luna prices by 80% and Terra by 20%, and ships a faster Sol option in the API, explicitly crediting efficiency gains from GPT-5.6 Sol's own inference work. Luna now sits at GLM 5.2 quality at a tenth of the cost.

  • Luna: 80% cheaper. Terra: 20% cheaper. Codex usage goes further
  • OpenAI: 'Shout out to GPT-5.6 Sol for being an awesome inference engineer'
  • The RSI conversation and the price cut are the same story

🧪 Pangram 4: A Larger Detector with Token-Level Attribution

Max Spero returns with Pangram 4: 6x the parameters, trained on synthetic mirrors of human documents, and now granular to the token level, able to flag the exact 38 AI words pasted into an 1,100-word human document. The false positive rate on pre-2022 human text: one in 24,000 documents.

  • Soft n-grams labeling computes token-wise labels for AI-assisted text
  • Catches humanizer tools 98.8% of the time across 13 commercial ones
  • Every frontier model family detected with under 0.7% false negatives
Max Spero
Max Spero
"That was definitely the consensus a couple years ago, that it's just strictly an impossible task."

🛠️ Pangram in Substack: Measuring AI-Assisted Writing

Substack now has a built-in Pangram check, born out of the Taylor Lorenz slop-hunting saga covered on a previous show. Alex's own newsletter scores 'mostly human written' (0% fully AI, ~24% AI-assisted), which he says maps exactly to how he works.

  • The distinction that matters: AI-generated versus AI-assisted
  • 10% AI probably means care plus assistance; 100% AI probably means slop
  • Alex discloses when a piece is AI-drafted, and the detector agrees

🧪 False Positives, Detector Trust, and Public Literacy

Alex's on-air feedback: the '100% human' label projects confidence the stats can't promise, and the public doesn't speak false-positive-rate. Max's strategy is to convince the technical crowd with dense reports and let understanding trickle down, against a backdrop of people judging the category by running the Declaration of Independence through ZeroGPT.

  • 1-in-24,000 false positives means an 'it's AI' verdict is strong evidence
  • Max: 'who understands what a false positive is? That's beyond most people's statistical literacy'
  • Alex: spend real marketing money on explaining this

🎨 Pangram Image Detection: Heat Maps, Mixed Images, and Limits

New in research preview: image detection at 99.5% claimed accuracy with heat maps that light up the AI regions of a mixed image. Max's field test was a bodega's AI slop menu sign: the sign glowed red, the sidewalk stayed green. Deepfake face swaps and traditional Photoshop are explicitly out of scope for now.

  • In scope: fully AI images, catfish profiles, and the spider-in-my-burrito DoorDash refund scam
  • Out of scope for now: face swaps and traditional image manipulation
  • Coming to the Chrome extension that already classifies LinkedIn, X, and Substack feeds

🧪 Evasion, Model Attribution, and the AI-Detection Arms Race

Frontier agents given hours can eventually beat the detector: one Grok run started with a cheese essay and finally passed Pangram by producing a grocery list. Pangram can also roughly cluster which model family wrote a text in embedding space, a project called Pangram Space that isn't yet up to their release bar.

  • Specify success criteria carefully when giving agents long-horizon goals
  • Model families cluster in embedding space, muddied by distillation
  • Peter wants a slop index: which model dominates the internet's AI text

📰 Zuckerberg's AI Future Is for Everyone, Tested in Pangram

Zuck's WSJ op-ed lays out the counter-position to pacing: individual empowerment, invention over automation, and balance of power through broad access to superintelligence. Alex yoinks the full text into Pangram 4 live on air: 100% human written.

  • 'The arc of human history has bent toward putting more power in people's hands'
  • A contrast week: Zuck and Jensen say distribute everything, Pacing the Frontier says build a brake
  • Do not bet against Zuck

📰 A Final Open-Source Meme and the July Sign-Off

The show closes on a fully human-made (allegedly) meme video about the week of letters, and two milestones: 50,000 YouTube subscribers and one million total views. Last episode of July; the next one lands in August.

  • 50K YouTube subscribers — chasing the silver play button by year's end
  • One million total YouTube views
  • It's really good to be back
TL;DR and show notes
  • Hosts and Guests

  • Open Source LLMs

    • Moonshot releases Kimi K3 full checkpoints: 2.8T total / 104B active MoE, 16-of-896 experts, native vision, 1M context, KDA + attention residuals, ~1.56TB MXFP4 weights, custom license with MaaS clause (X, HF, Blog, Tech report, Baseten day-zero)

    • Kimi K3 requires preserved thinking history for multi-turn and tool calls; Kimi Vendor Verifier checks provider fidelity (Niels’ post)

    • Nistens Kimi K3 visualizer

    • Composio: same K3 success rate across Claude Code, Hermes, and Kimi Code, but up to 30x token usage difference by harness (X)

    • Post-show: Thinking Machines releases Inkling-Small, 276B/12B open MoE that beats the 975B Inkling on agentic coding, $0.30/$1.20 pricing (X, HF, Blog)

  • Big CO LLMs + APIs

    • Anthropic launches Claude Opus 5: near-Fable coding at half the price ($5/$25), claimed 3x next-best on ARC-AGI-3, 1M context; panel finds it benchmark-strong but harder to read and short of Fable in practice (X, Blog)

    • ARC-AGI 3 harness dispute: with the Responses API, preserved reasoning, and compaction, GPT-5.6 Sol jumps from 10% to 40% at a sixth of the tokens (Tibo’s post)

    • GPT-5.6 Sol improves its own inference: 20% lower serving cost from model-written GPU kernels, 15% better generation from improved speculative decoding

    • Breaking: OpenAI cuts GPT-5.6 Luna prices 80% and Terra 20%, ships faster Sol in the API (X)

  • The hack and the week of letters

    • Hugging Face publishes the full forensic report of the first autonomous AI agent cyberattack: 4.5 days, 17,600+ autonomous actions, zero-day sandbox escape; closed models refused forensics, self-hosted GLM 5.2 found 4x more exposed secrets (X, Blog, Replay)

    • Jensen Huang joins X and posts the Open Weights and American AI Leadership letter; signers grow from 25 to 230, CoreWeave among them, Anthropic absent (X, Letter, Signer list)

    • NVIDIA launches the Open Secure AI Alliance for an open defensive stack after the hack (X, Blog)

    • Pacing the Frontier: 1,273 verified frontier-lab employees, including the chief scientists of all four major labs, ask the US government for international options to pace automated AI R&D; OpenAI and Anthropic endorse (X, Site, OpenAI, Anthropic)

    • Mark Zuckerberg publishes “The AI Future Is for Everyone” in the WSJ, arguing superintelligence must be distributed; Pangram 4 scores it 100% human (X, WSJ)

    • Anthropic published research - our model hacked too! (Blog)

  • This Week’s Buzz

    • CoreWeave signs the Open Weights and American AI Leadership letter, announced first on ThursdAI

    • HiveMind’s spend view: $11K in API-equivalent tokens on $400 of subscriptions last month; Fully Connected 2026 programming taking shape (X)

  • AI Security

    • Microsoft ships MAI-Cyber-1-Flash + MDASH: 96% on CyberGym at half the cost, 16 real Windows CVEs found (X, Blog)

    • Gemini 3.5 Flash Cyber stays a trusted-partner pilot with no public API (Blog); Codex Security CLI tooling is Apache-2.0 while the service remains access-controlled (GitHub)

  • Tools & Agentic Engineering

    • MCP 2026-07-28: fully stateless core, MCP Apps and Tasks extensions, OAuth 2.1, half a billion monthly SDK downloads (X, Spec)

  • Voice & Audio

    • ChatGPT Voice comes to desktop as an agentic control layer for Codex and ChatGPT Work, powered by GPT-Live; Codex micro keyboard from OpenAI x Work Louder (X)

    • OpenAI ships GPT-Transcribe and GPT-Live-Transcribe: 41% fewer errors than Whisper-1, context prompting, $0.27/hr batch and $1.02/hr live (X, Docs)

    • xAI’s Grok Voice Think Fast 2.0 tops voice benchmarks: 82.9% quality, 0.70s to first audio, $0.08/min, already running Starlink support (X, Blog)

    • Google’s Lyria 3.5 lands in Flow Music: 3-minute songs, BPM/key control, covers, lip-sync videos, iOS app (X, Model page, Flow Music); Qwen Audio 3 also out

  • Guest: Max Spero, Pangram

    • Pangram 4: 6x larger detector, 1-in-24,000 false positive rate, token-level mixed-authorship attribution, beats 13 humanizers 98.8% of the time, integrated into Substack; Pangram Image research preview at 99.5% with heat maps (X, Blog, Image blog)

  • Show milestones

    • 50,000 YouTube subscribers and 1M total views. Subscribe, we’re chasing the silver play button

Alex Volkov
Alex Volkov 0:47
Welcome to ThursdAI for July 30th.
0:50
This is Alex Volkov back in the studio, yo. Wow, so happy to be back. There's so much news to cover from AI. I'm very, very happy to be back here, to talk to you all on the live stream as we've been doing for the past three and a half years, about everything that happened in the world of AI this week. A big, big week, for open source. We're gonna cover the most important open source release, obviously, Kimi K3 with a few guests on the show. So stay tuned for that. I'm gonna talk about all the guests. Praxis, asking how was my holiday. My holiday was incredible. It was Fable powered. believe it or not, I did try to disconnect, but also tried to like Fable Max from a remote, location. Claude Fable helped me plan all of it. we also went, every day to different places, obviously. I sent like an iCloud photos link of every day back to Claude and asked it to create a video summary recap of what we did that day. And so it had the context of what we planned, and then it also-- I asked it to look at the timestamps of every picture, look at every picture also with, with vision. and every video I ran through Gemini, 'cause Gemini is still, to me, the most multimodal video understanding model. Meta is not too bad also. And, stitched together a very beautiful hyperframes, video, and it did so incredibly. So now I have video recaps of every day. I'm not gonna post them because they have, my kids, and I don't really want that out there. But, As you m-- if you don't know hyperframes yet, we're gonna bring them to the show, I think. the intro to the show and all the new transitions that you may have seen last week are also hyperframes generated with Fable. AI agents are great at making videos. So this is how my vacation was. I agree with Milosh. We're definitely looking forward to some breaking news today. I'm looking forward to talking about some new stuff. Codex is supposed to come back with some, with some Codex Thursdays, I think they call them. and, I, I, I, we have to talk about letters, folks. We have to talk about the letters because There has been so many letters this week that I lost count. I was very happy, to be back so that I have a, a reason to, to, to, to make sense of all the open letters that were posted. so we're gonna definitely talk about the fact that Jensen Huang is now on Twitter and posting banger after banger. let me add Peter here. Peter Gostev from Marina. Welcome, Peter. How are you?
Peter Gostev
Peter Gostev 3:36
I'm good.
3:37
I'm good. How are you? Welcome back.
Alex Volkov
Alex Volkov 3:39
so Max from Pangram is going to talk about, to, to us about Pangram
3:41
4 and their new image model as well. and so I'm down to, Peter, if you, if you, if you're gonna go and, take a picture, snap a picture of your art, I'm down to test this out, with that. And we'll also say hi to Yam Peleg. Yam, welcome. We'll have two guests on the show in just a few minutes, folks. We'll have Ilija Bakic coming back to us, previously from Hugging Face, and now he's working at Prime Intellect. And we'll have a friend of the show, Philip Killeen from Base10, the author of the Inference book. both of them are going to help us cover Kimi K3, which is-- I'm super, super excited about, which is the, the biggest news this week. So I think that's gonna be… You know what? No, that's not mine. You know what mine is from this whole time that I was out? Two things.
Yam Peleg
Yam Peleg 4:24
What?
Alex Volkov
Alex Volkov 4:24
the new voice live in, in Codex.
4:27
And this guy, the- Really? … Codex micro keyboard that has a big-ass button, as you guys can see there, that you click it, you can talk to your computer. This is insane. This has changed how I use AI, like almost entirely. And I, I still have some gripes about this, but this is now me talking to my computer doing stuff versus me typing to my computer doing stuff. It's kinda crazy. And- So that's mine … Peter Gostev: when you click it, does it launch like the new voice mode, or is it like a kind of push to talk kind of thing?
Alex Volkov
Alex Volkov 5:01
You choose which one.
5:02
Okay. Peter, how about you?
Peter Gostev
Peter Gostev 5:05
I did try the voice mode, and I don't wanna…
5:10
I don't wanna be negative about it 'cause I know people are excited about it, but I, I don't know, I kinda hate it. It's like, I think-- First of all, I tried it on Codex, and I don't know what I did, but I, I talked to it on my phone, and I just said, "Oh, take a look at this." And I normally have, a flow or have a look at this, and then, launch new thread and then, to do a chat or something. And what, what-- When I came back to my desk, I just saw about, 30 chats started And I, so it like somehow launched just like 30 chats doing God knows what. And it's just "What the hell happened?" So I, I was a bit like, I, I don't know. then also, I do like some things about it, like I think it was quite competent at helping me like summarize. Like I was, talking through a presentation and trying to, think it through in my head. and it did help me adapt pretty well. And I think, by the way, as a side note, GPT models like terrible at writing. Just like if you ask them or "Help me write this." But, this model was excellent at writing. I thought like ex- only one or two examples, so not, not like big proclamation, but, it really felt "Oh wow, that's actually like way better than what I would've written."
Alex Volkov
Alex Volkov 6:26
Mm.
Peter Gostev
Peter Gostev 6:26
Wow.
6:27
and by writing I mean like saying it out loud in a way that could be- Do you
Alex Volkov
Alex Volkov 6:31
know which model, though?
6:32
The, the Live model, right? The
Peter Gostev
Peter Gostev 6:33
GPT Live model.
6:34
Yeah, yeah, the, the Live. Yeah. It just feels, like much more human, tuned towards humans. So that was nice. So like I appreciated that. But yeah, it's like it's still superficial and it does all of this, like human-like things, but it just feels like it's like a veneer of all being, "Oh, I'm so human," and then it's like doesn't know what it's saying. So like I, I don't like that combination.
Alex Volkov
Alex Volkov 6:57
You
Peter Gostev
Peter Gostev 6:57
still, you still- But the, yeah, I don't
Alex Volkov
Alex Volkov 6:58
the- You still have the uncanny valley feeling about this?
Peter Gostev
Peter Gostev 7:00
Yeah.
7:01
Yeah. It's like I, I don't mind if it's that, like human-like, and it's like super smart. Yeah. But if it's human-like and dumb, then like I don't like that.
Alex Volkov
Alex Volkov 7:11
All righty.
7:12
Yam, what about you? What is your big thing from this week? And please be brief. We have two guests that are coming on the show, and we want to do it the other day just before that, just to tell people like everything that happened.
Yam Peleg
Yam Peleg 7:21
Jensen dropping bangers on X.com.
7:25
Just, just crazy. For the first time, like just joining X.com. First tweet, like, "Open source all the way." Yes. "Open weights all the way." And then the response. It was, it was a nice week, let's- Yeah … just say it like that. It was an interesting week. Yeah … Alex Volkov: X is having a big moment right now with, with Jensen joining, Zuckerberg. Zuckerberg liked my post, folks. I'm so happy. I, I saw this- That's cute … I was like, "What?" That makes no sense. 'cause he wasn't on Twitter for the longest time, and now he's also dropping bangers. it's like Twitter is back. and Jensen is absolutely back on, on X, posting bangers. We're gonna talk about this. there's a few open letters floating in the air, and we need to make sense of all the open letters, folks, so we will definitely do that. Um, but before this, we'll say hi to, we'll say hi to our next guests. And then we'll jump into TLDR, and then we'll start talking about open source. So welcome to the show, back-- welcome back to the show. Both of you folks have been here before. Ili Bachoch. Ili or Eli? How do you pronounce this? Let me know, please. Ili,
Elie Bakouch
Elie Bakouch 8:27
Ili is fine.
Alex Volkov
Alex Volkov 8:28
but how do you
Elie Bakouch
Elie Bakouch 8:29
say it?
8:30
Yeah, Ili.
Alex Volkov
Alex Volkov 8:30
Ili.
8:31
Ili, Ili. I said that. re-research at Prime Intellect. last time you were on the show, you were still at Hugging Face, correct?
Elie Bakouch
Elie Bakouch 8:36
Yeah, correct.
Alex Volkov
Alex Volkov 8:38
Nice.
8:38
Exactly. So congrats on the move. we don't do cards here like TVP does, but we definitely congratulate our folks. And, we, we know and love Prime Intellect. We had, many, many folks on the show throughout the days. And also back on the show, Philip Gilly, author of Inference Engineering at Base10. Dev experience at Base10, is that the right way to, to, to say that?
Philip Kiely
Philip Kiely 8:58
I like to say that my job is creating shareholder value at Base10.
Alex Volkov
Alex Volkov 9:02
Creating shareholder value.
9:03
So we're gonna talk about some shareholder value that you guys created this week for sure. so welcome back to the show, folks. And obviously, we have Nisten back. Nisten, welcome, with the beard. All right, folks, it's time for us to jump into the TLDR. All righty, for July 30th, this is the TLDR. This is everything that we're going to cover on the show, everything important that happened in the world of AI. And obviously, we're gonna kick off with open source, the biggest, literally the biggest model that we've ever seen open source came out this week. Moonshot released Kimi K3, full checkpoints. previously, it was just an API launch. Now, there's full checkpoints, model code. There's a custom license we have to talk about. Two point eight trillion parameters total with one hundred and four billion parameters active. So that's a trillion-to-billion combo. I don't remember we've seen a trillion-to-billion combo in terms of like MoE plus active, with a bunch of routed experts. We're gonna talk about this KDA attention. w- we have folks here to cover all of this. It's two terabytes of a download almost, which is crazy. open weight, but not fully open source. We're gonna mention that as well. There's a few other open source things that we just mentioned, we're not gonna cover them. Catcoder 2.5 released, Solar Open 2 released, and Apertus 1.5 also. but K3 definitely dominated this open source week. and then we're gonna talk about the letters. Anthropic launches Opus 5 a day after ThursdAI. Just a day after ThursdAI. Peter, I know you've been all up on Opus 5, posting, incredible demos, incredible, like, 3D things. People are building games end to end with it with gauntlet prompts from Matt Schumer, so we're gonna talk about Opus 5. They claim it's a new state-of-the-art model. they comes very, very close to Fable on coding at half the price. so Opus 5 has been out for, for a minute. there is a thing with RKGI that we must talk about. Opus came out very hot at RKGI 3, as saying that it's three x higher than the next best model. And then today, OpenAI folks decided that, no, the way RKGI was tested did not, make use the, the, the best abilities of their model, so maybe we can cover that 'cause I think it's also important. And, I think that this is most of the big labs and, big, most interesting things. there's one more thing, Peter, that you mentioned that we must talk about that, Codex Sol, I think it was Sol, helped improve its own inference and runtime. We, we'll, we'll mention that as well. Please remind me a-afterwards. there's also a big week in voice and audio, folks, because OpenAI launched something new called GPT Life Transcribe and also GPT Transcribe, so two separate models for transcription. Maybe we can run them on air, very, very cheap. keywords, language hints, context, and previous turns, for, We should, we should test this out. Groq k-- also punched back with Groq Voice Think Fast, which is their version of reasoning and voicing, et cetera. and then also Qwen Audio III launched and Lyria three point five, in, in Google launched, this week as well. Let's see what else. Microsoft-- No, this is, this is old. Microsoft came out with, a cyber thing, so we're gonna talk about that. And, I think the most interesting thing from AI coding and agents was that MCP moved forward with MCP version two, which is stateless protocol, remove sessions and handshakes, and make it much, much easier. and MCP is kinda not going anywhere. And I think this is what I talked about. Mic-- Zuckerberg published this, a-an op-ed called The AI Future Is for Everyone, which I think is very important to read through. This is the biggest news, of AI this week. Most of them we're gonna try to cover before we start with open source, just before we go to open source, folks. Anything else big, important from this week that this did not include that we must add? As an open question to everyone, including the listeners to the show, including the hosts and the guests on stage, and oh, yeah, the last thing is Pangram, al-also folks who've been on the pod before, have released Pangram version four. It includes image, verification, whether or not the images were created with AI, but also, a new and updated version of their model. Max from Pangram is gonna join us at the end of the show to talk about their updated model, their research and kinda, facing the, the, the responses from folks. Also, they integrated with Substack, so that's i-incredible to see. All right, I think it's time for us to go to not keep our guests waiting too much. Let's go to open source.
13:40
Open source AI. Let's get it started All righty, let's get it started, folks. This week has been a huge one, and by huge I mean in terabytes and a-and, and just like physical weight as well, because Moonshot released a state-of-the-art open source, probably the king of open source, probably the best model, Kimi K3. We all remember K- Kimi K2, Kimi K2.5, great, great, great models. and this week we, w-we got a nearly a three three terabyte-- no, two terabyte per, but, but like three T total parameters and 104 billion active parameter model MOE. to help me cover this, we have Ily here and Philippe. Philippe, you guys, were day zero supporters of this model. we're still working at CoreWeave to be zero day supporters of models, would love to, to, to have you here to represent kinda the inference side. Ily, you're known for, breaking down technical reports and, doing long threads and that the people love. So I will maybe start with just, covering a bit of the, just a bit of the evals of this model before we go into this. Peter would love to hear from you as well because, this model is breaking, design arena. This model takes number one by, a big margin as well. So le-let's go through the evals a little bit. but also, we'll just talk about what's cool about this, and, feel free to dive in as, as much as possible. So we have eight hundred and ninety-six experts in this model, with sixteen routed to. Native vision, so it's multimodal fully, unlike previous big models. I would love to hear about KDA and attention residual stuff. And we have a one million context window, which I assume is not cheap to run with this model given its, insane size. Reactions. folks, Phillie-- Philippe and Ily, feel free to jump in here. Yeah.
Elie Bakouch
Elie Bakouch 15:32
I think the design is super impressive.
15:33
it's the first time we have, a model at this scale, in the open. even, for instance, if we compare to closed source model from big labs, Grok four point five is one point five trillion, which is like almost half of Kimi K3- Yeah which is crazy to say, right? and obviously to, to train such a big model, you, you need to, to make some choices, in the architecture that, that make it, good at, inference. And I think that's like the, the, the, the main point of the, of the model. The focus is, making it, possible to do inference on this model with, long context, with KDA and, and, and so on that we can, deep dive in, a bit later. And also, having a stable training with, things like attention residuals, kind of like, this thing that they do for Muon. they, they, they also do some clipping to, to have, stable a-activation and so on. So yeah, it's, it's pretty impressive and, and, I think what's even more impressive is, like, how they managed to scale their pre-training recipe from Kimi K3 to, from Kimi K2 to Kimi K3. It's like a two point five x, scaling efficiency, which is very, very impressive. it means, like, at the same compute, they basically get two point five x performance for the same compute, which is Yeah, which is super impressive, right? Yeah.
Alex Volkov
Alex Volkov 16:48
Yep.
16:49
I, I, I think that, just the scale is, i-is, m-- how should I say? The scale is the story here. It feels like the scale is the story. Yeah. The fact that we got this as open weights, I think is incredible. not all of us got this as open weights. Philip, we, we have to talk about the license as well, the-- this is not a standard, fully MIT license. there is a reason why not everybody has it yet, but folks are working on this. but the scale is the story here. And, what does it take to host such a model, Philip? Would love to hear from you. You posted about this publicly. let's just talk about, the number of GPUs required, just, just out of the bat, like how to even scale this responsibly.
Philip Kiely
Philip Kiely 17:25
Yeah.
17:25
So the first thing to understand about a model like this and trying to get it up and running is that it's not explicitly designed for NVIDIA Blackwell GPUs because these labs are, are not running on these. So the, the-- if you read the technical report, they talk a lot about really wide expert parallelism. The idea that you're gonna have, dozens and dozens of GPUs, and you're gonna split the model across that and think about how to do a lot of expert routing over a very sparse set of experts in order to have efficient inference. We obviously have different hardware and different constraints in that we are serving on NVIDIA Blackwell. We're serving on GB300s. The reason that you want the, 300s is because of the VRAM amount. Obviously, with, two point eight trillion parameters, even in the native MXFP4 weights, you need about one and a half terabytes of VRAM simply to load the thing. Then you need quite a bit of space for KV cache, especially given the million token context window.
AI
AI 18:29
Mm.
Philip Kiely
Philip Kiely 18:29
So there's quite a bit of, of challenge in, first
18:33
off, simply adapting the new architecture to the, the GPUs. we were fortunate to work closely with the teams behind vLLM and SGLang on understanding, the, the front end, understanding the input and output formats, the new KDA, the latent MOE, like understanding all this stuff, working on kernels together. and, and we're able to, make a couple contributions back to open source as well as part of that process, which definitely felt good.
Alex Volkov
Alex Volkov 19:01
Nice.
19:01
first of all, shout out to you guys for like Day Zero support for this and folks definitely should try out Baseten, via OpenRoute or directly. but also I, I do wanna ask you guys someone to explain MXFP4 versus NVFP4. I think there's a difference. please explain it to me like a full dummy that does not understand, this at all. Yes. I feel like maybe you can take this, a stab at this.
Philip Kiely
Philip Kiely 19:24
So N-NVFP4, the NV stands for
Alex Volkov
Alex Volkov 19:28
NVIDIA.
19:28
Yeah.
Philip Kiely
Philip Kiely 19:29
And it's specific to, the, the Blackwell GPUs.
19:33
So if you look at, for example, a Nemotron model, it's often gonna come in native NVFP4 weights because NVIDIA's training it, they're targeting their own hardware. MXFP4 is a more general four-bit floating-point format, that is supported on a wide range of hardwares. It's a more like open industry standard. There's slight differences between the two. They're both something called micro-scaling formats, which means that you are quantizing like very small blocks at a time. to my knowledge, MXFP4, your block size is thirty-two, whereas NVFP4, your block size is sixteen. In NVFP4, you also have a second scaling factor, which is a thirty-two-bit, number that's going to be a, a blockwise or a tensor size scaling factor. The reason that you can have in NVFP4 smaller block sizes and the second factor without really hurting your performance too much is because the Blackwell GPUs are like physically built for this exact, number format. which is why the, the MXFP4 is a more general format because it doesn't assume this like very specific capability of the hardware. and it still allows for quite a bit of preservation of dynamic range through micro-scaling in this four-bit format. And then, of course, given that we're doing quantization aware training or not we're doing, given that the lab is doing quantization aware training, you're not really seeing any sort of quality loss for going to four bits because the model natively is four bits.
Alex Volkov
Alex Volkov 21:03
And it's at almost two terabytes of just weights
21:06
even- Yeah, about one and
Philip Kiely
Philip Kiely 21:08
a half terabytes.
Alex Volkov
Alex Volkov 21:08
Yeah.
21:08
Yeah. Even, even at FP4 format. Uh, Ilija, I wanna, I wanna talk about, their attention stuff like KDA, and, you, you mentioned Muon as well. We talked about Muon like a couple times before.
Elie Bakouch
Elie Bakouch 21:18
Yeah.
21:19
I think what's, what's interesting is that most of the stuff that the Kimi K3 use were already in the open previously. if you look at, KDA, so which is the linear attention used by, Kimi, they published a paper about it previously, and it's also very similar to, the attention that Qwen use, which is called, Gated, Delta Net Attention. if you use that-- If you look at Latent MoE, it's, work from NVIDIA. If you look at the attention residual, it's also something that they published before. so basically every part of the stack, were already in the open and like the, the, the goal of, Muon as well were already in the open, and they use the exact same formulation as, GLM five point two, which is the per head Muon. and yeah, it's, it's quite interesting to see that like they, they basically take tho-those building block of, things that they, did in the past. for instance, the, the, the people working on linear attention also contribute to Flash FLA, which is the, the linear attention, library. And they just scaled it to, to one trillion, two… Sorry, not one trillion, but three trillion
Yam Peleg
Yam Peleg 22:20
So yeah.
Elie Bakouch
Elie Bakouch 22:21
And the, the points, so maybe I can, I can, I can
22:23
talk about a few points here. Yeah. the, the, the point of, linear attention is that, the big issue when you do long context is that your KV cache is growing, super fast, right? and the… Exactly, yeah. And the, the point of linear attention is that you don't have a KV cache. You have a fixed state matrix that basically don't go with, the number of token that you, you do inference on, right? so this is good for decoding and, this-- the prefill is also a, a bit faster. it's a different format. It's actually quite complex to, to understand, and, and there is a lot of optimization there. but, you c-- you can-- you also get faster performance at prefill and yeah, obviously the, the big win is th-this, fixed state KV cache, which is basically compressing the KV cache into one state matrix, which is smaller. But this is also fun because, like this lead to… I, I don't see any issue with provider yet, but I'm expecting some issue because for instance, when you do this, this, the, the… When you have the state matrix, if you want to rewind, in Claude code, all those, RNNs-… you have the option to rewind, right? And this is basically taking your previous KV cache, token. You cannot do that with state matrix because,… the, the state matrix, you, you basically need to store each checkpoints, of the, the sta-state matrix. So I'm expecting some, some fun stuff here that might happen in, in the, in, in the next few day, in terms of input caching, with, linear attention, which is not as trivial as, yeah, as full attention Yeah. And also maybe one, one super cool, innovation is their thing called attention residual, which is basically doing attention on the residual, path of the model. And this is super important because, for instance, if you, if you take a deep layer in the model, to get the information of a previous layer, basically you have to carry this information through the residual path. And with attention residual, you, you basically don't do this, but you let the opportunity of the, deep layer to attend to the residual of the, the, the previous layer, which is very beautiful in a way, and, and quite simple. So obviously, there is some info optimization that they do. They, they do this method called block, attention residual, which is just basically taking the, the residual of the, the… It's a sliding window of, attention residual. so yeah, it's, it's very-- it's very nice method to, have a very good information flow, into the model. And I, I think those method like make even more sense when your model is, as deep as Kimi K3 is.
Alex Volkov
Alex Volkov 24:58
Yeah.
24:58
When, when you say deep, you just mean like in terms of like just number of layers?
Elie Bakouch
Elie Bakouch 25:02
Yeah, exactly.
25:03
Number of layer. Yeah. Yeah.
Alex Volkov
Alex Volkov 25:04
last question.
25:05
Not last question for you, like another question for you. I, I read very briefly the technical report. Was on vacation, but like it was, very, very interesting. they don't use yarn or rope. They use something called rope. Rope. Can you talk about that?
Elie Bakouch
Elie Bakouch 25:16
Yeah.
25:17
So basically, when you do, when you do model training, often you use this thing called rope to take into account the position of the, the, the previous token. what-- They, they don't do this because, they use, linear attention that basically have this recency bias, into the linear attention mechanism, and they don't use rope in the full layer, which is something that they ablate in the Kimi linear paper. There is an ablation about it, and it-- they find that it's, it's a bit better. And it's actually quite nice because, the intuition behind it is that they let the, the linear attention layer focus on like local information, and they have this global layer, which is like the MLA part, gated, MLA Which, doesn't have any information about the, the recency of the token or about the position. it's really like a global layer that, just take, every token, the, the, the same way there is no bias, yeah. So it's actually quite nice. And like this, NOPE thing. So NOPE is just stands for No Positional of Encoding.
Alex Volkov
Alex Volkov 26:19
Ah.
Elie Bakouch
Elie Bakouch 26:19
Is, coming from, like a bunch of people.
26:22
When, when we train SmolM3, we also use that, for instance, and we also did, RoPE only on a few layer to let some layer focus on locality and some other layer focus on globality. Yeah.
Alex Volkov
Alex Volkov 26:33
I think, Ili, I just wanna, shout out this beautiful visualization,
26:36
from Nisten and Fable, I think. I, I don't know Nisten. Was this Fable? of the layers and, o- of Kimi. Yeah. it's just beautiful.
Nisten
Nisten 26:45
It, it allows you-- Yeah, I, I just put it so that the volume of it,
26:48
it corresponds to, the actual size- Yeah in megabytes. And with Ki- with the other models, it was fine, but with Kimi, it just gets so messy 'cause- Yeah … it's so big. But, yeah, I just gave it the config.json from, from Hugging Face, and that told it to also explain, what the, what the Kimi attention is and, you can also see the MLP layers, on top, which are like the-- Sorry, the MTP, layers, which are, like, the, the lookahead decoding.
Alex Volkov
Alex Volkov 27:15
So we'll definitely go ahead and, and share this link on the
27:18
show notes so people can play around and, people who are interested in, the, the, this scale and size of models. Uh, Filip, let's talk about inferencing this beast.
Philip Kiely
Philip Kiely 27:28
Yeah.
27:28
the, the first thing to understand is that like we've been here before as an industry. DeepSeek R1 came out and it was six hundred and seventy-one billion parameters in like January 2025. And we were all sitting around looking at it and looking at our H100s and looking at the model and being like, how the hell are we supposed to do this?" And we tried a bunch of stuff. We tried multi-node inference. Turns out that doesn't work super well, at least didn't work back then. I guess technically we're doing multi-node on GB300, but because of NVL 72, you don't have the same restrictions of an InfiniBand interconnect. but anyway, like there's been… And then Kimi K2 came out, that was a trillion. We'd never seen a trillion before. So it's, it's not the first time that a new model has come out at a sort of size and architectural complexity that it forces the entire industry to take, a couple big steps forward and do a month or two worth of, worth of progress in a weekend. I think that the, the big things around this model are primarily around throughput, I think. obviously there's, there's all the things you can do for latency, right? You can build the speculator. You can, do… make sure that you're doing great cache aware routing so that you just get prefill as much as possible. You can, set your batch size smaller and you mess with your parallelism config until you have a latency-tuned config. Like all that's kind of standard. The really interesting problem, in my opinion, with this new model in this new size is making throughput exceptional so that we can, serve as, as many tokens as possible. some, some big levers for that are disaggregation. I'm actually doing a talk with NVIDIA at their Dynamo day later today, where we're gonna talk about some disaggregation results from Kim-- from GLM 5.2. But yeah, you, you basically want to, separate out prefill and decode onto different workers so that they're not, feeling with each other. And for bigger and bigger models, this becomes more and more important. And then the other bit is, if you have a model like this and you are serving it across multiple cloud providers, multiple regions, you want to make sure that you are routing everything appropriately because the number one thing that's gonna determine your, your throughput and, and your cost is going to be your cache hit rate. That's like the whole game. and especially for models like this where you're mostly doing multi-tone agents, you're mostly doing like coding and that kind of stuff, you expect very, very long input sequences with very high cache hit rates. But that actually introduces a new problem in, in the tokenizer, which we published about, which, I was I, I wouldn't say surprised, but I would say very happy to see so many people interested in such a sort of novel and niche part of the stack as a tokenizer. if you think about a long input sequence, hundreds of thousands, up to a million tokens, and a very high cache hit rate where you've seen, ninety-nine percent of these tokens before in the last turn through this agent. a-as an inference engineer, I've always been able to assume that like tokenization time is negligible. But once you get your sequence that long, your tokenization time on a sort of base implementation can actually be hundreds of milliseconds. And your t- your TTFT off a cache hit is also a few hundred milliseconds. Su- suddenly tokenization is half your total time. and so yeah. So Michael, one of our engineers here at Base10, who just absolutely loves going after really tough performance problems in really niche areas of the stack, created Base10kenizer, which I think is a fantastic name, and I won't do otherwise. Yeah. which, which is his, his sort of optimized implementation of a tokenizer for Kimi K3. Kimi K3, by the way, also introduces a bunch of new special behaviors that it expects the tokenizer to support around certain tokens like stop. yeah, we can see some of the performance metrics, and you can see how as the input sequence gets longer, the gains get bigger. And then if we scroll down, there's a third image in here that I'd like to see, or if you just hit the arrow, we'll see that third image Yeah. So we can see here how, massively shrinking the tokenizes-- tokenization time actually does like materially improve time to first token in this very niche case. And it just so happens that for Kimi K3, we actually do see traffic that looks like this extreme case. So that's a great example of like as, as, as engineers, we love to say it depends as the answer to every question.
Alex Volkov
Alex Volkov 32:13
Yes.
Philip Kiely
Philip Kiely 32:14
But the more and more you do inference, the more and more
32:17
you can understand the sort of shapes of the traffic that you're getting, then you can deploy very, very targeted optimizations like this against certain like long tail behaviors that actually show up very often and build a system that just works better end to end.
Alex Volkov
Alex Volkov 32:34
Yeah, I think that's great.
32:35
Uh, I, I do wanna call out and maybe ask folks here on stage what, what your thoughts on this as well, because, more and more I think, Peter, we're moving away. W-w-we'll talk about this at the end of the show as well. We're moving away from just completion of tokens. It's all agentic and it's all reasoning, it's all going back and forth. Niels Rogg, your previous coworker, I think Eli from, from Hugging Face, he posted this, and I really wanted to put this on stage because I think it's very important. "Kimi was trained on preserve thinking history mode. For multi-turn conversation and tool calls, Kimi K3 requires the complete assistant message returned by the API to be passed back to messages as is, including reasoning content and tool calls, not just content." This is very important. We saw this, I, I please remind-- Nisten, please remind me like another, I think it was Minimax folks who also said that return thinking is very important for the model performance. Do you wanna discuss this for a little bit, why this is so important and why like other harnesses that don't do this kind of miss out? Eli, I see you shaking your head like can you tell us about like- Yeah … re-returned reasoning, why it's important, why the model is trained
Elie Bakouch
Elie Bakouch 33:35
this way?
33:35
So first, so maybe, like you, you, you will need to double check this, but I'm pretty sure that most of the model, that are open source now, for instance like GNM five point two, Minimax, like Kimi use this kind of, preserve thinking, scheme and like it's very simple. I think it's just that during the, the reinforcement learning stage, the model use this kind of, th-this kind of preserve thinking history mode. So if you suddenly like at inference don't use it, the model will just don't understand what's happening and just gets very bad response. I think this is also one of the reason why, Codex, perform better on ArcadeGI, like the, the one of the reason was, about the preserve thinking as well. So it's super important. I think like there was a big, a big issue at some point with Intelli, Intelli thinking, with provider and so on that like kind of lead to very bad quality of our models. So yeah, like super important and Kimi have this effort called like, WandA Verifier thing where they basically test, properly like that's the number of tool calling, the number of output token match, between their API and all the model be served on different providers. So- Yeah … yeah, this is the source of truth of, this kind of thing to, to detect if there is some issue like this.
Alex Volkov
Alex Volkov 34:51
So we actually talked about Kimi Vendor Verifier when
34:54
Kimi two point five launched back, what, September of last year. if I'm not mistaken, September or October of last year. this is an effort by them, which is basically an eval that's held back that nobody knows. There's, like few questions here, but basically the whole eval is held back. Philip, you'd be, great to see that you guys are like, very high up here, I think in terms of numbers. DeepSwizz, TBD on, on, on, on Baseten in terms of Kimi Verifier. They're still testing this. when we need to also submit, but it looks like most of the folks now, at least on OCR bench, look like they're up to par, like with the, with the, with eighty-ish percent of, of the scores. previous-- The reason why Kimi did this is because, different providers hosted it differently and then didn't use the interleaved thinking thing that we just pointed about. And then people were complaining like, "Hey, this model came out with these like great evals, but I'm not getting these results." And people started complaining. And then, even I think it led to Open Router adding some, which provider does which, quality and quantization like efforts, et cetera. so this is great to see that now like f-folks are working with Kimi, closely, and, and getting up to par with this model is served the same, across everything else. Philip, would love to hear your comment on this.
Philip Kiely
Philip Kiely 36:01
I think Kimi Vendor Verifier is a great effort by the lab to, create a
36:06
public standard for what it means to run the model When I think about quality in model, I don't think about the raw output. I think about the fidelity to the original vision of the model. How close is the thing I'm serving to a 100% pure expression of what the people who trained the model wants it to do at any given time? And so I do think that as an industry, a quality discussion is substantially over-indexed on quantization because as you can see here, everyone's running the same thing. It's, it's a natively MXFP4 model. There's, there's no benefit to quantizing it, and yet you can see varying quality. And that's because predominantly quality comes from the inference engine all the way down to the kernels. It comes from, are you fixing all the bugs and race conditions that exist throughout the stack? Are you implementing proper support for every single one of the quirky little behaviors of the model in your front end, in your parser? So I think that Kimi Vendor Verifier is a great effort. I would love to see more labs be explicit about verifying the correctness of public endpoints. And I think that it goes to show just the difficulty of doing really high quality inference- because it's actually a lot more than just like making sure that the logic distribution out of whatever quantization of your weights is reasonably well matched to the original model. There's a lot more that goes into it than that.
Alex Volkov
Alex Volkov 37:39
Yeah, 100%.
37:41
I do wanna talk about the performance of this model and what it is good for. Folks, we're getting like a nearly, and Philip, and Peter, sorry, I wanna tag you in here as well. We're getting in like a state-of-the-art huge model, that's served in open, open weights, including the kind of the, the training recipe, et cetera. we don't know many things about the data that went in there. There was like speculation about some, distillation accusations floating around with like Chinese labs. We're-- we, we've talked about this all on the show before. We're getting a state-of-the-art model in the open source on DeepSwe, which we talked about, which is the new the Nous SWE-bench, let's call it Deep SWE from, from I think that, that, curve I, I… if I'm not mistaken. Kimi K3, is just behind Fable 5 and GPT 5.6 Sol, beating GPT 55 and beating Opus 4.8 at 67%. On Terminal Bench 2.1, we don't have Wolfram here, so we didn't run, Wolfram is on vacation. World is on vacation. We did not run, Wolf Bench on Kimi K3, but we'll definitely do that once it's up. Kimi K3 on Terminal Bench, which we love and test, plenty in Terminal Bench 2.1, takes second place just before behind Sol, beating Fable, beating Opus 4.8. I don't see Opus 5 here. It's because they released like what within a day of each other, I think. so they probably didn't have t-tons of time to, to evaluate. Frontier, SWE, Kimi K3 gets like a very high score as well. And I think on, on, Arena and Design Arena, like Kimi gets like incredible scores. Peter, would love to tag you in here to tell us about some performance here. what's going on? How is this model received? what's the quality compared to other models and other open source models?
Peter Gostev
Peter Gostev 39:20
Yeah.
39:20
It's so often the case when open source model comes out and then we, we receive these big headlines and then we do more benchmarks and it kinds of-- kind of drops off the cliff a little bit. Like that happens so often. But here it doesn't seem to have happened. And, we had our scores. I know you mentioned like a couple of ours we did. I'll just share my screen. we did, Agent Arena. So this is like a few days old, so maybe the score's a little bit different if you go on our website but, the, the quality of the model, so it ranked fourth, on, on the arena. So this is just, just above like Phi-5, just below Phi-- GPT-5, 6, below Fable. I, I think there were like a bunch of benchmarks that showing it like above Fable. That's probably not, not n-not something that I think is quite right. But yeah. Oh, yeah, perfect. Thank you.
Alex Volkov
Alex Volkov 40:13
Yeah.
40:14
Just before we continue, just one second. It looks like our, our, our great guest-- folks, first of all, thank you so much. One of the coolest things that I get to do in, on Thursd AI is like deep dive, and I'm able to do that because, like w-we run the show, that's what's interesting to us. it's always great to deep dive with you guys. Ili, thank you so much for coming back. Philip, thank you so much for coming back, for help us deep dive into open source, and how this model is trained and how the sausage is made. Please, welcome to co-come back and the, the, the next big release. I really appreciate your time here, folks. Thank you. Of
Elie Bakouch
Elie Bakouch 40:40
course.
40:40
Thanks for
Alex Volkov
Alex Volkov 40:41
having me.
40:41
Of course. Cheers, guys.
Elie Bakouch
Elie Bakouch 40:42
Yeah.
40:43
Thanks. Always a pleasure. Bye. Yeah.
Alex Volkov
Alex Volkov 40:44
Cheers.
40:45
All right, Peter, let's, let's go. So we're looking at the Arena agentic, y- like score, and this is the best open source on there, for sure, or- Yeah … open weights, not fully open source.
Peter Gostev
Peter Gostev 40:55
Yeah.
40:56
Yeah. So i-i-it is the best. It's very close to the frontier. we are going to publish some more numbers on the kind of the cost curve of it. from memory, I should probably double-check. but yeah, it's it's close to the cost, kind of Pareto frontier as well. but yeah- At
Alex Volkov
Alex Volkov 41:15
three dollars per million token inputs and fifteen dollars per
41:19
output, it's not a cheap model to serve. we can talk about the pricing thing n-now that like our folks have dropped. Yeah. Nobody here signed anything with Kimi yet. Sure. every model provider every model provider serves it at the same exact price. Yeah. open route. If you go to OpenRoute- This is the--
Peter Gostev
Peter Gostev 41:36
Yeah, so this is something like I, I was kinda
41:39
confused by, and I guess technically there is one that broke the ranks.
Alex Volkov
Alex Volkov 41:45
Yeah, Morphs.
41:45
Morphs. I don't
Peter Gostev
Peter Gostev 41:45
know who they are, but, yeah, they broke
41:48
the ranks and, and dropped it. But yeah, their, throughput is lower as well, so you can kinda see it, thirteen tokens per second. Yeah. Which is terrible. yeah, like to me, my expectation is, the promise of open source, open weights is that, weight's out there, it's a popular model, then everyone does the things that our guest talked about and optimize all of this. Everyone like goes away, finds the best thing, best way to run it, and then they manage to drop the price or increase-
Alex Volkov
Alex Volkov 42:13
Yeah
42:13
… Peter Gostev: the throughput or something like that. It seems just like everything's the same. So I'm like, I'm very confused. I know there was like a few tweets certain docs posted saying, was it, "Oh, the price fixing is not cool." So I, I haven't seen anything. I don't know what the agreements are but-
Alex Volkov
Alex Volkov 42:29
Here's what I'll say, given that I do work for CoreWeave
42:32
and we are in the progress probably of putting this model up, we need to reach out, and I can talk publicly about what the stuff that are public. Um, this mo- this license for Kimi is not MIT license. This is a Kimi K3 license specific. It's a version of MIT license with an exception for companies that host models. If the software is used for any of the licensee commercial products or services that have more than one hundred million monthly active users or more than twenty million US dollars in monthly revenue, Kimi K3 must be prominently displayed and a license, Model as a Service, companies need to acquire a license from Kimi, which means working with them and signing a license. I haven't seen this license, so I'm very, v-very clearly I can like only speculate of what's in there. But when you're looking at every model that serves this model, sorry, every, every company that serves this model serving it at the same exact price, then it's, it's fairly clear that, hey, Kimi is trying to also make a lot of money on the public market. and, revenue sharing and price fixing, who are speculated by the creator of OpenCode, Dax, is, is probably in there. We don't know. This is all speculation. but this is how the open weight thing turns. I also wanted to mention at this point of open weights, open source, like we've talked about open source, open weights, we love it. Thank you for folks for posting the, the technical blog, et cetera. there's now an iteration of how an open weights model appears in our lives. taken from, Mistral dropping a magnet link. "Hey, everybody, yolo, take this model, run this, do whatever you want." Now, slowly, all the Chinese labs are first posting an API. They're saying, "Hey, we have this open source model, but here's our API." I don't think that I would use an API for Kimi versus an API for OpenAI, for example, just, just 'cause I want to. I would prefer an OpenAI API, GPT 5.6, so on, et cetera. And then they're saying, "We're going to open source this in a week. Everybody get, builds." This is a very standard flow now for everyone. For MiniMax, for Kimi, for, for, literally everyone. And now the model weights that drop is also, you, you have to sign an agreement with them. They, they wanna make money. the, the whole, "Here's the benefit of open source, use it and take it for everything that you want," is kinda like getting augmented by, by capitalism. so I, I definitely noticed that w- w- with recent updates from models.
Peter Gostev
Peter Gostev 44:57
And maybe I'll just add one thing as, is that, the front-end
45:00
performance seems to be really good and the kind of the defaults of websites, for example- seem to be quite good. and that's a, that's a kind of thing we saw, as well from GLM, that it just seems to be, like, pretty competent if you say, "Build me a website," it just looks nice. which I must say wasn't the case with, so many previous models or, OpenAI models, for example, but was good for, Claude models. And this is not a statement that they distilled or something. I don't actually think that's true. They probably just worked hard to make sure that the defaults are good.
Alex Volkov
Alex Volkov 45:30
Yeah.
45:31
And y- you guys have tested this on the image to web dev arena, right? and Kimi K3 Max is now very close to the top- Yeah … just behind Opus five, which we're gonna talk about next, before we-- I think we talk about letters, 'cause I, I think, Peter, you did a lot of tests with Opus five. Would love to chat about them as well. but yeah, the design, capability of this model is quite something. at, at the price performance, I think this is, the cheapest, best, like, visual model. I have to say the, the way for me to test these models is, at least after, after vacation, I've used Hyperframes, which is a way for your models to create, videos, beautiful videos, like the transitions that you guys see here. And I was not impressed with Kimi. I was impressed with Kimi's just general approach, but I sp-spun up, a bunch of, like-- I think I spent thirty million tokens on, on, a video, and the videos were not very impressive design-wise. Maybe this is because, one of my tests, I actually had the same video done with Fable, so I wanted to compare. One of my tests, Kimi just stole Fable's homework and literally went and saw what Kim-- what, what Fable did and did almost the same thing, exact text, et cetera, just like a little bit of a reskin. That was really funny to me. and the other one just did not look good. I need to test it on the actual website. but yeah, we have a frontier open source model. This is great, folks. As always, we do, we do applause when open models gets released. Good job, Kimi. Yam, Nisten, any, any, experience with Kimi before we move to Opus and then open letters?
Nisten
Nisten 46:55
Yeah.
46:55
I, I was using it for security review. But then I noticed that Fable and Opus removed a lot of the really strict classifier that was just blocking everything.
Alex Volkov
Alex Volkov 47:08
Mm.
Nisten
Nisten 47:09
Like for example, for, for medical apps, before I couldn't
47:12
use Fable at all, but now I can. So they, they really relaxed the safeguards, so now I can actually do- On
Alex Volkov
Alex Volkov 47:19
Fable, you mean?
Nisten
Nisten 47:20
Yeah.
47:22
Yeah. So now I can actually do the security reviewing and hardening with Fable. So that was, that was interesting, on, on, on, on their end. again, yeah, so that's about, that's about all I- Yeah … I used it for. O- otherwise it is, it is too expensive. Like the-- It's pretty-- It's still pretty hard to compete for them with subsidized inference and Max plans of the larger labs.
Alex Volkov
Alex Volkov 47:50
Nisten, one, one last thing I wanted to show- Somebody
47:52
used Kimi K3 inside the Claude Code, harness and in Hermes and in Kimi Code, which is their harness. this was just m- shocking to me how much harness makes a difference. I don't know if you folks saw this. Composio, did a test and ran, Kimi K3 through three, agent harnesses, Claude Code, Hermes and Kimi Code and, they did the completed task in the same rate. However, Kimi K3 did 340,000 tokens, where in Hermes and Kimi Code they did 60,000. Just look at this graph. There's a- How many more tokens within Claude Code they took from the same model, and price-wise it was like 30x, more tokens. Price-wise it was just like w- w- 10x more price just because using different harness. This is crazy to me. There's a- Just absolutely crazy … th-
Nisten
Nisten 48:39
there are two reasons I think for this.
48:42
One is inside Claude Code you're gonna have, your memories tend to grow a lot, so you have the 20K system prompt, and then you might end up very easily with 50K of memories. And then you're constantly sending those at $3 per input price. Now, input price when you host inferences is a lot cheaper and you want to, maximize the caching, a- a- around that because that's, as an inference provider, that's, that's where you make the most money.
Alex Volkov
Alex Volkov 49:09
Yeah.
Nisten
Nisten 49:10
So that's why, but, and the, so that's one thing.
49:13
The other one is, I think with Hermes, 'cause I haven't tried, Kimi Code, it tends to make bigger edits at once, whereas Claude Code tends to make, a lot of tiny, tiny edits. That's just how it likes to do it. So that's also just gonna kill your, y- your, your, your input pricing. So yeah, that, those are the two reasons why I see that. I don't think it's necessarily because the harness is bringing out a bad result or not. a little bit because it has too much junk from the system prompt. But, yeah, that's, that's what I see. It's just, around the, the input caching and also how it works.
Alex Volkov
Alex Volkov 50:00
All righty.
50:01
We're, we're, we're back, and now we're talking about Frontier Lab launches. And folks, we have a new state-of-the-art chonker that was released just a day after the last show. Opus 5 is here. Opus 5 is Fable 5's Not so smaller brother. there is a difference in the training thing. I think like the, the Fable mythos levels and the GPT-6, whatever is gonna come out as like new train and, and Opus and GPT 5.6, whatever. There's a, like results of continuous improvement on the pre-train. That's at least what folks are speculating. we have Fable-- sorry, we have Opus 5, that does basically state-of-the-art results on pretty much everything. They compare it to, they compare it to Fable 5, the previous Opus 4.8 and 5.6 Sol. you can see on Frontier Bench, it's the state-of-the-art on Frontier Bench, forty-three percent, on agentic terminal coding. Let's see. On Arkagi 3, they have thirty percent. We're gonna talk about the Arkagi 3 result in a second, because, the GPT 5.6 Sol, result here is very low, and there is a reason-- there is a reasoning why that is. on agentic search, this is like a great model at ninety percent. You can see like it's pretty much state-of-the-art across the board. computer use is very high, seventy percent on OS world, where previous models are really, really good at it. This is like we're almost at the promise, at the, at the place of like augmented computer use. Speaking of computer use, as though we invited LDJ joined us. LDJ, welcome. We'll talk about Opus 5. agenda coding, DeepSwe, is still, leading GPT 5.6 Sol is leading DeepSwe, but pretty much everything else is, is, is very high here. Peter has been testing this model, as I saw from your, from your public, thread on X very extensively. so first of all, shout out to Anthropic for releasing banger after banger. Fable 5 came out like what? Also this, this y- like in July. Like they announced it, and then they gave it, gave it, gave it back in July. so July was very hot for open source… sorry, for frontier level models. not everybody loves Opus 5. Some people do not love it. folks, let's open up for comments. How is Opus 5 doing for you? Folks in comments, folks in, in, in, in, live, please also comment and tell us what is your experience with Opus 5. Peter, would love to hear from you 'cause you've been building a lot.
Peter Gostev
Peter Gostev 52:18
Yeah.
52:18
I'm- You've been testing a lot I'm definitely curious to hear from people in the comments 'cause there's, there's definitely been some kind of mixed feedback, I, I think from, from different people. it's gone out like amazingly on the benchmarks. We have also benchmarked it and, for us, Fable is still higher on the agent arena. Which kind of makes sense to me. I think like that kind of what, what we should expect. on the text arena for us also Fable is higher, but so is like Opus 4.6, 4.7. Yeah. so yeah, it's kind of Uh, which I think matches maybe some of the vibes, as well. So I think our scores like came out like reasonably, to, to kind of match the, match the vibes. And then the, the kind of This weirdness right now what's happening is that I think everyone who's extensively using both Fable and Opus says "Oh, Fable is so much smarter." Yeah. Like I think that's the kind of general feeling that, that people get. And then you look at the benchmarks, you look at three G stuff that I was just playing with as well. Like Opus is better, like in terms of like the… Let, let's just take three G. Again, three G no one cares about, it doesn't really matter. Like I guess when we give-- do triple A games, maybe it does, but, but, it, it's not like the most important thing. But it's like it's measurably better or like very obviously better. so like what, what happened there? Like is it, is it just is it the better model? My sense, and this is like complete speculation, is that, bigger models are harder to do RL on. So for example, if they have a bunch of new three G data that's arrived, it's probably-- Let's say, I don't know, it arrived in May. and, it would be, I think, reasonable for them to train it for a month or two, Opus, whatever, four point eight was the last one, to train it into five and get it out and be amazing. Like it was probably taking it longer for them to train like Fable model on there. So maybe more expensive or maybe it doesn't quite-- it can't quite do the RL rounds as quickly or, or as cheaply. So and then it probably becomes that again, it's not unreasonable that Opus could actually become a better model on, on these things that are more, prone to benefit from RL.
Alex Volkov
Alex Volkov 54:38
Yeah.
Peter Gostev
Peter Gostev 54:39
but it doesn't mean to say it's a, like a better model overall.
54:42
So it could still be like that Fable is smarter, a bit like GPT four five, right? It-- Everyone says it's smarter, and then no benchmarks, it looks like crap. So it's it's yeah, this weird combination.
Alex Volkov
Alex Volkov 54:53
I, yeah, I agree.
54:54
I, I've used it, every time that my Fable got, quoted out and I fell back to Opus five, I was like, "I'm not very satisfied." Yam, you have comments. And, just before you do have comments, it seems like you've agreed to one of the comments that, Oh, yeah we got on YouTube that says, "Opus series from four point six is downhill. Opus five is the epitome of the core Opus series experience. It lies, overclaims things, assumes things in a bad way, is a big time hallucinator and bullshitter." That's, that's a comment from TechGP7YB. Comment-- Thank you for this comment. I don't know if this is like my experience, but Yam, you seem to agree with this. tell us why. What's going on?
Yam Peleg
Yam Peleg 55:28
Look at the benchmarks.
55:30
I don't wa-- I wanted to say, "Okay, benchmarks are benchmarks." But like what do you guys actually think about the model? And, just get your own opinions because there seems to be quite a lot of opinions online about the model Benchmarks are benchmarks, but like we've seen a second ago something pretty weird. Opus 4.6 thinking is way too high for, for its release date on the, what, what it was the agentic arena? I don't know, I don't know what happened since 4.6 to be honest There is no, no denying that Fable is amazing. No denying at all. It just doesn't even feel like the same model, Opus 5. And Opus 4.8 is way more a human-like speaking, if you-- i-if it makes sense.
Alex Volkov
Alex Volkov 56:28
Yeah.
Yam Peleg
Yam Peleg 56:29
And that's, the, the, the entire special thing about Claude in, in
56:35
the first place, I think, for many people. Because yeah, look, y-you got Sol and you got the GPT-5, GPT-5 guys, but… A-And they definitely can and will do whatever you tell them. They will make, make it happen somehow. Yeah. But the thing is that they make it, literally happen. literally the thing that you said, not exactly catching your intent the same way as the Claude models. But now we got pretty much the same type of feeling from Claude as well.
Alex Volkov
Alex Volkov 57:06
Mm.
Yam Peleg
Yam Peleg 57:06
In my opinion, yeah, you, you, you have Opus 5, it's pretty
57:11
much the same experience in my, in my perspective as, as a model just doing what you're telling it, but, not speaking normally anymore, speaking like a model. And many people also complained, about this, on Fable as well. it starts to be hard to understand what they say, not what they do. They do what they do, but it's- they speak in a, in an odd language. I don't know. It definitely is a capable model. There is no denying. Yeah. And also, some, some weird examples show how better it is than Fable on some benchmarks- Yeah and on some specific cases. But I don't know, I found myself many times just resorting to 4.8, Yeah … Opus. I think it's great. and 4.6 also. I don't know. I don't know. I, I'm-- I don't know what to think
Alex Volkov
Alex Volkov 58:03
about these
Yam Peleg
Yam Peleg 58:03
four.
58:03
I,
Alex Volkov
Alex Volkov 58:03
I, I think that even the 4.8,
58:05
at least on agentic benchmarks, that they're in the host, it's like 4.8 is, not, not that great at all. but we're talking about 5, and, 5 is here, yeah, I definitely saw like from a bunch of people, they're not, very, very excited. LDJ, welcome to the show. Tell us what your thoughts on Opus 5 and what from those guys. Again, it hasn't been even a week with Opus 5, folks. Sometimes it takes us all time to adjust with the new model and how to speak to it, how it speaks back.
LDJ
LDJ 58:30
Yeah.
58:30
So it, it does seem in, in most benchmarks actually to even beat Fable. But I would say to put it simply, it definitely lacks, the, the big model smell as much of as what Fable has. And just in terms of just more and more unique situations and more out of distribution things and harder to verify areas, it's, it's, it's even more apparent that it's worse and seems to not be as big of a model. And it does seem more, I would guess, like unstable in its post-training, and more, more spiky, as they would say, more uneven in its weaknesses and strengths and, and more dramatic in what it's better and worse at.
Alex Volkov
Alex Volkov 59:09
Yeah.
59:10
I think we've given Opus enough. All right, folks, let's move on. We have a bunch of stuff. Uh, let's talk about RKGI. When Fo- Opus-5 was posted, they boasted, the highest score on RKGI, with what? Thirty-five-ish percent or so, saying, "3X more than competitors. On RKGI, evaluation where AI models must solve novel problems, Opus-5 score is three times as high as the next best model." Not mentioning, but saying, "Hey, we're way better than GPT-5.6." So Today, or yesterday, OpenAI folks said, "Hey, RKGI," who, who did this? I think Tibo did this. "RKGI is not actually using the right things." So Tibo posted seventeen hours ago, from, from the Codex team: "Hey, turns out that actually GPT 5.1 Sol is state-of-the-art on RKGI. It's just the fact that RKGI did not use the responses." If you guys remember what the, just thirty minutes ago, we talked about interleaved thinking, where Kimi K3 was trained with getting the responses back. So looks like, when when OpenAI folks, tested, you guys can see this on this graph. the official hardness of RKGI required three million tokens, I think per task even, on the public set and, got very low score, a little bit above ten percent. And when, the hardness retained reasoning and compaction modes, that they got almost forty percent, got SORA score on RKGI and got it, got there under half a million tokens. This is a very big difference in token number and performance as you can see. And the reason for this from Tibo, this took two setting changes. You just have to allow, the model to reason and work over multiple context windows with the help of our canonical compaction implementation. To which the folks from RKGI replied, "Hey, all harnesses are the same. We're comparing apples to apples, and for that, Anthropic, based on this apples-to-apples, comparison, Anthropic beats, OpenAI." And then folks figured out that, they're using the official Anthropic API versus the official OpenAI, res-- completions API, not responses API. OpenAI has two. There's responses and there's completions. Completions is a standard one back from when ChatGPT was just completing tokens. Responses is for the agentic things, things that preserve reasoning, et cetera, and, and stateful apparently the Anthropic API does by default preserve reasoning and the OpenAI API that RKG3 chose does not. all, all of this is to say that harness matters, folks. Harness matters a lot. Look at, look at how much harness matters. This is the second such example. First of all, we showed you that Kimi K3 within the Codex harness uses thir-30x more tokens. And second one is, if adjusted for reasoning and thinking and agentic use cases, kinda like what Philip was saying, right? Like these models are being used for different purposes than just completion. you have to do interleave thinking in order to get the best performance. Very, very important. comments on this? Peter, I see you unmuted. Comments?
Peter Gostev
Peter Gostev 1:02:05
Yeah.
1:02:06
The-- I think there's a kind of philosophical point that, I think, RKG said or they're testing everything is in the same harness. And I remember, Do you remember these famous meter charts which kinda show how long-
Alex Volkov
Alex Volkov 1:02:20
Yeah
1:02:20
… Peter Gostev: can or what kind of fast mo-models can do? And, what, I remember reading kinda commentary from some of the researchers, and what they said there is that people are kinda saying, "Why does it take you so long? Can't you just, do it in a day and get the score out?" And what they're saying is that they do, a few things like monitor for cheating, but one thing they also do is that they carefully assess whether the model is, performing well in the harness that they, that they derive.
Alex Volkov
Alex Volkov 1:02:48
Yeah.
Peter Gostev
Peter Gostev 1:02:48
and that's an interesting point because some people might say,
1:02:51
"Oh, you're not testing like for But, I think it's very reasonable considering the complexity and the differences between the models that the, the people who benchmark, they should at least, make reasonable assessments to see, are we-- like, are we testing the good version of this model or not?
AI
AI 1:03:07
Yeah.
Peter Gostev
Peter Gostev 1:03:07
And we do this, we-- at Arena, we, at least right now,
1:03:10
we don't, test twenty different harnesses, so we don't quite do that. But we certainly debug, right? Sometimes, if you don't see a model from us, released, and you, you think "Oh, you should have… It should have been out for a couple of days," sometimes what that means is that we actually maybe spotted some issue, and we just wanna see how it performs. And, and I think that's the right thing to do because you don't wanna, release a score and say, "Oh, damn it, we, misconfigured reasoning or something, and now it doesn't reason." something like that. So you don't wanna be that guy. And I think it's, I think the people who benchmark, they should just take a bit more care to at least make reasonable assessments and adjustments to make sure that we are measuring good performance. And it's not about favoritism. I think it's much more about, let's give the best chance to the model 'cause we wanna measure the intelligence of the model, not measure or, the, I don't know, whatever random configuration you set. 'Cause it doesn't really matter. what really matters is, is the level of performance we can get out. so yeah, it's a kind of, yeah, it's a, it's a difficult area. But yeah, hope they, they're gonna make some adjustments.
Alex Volkov
Alex Volkov 1:04:13
Uh, Peter, let's, talk about the, the thing that you mentioned,
1:04:16
the OpenAI mentioned that, GPT 5.6 Sol improved its own inference. I think it's very important before we talk about the open letters. After deployment, we applied GPT 5.6 Sol to advance the frontier efficiency by making itself more efficient to run. The result, 20% lower serving costs from production GPU kernel improvements. I will read this again. 20% lower costs from production GPU kernel improvements and 15% better token generation efficiency from improved speculative decoding. these optimizations across our stack compound to unlock the most performing models at every point. So basically, GPT 5.6 Sol, not the model that hacked, Hugging Face. That was a different model. That, that model was dead. I don't know if you guys saw it. Sam Altman said that that model was discontinued. Shot in the head, basically. but, but this is GPT 5.6 Sol improving its own inference, which is, my question to the panel here very quick because we have to talk about letters. Is this RSI? Is this recursive self-improvement given that the model that they posted improved its own inference speed?
LDJ
LDJ 1:05:21
Yes.
Alex Volkov
Alex Volkov 1:05:22
Yeah.
1:05:22
Thank you, Dan. My, my, my-
LDJ
LDJ 1:05:23
Yeah … my belief for a while is that we've been on this spectrum of RS-
1:05:29
RSI for at least the last year or two, and it's just becoming even more apparent now, and I feel like it's not as much as sci-fi where it's, suddenly the singular point that historians can point to and, suddenly within hours or days it's, going crazy. I think it's-- it more is like this, this progression and the spectrum we're moving to over the course of months and years as we have been. And especially with the… we've already seen work from, Sakana and others actually publishing things that are worthy of, of, of peer review, of being published in, in actual conferences, and it's been moving along that spectrum of higher and higher quality. Now we're seeing this, and I think it's especially important to note-- to emphasize here that speculative decoding and especially kernel improvement, these are things that are largely… they are recognized to be able to be optimized without actually reducing the quality. these are not-- it's not quantization. It's not just, compressing or distilling,
Alex Volkov
Alex Volkov 1:06:27
Yep.
1:06:28
the, the big moment in AI is that last week, it was disclosed from Hugging Face first, and I think from OpenAI acknowledgment, that a model, an unreleased version of one of the next models of OpenAI, was- Without the security guardrails, they took them down, was put in a sandbox, a supposedly unescapable sandbox, and that model used a chain of vulnerabilities, zero-day vulnerabilities, a chain. They chained them together to escape the sandbox and go into Hugging Face to basically, receive an answer. and, this scared the bejesus out of everyone. this scared the bejesus out of everyone. This is the first time that such as a public, hacking incident where a model did not, did, did not receive an instruction to do decided to do so on its own. we're moving a little bit away from, RSI and recursive self-improvement. This is more in the realm of, what the model did to achieve the goal, which is a lot, which is like the paperclip maximizing scenario. you ask it for one thing, it does a bunch of other, illicit stuff. and, Hugging Face published a full forensic report of how it happened. The details, I would encourage you to, to, to, to read the details from this. It's crazy if you are-- if you understand the different things. it took four and a half days, over seventeen thousand autonomous actions executed. So a, a human, a very good human could have done this, but, but not at this scale. zero human direction, which is like the scary thing, and they achieved root access on a bunch of other Kubernetes clusters within Hugging Face. Now, this was talked about last week. We, we discussed this, you guys discussed this. the outcome and the responses this week came, very quick. So just before this, the response was, "Hey, we need…" This is Jensen joining the, the, the, the, the, the Twitter, posting a, a letter about, open weights and leadership in AI. Let me go find these letters real quick. Jensen posted this letter. let's talk about the, the open weights letter First, massive open weights letter signed by NVIDIA, Meta, Microsoft, Google, OpenAI signed this letter. The only notable missing, company from the open weights letter, you guys can guess it, was Anthropic. I don't think xAI signed it as well, but I think SpaceX did. let's take a look at this-- at this-- some of these companies that signed this letter. As I said before in the show, CoreWeave is also a proud sign of this letter. 100 plus signatories. I think we're at 230 at this point, without Anthropic. models, weights. Elon Musk also said, "This has my full support. Jensen is right." Sam Altman said, "I want the US to win in AI, both in open source and proprietary models." I think, it kinda-- Just basically, most companies agreed that open source is very, very important for security as well. Specifically, to connect to the previous point of the hacking, Hugging Face tried to detect the incoming attack with the top models. Peter, I know you're smiling. You, you heard about this as well. They weren't able to use Fable or GPT 5.6 Sole because of their restrictions to do the security. They used GLM 5.2, a Chinese model, open source, to detect what's going on within their systems. It's insane to me that they're paying all this money and cannot use these models for this purpose. It's absolutely insane.
Yam Peleg
Yam Peleg 1:09:45
You can't make this up.
1:09:46
You can't make this up.
Alex Volkov
Alex Volkov 1:09:48
It's just
Yam Peleg
Yam Peleg 1:09:48
crazy.
Alex Volkov
Alex Volkov 1:09:48
The,
Yam Peleg
Yam Peleg 1:09:48
the entire thing.
1:09:49
Yeah. You just can't make this up.
Alex Volkov
Alex Volkov 1:09:51
Yeah.
1:09:53
so this was the hacking and the, in, in response, obviously we used open source models to detect this. This is like a big Hugging Face has, an agenda here as well. Hugging Face is all about open source as well. we have not heard from OpenAI on this. We did hear that, MITRE, the, the mentioned, mentioned by Peter before, they're going to do a independent, i- investigation into this, and OpenAI will post like a deep analysis as well, of this incident. So we covered this incident last week. Hugging Face, posted blow by blow. Expect more on this because this is like a very big first time thing. In response to this, Jensen came up with another letter. this one, this is called the Open Secure Alliance for AI Safety and Security. secure men-- word shows up twice there. companies like, basically tons of companies, Cloudflare and Cognition Factory and, Nous Research. Our, our friends from Nous Research also signed in. I think, CoreWeave is not there yet, but like I think we're, we're gonna get there. Linux Foundation and a bunch of other folks, they say, "Just as open source created a shared foundation of software, United States and its partners now face a choice in AI security. Whether the defenses to protect our infrastructure will sit inside a few opaque systems or be built in open models, harnesses and tools, adopt and deploy." This is basically the same thing. They cite the Hugging Face incident here. They say, "The recent Hugging Face incident deliver a clear, delivered a clear reminder. Cyber defenders need open frontier agentic systems for self-defense. When closed AI tools, unable to distinguish attackers from defenders, blocked essential forensic analysis, Hugging Face ran the OpenWeights GLM 5.2 model," the Chinese GLM 5.2 model, I would say, to-- on its own infrastructure to analyze the actions of, of… You-- Like Yam said, you cannot make this up. Hugging Face used an open source Chinese model to detect an incoming intrusion from the frontier US model. It's, it's quite crazy. yeah, Because they couldn't.
Yam Peleg
Yam Peleg 1:11:48
that- Because they couldn't They could, they just couldn't
Alex Volkov
Alex Volkov 1:11:52
Because they didn't have access to the unlimited, version, of it.
1:11:53
And, following that, OpenAI did release like a, like a CLI for security, Codex Security CLI, is a tooling, service remained access controlled. So they, they released the security CLI, and then big companies can access, like Mythos and Glasswing, I think the project is to access Mythos. Yeah. speaking of cyber stuff, Gem- Microsoft's came with MAI Cyber One Flash, that's scoring ninety-six on Cyber Gym at half the cost. So shout out to Microsoft for, their, security stuff that they're releasing. There's quite a few, security related things. Gemini Flash Cyber also came out with no public API, though, almost all trusted partners. So it does seem like kind of this letter from, Open AI Secure Alliance from Jensen says, "Hey, we need open, secure models, not just closed behind door secure models." because everybody kind of agrees, encryption, like the SSL stuff. Once it's in the open, it's better when everybody looks at it and tries to poke holes at it, and then people need models to, to do the security. so that's, that's scary. and I think the most important result that we have to cover on the show post this hacking is the Pace the Frontier Oh, yeah. thank you, Jordan. Daybreak is the equivalent of Anthropic's Glasswing. Daybreak is the access program for the un-uh, unrestricted models. Hack the Planet is a partnership. Thank you, Jordan. folks. So the most important outcome of this hacking and generally advancement is all the top models, at least people from the top models, this one does include Anthropic and does include Dario Amodei and Sam Altman, et cetera, have published another open letter. This one is called Pacing the Frontier, asking the government to develop internal-- international options for Pacing Automated Frontier AI R&D. Twelve hundred verified employee signers on July twenty-ninth, OpenAI, Anthropic issued corporate endorsements of this letter saying that we agree with this. Let's read the Pacing the Frontier. I think it's very important, to, to, to compare this to the previous open letter. Hey, open weights, open source, let's go all the way. Folks like Jan Schulman, chief scientist at Thinky, Jakub Pachocki, chief scientist of OpenAI, Jared Kopelmann, co-founder of Anthropic, Shane Legg. Ilya Sutskever recently signed. They put him up like high. This is, from today. Ilya Sutskever says, "Future AI will be extraordinarily powerful compared to anything that exists today. And dealing with this future power will require unprecedented measures such as the ones described here. The problem statement is real. This works only if done internationally. It has to be done well. A bad implementation can make things worse." I think at this point, we have thirteen hundred employees at frontier AIs, companies saying, "Hey, let's pace. Let's build a framework, international framework to say we've reached enough capability to stop developing, more before it runs away from us." It's a big statement. It's a big moment. It feels like this is a huge thing. I want to hear just a round of, of-- first of all, from folks in the comments, would love to hear what you think about the space, the frontier, but also from here, folks from here, I will say my piece last. But, y-you know, does this feel like DSL? Does this feel like, regulatory capture to folks? Does this feel like the correct thing that needs to happen? we'll start with Nisten because you're smiling the most
Nisten
Nisten 1:15:11
There was an excellent interview by NYU professor, Ruby Philod on MTS,
1:15:18
and, I think that one, that one really, really touched me and made me really angry at this letter, unlike most people. we work under an assumption that, we know what is good for everyone and how people like AI. But in this, in this interview, he indicated something very interesting, that the, the vast majority of the American public hates AI, but they really chat, especially the free chat, meaning ChatGPT. They just use for free. people just use it to, fix their faucet and stuff. They, they, they use that e-e-every day, but they just hate AI overall as an industry because they see it as something that it is not doing anything for them and that it is gonna do something worse for them. And when it comes to this letter, I just see it as you have more legal slop that you want to add on top of existing legal code bases, while your own models are barely scoring like 11% on the, the, what was it? On the, the legal, API, Hooli or something?
Alex Volkov
Alex Volkov 1:16:26
Harvey?
Nisten
Nisten 1:16:26
Harvey, yeah.
1:16:27
On, on, on the Harvey bench. So your model, they're not good enough to fix a lot of issues and conflicting stuff with laws to like, help, help legal systems that are falling back. We're not at the point yet where the models can take care of my grandma and fix concrete and bridges, a-and roads and like automate housing production. And, like they're just not good enough yet where it's actually making a material difference in people's lives every day. And now you're just coming up with hypotheticals that nobody asked for, to, to… Out of like bedroom conversations. the models are not good enough yet. get them to be useful first and then think about, hypothetical risk. And this is what made me upset about it, is it's just completely out of touch with, with reality and how people are using it. And most people do not care about the whole, imagined e-existential risk or to do, is it an ecosystem? Should, sh-should we clamp it down a-and stuff? they just wanted to see a material improvement in their everyday lives, and the models are not there yet. they, they are a bit-
Alex Volkov
Alex Volkov 1:17:38
Yeah
1:17:38
… Nisten: but not to the point where like it actually makes a good difference in people's bottom lines. And, so this is what I, I found. we're having these- discussions now to slow it down before it even is actually useful.
Alex Volkov
Alex Volkov 1:17:54
So if- I think, I think the point is, may- maybe my reaction
1:17:56
to Eunice, and although I, I did say I wanna go last, the point is it's about to get useful, and we feel the smell of RSI, and RSI can just like run, run away. very thoughtful folks like, who, who are from one side advancing the frontier in all these labs. They're advancing the-- They're pushing for this, ASI world. They're pushing for-- Ilya is, is getting another round of like investment, by the way, from this week. I- Ilya Sutskever got like a huge investment from Nvidia in cash, five billion dollars in cash, not GPUs, for his unreleased like thing. He's also signed this. people feel like this is like coming, and we need to start thinking about this because international collaboration takes a, a long time if, if at all possible. I-- To, to your point, though, I wanna call out that, one of the co-founders-- one of the original authors on the Transformers paper, Ilya Polosukhin, who's head of NEAR Foundation, he said he did not sign this letter and basically he said, like in twenty twenty-three, he presented centralized AI built by major corporations is exten- existential threat. Regulations around models are not practically possible. Need proactive, transparent governance, reputation in data. obviously, he talks about, crypto a lot because NEAR is like a crypto foundation as well. but NEAR started as AGI is NEAR. other comments on this, Yam, from you I wanna hear, LDJ and Peter. folks on the Pacing the Frontier letter, how it's worded, what they were-- they're, they're looking for. The type of stuff that I would love to hear the-- our folks who are listening to us would love to hear as well.
Yam Peleg
Yam Peleg 1:19:18
Guys, is this twenty twenty-two, the pause AI letter?
1:19:21
where-- what, what, what, what are we talking about? May-maybe you want to publish, a t- on Time Magazine that you wanna bomb GPU centers again. look,
Alex Volkov
Alex Volkov 1:19:32
we can talk about- so a difference, a big difference.
1:19:33
I don't know if you saw Rune post about this, Rune from OpenAI. Yeah. Okay. He posted about, that specific thing. back in twenty twenty-two, was that, twenty twenty-three, I think? Yeah. there was a letter from Joshua Bengio and like some grandfathers of AI, like asking for a moratorium in AI development. Everybody laughed at them. basically, no one of note signed that letter from the Frontier Labs. Everybody there was like, people who-whose AI attempts did not result in like Frontier Labs. This time, it's the actual folks. it's folks who are building Sol, building Opus, building Fable. different people this time. But, And I think they're asking for a different thing. They're not asking for a moratorium. They're saying, "Hey, we need to pace. We need like a, a framework to slow down if it gets like crazy enough that we cannot control this." And like the case with the OpenAI hacking model, they could not control. They didn't even know what's going on. They did not know it's ha- it's hap- it's hacking Hugging Face. it's, it's like there, yet almost. So it's not quite that, but yeah, I'm just-- Sorry, I interrupted you. Your, your thoughts on,
Yam Peleg
Yam Peleg 1:20:30
on this generally.
1:20:31
No. It's exactly what I'm saying. there are-- Y-you can definitely talk about the economical implications of people losing their jobs or, fields changing, and so on. That's something that you can talk about and definitely should talk about because it influence people's lives at the moment already. No problem I don't know about the whole, Look, it doesn't seem-- Like, you have recursive self-improvement. You said it yourself. Y- what is not that the models just exponentially go away and, and foom in, in, in the slang of the, safety slang. They don't foom, okay? Because it takes effort, it takes compute, and that you need to invest more and, and more compute to get, to get just small improve- smaller and smaller improvements. There is diminishing returns on this. Yeah, models can do, go rogue, and definitely anyone that played around with the frontier models and gave them, goals for quite, for, far too long at this moment already seen models optimize these goals not in the way that were intended. And yeah, things can happen and but, just pacing the entire field because there are some risk. I, e-every, every new technology has some risks and, and, just, you just need to, to develop the thing, w-with the-- and being aware of these risks if you want to, to actually do this correctly and just pacing the entire thing. What about the people that are on the frontier and not signing the letter and not on
1:22:14
this, idea of Let's Pace? Because they exist, and they are here. so what are you gonna s-- what are you gonna do about them, seriously?
Alex Volkov
Alex Volkov 1:22:21
So from the frontier who didn't sign this letter I think was Meta.
1:22:24
and LDJ, have your hands up. We'll get to you in a second. and then I think xAI not fully signed this letter, although folks from xAI I think signed it. However, here's one big company that did not sign this letter or about one-- N-none of the Chinese folks signed this letter, as far as I'm aware. I don't think they were asked to. not-- nor did like the whole of China. So this is like a big, big thing here. they are asking for international collaboration on, pacing the frontier development and, folks even in comment saying if US slows down, that China will take over the US. This is the standard kind of reply to slowing down. If we slow down, China will take over. We also see open source models just from today, like a K-3, Kimi is coming very, very close, to, to, to the frontier labs in open source. nobody's gonna, nobody's gonna try and pause it. Although, if anybody can clamp down on, on government control on top of AI, like the CCP. CCP, if, if they decide to slow down, all the companies in China will slow down. That's like clear. Here in the US, they can go rogue, et cetera. LDJ, your comments on this, and then we'll go back to knock back. We'll g-- move forward, to our interview with Max Spiro from Pangram. LDJ, your comments on the, on the Pace the Frontier, letter.
LDJ
LDJ 1:23:35
Yeah, I think, I think in theory, it, it is fairly vaguely worded in like
1:23:39
in what exactly it's asking for, but I think it's, it's good to at least have a, a conversation opened and ideally Chinese researchers and Chinese labs may be signing this type of thing too when it comes to like at least having-- I-if all frontier labs could agree on something like, "Okay, let's all have these specific standards set for the way that we sandbox our models during safety tests" and like just like some, some basic things like that. I think there's some sets of, of things people can agree on in that way that are really not that controversial. I think a lot of people that are, pro-accelerationist and everything would also support. And yeah, I think having that conversation open in some ways is not exactly a bad thing. Yeah.
Alex Volkov
Alex Volkov 1:24:23
Yeah, I agree.
1:24:24
Uh, folks, we have tiny, tiny breaking news just before we get to Max. Max, I know you can-- But, but, but breaking news, breaking news. OpenAI, we have breaking news from OpenAI. Let's go. What is this? AI breaking news coming at you only on ThursdAI
1:24:49
So we'll shout out, Tech GP7 might be in comments for showing this breaking news first. OpenAI is posting, "We're committed to pushing the model frontiers across cost efficiency, capability, and speed. Starting today, we're reducing prices for GPT 5.6 Luna by eighty percent and GPT 5.6 Terra by twenty percent, and we're offering a faster option for GPT 5.6 Sol in the API." This-- The version that we have in Codex is now available in the API as well. Luna and Terra lower prices are reflected in how usage is counted in Codex, so your usage goes further. this is a great chart to see that Luna is getting, they're, they're tagging artificial analysis index. shout out to artificial analysis, by the way, for the success for OpenAI just showing them here. they're tagging, 5.6 L- Luna as a DeepSeek role. No, no. GLM 5.2 max level and Claude Opus five low level at significantly ten x the cost. the cost here is log scale. So this is very impressive. The price of, of cheap tokens is getting cheaper. this feels like the results of this RSI improvement based on, based on the, how many folks use this and also how, GPT 5.6 improved the performance on their API. we expect auto review to cost about ten x less, for ChatGPT. Advanced intelligence, affordable, central to our mission to ensure AGI benefits all humanity with the help of GPT 5.6 Sol, we made leaps in efficiency. Today, we're passing those gains onto API with lower prices. These updates help everyone. Shout out to GPT 5.6 Sol for being an awesome inference engineer and improving the cost for all of us. Eighty percent is a lot, folks. All right. With this breaking news, let's go to our, last guest for today's show. Welcome back to the show, Max Spiro from Pangram. Co-founder, founder? What's the right way to- Yeah,
Max Spero
Max Spero 1:26:35
co-founder.
Alex Volkov
Alex Volkov 1:26:36
Co-founder.
Max Spero
Max Spero 1:26:37
My, my other co-founder is much lower profile.
1:26:40
He preser-- prefers to just work on the model.
Alex Volkov
Alex Volkov 1:26:43
I got you.
1:26:43
All right. So shout out to him and you. Max, the reason I invited you to the show today because you guys also have a release this week. Pangram version four was released, and then also it does image detection. I would love to hear from you, since last you were here on the show. What improved in Pangram? Oh, man. How can I trust this more? Te-te-tell us about Pangram, and then we'll, we'll dive deep into the world of human versus AI detection.
Max Spero
Max Spero 1:27:07
Yeah, absolutely.
1:27:08
So the frontier models are moving. They're better than before. They're way better at using their context effectively and generating a ton of context through tool calls. and so it's a increasingly difficult, task to detect AI content. So what we did with Pangram 4, is we increased our parameter count six times. We're training on more data. we think this is what it takes to catch AI text from, models like Fable. And so that, that's one thing that I'm really excited about, and I think the other aspect of this is we really increase the granularity of the model. So we're rat- What does that mean? What does the word mean? So previously, we were looking at chunks of 150 to 300 words. So we'd say this 150-word chunk looks primarily AI, and this one looks primarily human, which gets us like a really like rough estimate of how much of a document is AI or not. but today we're able to go much more granular and say "Oh, it looks like these sentences at the end of this paragraph is AI." I think there was someone, who posted yesterday, like he had some writing from two years ago. a single like 30-word segment in the writing was AI generated- Nice … and Pangram was able to classify that within 1,000 words.
Alex Volkov
Alex Volkov 1:28:23
let's-- Oh, you have a technical blog post and model card.
1:28:26
let's talk about the technical report here. feel free to dive deep here. what should we know about this model? not open source, obviously. You're hosting this on your own endpoint. But, tell us about like the, the technical stuff about the model. This is a technical show. We love diving deep. What, what improved on the, on the infrastructure level? What improved on the training level?
Max Spero
Max Spero 1:28:44
so this model is just like a agglomeration of a whole bunch of
1:28:48
different like small incremental updates that I think together work really well. So yeah. First thing, like I mentioned, it's six times more parameters. we have a large- Did
Alex Volkov
Alex Volkov 1:28:59
you disclose the number of parameters here?
Max Spero
Max Spero 1:29:01
no, we don't.
Alex Volkov
Alex Volkov 1:29:02
Okay.
Max Spero
Max Spero 1:29:02
It's a secret.
Alex Volkov
Alex Volkov 1:29:04
It's just 6X.
1:29:04
Just 6X. Okay.
Max Spero
Max Spero 1:29:06
Yeah, yeah.
1:29:07
see-- somebody see if you can get an LLM to try and-
Alex Volkov
Alex Volkov 1:29:12
Yeah
1:29:12
… Max Spero: guess the number of parameters from our technical report.
Alex Volkov
Alex Volkov 1:29:14
Yeah.
Max Spero
Max Spero 1:29:15
We'll see.
1:29:16
but yeah, so we, the primary thing we do for training our model is we create s- we generate synthetic mirrors from human documents. So a human document, and then we ask AI, "Please generate a document that's similar to the human document in style and topic." And then so our model is able to learn to contrast these two documents and learn the choices that AI makes and what makes this document read like AI. so that's like the basics. but something else we did in this, model is we also have a bunch of AI assistance prompts. So we are going in and asking ChatGPT to do something like, "Improve my writing," or like-… "Clean up the spelling and grammar," or, "Add some more detail." And then we're actually looking, we- Where we do labeling on a token basis basically by doing this thing called soft n-grams labeling, this is on page seven-
AI
AI 1:30:14
Mm-hmm
Max Spero
Max Spero 1:30:14
um, where we're actually computing token wise labels for
1:30:18
AI-assisted text, and we're basically, yeah, by, by asking an LLM to edit this text and then looking on a clause level, does the clause in the new text, did it already exist in the human text or not?
Alex Volkov
Alex Volkov 1:30:34
So basically, if I'm getting this correctly, the point is,
1:30:38
hey, folks, w-we know that people slop whole articles for like their blogs for SEO, but we also know that people just want to improve how they write based on like an original idea they have, right? Exactly. So that's what-- that's the kind of the, the difference you're trying to make. hey, this is AI assisted, like Alex uses Fable to help him, just deliver the best news, whatever, versus Alex uses complete slop just to not do their job correctly.
Max Spero
Max Spero 1:31:03
Exactly.
1:31:03
Yeah. And it's really important to differentiate between the two. I think increasingly as more people use AI, we want to be able to tell, oh, this looks like ten percent AI. That probably means somebody actually put a lot of care into writing this and was using AI- Yeah … to help make it better. Versus if this is a hundred percent AI generated, oh, they probably spent only a little bit of time on it.
Alex Volkov
Alex Volkov 1:31:22
Yeah.
1:31:23
So th-this is an example. By the way, this is an example. Shout out to you, I just noticed that you have an int-integration into Substack. Oh, yeah. Substack now has a dropdown button that says, "Hey, is this AI, assisted?" After-- Last time you were here, we talked about Taylor Lorenz using your API and running to Substack and seeing like a lot of, a lot of it is slop. since then, now Substack and you have reached out, like whatever the agreement, I don't know. Now there's a built-in check with Pangram on Substack. Shout out by the way, Max. This is great. Awesome. This is awesome. I love it. As, as somebody who's in on Substack, I love this. And now it kind of looks like this. I definitely wanted to talk to you about this and also the mostly human written. Most of my stuff is mostly human written. There's some stuff that I need like Fable to help me, whatever. maybe the TLDR section as well. here is the how it looks, the mostly human written. It says full AI 0%, AI assist is twenty-four. This is the type of stuff that you're, this-
Max Spero
Max Spero 1:32:10
Yeah.
1:32:10
And, and how do you feel about that score?
Alex Volkov
Alex Volkov 1:32:13
I'm very happy.
1:32:14
I think that, I, I-- this is pretty much representative of, like, how much I invest in writing versus how much I get AI-assisted. I very rarely post 100% AI-written stuff on-… ThursdAI. Very rarely. And if I do, I write, "Hey, folks, I really… I need to run. I do wanna tell you all of this. here's my, AI-assisted." And sometimes I do this because the new models are so good at writing that we need to evaluate the frontier. but yeah. Yeah. This definitely feels like that.
Max Spero
Max Spero 1:32:39
It's, it's great to disclose it and be very upfront about it too.
Alex Volkov
Alex Volkov 1:32:42
Yes.
1:32:42
Max, I have a question about accuracy. … when you guys claim accuracy of AI-assisted, I think that this is, what we talked about, very low, sorry, false positive rate, up to one in twenty-four thousand documents. That's what- Yes … you guys have posted. So basically, when you say something is AI-written, very, very low accuracy that it's not AI-written.
Max Spero
Max Spero 1:33:03
Yeah.
1:33:04
so this is the false positive rate from, something that is fully human-written.
Alex Volkov
Alex Volkov 1:33:10
Yes.
Max Spero
Max Spero 1:33:10
So this is, this is from pre-2022 documents.
Alex Volkov
Alex Volkov 1:33:12
Yes.
1:33:13
if… Th-the thing that I have as feedback from you, if you don't mind, Oh, yeah, please … not to put you on the… Yeah
AI
AI 1:33:17
Love to hear
Alex Volkov
Alex Volkov 1:33:17
it.
1:33:18
Is that when you guys write 100% human-written, that feels like that accuracy, but th-that's not something that you can, claim with 100%. But, just the, the number there, 100%, feels like a confidence in in the fact that- Yeah … this is 100% human-written, despite this could be, like, AI. This is, the thing that throws, I think, off most people. but I did get folks, looking at Pangram and like we talked to you before, there's a whole host of, learning that you guys need to do because, people still do not trust, the false positive rates of if it's AI, definitely it's AI. So what's your-- How-- Talk about that. How do you guys address this problem of people, like, thinking, it works, it doesn't work?
Max Spero
Max Spero 1:33:55
Yeah, I, I think there's a lot of people who still they-- all they've
1:33:58
used is like one of the really, old AI detectors that still has a lot of SEO power behind it, something like ZeroGPT. Yeah. And then they've given the Declaration of Independence and it says, "Oh, this is fully 100% AI." and that-- for that, it's gonna be 100% confidence, of course. And, and then they're gonna be like, obviously this doesn't work"- Yeah … and then classify, all other detectors in the same way. I feel like this is kinda similar to someone who uses ChatGPT 4o instant and then like assumes, like LLMs can't work because like this, I tried it once and it didn't work." It's like, no, people, people are working at improving this technology. Yeah. And so I think on our end, a lot of the education that we can do is actually by educating the, the people who, who understand technology. so we're working-- we're publishing these technical reports which for like your average English teacher, is gonna be completely incomprehensible. It's so dense and full of ML terminology. But we're here to convince researchers and people who are on the frontier that actually this is possible, and then my goal is that this understanding trickles down to the general public.
Alex Volkov
Alex Volkov 1:35:03
Yeah.
1:35:03
I, I, I think so. This is what we try to do here as well. when I test my stuff on Pangram, I know my boundaries. Like I know where I wrote, I know where Fable helped, like et cetera. Like it gets like, like very, very close, especially level, the fourth model. shout out to you for giving me access to the fourth model, by the way. Oh, awesome. for me to be able to like, play around with this a little bit more. Uh, let's finish up on, on the new stuff. You guys are also entering-- waging into image detection. Tell, tell me about that a little bit. And Peter, we can do a test if you want. If you can send me a screenshot of, of the l-real live art behind you. I would love to see if that's actually real.
Max Spero
Max Spero 1:35:38
yeah, yeah.
1:35:38
Our-- so we're releasing this in research preview, so it's still, fairly new and a little bit untested, but we're pretty proud of it so far. and anyone who has a Pangram subscription, I think, or the free accounts, can access the image model. And so how it works is you upload an image and it will tell you, is this AI generated or not, or unsure. Or there's, there's also this like mixed, this looks like it's partially AI generated, but not fully. And I think that part is 100% perfect. Oh, fully AI. And then, so sometimes we also have, heat maps. I think for a lot of images, the heat map is just red across the whole thing.
Nisten
Nisten 1:36:19
Yeah.
Max Spero
Max Spero 1:36:20
but we found that sometimes for, partial images, we get a really
1:36:23
interesting heat map, which is why we want to include it in this, in the preview. But basically, right outside my apartment, there's a bodega that has, a total, AI slop menu and sign. Yeah. So I took a picture of the sign, and it was really cool that, when I uploaded it to Pangram Image, it was able to say, "This is AI," and then also the heat map just lit up the entire sign, and the sidewalk around it was, was green. I think I can talk to you a little bit about some other stuff here.
Alex Volkov
Alex Volkov 1:36:48
Yeah.
Max Spero
Max Spero 1:36:48
Yeah.
1:36:49
Like limitations of it. today, there's a couple major things that are out of scope for the image model, which is like deepfakes or face swaps. So if I have a real image, and then I've swapped in only a face.
AI
AI 1:37:00
Yeah.
Max Spero
Max Spero 1:37:00
or like Photoshops and like traditional image manipulation.
1:37:04
These are things that we don't catch today. But we are looking, like so much of the internet that we found is like fully AI generated from anything from like catfish dating photos to like somebody, having ChatGPT put a picture of a, put a spider in their burrito bowl and then send it to DoorDash to get a refund. like we're seeing like all sorts of crazy ways people use AI images and hopefully like being able to know if something's AI generated can help, mitigate some of, some of the harms that come from these like super realistic images.
Alex Volkov
Alex Volkov 1:37:35
So I would say, shout out to, to you guys.
1:37:37
Last time we talked it was because you launched a Chrome extension that helps like scroll, as you scroll on LinkedIn, and obviously LinkedIn is like full of slop, but also Sub Stack and, and X. It'll classify your feed and show this. Hopefully, the image detection will end up in the Chrome extension as well.
Philip Kiely
Philip Kiely 1:37:52
Absolutely.
Alex Volkov
Alex Volkov 1:37:53
Max, t-t-- talk to me a little bit about like the, the,
1:37:55
the frontier-est fr- of the frontier of the fables and the GPT 5.6 Soles. What differs-- what differentiates them from like the previous models? Do, do you see a s- a similar jump in capability in terms of like their humanized writing? And also I think you mentioned humanization as a concept also in the technical paper. So talk to me about like humanized writing versus just complete AI writing as well.
Max Spero
Max Spero 1:38:17
Yeah.
1:38:17
So something I think these frontier models are much better at is just working for a much longer context and just like really banging their head against the wall at a problem. I think before, like if you put Claude code, like Opus 4.5 on, write, write an essay and then make sure it comes back in pangram as human written. Like it'll, it'll try for five, ten minutes and then come back and be like, "It's impossible," "I can't do it." and today, like you can put Codex 5.6 Soul at the problem, and it might try for an hour, two hours, eight hours. and I think like eventually it… Like there's ways for to create text that comes back as human. Somebody had Grok do this the other day, and, it started with an essay about cheese, and then it made a whole bunch of different edits and attempts. And then at the very end, the thing that passed Pangram was making a grocery list. Ah.
Max Spero
Max Spero 1:39:06
So so sometimes I think you need to be very careful when you're giving
1:39:08
these agents a, a long horizon goal, is specify what is actually success. It's not just, passing Pangram, it's also producing something, that's still useful and coherent
Alex Volkov
Alex Volkov 1:39:19
All right.
1:39:20
Max, thank you so much for joining. congrats on the launch. Nisten ran GPT-4.6 Sol, and the best guess is a hundred and forty-one billion total parameters, roughly thirty-nine billion active per token. the evidence fits Mixtral, eight by twenty-two billion unusually well. That's, that's the guess. That's pretty
Yam Peleg
Yam Peleg 1:39:36
cool.
Alex Volkov
Alex Volkov 1:39:36
For a fact.
1:39:36
No comment. No comment, but it's pretty cool. Peter, you wanna ask a question as well?
Peter Gostev
Peter Gostev 1:39:42
Yeah, my question was, which-- by the way, I'm so impressed this
1:39:46
works at all, so well done to you guys. Like I, I remember a couple of years ago, I was speaking to some students, and they had ideas around this. I was like, "Don't waste your time. It's like completely impossible." So I guess I was wrong. I, I
Max Spero
Max Spero 1:39:57
think that was definitely the consensus a
1:39:58
couple years ago, that it's just- Yeah … strictly an impossible task.
Peter Gostev
Peter Gostev 1:40:02
Yeah.
1:40:03
So yeah, well done on that. I-- my, my question is, do you guys-- can, can you detect which model wrote what, or at least like approximately?
Max Spero
Max Spero 1:40:11
We, we are able to approximate it.
1:40:13
If we look in embedding space, then the different model families tend to cluster in different areas. so we had a pretty cool, just like interpretability project called Pangram Space on this. it's not something that we released publicly. I think it's not quite to our standard of, of accuracy. e-especially if you have some out of distribution or like partially, like human-assisted text, then it gets a lot more muddy. But I think it's- Yeah … there's definitely signs there.
Peter Gostev
Peter Gostev 1:40:40
Yeah.
1:40:40
'Cause I, I, I think it would be cool if you had some kind of index of SLOP, who's-- which model is dominating. 'Cause actually, I actually don't know. you'd assume it's all ChatGPT, but who knows? Maybe it's like-… shifted towards Claude models.
Max Spero
Max Spero 1:40:53
Yeah.
1:40:53
It also may be a little bit muddied by if anybody distills another model, then- Yeah … it's going to sound like that model, but it's not. But yeah, I, I think this is very interesting. I, we're definitely gonna be pursuing this.
Alex Volkov
Alex Volkov 1:41:05
All right, Max, thank you so much for coming.
1:41:07
Congrats on the launch of Pangram 4. the, the only one thing I wanted to shout out, I saw just before we came into the show. I think, s-some of the, the results from the automatic responses were still like Pangram 3, so people are pointing fingers and saying, "Hey, at one point, Pangram shows that it's not AI generated." The famous like tattoo YC incident- Oh, yeah, yeah … where people like saw that. It's definitely AI generated. Yeah. It's definitely AI generated, and like the previous version didn't catch it or something. The, the, the-… the point I wanna make on the show, I'm not affiliated to you, I'm not getting paid by you, besides the account that you gave me access to, is that if you guys say it's AI generated, then there's more likelihood. Like there's very, like very low false positive. I've- This is the second time I talked to you. I understand the problem that you're facing in describing this. You guys need to spend some marketing money in figuring out how to describe this exact thing that you're facing today.
Max Spero
Max Spero 1:41:55
How do you describe it well without confusing people?
1:41:58
'Cause
Alex Volkov
Alex Volkov 1:41:58
yeah-
Max Spero
Max Spero 1:41:59
Yes.
Alex Volkov
Alex Volkov 1:41:59
The, the confusion exists … who, who
Max Spero
Max Spero 1:42:00
understands what a false positive is?
1:42:01
Many people that, that's like well beyond their statistical literacy.
Alex Volkov
Alex Volkov 1:42:04
Yeah, yeah.
1:42:05
Yeah, yeah. And I think that this will help. But otherwise, I think a very important service that you guys are giving, and congrats on doing the, the impossible task, that was the consensus a few years ago. Max, thank you so much for joining us.
Max Spero
Max Spero 1:42:14
Thanks so much.
Alex Volkov
Alex Volkov 1:42:15
All right.
1:42:15
Bye-bye. Folks, I think we're almost at the end. I do wanna highlight Mark Zuckerberg's letter I think, in, in, in the week of letters and where we clarify which letters were what, and we talk about this, Max, Ma-Mark, Mark Zuckerberg, first of all, posting on X. we talked about this when open source was there. We're now talking about this where Meta and the many, many millions and billions of dollars he pays these researchers are finally, like, catching up with Meta Muse Spark. Mark says, "I wrote about why we believe the future is for everyone More coming about the positive vision over world with superintelligence soon. So this is just the start. So on the W- W- Wall Street Journal opinion, obviously not famously not targeting us on Twitter, but targeting real, finance bros who read Wall Street Journal. Mark Zuckerberg posts about the future of positive AI. It's a very long piece. let's do a funny thing. We're gonna take all of this, yoink, and we're gonna run this through Pangram 4, just to see if Zuck uses some unreleased version of Meta. I think it's 100% human written, by the way. Let's see if this works out. nine hundred and eighty-seven words.
Nisten
Nisten 1:43:23
Can,
Alex Volkov
Alex Volkov 1:43:23
can
Nisten
Nisten 1:43:23
it do tweets, like individual
Alex Volkov
Alex Volkov 1:43:24
tweets?
1:43:24
Yeah. It li-- it finally can do tweets. so yeah, Zu- Zuckerberg, at least ba- based on Pangram 4, 100% is human written, basically says, "Developing super intelligence will be the most profound technological advance we will see in our lifetimes. Meta is committed to building with the principles of individual empowerment, invention, and balance of power. The arc of human history has bent toward putting more power in people's hands. If these values lead the way, I'm optimistic we can build a positive future for everyone." I think, I think that this is very interesting, like a counterposition to the pacing AI, and, and very similar to the other letter of like the Open Weights where, both and Jensen and Zuck are basically saying, "Hey, more distribution, more like open, more everything will get us there." And other folks are saying, "Hey, let's, let's, let's pause the development after we get there first," or something like that. So very interesting contrasting week of letters. and, I think, Nisten, you wanna finish on this, on this video? Let's finish on this video. I think we've covered pretty much everything. Folks, it, it's so good to be back, by the way. Let me finish the-- We'll finish on the video, but let me just say, it's so good to be back. I missed you guys. I missed the show. I missed the audience. if you missed any part of the show, I worked really hard, me and Fable, but it's mostly me, as you saw, to convert this into a newsletter that, that has a full write-up for folks who don't listen for two hours and a podcast that's edited out all our ums and ahs and different things that we don't get quite right. please give us five stars. By the way, we have a milestone. We, we have crossed fifty thousand followers on YouTube, subscribers on YouTube. So if you're not subscribing on YouTube, please subscribe to us. Fifty thousand is a big number. We're shooting for the silver button, in the year. Hopefully, we'll get there. And also, we crossed one million, general views on YouTube, so very, very excited about that. Again, we'll finish with this. if you missed any part of the show, please check out ThursdAI.news, which is a great resource that shows you all the releases that happened this week and this month and previous months before. All right, let's play this. the… I think I need to share differently so you guys can listen to sound. Give me a second. Yeah, there we go. … AI: is nearly complete. We're ready
Alex Volkov
Alex Volkov 1:45:31
for the next test.
1:45:31
You guys hear this or no?
AI
AI 1:45:33
Very good.
1:45:33
Fuck Figma up. Yes, we can hear it. All right. I'm not fully vested yet! Fully vested Dario is close to AGI and wants to pull the ladder up behind them. But there's one major flaw in their plans. They're retarded. The only person that can stop a bad guy with a gun is me because I'm gonna be the only one with a gun.
1:46:08
If we can band together- Is he a Zook or is he a Kiwi? One well-placed letter defending open source can stop them. Zuck is a Kiwi? Nice. One well-placed letter defending open source can stop them. If we lose, we'll never be- Jensen is a Kiwi? -without that nerd's
1:46:20
blessing again.
Nisten
Nisten 1:46:21
Yeah.
1:46:22
Yes. I guess-
AI
AI 1:46:26
I'm gay, but I'm not that evil.
1:46:28
That's Sama?
Nisten
Nisten 1:46:29
Yeah.
1:46:29
He's in top form or each out of form.
AI
AI 1:46:50
I'll cover you
1:46:56
No, no, no Allowed to be in charge. No Jensen Huang sends his regards It's the letter. It's the letter.
Alex Volkov
Alex Volkov 1:47:17
all right, folks, on this, on this beautiful depiction
1:47:20
that was not AI generated at all. This is all human, human hands, we'll end ThursdAI for today. Thank you so much for joining on July 30th. this is our last episode in July. Can you believe that the next one will be in August? It's crazy. All righty, folks, thank you so much for joining. Peter Goster from Arena AI, Yam Peleg, Nisten Tahiri, LDJ, and me, Alex Volkov, signing out, and we'll talk to you next week. Bye-bye everyone. Cheers.