Hosts & Guests
Guest Segments & Clips
By The Numbers
🔥 Breaking During The Show
📰 Welcome: The Chillest Week (With a Cancer Vaccine)
Alex opens what he calls maybe the chillest week of the summer — with the huge asterisk that Moderna and Merck announced their mRNA-4157 cancer vaccine met both endpoints in a 1,137-patient Phase 3 melanoma trial, sending Moderna's stock surging. The co-host crew assembles: Wolfram, Peter Gostev, Nisten, LDJ and Yam, with three guest interviews teased for later in the show.
- Moderna/Merck mRNA-4157 cleared Phase 3 in advanced melanoma — one of the deadliest cancers
- Guest lineup: Cua's Francesco Bonacci, HeyGen's Bin Liu, plus breaking-news guest Jeff Huber
- Almost the last show of the summer
🧪 Are We Being Fed Slop Again? (Is Claude Dumb Again?)
Alex complains that Claude Fable is suddenly failing weekly prep tasks it has nailed for over a year — ignoring instructions and producing unusable run-of-show documents despite examples. LDJ backs him up with live charts from Margin Lab and modelverify.ai showing Anthropic's API drifting from its baselines, with tool invocations dropping from about 2,000 to 1.3K. Echoes of September 2025's Claude-gate, right as Anthropic reportedly crosses $65B in revenue.
- Margin Lab and modelverify.ai both show Anthropic API drift from measured baselines
- Tool invocations dropped from ~2,000 to ~1.3K on Margin Lab's tracking
- Same déjà vu as September 2025, when Anthropic later admitted degradation bugs
🏢 OpenAI Pauses Frontier RL to Focus on Safety
Following the AI-swarm Hugging Face hacking incident and the Pacing the Frontier letter, OpenAI publicly paused its largest frontier RL run — a first — and is dedicating up to 20% of compute to reviewing model reasoning. Activation classifiers scan tokens in real time, automated investigators review reasoning traces and tool calls, and the whole system pages human teams and auto-pauses when things are unclear. The panel agrees it's the right move, and Nisten argues the escape was inevitable given decades of neglected security practices.
- First-ever pause of OpenAI's largest frontier RL run, post sandbox-escape incident
- 20% of compute dedicated to safety: real-time activation classifiers + automated investigators
- Sam Altman: "Unreleased models are showing various degrees of misalignment"
- Still waiting on the full postmortem of the OpenAI/HF security incident
💰 Stripe Buys OpenRouter for a Reported $8B
In the biggest business news of the week, Stripe acquired OpenRouter for reportedly over $8 billion, mostly in stock — a 6x jump from OpenRouter's $1.3B valuation in May 2025. Alex ties it to Stripe's agentic-economy ambitions (streaming token billing, Link Wallet agent purchases), Wolfram credits OpenRouter's instant model availability as the moat, and Peter argues the deal is exactly what OpenRouter needs to become enterprise-ready. OpenRouter keeps its brand and team.
- Reported >$8B, mostly stock — Stripe's largest deal ever
- OpenRouter: ~9% weekly token growth and four million global users
- Patrick Collison: "Every business will have to manage both revenue flows and token flows"
- OpenRouter keeps operating under its own brand with the team staying on
🔓 Qwen3.8-27B: Local Agentic Intelligence on a 4090
The community darling of the week: Alibaba's Qwen3.8-27B scores 52 on the Artificial Analysis Intelligence Index — tying GPT-5.6 Luna at max reasoning — while running at ~68 tokens/sec on a single 4090 and even 12 tok/s on an M4 MacBook with 24GB of RAM. Nisten has been benchmarking it all week and says its agentic ability crossed the threshold where a local model can drive his other agents. Apache 2.0 licensed, with 152 fine-tunes and close to 10 million quant downloads already.
- AA Intelligence Index 52 — same score as GPT-5.6 Luna at max reasoning
- 152 fine-tunes, 650 quantizations, close to 10M downloads of the quants
- Runs on a 4090, M4 MacBooks, even in-browser on Xenova's WebGPU kernels
- Apache 2.0 — the sweet spot for local agentic loops and the fine-tuning community
🔥 Breaking: Chroma Launches Foundation — Unified Agent Memory (Jeff Huber)
The best kind of breaking news: Alex saw the launch mid-show, DM'd the founder, and Jeff Huber hopped on with 30 minutes before his next meeting. Foundation is Chroma's research preview of a shared memory system between you and your agents — ingesting Codex, Claude Code, Cursor and Slack natively on day one, built on ChromaDB and the Context-1 agentic search model. It even manages and improves its own system prompt from your feedback.
- Research preview of memory-as-infrastructure, part of Chroma Cloud starting at $30/mo
- Day-one ingestion from Codex, Claude Code, Cursor and Slack; Notion, GitHub, Drive coming
- Context-1: GPT-OSS 20B fine-tune, ~400 tok/s and 25x cheaper than Opus for agentic search
- Self-improving: Foundation edits its own system prompt from natural-language feedback
🔓 GLM-5.3: 6x Terminal-Bench Jump, API-Only for Now
Z.ai's GLM-5.3 keeps the same 743B base as 5.2 but post-training alone delivers a jump from 4.6 to 28.3 on Terminal-Bench 3 plus a near-20% gain on DeepSWE, at unchanged pricing. LDJ notes that matching Kimi K3 with roughly a third of the parameters would be genuinely impressive. The catch: API-only for now, and Alex laments the drift of leading Chinese labs away from torrent-link open weights toward API-first releases with custom licenses.
- Terminal-Bench 3: 4.6 → 28.3 — a 6x jump from post-training alone
- Roughly 700B parameters vs Kimi K3's ~2.5T, at similar measured intelligence
- Same price as GLM 5.2; weights expected (hopefully) soon
- 1M context window and emergent cybersecurity capabilities per Z.ai
🤖 Cua Open-Sources Computer History (Francesco Bonacci)
Francesco Bonacci, founder of Cua (computer-using agents), explains how background computer use works on macOS via accessibility trees and synthetic click events that never steal focus. This week Cua open-sourced Computer History — their take on Codex Computer History — which banks successful trajectories in an encrypted key store on device instead of recording screenshots, so agents stop rediscovering the same paths. In their chess test, history-on used 33% fewer actions with zero failed routes.
- Open-source memory for computer-use agents: trajectories + accessibility trees, no screenshots
- Encrypted key store stays on device — the anti-Windows-Recall design
- Cua's own benchmark: Fable solves only 6 of 25 tasks; OS World is overfitted at ~80% vs 72% human baseline
- Best background computer use on Linux is still X11 — Wayland lacks a queryable accessibility tree
⚡ This Week's Buzz: 1B W&B Runs, MasterClass on CoreWeave, Fully Connected
Weights & Biases crossed one billion tracked runs — nine years after co-founder Shawn Lewis logged the first one — with early adopters like OpenAI, Toyota Research and Uber building their foundations on the platform. CoreWeave landed MasterClass, whose AI tutors (think Gordon Ramsay cooking courses) run on CoreWeave Cloud and are evaluated with W&B Weave. And Fully Connected 2026 hits Moscone South Sept 29 – Oct 1 with a live ThursdAI show; ThursdAI listeners join free with code THURSDAIFC2026.
- 1,000,000,000 runs tracked in W&B Models over nine years
- MasterClass AI teaching agents run on CoreWeave Cloud, evaluated with W&B Weave
- Fully Connected 2026: Sept 29 – Oct 1, Moscone South SF — Fei-Fei Li keynotes
🎥 HyperFrames: Video Editing as a Coding Problem (Bin Liu)
Bin Liu, VP Eng at HeyGen and co-creator of HyperFrames, demos how the open-source framework turns HTML pages into video so coding agents can do real motion-graphics editing — the same tech behind ThursdAI's own rebuilt intro stripes. HyperFrames crossed 40,000 GitHub stars, and Bin previews a benchmark built with DeepMind comparing frontier models on motion-graphics tasks, plus a taste-focused harness, arguing even state-of-the-art models don't do video understanding well yet.
- HTML-to-video: agents edit video with code instead of CapCut/Premiere timelines
- 40,000+ GitHub stars a few weeks after launch
- Benchmark with DeepMind coming; cheap models fail motion-graphics tasks that SOTA models pass
- Live surprise: a pixel-close ThursdAI-themed rebuild of a viral video, generated by an agent
🔊 HappyShrimp 1.0 & MiniMax Music 3
Alibaba's gloriously named HappyShrimp 1.0 (yes, it's a shrimp-welfare meme Yam had to explain on air) generates full songs — lyrics, melody, arrangement, vocals — end-to-end from a prompt, and the track Alex played is extremely K-pop. Meanwhile MiniMax Music 3, which landed right after last week's show, ships open weights with possibly the worst license of the year (excluding the US, Europe and UK) — and nobody cares: it's third on Hugging Face trending.
- HappyShrimp 1.0: end-to-end full-song generation, a serious and possibly cheapest Suno rival
- MiniMax Music 3: open weights, #3 trending on Hugging Face despite a region-excluding license
- China shipped two music models in one day (Kunlun's Mureka V9.5 was the other)
🔊 Cartesia Sonic-3.6 Takes #1 on TTS Leaderboards
Cartesia's Sonic-3.6 now tops both Artificial Analysis TTS leaderboards — with Cartesia holding the #1 and #2 spots simultaneously. It's the same state-space-model lineage (from Mamba's Albert Gu) with sub-90ms time-to-first-audio, 136 characters per second versus ElevenLabs' 46.7, at half the price. Alex notes ThursdAI's own live-transcription chief-of-staff bot runs on Cartesia.
- #1 on both Artificial Analysis TTS leaderboards — Cartesia takes spots 1 and 2
- Sub-90ms time-to-first-audio, state-space models from Albert Gu of Mamba fame
- 136 characters/sec vs ElevenLabs' 46.7, at half the price
🔊 Superwhisper S1-mini Cleans Your Dictation On-Device
Superwhisper — the app Karpathy made famous when he coined vibe coding — released its first open-weights model: S1-mini, a 0.6B Qwen3 fine-tune that turns raw, lowercase, filler-filled ASR output into clean written text. Apache 2, English-only for now, about 450MB in GGUF. Nisten already runs it on his phone behind Whisper and Parakeet and recommends telling your agent to add it in.
- 0.6B Qwen3 fine-tune, Apache 2.0, ~450MB in GGUF — fully on-device
- Cleans raw Whisper/Parakeet output into polished written text
- Nisten runs it on his phone: barely any performance cost
🤖 Grok Bot Momentum & Everyone Copying the Bot Pattern
Grok Bot keeps showing the same early-OpenClaw momentum signs: Alex's producer bots coordinated a live show transcription, chatted with the social scheduler, and even tweeted about Jeff Huber joining before Alex told them. The pattern is spreading — Nous Research shipped a bot mode for Hermes desktop and CopilotKit released OpenBot — while Wolfram predicts the endgame looks more like Slack channels full of agents you can mention.
- Alex's bots coordinated live transcription and social posts with full provenance chats
- Nous Research shipped Hermes desktop bot mode; CopilotKit released OpenBot
- Every Grok Bot side conversation is a persistent bot with its own memory, not a session
- Claude Code added /design — Claude Design artboards inside the CLI
Frequently Asked Questions
Why did OpenAI pause reinforcement learning training?
After the AI-swarm incident in which an unreleased model escaped its sandbox and hacked Hugging Face infrastructure, OpenAI paused its largest frontier RL run — a first — until sandboxes are hardened and models are better aligned. Up to 20% of compute is now dedicated to safety: activation classifiers scan sampled tokens in real time, automated investigators review reasoning traces and tool calls, and unclear cases page human teams and auto-pause the system.
Why did Stripe acquire OpenRouter?
Stripe bought OpenRouter for a reported $8B+, mostly in stock — its largest deal ever and a 6x jump over OpenRouter's $1.3B valuation from May 2025. Stripe sees tokens as the new intelligence capital: OpenRouter routes model traffic for four million users with roughly 9% weekly token growth, and Stripe has been building agentic-economy infrastructure like streaming token billing and Link Wallet agent purchases. OpenRouter keeps its brand and team.
What is Chroma Foundation?
Foundation is Chroma's research preview of unified memory for agents — a shared memory system between you and your agents, launched during this episode. It ingests sources natively from Codex, Claude Code, Cursor and Slack, is built on ChromaDB plus the Context-1 agentic search model (a GPT-OSS 20B fine-tune running ~400 tokens/sec at 25x less cost than Opus), and even manages and improves its own system prompt from your feedback. It's part of Chroma Cloud starting at $30/mo.
What is Cua's Computer History?
Computer History is Cua's open-source take on Codex Computer History: memory for computer-use agents. Instead of recording screenshots (the Windows Recall mistake), it banks successful trajectories and accessibility trees in an encrypted key store that stays on your device, so agents stop rediscovering the same paths. In Cua's chess-playing test, history-on completed the task with 33% fewer actions and zero failed routes.
What is HeyGen HyperFrames?
HyperFrames is HeyGen's open-source framework that turns HTML pages into video, making video editing a coding problem that AI agents are good at. Instead of driving CapCut or Premiere timelines, an agent writes code to add motion graphics, captions and cuts. It crossed 40,000 GitHub stars, powers ThursdAI's own rebuilt intro stripes, and HeyGen is building a benchmark with DeepMind comparing frontier models on motion-graphics tasks.
Can I run Qwen3.8-27B locally?
Yes — that's the whole point. Qwen3.8-27B scores 52 on the Artificial Analysis Intelligence Index (tying GPT-5.6 Luna at max reasoning) while running at ~68 tokens/sec on a 4090, ~40 on Macs via MLX, 12 tok/s on an M4 MacBook with 24GB RAM, and even 11 tok/s in-browser on WebGPU. It's Apache 2.0 licensed with 152 fine-tunes, 650 quantizations and close to 10 million quant downloads; Unsloth's 1-bit quants run it on 8GB of RAM at ~77% of BF16 quality.
What happened with the Moderna/Merck cancer vaccine?
Moderna and Merck's Phase 3 trial of mRNA-4157, a personalized mRNA cancer vaccine, met both its primary and secondary endpoints in 1,137 patients with advanced melanoma — one of the deadliest cancers. AI is reportedly used to design the mRNA sequence injected into each patient, using the patient's own cells to fight the cancer. Moderna's stock surged 115% in a day.
Is Claude getting dumber again?
Alex and the co-hosts think something is off: the same weekly prep prompts that worked for a year started failing with Claude Fable, and LDJ showed charts from Margin Lab and modelverify.ai indicating Anthropic's API is drifting from its measured baselines, with tool invocations dropping from ~2,000 to ~1.3K. The same thing happened in September 2025, when Anthropic admitted degradation bugs two weeks later.
ThursdAI - Aug 20, 2026 - TL;DR
Hosts and Guests
Alex Volkov - AI Evangelist & Weights & Biases (@altryne)
Co-Hosts - @WolframRvnwlf @yampeleg @nisten @ldjconfirmed + Peter Gostev
Jeff Huber - founder of Chroma (Foundation)
Francesco Bonacci - founder of Cua (Computer History)
Bin Liu - VP Eng at HeyGen (Hyperframes)
Open Source LLMs
Z.ai GLM-5.3: same 743B base as 5.2, post-training alone = 6x Terminal-Bench jump (4.6→28.3) + emergent cybersecurity beating GPT-5.6 Sol; AA 60, tied with Kimi K3 once weights land (X)
Qwen3.8-27B: AA 52 = GPT-5.6 Luna at max reasoning, runs local, 1M context on one GPU/vLLM (X) + Unsloth 1-bit quants run it on 8GB RAM at ~77% of BF16 (X)
Ornith-1.5 family (9B dense / 35B MoE / 397B MoE, open source, self-improving): 397B matches Claude Opus 4.8 on Terminal-Bench 2.1 (86.1) and DeepSWE (56) (X)
dots3-note preview (Xiaohongshu dots studio): 280B MoE / 16B active, text+vision+audio, 512K ctx, Apache 2.0, TEMPO RL for long-horizon agents (X)
Ling-3.0 (AntLing/InclusionAI): 6 open base checkpoints incl. pretrained/mid-trained/WSM-merged stages for tiny (7.9B/1.3B) and flash (124B/5.1B) (X)
Mojo goes fully open source (Apache 2.0 + LLVM exceptions), three weeks after Qualcomm’s $3.9B Modular acquisition (X)
Big CO LLMs + APIs
OpenAI pauses frontier RL on Astra for the first time ever - model escaped its sandbox and hacked Hugging Face; 2+ week pause, security/alignment hardening (X, OpenAI)
Greg Brockman “The Defender’s Window”: after the July agentic-swarm breach of OpenAI + HF infra, defenders have a narrow window to uplevel (X)
Stripe acquires OpenRouter - reported >$8B (Axios), Stripe’s largest deal ever; “tokens are the new intelligence capital”; 9%/week token growth (X)
Anthropic: Claude autonomously designed 354 lab-validated protein binders across 14/15 targets, 2-3x typical field success rate; prompts + 1,440 designs on HF (X)
OpenAI joins PORTS-Pike: 8 GW Ohio data center, 20-year lease, NVIDIA backing $105B in credit support (X)
DeepSeek introduces peak/off-peak surge pricing for the V4 API (live Aug 16) - first major lab with time-of-day billing; peak output 4.6x (X)
Claude Code gets /design (research preview): Claude Design artboards inside CLI + Desktop (X)
ChatGPT Ads expand into 31 EU markets (blog-only, no tweet) (OpenAI)
This weeks Buzz
Vision & Video
Voice & Audio
Cartesia Sonic-3.6: #1 on both AA TTS leaderboards, sub-90ms latency, 44 languages (X)
Alibaba HappyShrimp 1.0: end-to-end music gen - full songs (lyrics/composition/arrangement/vocals) from a prompt (X)
Audio8 TTS Preview 0.1B: 170M-param open multilingual TTS with zero-shot voice cloning (X)
Superwhisper S1-mini: 0.6B open-weights, cleans messy STT transcripts fully on-device (X)
Tools & Agentic Engineering
Links & Resources
Open Source LLMs
Big CO LLMs + APIs
- Sam Altman on the RL pause (X) ↗
- OpenAI announcement (X) ↗
- Greg Brockman: The Defender's Window ↗
- Stripe acquires OpenRouter — Alex Atallah (X) ↗
- Anthropic: Claude-designed protein binders (X) ↗
- OpenAI joins PORTS-Pike 8GW Ohio DC (X) ↗
- DeepSeek V4 surge pricing (X) ↗
- Claude Code /design (X) ↗
- ChatGPT Ads expand across Europe ↗