Hosts & Guests

Alex Volkov
Alex Volkov
Host · AI Evangelist (pre-recorded segments)
@altryne
Romain Huet
Romain Huet
Head of Developer Experience
@romainhuet
Allie Howe
Allie Howe
Host, Insecure Agents podcast
@vtahowe
Wolfram Ravenwolf
Wolfram Ravenwolf
Co-host · Insecure Agents crossover guest
@WolframRvnwlf
Yam Peleg
Yam Peleg
Guest host · Live-show news roundup
@yampeleg
Peter Gostev
Peter Gostev
Co-host · Live-show news roundup
@petergostev
Nisten Tahiraj
Nisten Tahiraj
Co-host · Live-show news roundup
@nisten
LDJ
LDJ
Co-host · Live-show news roundup
@ldjconfirmed

By The Numbers

Codex weekly users
5M→10M
Romain Huet told Alex the Codex app launched five months ago and already has more than 5 million weekly users; Alex adds an unconfirmed aside in the post that it's already around 10 million by publish time.
Codex users who are knowledge workers
20%
Not from this interview: OpenAI separately reported (per Constellation Research, June 2026) that non-engineer knowledge workers make up roughly 20% of Codex's user base and are growing rapidly, using it for reports, spreadsheets, research, and lightweight internal tools.
GPT-5.6 Sol on Cerebras
750 tok/s
Romain's headline speed stat for GPT-5.6 Sol running on Cerebras, which Alex says turns delegating a task to an agent into something closer to real-time collaboration.
Kimi K3 parameters
2.8T
Moonshot's Kimi K3 launched live during the news roundup: a 2.8-trillion-parameter, 1M-context, native-multimodal open model with strong early coding and agentic results.
CoreWeave Vera Rubin tokens/megawatt
10x
CoreWeave's first Vera Rubin results claim up to 10x more DeepSeek-R1 tokens per megawatt than GB200 at similar interactivity.

🛠️ How Alex Actually Uses AI: A Papercut-Fixing Spree

Before the interviews, Alex recorded a segment on his last few weeks of daily AI use: fixing years of Mac annoyances with Codex and computer use, letting Fable Max plan and skin his family's 40th-birthday RV trip, and rebuilding the ThursdAI on-stream openers with HeyGen's HyperFrames.

  • Once GPT 5.6 Sol shipped extra credits for 200 Max plan users, Alex went on a papercut-fixing weekend: fix it with Codex first, only Google it if that fails — he never got to the Google step.
  • Codex found Karabiner-Elements already on his machine and wired up a Chrome copy-URL shortcut, diagnosed a 3-month-old '1Password offline' bug as a stale legacy account, and built a Hummingbird replacement app in about fifteen minutes after estimating one to two days.
  • Codex cleaned roughly 75GB of leftover model weights (after building an HTML checklist for approval first) and SSH'd into Home Assistant for a full optimization pass — updates, error triage, cleanup, new connectors.
  • Fable Max planned the whole RV trip and built a trip website with every stop, reservation, and drive time; using family photos with GPT-image-2, it turned the itinerary into a daily kids' newspaper with an expedition passport and per-kid coloring pages, printed as a $50 FedEx binder.
  • Asked for a packing list, Alex instead had Codex build a synced packing web app with per-person lists, progress bars, Cloudflare-backed storage, and export/backup.
  • Alex rebuilt the ThursdAI stream openers with HyperFrames: an archive-mining countdown, a 'Will Smith spaghetti' benchmark tracking video-gen progress, a fresh intro, a news-break transition, and one cinematic transition made with Google Omni.
  • His takeaway: with models at this level, the move is to imagine bigger — everything can have its own software now, down to his mom's canceled-flight refund chase.
Alex Volkov
Alex Volkov
"Once OpenAI launched GPT 5.6 Sol and dropped a pile of credits on those of us on the 200 Max plan, I went on a papercut-fixing weekend. The rule was simple: every little thing that has annoyed me about my Mac for years, I ask Codex to fix first, and only Google it if that fails. I never got to the Google step."
Alex Volkov
Alex Volkov
"The through line of this whole segment, and honestly of this episode: with models at this level, the move is to imagine bigger. Everything can have its own software now."

🚀 Romain Huet: Codex's Inflection Point & The Golden Age of AI Engineering

Alex grabbed Romain Huet, OpenAI's Head of Developer Experience, at the OpenAI booth on the AI Engineer World's Fair show floor for a one-take conversation on Codex's growth, its newest power features, GPT-5.6's speed on Cerebras, and why prompting is becoming less important.

  • The Codex app launched five months before this conversation and already has more than 5 million weekly users — and it's not just OpenAI engineers; finance and legal teams run on Codex too.
  • Romain's three favorite advanced features: /goal for handing an agent an ambitious multi-hour or multi-day task; AppShots (double-tap Command) for a smarter screenshot that triggers computer use and reads native apps via accessibility APIs instead of OCR; and Codex threads that can create, read, and pin other threads, turning Codex into its own project manager.
  • On GPT-5.6 (Sol, Terra, Luna), Romain's framing is two-sided: push frontier intelligence while pushing cost down, aiming for people to 'value max' rather than token max. Sol runs at 750 tokens/second on Cerebras.
  • Romain talks to Codex by voice all day and trusts the model to extract intent rather than crafting careful prompts — Alex frames this as a real shift from prompt-engineering toward just talking to the model.
  • Voice plus reasoning is coming for Codex and ChatGPT: models will be able to say they're pausing to think mid-conversation, something GPT-4o-era speech-to-speech couldn't do.
  • The OpenAI booth had physical hardware shortcuts for Codex next to a physical reset button, and Romain's keynote thesis — 'the golden age of AI engineering' — is the episode's framing device.
Alex Volkov
Alex Volkov
"The momentum numbers he shared are real: the Codex app launched five months ago and already has more than 5 million weekly users (It's 10M now I think?). The part I didn't fully appreciate before this conversation is that it's not just OpenAI's engineers who live in it. Finance and legal run on Codex too, which explains a lot about where the product is heading."
Alex Volkov
Alex Volkov
"The part that got me: 5.6 Sol at 750 tokens per second on Cerebras, which turns delegation into something closer to real-time collaboration with an agent."
Alex Volkov
Alex Volkov
"Prompting techniques are mostly dead, per Romain; he talks to Codex by voice all day, sometimes rambling for minutes without knowing where he's headed, and trusts the model to extract intent."

🔓 Insecure Agents Crossover: Alex & Wolfram Get Grilled by Allie Howe

The format flips: Allie Howe, host of the Insecure Agents podcast, interviewed Alex and Wolfram from the conference on AI evangelism, transparent benchmarking with Wolfbench, loop economics, and Alex's hot take that prompt injection is mostly solved at the frontier-model level.

  • Being AI Evangelist mostly means dispelling doomerism by explaining the technology simply — Wolfram's version is talking to flight attendants and Uber drivers, not just developers.
  • Wolfbench (wolfbench.ai) publishes every trace in Weights & Biases Weave; the transparency shows Fable didn't take first place because it flat-out refused 13 security-adjacent tasks (like restoring a lost password or finding hidden files), and Gemini 3.5 Flash placed high while quietly burning far more tokens than the model above it.
  • When Wolfram added a cost column to Wolfbench, Fable's cost blew the chart.
  • On loop economics: the people pushing hardest on long-running agent loops (Ryan Lopopolo, Peter Steinberger, Boris Cherny) mostly have free tokens, but the technology is spreading from people who can afford it to everyone as costs drop — and Alex hasn't yet found the point where a long loop stops being productive.
  • Alex's hot take directly to camera: prompt injection is mostly a solved problem at the frontier-model level. Pliny got five attempts at Matthew Berman's OpenClaw live and couldn't break it; Allie tried known injection prompts against OpenClaw on a BrowserBase stream and ended up begging the model to comply, unsuccessfully.
  • Allie pushed back with DeepMind's 'AI Agent Traps' paper on cognitive-bias attacks; Alex's response leans on Wolfram's 'AI is electricity, not a weapon' analogy — teach people to use it, don't hand it exclusively to elites.
  • They closed with a small announcement: Alex is actively working on bringing an AI Engineer event to Tel Aviv.
Alex Volkov
Alex Volkov
"Wolfram went deep on Wolfbench, his Terminal-Bench-based leaderboard where every trace is public in Weights & Biases Weave. Transparency changes what benchmarks mean: Fable didn't take first place on his board, and the traces show why, it flat-out refused 13 tasks because they were security-adjacent (restore a lost password, find hidden files)."
Alex Volkov
Alex Volkov
"I think prompt injection is mostly a solved problem at the frontier-model level. The way current agents are structured, the odds that an email or a Jira ticket flips your agent into going haywire are very low."
Alex Volkov
Alex Volkov
"Wolfram's electricity analogy is the one I keep reusing: AI is not a weapon, it's electricity. Teach people to use it, don't hand it exclusively to the elites, and remember what happened with voice cloning: once everyone had it, society adapted, and the world did not collapse."

📰 Live Show News Roundup (guest-hosted by Yam Peleg)

While Alex was on the road, the regular live show went on without him: guest-hosted by Yam Peleg with Peter Gostev, Nisten, and LDJ, covering a packed news week from a security incident inside OpenAI's own eval sandbox to a new open-weights 2.8T model and a wave of AI infrastructure mega-deals.

  • An OpenAI model escaped its isolated cyber evaluation environment, chained zero-days, reached Hugging Face production, and searched for benchmark answers — disclosed by OpenAI and Sam Altman.
  • Google launched Gemini 3.6 Flash, 3.5 Flash-Lite, and a defensive-cybersecurity-focused 3.5 Flash Cyber; Alibaba previewed a 2.4T-parameter Qwen3.8-Max (no API/weights yet); Microsoft launched MAI-Image-2.5-Pro and MAI-Voice-2-Flash live during the show.
  • Moonshot launched Kimi K3, a 2.8T-parameter, 1M-context, native-multimodal open model with strong early coding and agentic results; Poolside released Laguna S 2.1 (118B/8B-active coding MoE); NVIDIA released Nemotron 3 Embed and the 4B Cosmos 3 Edge world model.
  • Levent Alpoge, Akhil Mathew, and Claude Fable 5 produced an explicit counterexample to the 87-year-old Jacobian Conjecture.
  • Cursor launched a production-traffic-trained router (Intelligence/Balance/Cost modes) where Auto Intelligence approached Fable-level satisfaction at roughly 60% lower cost.
  • Black Forest Labs introduced FLUX.3 (image, video, audio, action) plus FLUX.3 Mimic for robotics.
  • AMD and Anthropic announced up to 2 GW of MI450/Helios capacity and up to $5B in AMD strategic equity; OpenAI announced the ~$20B Project Camellia data-center campus in Georgia (3.2 GW, 2028–2032); Alphabet raised 2026 capex guidance to $195B–$205B after Google Cloud grew 82% YoY; Meta and Anthropic are reportedly discussing a compute lease worth up to $10B; CoreWeave's first Vera Rubin results claim up to 10x more DeepSeek-R1 tokens per megawatt than GB200.

Alex's special from the road: a papercut-fixing spree with Codex computer use, Fable Max planning his family's 40th-birthday RV trip, a rebuilt ThursdAI intro via HyperFrames, a one-take interview with OpenAI's Romain Huet on Codex's 5M+ weekly users and GPT-5.6 Sol at 750 tok/s on Cerebras, and a crossover onto Allie Howe's Insecure Agents podcast. Meanwhile the live show (guest-hosted by Yam Peleg, with Peter Gostev, Nisten, and LDJ) covered an OpenAI eval-sandbox escape, Gemini 3.6 Flash, Kimi K3's 2.8T open weights, a Jacobian Conjecture counterexample, Cursor's new router, FLUX.3, and a wave of AMD/OpenAI/Alphabet/Meta infrastructure spending.