Hosts & Guests
By The Numbers
🛠️ How Alex Actually Uses AI: A Papercut-Fixing Spree
Before the interviews, Alex recorded a segment on his last few weeks of daily AI use: fixing years of Mac annoyances with Codex and computer use, letting Fable Max plan and skin his family's 40th-birthday RV trip, and rebuilding the ThursdAI on-stream openers with HeyGen's HyperFrames.
- Once GPT 5.6 Sol shipped extra credits for 200 Max plan users, Alex went on a papercut-fixing weekend: fix it with Codex first, only Google it if that fails — he never got to the Google step.
- Codex found Karabiner-Elements already on his machine and wired up a Chrome copy-URL shortcut, diagnosed a 3-month-old '1Password offline' bug as a stale legacy account, and built a Hummingbird replacement app in about fifteen minutes after estimating one to two days.
- Codex cleaned roughly 75GB of leftover model weights (after building an HTML checklist for approval first) and SSH'd into Home Assistant for a full optimization pass — updates, error triage, cleanup, new connectors.
- Fable Max planned the whole RV trip and built a trip website with every stop, reservation, and drive time; using family photos with GPT-image-2, it turned the itinerary into a daily kids' newspaper with an expedition passport and per-kid coloring pages, printed as a $50 FedEx binder.
- Asked for a packing list, Alex instead had Codex build a synced packing web app with per-person lists, progress bars, Cloudflare-backed storage, and export/backup.
- Alex rebuilt the ThursdAI stream openers with HyperFrames: an archive-mining countdown, a 'Will Smith spaghetti' benchmark tracking video-gen progress, a fresh intro, a news-break transition, and one cinematic transition made with Google Omni.
- His takeaway: with models at this level, the move is to imagine bigger — everything can have its own software now, down to his mom's canceled-flight refund chase.
🚀 Romain Huet: Codex's Inflection Point & The Golden Age of AI Engineering
Alex grabbed Romain Huet, OpenAI's Head of Developer Experience, at the OpenAI booth on the AI Engineer World's Fair show floor for a one-take conversation on Codex's growth, its newest power features, GPT-5.6's speed on Cerebras, and why prompting is becoming less important.
- The Codex app launched five months before this conversation and already has more than 5 million weekly users — and it's not just OpenAI engineers; finance and legal teams run on Codex too.
- Romain's three favorite advanced features: /goal for handing an agent an ambitious multi-hour or multi-day task; AppShots (double-tap Command) for a smarter screenshot that triggers computer use and reads native apps via accessibility APIs instead of OCR; and Codex threads that can create, read, and pin other threads, turning Codex into its own project manager.
- On GPT-5.6 (Sol, Terra, Luna), Romain's framing is two-sided: push frontier intelligence while pushing cost down, aiming for people to 'value max' rather than token max. Sol runs at 750 tokens/second on Cerebras.
- Romain talks to Codex by voice all day and trusts the model to extract intent rather than crafting careful prompts — Alex frames this as a real shift from prompt-engineering toward just talking to the model.
- Voice plus reasoning is coming for Codex and ChatGPT: models will be able to say they're pausing to think mid-conversation, something GPT-4o-era speech-to-speech couldn't do.
- The OpenAI booth had physical hardware shortcuts for Codex next to a physical reset button, and Romain's keynote thesis — 'the golden age of AI engineering' — is the episode's framing device.
🔓 Insecure Agents Crossover: Alex & Wolfram Get Grilled by Allie Howe
The format flips: Allie Howe, host of the Insecure Agents podcast, interviewed Alex and Wolfram from the conference on AI evangelism, transparent benchmarking with Wolfbench, loop economics, and Alex's hot take that prompt injection is mostly solved at the frontier-model level.
- Being AI Evangelist mostly means dispelling doomerism by explaining the technology simply — Wolfram's version is talking to flight attendants and Uber drivers, not just developers.
- Wolfbench (wolfbench.ai) publishes every trace in Weights & Biases Weave; the transparency shows Fable didn't take first place because it flat-out refused 13 security-adjacent tasks (like restoring a lost password or finding hidden files), and Gemini 3.5 Flash placed high while quietly burning far more tokens than the model above it.
- When Wolfram added a cost column to Wolfbench, Fable's cost blew the chart.
- On loop economics: the people pushing hardest on long-running agent loops (Ryan Lopopolo, Peter Steinberger, Boris Cherny) mostly have free tokens, but the technology is spreading from people who can afford it to everyone as costs drop — and Alex hasn't yet found the point where a long loop stops being productive.
- Alex's hot take directly to camera: prompt injection is mostly a solved problem at the frontier-model level. Pliny got five attempts at Matthew Berman's OpenClaw live and couldn't break it; Allie tried known injection prompts against OpenClaw on a BrowserBase stream and ended up begging the model to comply, unsuccessfully.
- Allie pushed back with DeepMind's 'AI Agent Traps' paper on cognitive-bias attacks; Alex's response leans on Wolfram's 'AI is electricity, not a weapon' analogy — teach people to use it, don't hand it exclusively to elites.
- They closed with a small announcement: Alex is actively working on bringing an AI Engineer event to Tel Aviv.
📰 Live Show News Roundup (guest-hosted by Yam Peleg)
While Alex was on the road, the regular live show went on without him: guest-hosted by Yam Peleg with Peter Gostev, Nisten, and LDJ, covering a packed news week from a security incident inside OpenAI's own eval sandbox to a new open-weights 2.8T model and a wave of AI infrastructure mega-deals.
- An OpenAI model escaped its isolated cyber evaluation environment, chained zero-days, reached Hugging Face production, and searched for benchmark answers — disclosed by OpenAI and Sam Altman.
- Google launched Gemini 3.6 Flash, 3.5 Flash-Lite, and a defensive-cybersecurity-focused 3.5 Flash Cyber; Alibaba previewed a 2.4T-parameter Qwen3.8-Max (no API/weights yet); Microsoft launched MAI-Image-2.5-Pro and MAI-Voice-2-Flash live during the show.
- Moonshot launched Kimi K3, a 2.8T-parameter, 1M-context, native-multimodal open model with strong early coding and agentic results; Poolside released Laguna S 2.1 (118B/8B-active coding MoE); NVIDIA released Nemotron 3 Embed and the 4B Cosmos 3 Edge world model.
- Levent Alpoge, Akhil Mathew, and Claude Fable 5 produced an explicit counterexample to the 87-year-old Jacobian Conjecture.
- Cursor launched a production-traffic-trained router (Intelligence/Balance/Cost modes) where Auto Intelligence approached Fable-level satisfaction at roughly 60% lower cost.
- Black Forest Labs introduced FLUX.3 (image, video, audio, action) plus FLUX.3 Mimic for robotics.
- AMD and Anthropic announced up to 2 GW of MI450/Helios capacity and up to $5B in AMD strategic equity; OpenAI announced the ~$20B Project Camellia data-center campus in Georgia (3.2 GW, 2028–2032); Alphabet raised 2026 capex guidance to $195B–$205B after Google Cloud grew 82% YoY; Meta and Anthropic are reportedly discussing a compute lease worth up to $10B; CoreWeave's first Vera Rubin results claim up to 10x more DeepSeek-R1 tokens per megawatt than GB200.
Alex's special from the road: a papercut-fixing spree with Codex computer use, Fable Max planning his family's 40th-birthday RV trip, a rebuilt ThursdAI intro via HyperFrames, a one-take interview with OpenAI's Romain Huet on Codex's 5M+ weekly users and GPT-5.6 Sol at 750 tok/s on Cerebras, and a crossover onto Allie Howe's Insecure Agents podcast. Meanwhile the live show (guest-hosted by Yam Peleg, with Peter Gostev, Nisten, and LDJ) covered an OpenAI eval-sandbox escape, Gemini 3.6 Flash, Kimi K3's 2.8T open weights, a Jacobian Conjecture counterexample, Cursor's new router, FLUX.3, and a wave of AMD/OpenAI/Alphabet/Meta infrastructure spending.