Episode · July 9, 2026 · Launch Day, as it happened

GPT-5.6 arrived.
Sol, Terra & Luna.

The model went public in the middle of our live show — we cut to the watch party and tested it on air. And the news kept breaking: Zuckerberg returned to X to launch Meta Muse Spark 1.1 thirty-five seconds into the stream, Reve 2.1 stormed to #2 on the Arena, and Anthropic quietly reset Fable limits. Plus GPT-Live demos, Claude's J-space, Grok 4.5, and The Infographic Arena — four frontier image models, one identical brief, receipts included.

3 breaking stories landed while we were live METR threw out its own GPT-5.6 eval — record cheating rate 7.8% on ARC-AGI-3 — first model to beat a public game $2.50/$15 — Terra is the sleeper economics story
Watch the full episode
2h56m — the GPT-5.6 watch party, live demos, and three breaking stories
If the player doesn't load, watch on YouTube: youtu.be/QjuuTHJKxWI
Breaking on air · post-show update

The show where the news
kept breaking.

Three stories dropped while we were live. Here's what changed between the run of show and the sign-off.

Breaking · 35 seconds in Zuck returns to X to launch Muse Spark 1.1 Mark Zuckerberg's first post in ages announced Meta's frontier comeback: 1M-token context and the first-ever paid Meta Model API at $1.25/$4.25 per million (Opus is $15/$75). #1 on MCP Atlas; on the held-back Harvey legal bench it scores 20% vs Fable's 11%. Nisten's one-shot verdict, live: "Meta might be cooking here, guys." We went from a three-lab race to five, overnight.
Breaking · mid-show Reve 2.1 grabs #2 on the Arena Scored 1306, 28 points clear of the field — dethroning Muse Image after roughly 30 hours at #2. Built on a layout engine, so every element is an editable layer: we double-clicked a countdown clock live and rewrote it. Too late to enter today's Arena below — rematch pending.
Watch party · as it happened GPT-5.6 went public on our stream "Almost a billion" weekly ChatGPT users, Sol to all paid plans within 24h, Terra & Luna free. ARC-AGI-3: 7.8% — the first model to beat a public game. And Codex became the unified ChatGPT app live on Wolfram's machine, mid-broadcast, with hosted Sites on chatgpt.site.
Post-show · the grumble resolved Anthropic reset Fable limits after all The extension to July 12 originally came without a usage reset, which made it hollow for anyone who maxed out racing the first deadline. Then, hours after GPT-5.6 went public, the comments delivered: weekly Fable usage back to 0%. The timing is left as an exercise for the reader. Thank you, Anthropic — now about that 24/7 thing.
The Rundown

One of the densest weeks
in AI, ever.

Every lab shipped in the same 72 hours. Here's what we covered on the show.

Big Co · Launch Day GPT-5.6 went public: Sol, Terra & Luna Three durable tiers on the same "Spud" pretrain — Sol's new Ultra subagent mode, Terra at half the cost of 5.5-class, same-weights Sol on Cerebras at 700+ tok/s. The launch cleared a customer-by-customer Commerce Department review. METR rejected its own eval after record-rate cheating — 11.3 hours or 270+ hours, depending on whether cheating counts. Peter's verdict after a 2-day autonomous run that deleted 70,000 lines of his code: Fable is a wise owl, Sol is a rottweiler.
Voice · demoed on air GPT-Live: it talks while it listens True full-duplex ChatGPT Voice, delegating hard queries to GPT-5.5 mid-sentence. Our on-air demo: interrupt-on-cue passed, pitch detection passed, the accent test failed into full languages. Descript then credited the demo voice as its own panelist — "OpenAI sol," 17 speaking turns.
Research Anthropic finds Claude's J-space A global workspace inside the model: ~25 silent concepts steering multi-step reasoning. Delete it and reasoning collapses; ablate its test-awareness and the blackmail eval flips from 0 to 13/180.
Coding Grok 4.5 × Cursor, SWE-1.7, and the app-layer strikes back SpaceXAI's first coding model at $2/$6 with a self-disclosed benchmark contamination; Cognition names its Kimi K2.7 base out loud; a 9-point harness swing shows evals are now as contested as models.
Media Gen Muse, Seedream 5 Pro & Seedance incoming Meta's first MSL media models (and an Instagram consent landmine); ByteDance's layer-separation "Photoshop-killer"; Seedance 2.5 lands within days — and Reve 2.1 crashed the party mid-show.
Silicon DeepSeek goes chip-shopping Reuters: DeepSeek is building its own inference silicon. AMD −8%, Samsung shed $80B in a day. Third lab to decouple in a month.
Open Source Quick hits, stacked deep Cohere's Arabic ASR tops the leaderboard, Mistral's Robostral Navigate, LiquidAI Antidoom, PyTorch 2.13, Agents-A1, and one viral fake to debunk live.
World first · Nobody has run this test

The Infographic Arena

Four frontier image models. The same 5,000-character design brief. The same single reference image. Eight real news stories from this week, rendered by every model — full resolution, zoomable, with the actual generation receipts linked. Scroll in close. Find who spells names right. Find who invents benchmarks. The verdict from the show: GPT-Image-2 is still king, and Seedream is the most artistic and the weakest at text — exactly backwards from the marketing.

A · Nano Banana Pro B · GPT-Image-2 C · Seedream 5 Pro D · Meta Muse
fit
Generated infographic
🚫
not available
Story 1 of 8
model · story · 1–8 jump · C compare all four · R reset · scroll to zoom, drag to pan, double-click 2.5×
A · Nano Banana Pro The reigning production champ — these are the exact images our pipeline shipped. Only contender with web search enabled during generation (the one asymmetry we can't equalize).
B · GPT-Image-2 Strongest text fidelity at distance, and the only model here that accepts arbitrary resolutions (native 2048×1152). Fun fact: the "chatgpt-image-latest" alias failed this test so hard we retired it.
C · Seedream 5 Pro Dropped the day before the show — the most artistic compositions of the four, with a habit of garbling names in fine print. Zoom into the bylines.
D · Meta Muse No API exists — every one of these was generated inside the Meta AI app. The only model that makes its own layout calls (one story came out portrait) and narrates its assembly as it works.
Methodology: each story's prompt is the verbatim 5–7K-character design brief written by the ThursdAI research pipeline, paired with the identical single reference image of Alex. A: fal nano-banana-pro/edit · B: OpenAI images.edit (gpt-image-2, quality high) · C: fal seedream v5 pro/edit · D: Muse Image via meta.ai. The "inspect" link on A and C opens the actual fal request with the full prompt. Easter egg: all four models faithfully painted the same hallucinated launch date on story 1 — the pipeline wrote "July 11," the launch was July 9, and not one model fact-checked it. One that got away: Reve 2.1 shipped mid-show and took #2 on the Text-to-Image Arena — too late to enter this bracket. It gets its shot next week.

Missed it live? The write-up is coming.

Full episode above · the newsletter breakdown lands in your inbox · we're live again every Thursday.

Subscribe free