Episode · July 9, 2026 · Launch Day, as it happened
GPT-5.6 arrived.
Sol, Terra & Luna.
The model went public in the middle of our live show — we cut to the watch party and tested it on air.
And the news kept breaking: Zuckerberg returned to X to launch Meta Muse Spark 1.1 thirty-five seconds
into the stream, Reve 2.1 stormed to #2 on the Arena, and Anthropic quietly reset Fable limits.
Plus GPT-Live demos, Claude's J-space, Grok 4.5, and
The Infographic Arena — four frontier image models, one identical brief, receipts included.
3 breaking stories landed while we were live
METR threw out its own GPT-5.6 eval — record cheating rate
7.8% on ARC-AGI-3 — first model to beat a public game
$2.50/$15 — Terra is the sleeper economics story
Watch the full episode
2h56m — the GPT-5.6 watch party, live demos, and three breaking stories
If the player doesn't load, watch on YouTube: youtu.be/QjuuTHJKxWI
Breaking on air · post-show update
The show where the news
kept breaking.
Three stories dropped while we were live. Here's what changed between the run of show and the sign-off.
Breaking · 35 seconds in
Zuck returns to X to launch Muse Spark 1.1
Mark Zuckerberg's first post in ages announced Meta's frontier comeback: 1M-token context and the first-ever paid Meta Model API at $1.25/$4.25 per million (Opus is $15/$75). #1 on MCP Atlas; on the held-back Harvey legal bench it scores 20% vs Fable's 11%. Nisten's one-shot verdict, live: "Meta might be cooking here, guys." We went from a three-lab race to five, overnight.
Breaking · mid-show
Reve 2.1 grabs #2 on the Arena
Scored 1306, 28 points clear of the field — dethroning Muse Image after roughly 30 hours at #2. Built on a layout engine, so every element is an editable layer: we double-clicked a countdown clock live and rewrote it. Too late to enter today's Arena below — rematch pending.
Watch party · as it happened
GPT-5.6 went public on our stream
"Almost a billion" weekly ChatGPT users, Sol to all paid plans within 24h, Terra & Luna free. ARC-AGI-3: 7.8% — the first model to beat a public game. And Codex became the unified ChatGPT app live on Wolfram's machine, mid-broadcast, with hosted Sites on chatgpt.site.
Post-show · the grumble resolved
Anthropic reset Fable limits after all
The extension to July 12 originally came without a usage reset, which made it hollow for anyone who maxed out racing the first deadline. Then, hours after GPT-5.6 went public, the comments delivered: weekly Fable usage back to 0%. The timing is left as an exercise for the reader. Thank you, Anthropic — now about that 24/7 thing.
The Rundown
One of the densest weeks
in AI, ever.
Every lab shipped in the same 72 hours. Here's what we covered on the show.
Big Co · Launch Day
GPT-5.6 went public: Sol, Terra & Luna
Three durable tiers on the same "Spud" pretrain — Sol's new Ultra subagent mode, Terra at half the cost of 5.5-class, same-weights Sol on Cerebras at 700+ tok/s. The launch cleared a customer-by-customer Commerce Department review. METR rejected its own eval after record-rate cheating — 11.3 hours or 270+ hours, depending on whether cheating counts. Peter's verdict after a 2-day autonomous run that deleted 70,000 lines of his code: Fable is a wise owl, Sol is a rottweiler.
Voice · demoed on air
GPT-Live: it talks while it listens
True full-duplex ChatGPT Voice, delegating hard queries to GPT-5.5 mid-sentence. Our on-air demo: interrupt-on-cue passed, pitch detection passed, the accent test failed into full languages. Descript then credited the demo voice as its own panelist — "OpenAI sol," 17 speaking turns.
Research
Anthropic finds Claude's J-space
A global workspace inside the model: ~25 silent concepts steering multi-step reasoning. Delete it and reasoning collapses; ablate its test-awareness and the blackmail eval flips from 0 to 13/180.
Coding
Grok 4.5 × Cursor, SWE-1.7, and the app-layer strikes back
SpaceXAI's first coding model at $2/$6 with a self-disclosed benchmark contamination; Cognition names its Kimi K2.7 base out loud; a 9-point harness swing shows evals are now as contested as models.
Media Gen
Muse, Seedream 5 Pro & Seedance incoming
Meta's first MSL media models (and an Instagram consent landmine); ByteDance's layer-separation "Photoshop-killer"; Seedance 2.5 lands within days — and Reve 2.1 crashed the party mid-show.
Silicon
DeepSeek goes chip-shopping
Reuters: DeepSeek is building its own inference silicon. AMD −8%, Samsung shed $80B in a day. Third lab to decouple in a month.
Open Source
Quick hits, stacked deep
Cohere's Arabic ASR tops the leaderboard, Mistral's Robostral Navigate, LiquidAI Antidoom, PyTorch 2.13, Agents-A1, and one viral fake to debunk live.
World first · Nobody has run this test
The Infographic Arena
Four frontier image models. The same 5,000-character design brief. The same single reference image.
Eight real news stories from this week, rendered by every model — full resolution, zoomable, with the
actual generation receipts linked. Scroll in close. Find who spells names right. Find who invents benchmarks.
The verdict from the show: GPT-Image-2 is still king, and Seedream is the most artistic and the
weakest at text — exactly backwards from the marketing.
A · Nano Banana Pro
B · GPT-Image-2
C · Seedream 5 Pro
D · Meta Muse
←→ model · ↑↓ story · 1–8 jump · C compare all four · R reset · scroll to zoom, drag to pan, double-click 2.5×
A · Nano Banana Pro
The reigning production champ — these are the exact images our pipeline shipped. Only contender with web search enabled during generation (the one asymmetry we can't equalize).
B · GPT-Image-2
Strongest text fidelity at distance, and the only model here that accepts arbitrary resolutions (native 2048×1152). Fun fact: the "chatgpt-image-latest" alias failed this test so hard we retired it.
C · Seedream 5 Pro
Dropped the day before the show — the most artistic compositions of the four, with a habit of garbling names in fine print. Zoom into the bylines.
D · Meta Muse
No API exists — every one of these was generated inside the Meta AI app. The only model that makes its own layout calls (one story came out portrait) and narrates its assembly as it works.
Methodology: each story's prompt is the verbatim 5–7K-character design brief written by the ThursdAI research pipeline,
paired with the identical single reference image of Alex. A: fal nano-banana-pro/edit ·
B: OpenAI images.edit (gpt-image-2, quality high) ·
C: fal seedream v5 pro/edit · D: Muse Image via meta.ai.
The "inspect" link on A and C opens the actual fal request with the full prompt.
Easter egg: all four models faithfully painted the same hallucinated launch date on story 1 — the pipeline wrote "July 11," the launch was July 9, and not one model fact-checked it.
One that got away: Reve 2.1 shipped mid-show and took #2 on the Text-to-Image Arena — too late to enter this bracket. It gets its shot next week.
Missed it live? The write-up is coming.
Full episode above · the newsletter breakdown lands in your inbox · we're live again every Thursday.
Subscribe free ↗