APIs & Platforms
GPT-5.6 API pricing
Breaking on the show: OpenAI cuts GPT-5.6 Luna prices 80% and Terra 20%, crediting Sol's self-optimization
Dropped live during the episode: Luna prices fall 80%, Terra 20%, and a faster GPT-5.6 Sol option lands in the API, with lower prices reflected in Codex usage metering. OpenAI explicitly credits efficiency work GPT-5.6 Sol performed on its own serving stack — 20% lower serving costs from production GPU kernel improvements and 15% better token generation from improved speculative decoding — prompting the panel's on-air debate about whether recursive self-improvement is already here as a gradual spectrum.
-80% / -20% Luna / Terra price cuts20% + 15% serving-cost and token-generation gains, model-authored
Also ReleasedOpen weights
MCP 2026-07-28 spec
MCP's biggest update ever: fully stateless core, MCP Apps and Tasks extensions, OAuth 2.1
The 2026-07-28 spec makes MCP fully stateless — no handshakes, no sessions, every request self-describing — enabling serverless deployment behind plain round-robin load balancers (GitHub dropped its Redis session store). The extensions framework formalizes Tasks for long-running async work and MCP Apps for interactive UIs rendered in sandboxed iframes inside conversations. Monthly SDK downloads hit half a billion, up from 97 million in March, and Amazon Bedrock supports the new spec day one.
500M monthly SDK downloads, up from 97M in March
New Models
Ling-3.0-Flash
Ant Group's Ling-3.0-Flash: a 124B MoE that claims to match its 1T flagship on ~5B active parameters
Ant Group's inclusionAI lab released Ling-3.0-Flash, a 124B-parameter MoE activating only ~5.1B parameters per token, which Ant says matches or beats its own trillion-parameter flagship on most published benchmarks. It pairs native hybrid-linear attention with cluster-level hierarchical caching that Ant claims cuts time-to-first-token on long inputs by 60-80%. Free via API on OpenRouter and Vercel AI Gateway through August 3, with weights promised as open source afterward — an efficiency play aimed squarely at models 2-3x its scale.
124B / ~5.1B total / active parameters (MoE)60-80% claimed TTFT reduction on long inputs
Benchmarks & Evals
NVIDIA Vera Rubin NVL72 results
CoreWeave posts first Vera Rubin results: up to 10x more tokens per megawatt than Blackwell
CoreWeave published the first customer results for NVIDIA's Vera Rubin NVL72 platform, claiming up to 10x more DeepSeek-R1 tokens per megawatt than GB200 at similar interactivity — an efficiency leap that lands directly on the industry's power-constrained bottleneck. Covered on the Jul 23 live show's infrastructure block.
10x DeepSeek-R1 tokens per megawatt vs GB200 (up to)
Also Released
AMD MI450 / Helios capacity deal
Anthropic signs an AMD compute deal: up to 2GW of MI450/Helios capacity, with up to $5B in AMD equity
Covered on the Jul 23 live show: Anthropic and AMD struck a capacity agreement giving Anthropic access to up to 2 gigawatts of MI450/Helios-generation compute, a deal that includes up to $5 billion in AMD equity — landing the same week AMD's Advancing AI 2026 event launched the Helios/MI400 platform and reports surfaced of separate Meta/Anthropic compute-lease talks (~$10B over two years). A diversification play away from single-vendor GPU dependence at frontier scale.
2GW MI450/Helios compute capacity (up to)$5B AMD equity component (up to)
Major Features & UpdatesOpen weights
Gemma 4 (stealth update)
Google quietly patches Gemma 4 with Flash Attention 4 and tool-calling fixes — no version bump
Google shipped a stealth update to the Gemma 4 family: Flash Attention 4 support on Hopper-class GPUs (a reported 25-70% prefill throughput speedup), tool-calling bug fixes, reduced model 'laziness,' and configurable vision resolution. The ThursdAI panel criticized shipping new weights under the same Gemma 4 name with no version bump, leaving users unsure which checkpoint they're actually running.
25-70% prefill throughput speedup (Flash Attention 4)
Dev ToolsOpen weights
PyTorch 2.13
PyTorch 2.13 lands FlexAttention on Apple Silicon and big memory wins
3,328 commits from 526 contributors: FlexAttention on Apple Silicon at roughly 12x over SDPA for sparse patterns, a deterministic CUDA backward path, nn.LinearCrossEntropyLoss with up to 4x peak-memory reduction, torchcomms for large-cluster training, and expanded ROCm/Arm/XPU support.
~12x FlexAttention on Apple Silicon vs SDPA3,328 Commits from 526 contributors
Products & Apps
local.ai
Exo Labs launches local.ai to track the local-AI frontier
Announced live on ThursdAI at AI Engineer: local.ai tracks the best model for your hardware, the performance trade versus the cloud, and whether running local beats API-token pricing. Early access is live with signup codes, and the Exo CLI — 'vLLM for consumer devices, with the configs figured out for you' — ships in the coming weeks.
71% Terminal Bench 2.1, REAP-pruned GLM 5.2550B Nemotron-3 Ultra running on 4 NVIDIA Sparks
Funding
Series C
Together AI raises $800M Series C at an $8.3B valuation
Aramco Ventures led the round with NVIDIA, Vista Equity and General Catalyst participating. The open-model cloud reports over $1B in annual bookings, says open-model usage on the platform tripled year over year, and plans roughly 50x infrastructure growth over five years.
$800M Series C$8.3B Valuation>$1B Annual bookings