APIs & Platforms

Model APIs, developer platforms, pricing, and model routing. — 46 releases covered on the show.

August 2026

Meta AI
Products & Apps

Muse Code + Muse Spark 1.2

Meta ships Muse Code, a terminal coding agent that's 12-21x cheaper if you feed Meta your data

Meta Superintelligence Labs released Muse Code in beta, a terminal coding agent on Muse Spark 1.2 that plans, writes, and validates changes across large repos, now available globally. The story is the pricing: $1.25/$4.25 per million tokens standard, or $0.10/$0.20 on the 'contributor' tier where Meta trains on your data, with cached input at $0.002 per million. Early testing puts Spark 1.2 around Grok 4.5 level using ~50% more tokens. Wolfram made the case for universal open harnesses instead; Nisten flagged it as a data-generation gift for open source maintainers.

$0.10/$0.20 contributor-tier price per 1M tokens in/out$1.25/$4.25 standard price per 1M tokens1M token context window

July 2026

DeepSeek
APIs & Platforms

DeepSeek V4-Flash

DeepSeek V4-Flash enters public beta, beating its bigger sibling on agent benchmarks at $0.14/$0.28

Same 284B/13B-active architecture as the preview with all gains from post-training: 82.7 on Terminal Bench 2.1 (above V4-Pro-Preview), DeepSWE up 7.3 to 54.4, CyberGym 76.7, at $0.14/$0.28 per million tokens with a 1M context. It natively speaks the Responses API protocol with one-click Codex CLI setup. The honest caveat the panel kept: API-only, no weights, no license, so the 'open source DeepSeek' habit doesn't apply yet. Wolfram places it 'Terra level,' second on his Wolfbench; Nisten reports devs delegating 90-95% of tasks to it.

82.7 Terminal Bench 2.1, above V4-Pro-Preview7.3 → 54.4 DeepSWE jump from post-training alone$0.14 / $0.28 per 1M tokens in/out
OpenAI
APIs & Platforms

GPT-5.6 API pricing

Breaking on the show: OpenAI cuts GPT-5.6 Luna prices 80% and Terra 20%, crediting Sol's self-optimization

Dropped live during the episode: Luna prices fall 80%, Terra 20%, and a faster GPT-5.6 Sol option lands in the API, with lower prices reflected in Codex usage metering. OpenAI explicitly credits efficiency work GPT-5.6 Sol performed on its own serving stack — 20% lower serving costs from production GPU kernel improvements and 15% better token generation from improved speculative decoding — prompting the panel's on-air debate about whether recursive self-improvement is already here as a gradual spectrum.

-80% / -20% Luna / Terra price cuts20% + 15% serving-cost and token-generation gains, model-authored
xAI
New Models

Grok Voice Think Fast 2.0

Grok Voice Think Fast 2.0 tops speech-to-speech benchmarks at $0.08/minute, already answering Starlink support

xAI's next-gen voice model scores 82.9% on Artificial Analysis' Speech-to-Speech Quality Index (ahead of GPT-Realtime-2.1's 79.1%) and leads the tau-Voice agentic benchmark at 56.5% versus 45.7%. Time to first audio dropped to 0.70 seconds, reasoning tokens fell 60% with tool calls firing before the first sentence finishes, and it's already in production on Starlink customer support lines with measured conversion gains. It becomes the grok-voice-latest default on August 5.

82.9% AA Speech-to-Speech Quality Index0.70s time to first audio, down from 1.25s$0.08 per minute of audio
Model Context Protocol
Also ReleasedOpen weights

MCP 2026-07-28 spec

MCP's biggest update ever: fully stateless core, MCP Apps and Tasks extensions, OAuth 2.1

The 2026-07-28 spec makes MCP fully stateless — no handshakes, no sessions, every request self-describing — enabling serverless deployment behind plain round-robin load balancers (GitHub dropped its Redis session store). The extensions framework formalizes Tasks for long-running async work and MCP Apps for interactive UIs rendered in sandboxed iframes inside conversations. Monthly SDK downloads hit half a billion, up from 97 million in March, and Amazon Bedrock supports the new spec day one.

500M monthly SDK downloads, up from 97M in March
OpenAI
New Models

GPT-Transcribe + GPT-Live-Transcribe

OpenAI ships two context-aware transcription models with 41% fewer errors than Whisper-1

GPT-Transcribe (batch, $0.27/hour) and GPT-Live-Transcribe (streaming, $1.02/hour) replace the 4o-era ASR models: 8.98% word error rate versus Whisper-1's 15.21%, an 18% improvement for the live variant, and multilingual errors roughly halved across 22 languages. The standout feature is context prompting — keywords, language hints, and prior conversation turns measurably lift semantic accuracy, especially for names, numbers, and technical terms in noisy audio.

8.98% WER vs 15.21% for Whisper-1 (-41%)$0.27 / $1.02 per hour, batch / live
Cursor
Major Features & Updates

Production-trained router

Cursor launches a router trained on production traffic: near-Fable satisfaction at ~60% lower cost

Cursor shipped a model router trained on its own production traffic, with Intelligence, Balance, and Cost modes. In Cursor's numbers, Auto Intelligence mode approached Fable-level user satisfaction at roughly 60% lower cost — routing between frontier and cheaper models per request. Covered on the Jul 23 live show.

~60% cost reduction at near-Fable satisfaction (Cursor's numbers)
Google DeepMind
New Models

Gemini 3.6 Flash / 3.5 Flash-Lite / 3.5 Flash Cyber

Google ships a three-model Gemini Flash refresh — but still no Gemini 3.5 Pro

Google DeepMind released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. Flash 3.6 is the 'workhorse': 17% lower output-token usage than 3.5 Flash (per Artificial Analysis), output pricing cut from $9.00 to $7.50 per million tokens (input steady at $1.50), 1M-token context with 64K max output, and 58.7% on SWE-Bench Pro. Flash-Lite is the budget tier; Flash Cyber is a vulnerability-hunting model piloted only with governments and trusted partners. Conspicuously absent: Gemini 3.5 Pro, delayed for an architectural rebuild even as Google confirms Gemini 4 is in pre-training.

$1.50 / $7.50 per 1M tokens in/out (output down from $9.00)17% output-token reduction vs 3.5 Flash58.7% SWE-Bench Pro
Meta AI
New Models

Muse Spark 1.1 & Meta Model API

Meta launches Muse Spark 1.1 and its first paid Meta Model API

Mark Zuckerberg returned to X (35 seconds into the ThursdAI live show) to announce Muse Spark 1.1: a 1M-token-context agentic model that rivals GPT-5.5 and Opus 4.8 on agentic evals, claiming #1 on MCP Atlas, JobBench, Humanity's Last Exam and Finance Agent V2. It ships with Meta's first-ever paid developer API in public preview ($20 free credits, US-only at launch), computer use across desktop, browser and mobile, and parallel subagent delegation. On the held-back Vals AI Harvey legal-agent benchmark it scores 20% against Fable's 11%. Replit, Cline and Box are early partners. No open weights.

$1.25/$4.25 Per 1M tokens (in/out)1M Token context window20% vs 11% Harvey Legal Agent Bench vs Fable
Google DeepMind
APIs & Platforms

Gemini API Managed Agents

Gemini API Managed Agents add background tasks and remote MCP

Google expanded Managed Agents in the Gemini API with background task support, remote MCP and function calling, and network credential refresh — available on the free tier, positioning Gemini's agent infrastructure directly against OpenAI's agent primitives.

Free tier Availability
OpenAI
APIs & Platforms

GPT-Realtime-2.1-mini

GPT-Realtime-2.1-mini brings reasoning and tool use to the Realtime API mini tier

Two days before GPT-Live, OpenAI upgraded the Realtime API mini lineup with reasoning and tool use at unchanged pricing, plus a 25%+ p95 latency cut from improved caching. Notably it does not include GPT-Live's full-duplex capability, which remains app-exclusive.

≥25% p95 latency reduction
OpenAI
New Models

GPT-5.6

OpenAI ships GPT-5.6 as a three-model family: Sol, Terra and Luna

GPT-5.6 arrives as three models — Sol (frontier), Terra (~5.5-level intelligence at half the cost) and Luna (small and fast) — plus a new Ultra mode with a Max reasoning level and heavier sub-agent use. Dominik Kundel confirmed on ThursdAI that 5.6 Sol is coming to Cerebras at extreme speed running the same weights as the API model, not a distill.

3 models: Sol / Terra / Luna50% Terra cost vs GPT-5.5-level intelligence

June 2026

OpenRouter
APIs & Platforms

Fusion API

OpenRouter launches Fusion API, a panel of budget models competing with frontier models

OpenRouter launched Fusion API, which routes or ensembles a panel of lower-cost models to reach near-frontier results. The episode notes framed it as beating GPT-5.5 and Opus 4.8 in some comparisons while landing within roughly 1% of Claude Fable 5 at half the price.

~1% from Fable 5 in episode notes
APIs & Platforms

Kimi K2.7 Code on CoreWeave Inference

Kimi K2.7 Code goes live on W&B/CoreWeave Inference

Kimi K2.7 Code became available on W&B/CoreWeave Inference, with the episode notes calling out Blackwell NVFP4 serving, speculative decoding, and 289 tokens per second near the top of Artificial Analysis speed and price-performance charts.

289 tok/s reported throughput

May 2026

Anthropic
Major Features & Updates

Claude off-peak usage boost

Anthropic doubles Claude usage limits outside peak hours for a limited time

Anthropic doubled Claude usage outside peak hours for a limited period, covering Claude Code and other Claude surfaces. The move gives heavy users substantially more agentic and coding throughput during off-peak windows.

Google DeepMind
APIs & Platforms

Managed Agents (Gemini API)

Gemini API gets Managed Agents with hosted sandboxes and the Interactions API

Google launched Managed Agents in the Gemini API, letting developers spin up hosted Antigravity agents with Linux sandboxes and persistent state. It ships alongside the next-generation Interactions API, which Logan Kilpatrick described as designed for agentic systems rather than the old tokens-in, tokens-out model interaction pattern.

April 2026

March 2026

February 2026

Weights & Biases
Major Features & Updates

W&B Inference: MiniMax 2.5 & Kimi K2.5

W&B Inference adds MiniMax 2.5 and Kimi K2.5

Weights & Biases added MiniMax M2.5 and Kimi K2.5 to its CoreWeave-backed Inference service. The panel emphasized price/performance, with MiniMax 2.5 presented as roughly 10x cheaper than premium alternatives in some tiers and Kimi K2.5 praised for practical function calling and image-in-loop use cases.

xAI
New Models

Grok 4.20

xAI silently drops Grok 4.20 with four 500B-param collaborating agents

xAI released Grok 4.20, a multi-agent system where four 500B-parameter agents collaborate in a multi-agent UI, with a $300/month Heavy tier scaling to 16 agents. No benchmarks or evals were released with the drop. The panel found it underwhelming for coding and day-to-day agent work but still top tier for deep research thanks to xAI's RAG over X data; Grok 4.1 Fast remains #8 on OpenRouter by API usage.

500B×4 Grok 4 20 Architecture
OpenAI
New Models

GPT-5.3-Codex

OpenAI answers Opus with GPT-5.3-Codex, first model that helped build itself

One hour after Opus 4.6, OpenAI released GPT-5.3-Codex, billed as the first model instrumental in developing itself — the Codex team used early versions to debug its own training and manage its own deployment. It scores 73% on Terminal Bench 2.0, a 10-point gap over Opus 4.6, while running queries 25% faster and more token-efficiently than its predecessor, with improved mid-task steerability.

73% Terminal Bench 2.025% Speed improvement

January 2026

December 2025

OpenAI
Products & Apps

ChatGPT App Store

ChatGPT App Store opens submissions via MCP app model

OpenAI opened app submissions for the ChatGPT App Store, built on the MCP-powered apps model. Developers can now submit apps that run inside ChatGPT, signaling OpenAI's platform play for distribution of agentic apps.

xAI
APIs & Platforms

Grok Voice Agent API

xAI Grok Voice Agent API ships at $0.05/min flat rate, powers Tesla

xAI launched the Grok Voice Agent API with flat-rate pricing of $0.05 per minute and integration into Tesla vehicles. xAI claims the #1 spot on Big Bench Audio at 92.3%, tightening competition in the rapidly commoditizing real-time voice stack.

$0.05/min Grok Voice Agent API

November 2025

xAI
APIs & Platforms

Grok 4.1 Fast + Agent Tools API

Grok 4.1 Fast: 2M context and Agent Tools API at 10x lower cost

Launched as breaking news during the show, Grok 4.1 Fast pairs a 2 million token context window with a new Agent Tools API offering native X search, Reddit search, web browsing, and code execution. Benchmarks are striking: 93-100% on tau2-Bench Telecom and 72% on Berkeley Function Calling v4 (top of the leaderboard) at $0.20/$0.50 per million tokens — roughly 10x cheaper than competitors, and free for the first two weeks on the xAI API and OpenRouter.

93–100% τ²-Bench Telecom72% Berkeley Function Calling v42M Token context window

September 2025

OpenAI
New Models

gpt-realtime

OpenAI ships gpt-realtime and takes the Realtime API to GA

OpenAI shipped the gpt-realtime speech-to-speech model and moved the Realtime API to general availability. The GA release adds remote MCP tool support, image input, and SIP phone calling, making it a full production stack for voice agents and tying into the episode's voice-agents discussion with Kwindla Kramer.

May 2025

Mistral AI
APIs & Platforms

Mistral Agents API

Mistral launches Agents API for building tool-using agents

Mistral released an Agents API, a framework for building custom tool-using agents on top of Mistral models. It joins the wave of big-lab agent frameworks, letting developers wire up tools and orchestrate agentic workflows through Mistral's platform.

Mistral AI
APIs & Platforms

Mistral Embed

Mistral ships new state-of-the-art embedding API

Mistral announced a new state-of-the-art embedding API. The release gives developers a SOTA option for retrieval and semantic search workloads served through Mistral's platform.

Anthropic
APIs & Platforms

Web Search API

Anthropic launches Web Search API for real-time retrieval in Claude

Anthropic released a Web Search API that gives Claude models real-time web retrieval, letting developers ground responses in current information directly through the API. It was covered among the week's big-company API updates.

April 2025

OpenAI
APIs & Platforms

gpt-image-1

OpenAI's GPT Image generation lands in the API as gpt-image-1

OpenAI's powerful image generation capabilities, previously locked inside ChatGPT, are now available to developers via API under the official name gpt-image-1. This was the big one many developers were waiting for, opening up the viral image generation and editing capabilities for building AI art and image editing applications.

Mistral AI
Products & Apps

Classifiers Factory

Mistral releases Classifiers Factory

Mistral announced Classifiers Factory, a service for building and training custom text classifiers on its platform. Covered as a quick item in the Big CO LLMs + APIs section of the show.

Anthropic
Products & Apps

Claude Max plan

Anthropic launches Max plan at $200/mo with higher usage quotas

Anthropic introduced a new Max subscription tier priced at $200 per month, offering significantly more usage quota than the standard Pro plan. It mirrors OpenAI's Pro-tier pricing strategy for power users.

$200/mo Max plan price

March 2025

Arcee AI
Products & Apps

Arcee Conductor

Arcee AI announces Conductor, an intelligent model router

Arcee AI's Lucas Atkins joined the show to announce Conductor, a model router that picks the best model (including Arcee's small specialized models) for each query. It targets cost and quality optimization by routing requests instead of sending everything to one large model.

OpenAI
APIs & Platforms

o1-pro API

OpenAI makes o1-pro available via API at $600 per 1M output tokens

OpenAI exposed its o1-pro reasoning model through the API for the first time, priced at $600 per million output tokens. The show jokingly framed the pricing as 'for oligarchs', but it makes OpenAI's highest-compute reasoning tier programmatically accessible.

Nous Research
APIs & Platforms

Portal

Nous Research opens Portal, an inference API for Hermes models

Nous Research launched Portal, its new inference API service offering access to models like Hermes 3 Llama 70B and DeepHermes 3 8B directly via API. It marks another open-source lab standing up hosted API access to make its models more accessible.

OpenAI
APIs & Platforms

Responses API + Web Search, File Search, Computer Use tools

OpenAI launches Responses API with Web Search, File Search, and Computer Use

OpenAI announced a new agent-focused developer stack at a livestream: the Responses API, a new way to build with OpenAI designed for agentic workloads, plus an Agents SDK. It ships with three built-in tools: Web Search, a File Search tool providing built-in RAG over your files, and a Computer Use tool for agents that operate computer interfaces.

Mistral AI
APIs & Platforms

Mistral OCR

Mistral announces state-of-the-art OCR API

Mistral AI announced Mistral OCR, a document-understanding API the company claims is state of the art at extracting text, tables, and equations from complex documents. It targets RAG and document-processing pipelines with structured markdown output.

February 2025

Google DeepMind
APIs & Platforms

Veo 2 (via FAL API)

Google's Veo 2 video model becomes available via FAL API

Google DeepMind's Veo 2 video generation model became accessible to developers through FAL's inference API. This was the first broadly available API access to Veo 2, letting builders generate high-quality video from text prompts without waiting on Google's own product surfaces.

January 2025

Anthropic
APIs & Platforms

Citations (Claude API)

Anthropic adds Citations to the Claude API

Anthropic launched a Citations capability in the Claude API, letting Claude ground its answers in provided source documents and return precise citations. It targets RAG and document-QA use cases where verifiable sourcing matters.

Perplexity
APIs & Platforms

Sonar Pro Search API

Perplexity ships Sonar Pro search API and an Android AI assistant

Perplexity released its Sonar Pro search-grounded API, giving developers programmatic access to Perplexity-style web-grounded answers, and also launched an AI assistant for Android. Two shipping moves that push Perplexity beyond its consumer answer engine.