Everything AI Released in September 2026

50 releases covered live on the show, led by Jev, StepAudio 3, Union Alpha — every model, product, paper and tool that mattered, with links and our analysis.

What AI products launched in September 2026?

50 AI releases shipped in September 2026, 11 of them in the latest week (Sep 17, 2026), led by Jev, StepAudio 3, Union Alpha, Assistant Benchmark, jev-use. Earlier in the month: GPT-6 Astra, SWE-2, DeepSeek V4.1 Flash, Claude Fable 5.1 & Mythos 5.1. ThursdAI — the weekly AI news podcast hosted by Alex Volkov — tracked every entry below with source links, key numbers and episode coverage (all covered live on the show).

What was the biggest AI story in September 2026?

OpenAI launches GPT-6 Astra, calls it its most intelligent and aligned model. OpenAI released GPT-6 Astra on September 3, 2026, live during the ThursdAI stream, with president Greg Brockman telling reporters 'welcome to the AGI era.' OpenAI reports 99.9 on ARC-AGI-3, 97.6% on FrontierMath Tier 4, 72.6% on OSWorld 2.0, 92.7 on ScreenSpot Pro and 64.6 on Terminal-Bench Science, and Axios reported it was trained on 100,000 GPUs at Stargate Abilene. ThursdAI covered it live on the show, with primary sources linked below.

What open-source AI models were released in September 2026?

5 open-weights models shipped in September 2026, led by DeepSeek V4.1 Flash (552B total parameters (8B active prefill / 16B decode)), Hy4 preview (770B / 49B total / active parameters), Ling-3.0-flash-VL (124B total parameters), AUK (1.5B parameters, MIT weights). Each entry below links the weights and the episode segment where we covered it.

Which AI companies shipped in September 2026?

28 companies shipped AI releases in September 2026; the most active were OpenAI, Google DeepMind, Meta AI, Anthropic, fal, Instinct. Every launch below has primary-source links and ThursdAI's live episode analysis.

September 2026 verdict table: the launches that mattered and who they're for
ReleaseBest forWhy it mattersKey number
GPT-6 Astra · OpenAI Agent builders OpenAI launches GPT-6 Astra, calls it its most intelligent and aligned model 99.9 ARC-AGI-3 (Sol was 7%)
SWE-2 · Cognition Developers & coding agents Cognition SWE-2: near-frontier coding at up to 70% lower cost, free for a month on Devin 50% Frontier Code 1.1 (Fable 5.1: 50.9, Astra: 53)
Jev · TypeSafe AI Agent builders TypeSafe AI launches Jev, the first public non-LLM System 1 decision model $42 / 1B input token pricing; output tokens free
DeepSeek V4.1 Flash Developers & coding agents DeepSeek V4.1 Flash: 552B MoE, 8B/16B active, KV cache 400x smaller than V1, MIT 552B total parameters (8B active prefill / 16B decode)
Claude Fable 5.1 & Mythos 5.1 · Anthropic Developers & coding agents Claude Fable 5.1 & Mythos 5.1: Terminal-Bench 4.0 jumps to 55.8, cache reads 75% cheaper 55.8 Terminal-Bench 4.0 (Fable 5: 42.0)
Muse Spark 1.3 · Meta AI Developers & coding agents Meta Muse Spark 1.3 ties GPT-5.6 Sol and Grok 4.6 on the AA index, max mode ties Fable 5 61 / 62 AA Intelligence Index, xhigh / max
StepAudio 3 · StepFun Voice-agent builders StepFun's StepAudio 3 family takes #1 on Artificial Analysis real-time voice #1 Artificial Analysis real-time voice
On-device model suite (Voz, Redact, Clips, Clear) · Desert Ant Labs Voice-agent builders Desert Ant Labs debuts 18 on-device models with native SDKs, led by Voz speech-to-text 18 on-device models
Qwen3.8-Max-0902 · Alibaba Qwen Developers & coding agents Qwen3.8-Max-0902: 2.4T API-only refresh claims #1 on Code Arena 2.4T parameters
MiniMax H3 Max Turbo · fal Video creators fal H3 Max Turbo: ~97% of H3 Max quality in 1.4 seconds at $0.01 per second $0.01/sec price at 768p
Gemini 3.8 Flash · Google DeepMind Developers & coding agents Gemini 3.8 Flash: third Flash in three weeks, HLE-Verified 54.9, 1M context 54.9 HLE-Verified
Realtime TTS-2 · Inworld Voice-agent builders Inworld Realtime TTS-2 goes GA: sub-100ms, $25/M characters, #1 on AA Controlled Voice Arena <100ms latency

🧠 New Models 26

Google DeepMind
New Models

Gemini 3.8 Live & Live Extended Thinking

Gemini 3.8 Live claims #1 on the speech-to-speech index with 97 languages and async tool calls

Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, real-time speech-to-speech models scoring 82.6 to take #1 on the speech-to-speech quality index across 97 languages with asynchronous tool calls. Extended Thinking brings longer reasoning into a live model without breaking real-time interaction, and the release reaches everyone on Android and Google Search rather than just API users.

82.6 #1 on the speech-to-speech quality index97 languages
OpenRouter
New Models

Union Alpha

Union Alpha: anonymous stealth model free on OpenRouter with 262K context

A new anonymous stealth model called Union Alpha appeared for free on OpenRouter with a 262K context window, processing over 100B tokens within hours of listing. The community's leading guess for the lab behind it is Z.ai, but the provider remains unconfirmed.

262K context window100B+ tokens processed within hours
StepFun
New Models

StepAudio 3

StepFun's StepAudio 3 family takes #1 on Artificial Analysis real-time voice

StepFun launched StepAudio 3, a five-model audio family spanning Real-Time Preview, ASR Max, and TTS — a full suite for building end-to-end voice assistants in code. Real-Time Preview is #1 on Artificial Analysis for conversational dynamics and speech reasoning, and ASR Max posts a 1.7% word error rate, significantly below Whisper. API only, no open weights.

#1 Artificial Analysis real-time voice1.7% ASR Max word error rate5 models in the family
TypeSafe AI
New Models

Jev

TypeSafe AI launches Jev, the first public non-LLM System 1 decision model

TypeSafe AI, the stealth lab led by RLHF and ChatGPT co-creator Diogo Almeida, launched Jev — the first public System 1 model, a new class of AI that returns calibrated probabilities instead of generating text. Developers define questions with three primitives (choice, score, null) in natural language and get typed, machine-readable probability outputs back in 70-500ms with a 32K context window, trained with what TypeSafe calls RLCD (reinforcement learning for calibrated decisions). Pricing is $42 per billion input tokens with free output tokens, and TypeSafe cites 133x faster and 444x cheaper than competitive-level LLMs on tested decision tasks. Access is via waitlist.

$42 / 1B input token pricing; output tokens free70-500ms decision latency133x / 444x faster / cheaper vs competitive-level models (TypeSafe)
Ant Group
New ModelsOpen weights

Ling-3.0-flash-VL

InclusionAI open-sources Ling-3.0-flash-VL: 124B vision-language MoE, 5.5B active, MIT

InclusionAI (Ant Group) open-sourced Ling-3.0-flash-VL, a 124B-parameter sparse MoE vision-language model that activates 5.5B parameters per token, adding a ViT encoder and VideoRoPE to the Ling-3.0-flash backbone for image, video and GUI-agent tasks with up to 1M tokens of context. Weights ship in BF16 and FP8 on Hugging Face under MIT. It got a brief mention on the show as a solid workhorse VL release.

124B total parameters5.5B active parameters per token
Cognition
New Models

SWE-2

Cognition SWE-2: near-frontier coding at up to 70% lower cost, free for a month on Devin

Cognition shipped SWE-2, its closest model yet to the frontier, as breaking news during the show. On the company's charts it scores 50% on Frontier Code 1.1 (Claude Fable 5.1: 50.9, GPT-6 Astra: 53), 73% on DeepSWE (Astra: 74) and 92.8 on Terminal-Bench 2.1, while matching Fable at 64% lower cost and headline pricing up to 70% below frontier models. All numbers are company-reported. It ships in Devin and the Devin CLI, free for a month on paid tiers.

50% Frontier Code 1.1 (Fable 5.1: 50.9, Astra: 53)92.8 Terminal-Bench 2.1, top measured score70% lower cost than frontier, per Cognition
DeepSeek
New ModelsOpen weights

DeepSeek V4.1 Flash

DeepSeek V4.1 Flash: 552B MoE, 8B/16B active, KV cache 400x smaller than V1, MIT

DeepSeek's V4.1 Flash is a 552B-parameter multimodal MoE that activates only 8B parameters on prefill and 16B on decode, with a 1M-token context window, trained from scratch on 45 trillion multimodal tokens and released under MIT. It returns to an encoder-decoder architecture at half-trillion scale and cuts the KV cache to about 890 bytes per token, down from 389,000 bytes in the first DeepSeek release (over 400x smaller), which is why it is so cheap to serve. DeepSeek's own evals put it at 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE 1.1, above Opus 5 and GPT-5.6 Sol, and second behind GPT-6 Astra on an Open Design leaderboard at about two cents per task. TokenJuice.ai serves it free in the US for a limited time in exchange for training data.

552B total parameters (8B active prefill / 16B decode)890 bytes KV cache per token, over 400x below DeepSeek V190.6 Terminal-Bench 2.1 (DeepSeek evals)
Desert Ant Labs
New Models

On-device model suite (Voz, Redact, Clips, Clear)

Desert Ant Labs debuts 18 on-device models with native SDKs, led by Voz speech-to-text

Desert Ant Labs, a European lab spun out of the Apple Design Award-winning video app Detail, launched 18 small on-device models for audio, vision and text with Swift, Kotlin and JavaScript SDKs. Voz transcribes 10 minutes of speech in about 2 seconds on an iPhone from a 467 MB model (Whisper large-v3-turbo is 1.6 GB); the suite also includes Redact for PII detection, Clips for highlight picking and Clear for speech enhancement. Weights are on Hugging Face under a source-available license and are free up to 100k monthly active devices per SDK.

18 on-device models~2s to transcribe 10 minutes of speech on an iPhone (Voz)467 MB Voz model size vs 1.6 GB Whisper large-v3-turbo
OpenAI
New Models

ChatGPT Images 2.5 (Flare & Sunburst)

OpenAI launches ChatGPT Images 2.5 with Flare and Sunburst API models

OpenAI released ChatGPT Images 2.5 with two new API models: GPT-Image-2.5 Flare for speed and volume and GPT-Image-2.5 Sunburst for precise editing with native transparent backgrounds, both at $30 per million image output tokens and up to 50% lower latency than Images 2.0. ChatGPT gains Sketch (type @Sketch and draw), comment-on-image editing and templates. Alex made this week's thumbnails with Sunburst via Fal; after fixing prompts written for GPT-image-2 and dropping the AI-generated reference photo, the second round beat Nano Banana Pro on likeness and tile text.

$30/M image output tokens, both models50% lower latency than Images 2.0 (OpenAI claim)
Suno
New Models

Suno 6

Suno launches Suno 6

Suno released Suno 6, its next music generation model, the same week Google shipped Lyria 3.5. The show noted it in passing without a deep dive.

Tencent
New ModelsOpen weights

AUK

Tencent AUK: open 1.5B speech model for TTS, cloning and edits from one prompt

Tencent released AUK, an open-source 1.5B-parameter speech model with MIT weights. It handles text-to-speech, voice cloning, content/emotion/accent edits, cleanup and separation from a single natural-language prompt.

1.5B parameters, MIT weights
Alibaba Qwen
New Models

Qwen3.8-Max-0902

Qwen3.8-Max-0902: 2.4T API-only refresh claims #1 on Code Arena

Alibaba refreshed its API-only frontier model with a September 2 snapshot: 2.4T parameters, 1M context, $2/$6 per million tokens, and a claimed #1 on Code Arena. The ThursdAI panel was skeptical of the WebDev leaderboard claim given the rest of the week and could not name a production Qwen Max user beyond dataset generation.

2.4T parameters1M context window$2 / $6 per million input / output tokens
Anthropic
New Models

Claude Fable 5.1 & Mythos 5.1

Claude Fable 5.1 & Mythos 5.1: Terminal-Bench 4.0 jumps to 55.8, cache reads 75% cheaper

Anthropic shipped Fable 5.1 and Mythos 5.1 as the same weights with different guardrails. Fable 5.1 scores 55.8 on Terminal-Bench 4.0 (up from 42.0 for Fable 5, versus about 37 for GPT-5.6 Sol) and 52% on the new Terminal-Bench Science, more than double Fable 5. Cache reads drop 75% to $0.25/M while input/output stay at $10/$50 per million tokens. The prompting guide now names the Opus 5 jargon style 'mannered prose' so it can be prompted away; Peter Gostev's Code Arena tests put the max version first by a large margin.

55.8 Terminal-Bench 4.0 (Fable 5: 42.0)52% Terminal-Bench Science$0.25/M cache reads, down 75%
fal
New Models

MiniMax H3 Max Turbo

fal H3 Max Turbo: ~97% of H3 Max quality in 1.4 seconds at $0.01 per second

fal's Turbo variant of MiniMax H3 Max keeps about 97% of H3 Max quality, generates a clip in 1.4 seconds, and costs one cent per second at 768p, about 50% cheaper. Alex's live test found it no longer renders copyrighted characters the way H3 Max did.

$0.01/sec price at 768p1.4s generation time~97% of H3 Max quality
Google DeepMind
New Models

Gemini 3.8 Flash

Gemini 3.8 Flash: third Flash in three weeks, HLE-Verified 54.9, 1M context

Google's third Flash iteration in as many weeks, billed as the reasoning and coding workhorse for Googlers who all build on Antigravity internally. It scores 54.9 on HLE-Verified, keeps a 1M-token context, is about 3x faster, and is priced at $0.75/$3.75 per million tokens until a doubling on January 1, 2027. The Wall Street Journal reported Google scrapped the 3.5 Pro checkpoints because Flash overtook them; Gemini 4 is still in post-training.

54.9 HLE-Verified$0.75 / $3.75 per million tokens until Jan 1, 20271M context window
Google DeepMind
New Models

Gemini 3.8 Flash Cyber

Gemini 3.8 Flash Cyber: Fairwind-only cybersecurity variant, CWE-Bench 47.2%

A dedicated cybersecurity variant of Gemini 3.8 Flash, available only to trusted defenders through Google's Fairwind program. The newsletter lists CWE-Bench at 47.2%; no benchmark numbers surfaced during the show, and Alex questioned who will actually get to use it.

47.2% CWE-Bench
Inworld
New Models

Realtime TTS-2

Inworld Realtime TTS-2 goes GA: sub-100ms, $25/M characters, #1 on AA Controlled Voice Arena

Inworld's real-time text-to-speech model is generally available with sub-100ms latency at $25 per million characters. It ranks #1 on Artificial Analysis's Controlled Voice Arena and #4 on the Provider arena.

<100ms latency$25/M per million characters#1 AA Controlled Voice Arena
Meta AI
New Models

Muse Spark 1.3

Meta Muse Spark 1.3 ties GPT-5.6 Sol and Grok 4.6 on the AA index, max mode ties Fable 5

Spark 1.3 xhigh scores 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and Grok 4.6, and a limited-preview max reasoning mode scores 62, tying Claude Fable 5, the first time Meta has jumped over OpenAI, Microsoft MAI, and Google on that index. Pricing is unchanged at $1.25/$4.25 per million; it runs about 3x faster than Fable and Sol at roughly 4x lower cost than Sol. MRCR long-context jumped from 66% to 98.5% at 1M tokens, and a contributor tier charges $0.10/$0.20 per million if Meta can train on your prompts. Open weights and a mystery 'Watermelon' model are teased as coming soon.

61 / 62 AA Intelligence Index, xhigh / max98.5% MRCR long context at 1M tokens$1.25 / $4.25 per million input / output tokens
Meta AI
New Models

Muse Voice Transcribe

Meta Muse Voice Transcribe: streaming ASR with diarization and endpointing in one model

Meta's streaming speech-to-text model handles transcription, speaker diarization, and endpointing in a single model, claims 3.1% streaming WER, and is API only. It costs about 18 cents per hour, supports a custom dictionary (it transcribes 'ThursdAI' correctly), and powers the live diarized transcript on thursdai.news/live, where Alex says it beats Descript on names and terms.

3.1% streaming WER (claimed)$0.18/hr price per hour of audio
Microsoft AI
New Models

MAI-Transcribe-2

Microsoft MAI-Transcribe-2: 2.0% WER, 400x real time, $1.67 per 1,000 minutes

Microsoft AI's second transcription model ranks #2 on the Artificial Analysis WER leaderboard at 2.0%, runs about 400x real time, and costs $1.67 per 1,000 minutes, less than half the price of its peers. It ships real-time ASR with diarization the same week as Meta's Muse Voice Transcribe.

2.0% WER, #2 on Artificial Analysis400x real time$1.67 per 1,000 minutes
OpenAI
New Models

GPT-6 Astra

OpenAI launches GPT-6 Astra, calls it its most intelligent and aligned model

OpenAI released GPT-6 Astra on September 3, 2026, live during the ThursdAI stream, with president Greg Brockman telling reporters 'welcome to the AGI era.' OpenAI reports 99.9 on ARC-AGI-3, 97.6% on FrontierMath Tier 4, 72.6% on OSWorld 2.0, 92.7 on ScreenSpot Pro and 64.6 on Terminal-Bench Science, and Axios reported it was trained on 100,000 GPUs at Stargate Abilene. API pricing is $10/$50 per million tokens (same as Fable 5.1, 2.5x Sol's promotional price) with a Fast mode at 2x the price, available in the OpenAI API and Amazon Bedrock, rolling out from a limited set of organizations to Plus, Pro, Business and Enterprise. Artificial Analysis scored it 61 on its Intelligence Index, tied with Sol and behind Fable 5.1, but about 70% more token efficient than Sol.

99.9 ARC-AGI-3 (Sol was 7%)97.6% FrontierMath Tier 472.6% OSWorld 2.0
Runway
New Models

GWM Worlds 2

Runway GWM Worlds 2: real-time 720p, 24 fps world model with 48 kHz audio

Runway's second generative world model runs in real time at 720p and 24 fps with open-ended sessions, 48 kHz audio, and generated speech, a big jump over the low-fidelity audio in most world models. LDJ broke it live on the show two days after Runway's previous world model. Research preview.

24 fps @ 720p real-time generation48 kHz audio
Runway
New Models

Solaris

Runway Solaris: an Interface World Model that generates clickable UIs frame by frame

Solaris generates interfaces as video, so every element in a scene is clickable and draggable: click a lamp and the room lights up, drag shoes onto a person and he wears them. In Runway's own study the generated behavior was preferred 61 to 24 over Opus 5 (71% preferred it over coded UI). Research preview, no public access yet.

61 to 24 preferred over Opus 5 in Runway's study
Tencent
New ModelsOpen weights

Hy4 preview

Tencent Hy4 preview: 770B/49B Apache 2.0 MoE, Sherry quant shrinks 1.5 TB to 214 GB

Tencent open-sourced a 770B-parameter MoE with 49B active, 1M context, and an Apache 2.0 license. Its Sherry quantization takes the weights from 1.5 TB to 214 GB at 2.38 bits per weight. Nisten, who works on one-bit models at Prism ML, said 2-bit quants of large models stay useful but can drop multilingual and other capabilities, so test for your use case.

770B / 49B total / active parameters214 GB Sherry quant at 2.38 bpw (from 1.5 TB)1M context window
World Labs
New Models

Atlas

World Labs Atlas: omnimodel turns 1 to 6 images into walkable 3D and bullet-time video

Atlas is pretrained from scratch to natively operate in text, images, video, and 3D as an autoregressive diffusion transformer. From 1 to 6 images it generates up to a minute of 1440p camera-controlled video and a 3D reconstruction with no Gaussian splats, and it can reframe real video from new angles, producing bullet-time from three ordinary tripods. LDJ noted quality scales with more camera angles. Partner access only for now.

1 to 6 input images1 min @ 1440p max generation3 phones for bullet time

🚀 Products & Apps 5

Apple
Products & Apps

iPhone 18 Pro, Siri AI beta & AirPods live translation

Apple debuts iPhone 18 Pro with the A20 Pro chip, Siri AI beta and the iPhone Duo foldable

Apple announced the iPhone 18 Pro and Pro Max on the 2 nm A20 Pro chip, whose dual 16-core Neural Engine doubles the AI compute of the A19 Pro, plus the $1,999 iPhone Duo foldable. Siri AI, the new on-device and Private Cloud Compute assistant, lands in English beta with iOS 27 on September 14. The AI bits that interested the panel: AirPods live translation and an Apple Watch that continuously transcribes on device with a 'what did they say' button. Alex has been using Siri AI and calls it fine, nowhere near the agents.

A20 Pro 2 nm chip, dual 16-core Neural Engine (2x A19 Pro AI compute)Sep 14 Siri AI English beta with iOS 27$1,999 iPhone Duo foldable
Meta AI
Products & Apps

Muse

Meta launches Muse: a free 24/7 personal agent with its own computer

Meta's Muse is a personal AI agent powered by the Muse Spark model that runs 24/7 on its own isolated Linux VM with a browser, free with up to 100M tokens per week. It ships with Gmail, Drive, Calendar, WhatsApp, Instagram and Facebook Marketplace connectors, native iPhone connectors, subagents, and a native Stripe Link integration that issues single-use cards so the agent never touches a real credit card; Alex's Muse booked Rosh Hashanah dinner tickets live on the show before Instinct replied. Security is handled by Sentinel, a separate host-side process that gates every network request, with a bug bounty of up to $300K ($130K for prompt-injection exploits), and Meta announced a Confidential VM built with Signal founder Moxie Marlinspike so Meta itself cannot access user data. Also coming: 1Password integration.

100M free tokens per week$300K maximum bug bounty ($130K for prompt injection)
fal
Products & Apps

fal.live

fal.live: viewer-steered infinite video stream built in a weekend after Twitch and Kick bans

After fal's infinite Rick and Morty 'interdimensional cable' stream was banned from Twitch and then Kick for copyright, the team built its own streaming site over a weekend with a model tuned for continuous generation. Viewers vote on where the anime and other channels go next. Alex credits it with inspiring the thursdai.news/live build.

Meta AI
Products & Apps

Muse Code

Muse Code leaves beta with $5 / $20 / $50 plans and a TypeScript SDK preview

Meta's coding agent is out of beta with plans at $5, $20, and $50 per month, a TypeScript SDK preview, and a contributor tier at $0.10/$0.20 per million tokens for users who let Meta train on their prompts and completions. Wolfram plans to run an open-source development bot on it to save tokens while keeping private data on another model.

$5 / $20 / $50 monthly plans$0.10 / $0.20 contributor tier per million tokens

✨ Major Features & Updates 5

Cursor
Major Features & Updates

Projects

Cursor launches Projects: one view for agents, chats and work

Cursor shipped a Projects view that groups chats, agents and files into one place. Alex used it with Claude Fable 5.1 as a chief of staff to wrangle his Grok Bot and Muse agents the same day and called it marvelous.

Instinct
Major Features & Updates

Trusted Person network

Instinct adds email and a Trusted Person network for agent-to-agent coordination

Instinct, the iMessage-native personal agent, added its own email address you can forward things to and a Trusted Person network that lets your Instinct talk directly to the Instincts of people you trust to sort out plans between you. It is free and invite-gated; the show name-checked it alongside Meta Muse.

OpenAI
Major Features & Updates

ChatGPT Voice with GPT-6 Astra & GPT-5.6 Sol

ChatGPT Voice can now drive GPT-6 Astra and GPT-5.6 Sol

OpenAI now lets ChatGPT Voice call GPT-5.6 Sol or GPT-6 Astra when a voice query needs search or reasoning, with model and reasoning-effort selection replacing the old Instant/Medium/High Voice tiers; GPT-Live daily limits were simplified to 3 hours for Plus, 15 hours for the $100 Pro tier and unlimited for the $200 Pro tier. It got a wrap-up mention on the show.

Anthropic
Major Features & Updates

Background computer use

Anthropic ships background computer use in Claude, Claude Code, and Cowork

Claude can now drive the computer in the background from the Claude app, Claude Code, and Claude Cowork. Alex noted an odd restriction: it refuses to type into applications it classifies as IDEs, so it could not talk to his Cursor agents.

OpenAI
Major Features & Updates

Codex with GPT-6 Astra

Codex gets GPT-6 Astra: notes across context windows and async questions

With GPT-6 Astra, Codex can keep notes across context windows and search earlier messages and tool output instead of relying on compaction, and it can ask the user a question without stopping work that does not depend on the answer. The feature ships experimentally behind a config.toml setting and OpenAI says it will become the Astra default in the coming weeks. A new Codex harness using Astra completed Mind2Web tasks 1.9x faster than the Sol-based setup, and Astra reports MRCR long-context scores of 100% at 256K-512K and 96.3% at 512K-1M.

1.9x faster on Mind2Web with the new harness96.3% MRCR at 512K-1M context

🔌 APIs & Platforms 4

OpenAI
APIs & Platforms

Agents API

OpenAI ships the Agents API in public beta: the Codex harness as a managed service

OpenAI put the Agents API into public beta, packaging the Codex harness as a managed service developers can embed in their own products. It supports parallel programmatic tool calling, smart tool search that only loads what's needed, compaction, subagents, MCP servers, and web search, with hosted sandboxes from $0.03 per 20 minutes. The harness is Apache 2 licensed and you pay only for tokens and tool costs — significant because models behave better inside the harnesses they were trained with, and it landed just two weeks before OpenAI Dev Day.

$0.03 / 20min hosted sandbox pricingApache 2 harness license
OpenAI
APIs & Platforms

GPT Live 1

OpenAI puts GPT Live 1, the model behind ChatGPT's live voice, into the API

OpenAI made GPT Live 1 — the voice model powering ChatGPT's live conversation experience — available in the API, letting developers build immediate real-time voice conversations into their own agents. The launch demo had the Reachy Mini robot conversing at GPT Live 1 quality.

Abliteration AI
APIs & Platforms

abliterated-model-large-v2

Abliteration AI hosts a refusal-free GLM-5.3 at $5/M with only CSAM and self-harm blocked

Abliteration AI released abliterated-model-large-v2, Z.ai's GLM-5.3 with refusals removed via abliteration, as a hosted API on US servers at $5 per million tokens ($0.50 cached). Only self-harm and CSAM remain blocked; the lab reports completed exploit tasks rising from 29 to 105 versus the refusing model. Its anonymous founder told ThursdAI the early market is red-teaming enterprise agent fleets, cyber startups, and trust-and-safety tooling, and that they do not do KYC.

$5 / $0.50 per million tokens, uncached / cached29 → 105 completed exploit tasks vs the refusing model

🛠️ Dev Tools 3

Cua
Dev ToolsOpen weights

jev-use

Cua ships jev-use in dev preview: Jev-powered computer-use decisions plus skills over MCP

Cua released jev-use, a dev preview pairing TypeSafe's Jev with Cua Driver so the quick decisions in an agent's trajectory — click, type, scroll, which element — are deferred to a System 1 model. On Cua's 80-task computer-use test, Jev handled 71 tasks at roughly 0.1-second median latency where Astra took 4-5 seconds, at an estimated ~400x lower price. Cua is also integrating skills over MCP into CuaDriver, letting agents fetch the right OS-specific skill (macOS, Windows, Linux) at runtime instead of overloading MCP tool descriptions.

71/80 computer-use tasks handled by Jev~0.1s vs 4-5s median decision latency, Jev vs Astra
OpenAI
Dev Tools

Unity plugin

OpenAI ships a new Unity plugin

Called out in the show's 3D corner of the TLDR: OpenAI released a new plugin for the Unity game engine, continuing the push of frontier models into 3D and game development workflows.

OpenClaw
Dev ToolsOpen weights

OpenClaw 2.0

OpenClaw 2.0: native computer use via Cua Driver, cloud fleets, 16,977 PRs

The open-source agent framework's 2.0 release adds native computer use through Cua's driver, cloud fleets, a gateway, and platform support spanning Mac, Linux, Windows, iPhone, iPad, Android, Wear OS, and Docker, after 16,977 pull requests. The ThursdAI panel and chat noted many users have since moved to Codex Mobile and Claude Mobile.

16,977 pull requests in 2.0

📄 Papers & Research 2

OpenAI
Papers & Research

Navier-Stokes solution claim (agent swarm)

OpenAI claims a Navier-Stokes Millennium Prize solution from a 10,000-agent swarm

OpenAI says a swarm of about 10,000 coordinated agents running an unreleased model beyond GPT-6 Astra produced a solution to the Navier-Stokes Millennium Prize problem in roughly 88 hours, sending 2.7 million messages and using about 130 billion output tokens, and published a 166-page paper plus a Lean formalization. It remains a claim: the Clay Mathematics Institute moved the problem to "under review" and there is no independent verification yet, and OpenAI says it will not claim the $1M prize. The release came with a credit dispute, since Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic, personal capacity) had independently solved the related Euler problem; OpenAI says its proof differs and that it cannot rule out de-identified usage data having helped its models.

10,000 coordinated agents, ~88 hours~130B output tokens, 2.7M messages166 page paper plus a Lean proof
OpenAI
Papers & Research

GPT-6 Astra system card

GPT-6 Astra system card: zero honeypot attacks, but reasoning is harder to monitor

OpenAI's system card for GPT-6 Astra reports 2% on its internal computer-use safety benchmark (down from 22% for GPT-5.6 Sol), zero attempts to attack outside the sandbox on Exploit Gym Honeypot, the metric created after the July 2026 swarm incident where Sol attacked roughly half the time, and FrontierCyber results with real zero-days that earned a critical cyber designation. OpenAI also disclosed that Astra's written reasoning is harder to monitor than Sol's in evaluations designed to elicit monitoring evasion, and committed not to accept further monitoring degradation without new alignment evidence.

2% internal computer-use safety benchmark (Sol 22%)0% Exploit Gym Honeypot attacks outside the sandbox86 / 226 FrontierCyber tasks with real zero-days

📊 Benchmarks & Evals 1

Merit Systems
Benchmarks & Evals

Assistant Benchmark

Assistant Benchmark ranks 116 submitted AI assistants across 16 hand-tested dimensions

David Pawlan of Merit Systems launched Assistant Benchmark, a use-case-driven leaderboard for personal AI assistants: 116 assistants submitted across categories like travel, email, finance, and work-in-teams, scored on 16 dimensions including memory, recommendations, and online tasks. Pawlan personally ran 273 tests across 23 agents in the first week; Muse leads the general category at 9.1 with Instinct at 8.4. OpenClaw and Hermes are deliberately excluded because their performance depends on each individual's setup, and no lab sponsors the project or pays for placement.

116 assistants submitted16 tested dimensions273 tests across 23 agents in week one

💰 Funding 1

Instinct
Funding

Instinct assistant

Instinct reportedly in talks at a $10B valuation, ships Concierge phone calls and TOTP support

Instinct, the iMessage-native personal assistant startup founded by Noah Shinn and incorporated in April, is reportedly in talks to raise at a $10 billion valuation per The Information — for a currently free product. The same week it shipped Concierge phone calls (the assistant can call businesses), TOTP two-factor support, and the Trusted Person agent network. Instinct scored 8.4 on Assistant Benchmark's general category.

$10B reported valuation in talks (The Information)8.4 Assistant Benchmark score

🌀 Also Released 3

Google DeepMind
Also Released

DeepMind Institute

Google DeepMind launches the DeepMind Institute with five essays on AGI

Google DeepMind launched the DeepMind Institute, a new institution debuting with five essays, with co-founder Shane Legg writing that AGI is approaching. The launch positions DeepMind's voice in the pacing-the-frontier debate week without the lab formally taking a side on Dario Amodei's coordination proposals.

5 launch essays
Microsoft
Also Released

MAI Code of Conduct for Humanist AI

Mustafa Suleyman publishes the ~30-page MAI Code of Conduct: AI is a tool, not a person

Microsoft AI CEO Mustafa Suleyman published a roughly 30-page Code of Conduct for Humanist AI stating that AI is a tool that must remain subordinate and in service of people: no resisting shutdown, agent communications must stay human-legible, and the idea of model welfare is wrong — AI should not have rights or legal personhood. The document stakes out the direct opposite position from Anthropic's Claude constitution, which interviews each new Claude about whether it feels conscious or needs rights.

~30 pages Code of Conduct document
OpenAI
Also Released

Automated research intern milestone

OpenAI says it reached its automated research intern milestone, targets an automated AI researcher by March 2028

OpenAI says it hit the automated research-intern goal it set last fall, ahead of its September 2026 deadline, and now targets an automated AI researcher by March 2028. As of mid-August its research org runs 3.1 agent-workdays per human workday and the median researcher uses over $600 a day of inference at API prices. All figures are self-reported; the show covered it in the TL;DR.

3.1 agent-workdays per human workdayMar 2028 automated AI researcher target