New ModelsOpen weights
Qwen3.8-Flash-Next
Qwen3.8-Flash-Next previews the Qwen4 architecture in open weights
Alibaba released open weights for Qwen3.8-Flash-Next: 125B parameters plus a 51B N-gram embedding table with only 6B active, trained at roughly 1/9 the cost of Qwen3.7-Plus while self-reporting DeepSWE 58.7 and SWE-bench Pro 62.5. It previews the Qwen4 architecture: Qwen Sparse Attention (QSA) replaces full attention layers, and the low-bandwidth N-gram embedding table can be offloaded to slow RAM or even an SSD.
125B+51B parameters + N-gram embeddings, 6B active1/9 training cost vs Qwen3.7-Plus58.7 DeepSWE (self-reported); SWE-bench Pro 62.5
New ModelsOpen weights
Apodex 1.1 + FrontierAgent
Apodex 1.1 agentic model family with open-weight 35B mini and FrontierAgent harness
Apodex released the Apodex 1.1 agentic model family, including an open-weight 35B mini model and the Apache 2.0 FrontierAgent harness. Benchmarks are self-reported, and the company was new to the ThursdAI crew.
35B open-weight mini model
New ModelsOpen weights
Breeze TTS 2
Breeze TTS 2 takes #1 open-weight TTS on Artificial Analysis
Breeze TTS 2 released open weights and immediately took the #1 open-weight spot on Artificial Analysis' Provider Voices arena at 1,215 Elo, surpassing Fish Audio by about 90 points. It supports emotion, pausing, and natural disfluencies. Weights ship under a non-commercial license.
1,215 Elo — #1 open-weight TTS on AA Provider Voices
New ModelsOpen weights
PhoneLLM Alpha 1
PhoneLLM Alpha 1: open-weights voice-agent model at a quarter cent per minute
Daily/Pipecat released PhoneLLM Alpha 1, an open-weights post-train of NVIDIA's Nemotron 3 Nano (30B, ~3B active) built for production voice agents where thinking-token latency ruins conversations. Post-training took PhoneBench v1 from a base 28% to 72% — beating GPT-5.6 Tera at a third the latency and one-eighteenth the price — and it runs 80+ concurrent agents on a single B200 (NVFP4, with Modal), landing at about a quarter of a cent per minute. Weights are on Hugging Face.
28%→72% PhoneBench v1, from post-training alone80+ concurrent agents per B200~$0.0025 per minute runtime cost
New Models
MiniMax H3 Max
fal's MiniMax H3 Max generates 5-second video in under 3 seconds
fal Research debuted MiniMax H3 Max, a post-train of the open-weight MiniMax H3 that ranks #1 on image-to-video and #3 on text-to-video on Artificial Analysis while generating five-second clips in under three seconds — 2.53s in the live on-air test. Priced at $0.04/second at 768p (promo until Sept 1), with a weights release planned. The speed puts it alone on the speed-versus-quality Pareto frontier.
2.53s to generate a 5-second clip, live on the show#1 image-to-video on Artificial Analysis (#3 T2V)$0.04/s at 768p until Sept 1
New Models
Gemini 3.5 Transcribe
Gemini 3.5 Transcribe launches in live and batch modes, replacing Chirp 3
Google launched Gemini 3.5 Transcribe in public preview with both live (sub-second streaming via the Live API, with WebSocket multi-turn support) and batch modes, reporting 2.6%/4.0% WER per Artificial Analysis. It replaces Chirp 3, shipped with launch-day Pipecat support, and was transcribing the show itself in real time via Alex's GrokBot producer.
2.6%/4.0% WER live/batch per Artificial Analysis
New Models
Gemini Omni 1.1 Flash
Gemini Omni 1.1 Flash tops Arena text-to-video with voice-consistent scene extension
Google's new video model dropped during the show: #1 on Arena's text-to-video leaderboard and #2 on image-to-video. It analyzes up to 10 seconds of previous footage to extend scenes while keeping character identity, voice, and lighting locked; adds first/last-frame control and infinite loops; and offers 360p draft generations with built-in upscaling. Rolling out in Google AI Studio, Flow, and Gemini Enterprise with API access.
#1 Arena text-to-video (#2 image-to-video)10s of prior footage analyzed for scene extension
New ModelsOpen weights
Granite Speech 5.0 Turbo CTC
IBM Granite Speech 5.0 Turbo CTC: 470M encoder-only ASR at 12,600+ RTFx
IBM released Granite Speech 5.0 Turbo CTC, a 470M-parameter encoder-only English ASR model under Apache 2.0. It reports 4.85% WER with throughput above 12,600 RTFx on an H200 — built for the fast, cheap end of the transcription spectrum.
4.85% WER, 470M encoder-only English ASR12,600+ RTFx on H200
New Models
Navigator n2
Yutori Navigator n2: 27B computer-use model at a fraction of frontier cost
Yutori (the Scouts team) announced Navigator n2, a 27B computer-use model scoring a self-reported 65.2% on OSWorld 2.0 — close to Fable at desktop control while significantly cheaper ($0.50/M input, $4/M output, API only). It can switch between Chrome tooling and full computer use depending on which is cheaper for the task. Not open source.
65.2% OSWorld 2.0 (self-reported)27B parameters$0.50/$4 per M tokens in/out
New ModelsOpen weights
GLM-5.3-Flash
GLM-5.3-Flash: the OX Alpha mystery model, open-sourced under MIT
Z.ai open-sourced GLM-5.3-Flash, a 320B-parameter MoE with 18B active under MIT license, after stealth-testing it for about six days as 'OX Alpha' with effectively unlimited free traffic on OpenRouter — all served on Chinese chips. Company-reported DeepSWE is 63.4 with Claude Opus 4.8-level coding claims, it's natively multimodal, and a hybrid sparse/linear attention architecture cuts KV cache size 4x versus GLM 5.3 with 3x serving performance.
320B-A18B parameters (total / active), MIT license63.4 DeepSWE (company-reported)4x smaller KV cache vs GLM 5.3