New ModelsOpen weights
Inkling-Small
Thinking Machines releases Inkling-Small: 276B/12B open MoE that beats its 975B sibling on agentic coding
The efficient sibling previewed alongside Inkling ships as open weights: 276B total with 12B active, natively multimodal with an encoder-free architecture (images via hierarchical patch encoding, audio via dMel spectrograms straight into the decoder), and variable thinking effort. On-policy distillation from Inkling plus two extra weeks of agentic-coding RL let it beat the 975B teacher on SWE-Bench Verified (80.2% vs 77.6%) and ARC-AGI-2 (40.1% vs 36.5%), though factual recall regressed hard (SimpleQA 20.6% vs 43.9%). Priced at $0.30/$1.20 per million tokens, roughly 3-4x cheaper than Inkling, with day-zero SGLang, Unsloth GGUF, and Baseten support. Dropped just after the July 30 show aired.
276B / 12B total / active parameters80.2% SWE-Bench Verified, beating the 975B Inkling's 77.6%$0.30 / $1.20 per 1M tokens in/out
Papers & Research
The AI Future Is for Everyone
Zuckerberg's WSJ op-ed: superintelligence must be distributed, not centralized
Mark Zuckerberg laid out Meta's three principles for the superintelligence era — individual empowerment, invention over automation, and balance of power through broad access — arguing the defining question is who gets access to superintelligence, not whether it arrives. Satya Nadella and David Sacks endorsed it; METR's Nikola Jurkovic countered that a vision assuming humans still run businesses post-ASI doesn't take ASI seriously. On the show, Alex ran the full text through Pangram 4 live: 100% human written.
New ModelsOpen weights
Kimi K3 open weights
Moonshot releases Kimi K3's full open-weight checkpoints — 2.8T parameters, the largest open model ever
Two weeks after the API launch, Moonshot published Kimi K3's full checkpoints, model code, and technical report: 2.8T total parameters with 104B active (16 of 896 experts), native vision, a 1M-token context window, and roughly 1.56TB of MXFP4 weights. The report details KDA linear attention, attention residuals, NoPE, and a claimed 2.5x scaling-efficiency jump over K2. On the show, Elie Bakouch called it public building blocks scaled superbly, and Baseten's Philip Kiely described serving it day-zero on eight GB300s. The custom license requires branding above 100M MAU or $20M monthly revenue and a signed agreement for model-as-a-service providers — every provider lists the identical $3/$15 price.
2.8T / 104B total / active parameters1.56TB MXFP4 weights — eight GB300s to serve2.5x claimed scaling efficiency over Kimi K2
Also Released
Open Secure AI Alliance
NVIDIA launches the Open Secure AI Alliance: an open defensive stack, born from the Hugging Face hack
Jensen Huang's second letter of the week proposes an open defensive stack — identity, permissions, isolation, harnesses, logs, and evals — under Linux Foundation stewardship, with launch partners including Microsoft, Hugging Face, CrowdStrike, Mistral, Cloudflare, and Nous Research. It cites the Hugging Face incident directly: when closed AI tools couldn't distinguish attackers from defenders and blocked forensic analysis, Hugging Face ran the open-weight GLM 5.2 on its own infrastructure to contain the intrusion. OpenAI and Anthropic are absent.
37 → 52 partners from launch day to the July 29 page
Also Released
Open Weights and American AI Leadership
Jensen Huang joins X and publishes the Open Weights and American AI Leadership letter
Jensen Huang's first-ever X post published a coalition letter arguing open-weight models are the path to AI diffusion and security, signed at launch by NVIDIA, Microsoft, Meta, Google, and OpenAI and growing from 25 to 230 signatories within a week — CoreWeave among them, announced first on ThursdAI. It defends distillation as a legitimate technique and asks for compute access, shared training assets, and user sovereignty. Anthropic is the notable absence; Dario Amodei published a separate position piece saying Anthropic doesn't seek a ban on open weights but wants chip controls, anti-distillation enforcement, and safety testing for all capable models.
25 → 230 signatories in the first week
New Models
Ling-3.0-Flash
Ant Group's Ling-3.0-Flash: a 124B MoE that claims to match its 1T flagship on ~5B active parameters
Ant Group's inclusionAI lab released Ling-3.0-Flash, a 124B-parameter MoE activating only ~5.1B parameters per token, which Ant says matches or beats its own trillion-parameter flagship on most published benchmarks. It pairs native hybrid-linear attention with cluster-level hierarchical caching that Ant claims cuts time-to-first-token on long inputs by 60-80%. Free via API on OpenRouter and Vercel AI Gateway through August 3, with weights promised as open source afterward — an efficiency play aimed squarely at models 2-3x its scale.
124B / ~5.1B total / active parameters (MoE)60-80% claimed TTFT reduction on long inputs
New ModelsOpen weights
Laguna S 2.1
Poolside open-sources Laguna S 2.1, a 118B agentic coding model with a 1M-token context window
Poolside released Laguna S 2.1, an open-weight 118B-parameter MoE (8B active) built for agentic coding, following the smaller Laguna XS 2.1 (33B/3B) from earlier in July. It runs up to a 1M-token context and scores 70.2 pass@1 on Terminal-Bench 2.1 — matching or beating open models several times its size, including DeepSeek-V4-Flash and Nemotron 3 Ultra, on agentic coding benchmarks. Both Laguna models are free for a limited time on Hugging Face under the permissive OpenMDW license.
118B / 8B total / active parameters (MoE)1M context window (tokens)70.2 Terminal-Bench 2.1 pass@1
New ModelsOpen weights
Kimi K3
Moonshot's Kimi K3 — 2.8T parameters — launches its API live mid-show, with full open weights following July 27
Kimi K3 went from rumor to released API in the middle of the ThursdAI broadcast: a 2.8-trillion-parameter MoE (16 of 896 experts active, ~60-75B per LDJ's estimate) with Kimi Delta Attention and attention residuals for roughly 2.5x the scaling efficiency of K2 (Moonshot's technical report later confirmed ~104B active), native vision, a 1M-token context window, and pricing around half of Opus 4.8 or GPT-5.6 Sol. It debuted #1 on the Frontend Code Arena above Claude Fable 5 and #3 on Artificial Analysis's Intelligence Index; demand forced Moonshot to pause new API subscriptions. The full weights shipped July 27 under a bespoke open-weight (not OSI) Kimi K3 license — the first open 3T-class model.
2.8T total parameters (16 of 896 experts active)1M context window (tokens)#3 / #1 AA Intelligence Index / Frontend Code Arena debut
Major Features & UpdatesOpen weights
Gemma 4 (stealth update)
Google quietly patches Gemma 4 with Flash Attention 4 and tool-calling fixes — no version bump
Google shipped a stealth update to the Gemma 4 family: Flash Attention 4 support on Hopper-class GPUs (a reported 25-70% prefill throughput speedup), tool-calling bug fixes, reduced model 'laziness,' and configurable vision resolution. The ThursdAI panel criticized shipping new weights under the same Gemma 4 name with no version bump, leaving users unsure which checkpoint they're actually running.
25-70% prefill throughput speedup (Flash Attention 4)
New ModelsOpen weights
MOSS-VL-Realtime
OpenMOSS open-sources MOSS-VL-Realtime, an 11B VLM that decides when to speak — and when to stay silent
OpenMOSS released MOSS-VL-Realtime, an open-source 11B-parameter vision-language model (~22.7GB) built for real-time streaming video that proactively speaks up or deliberately stays silent instead of only answering prompts. ThursdAI reported it as state-of-the-art on all three open proactivity benchmarks, with the base model included in the release.
11B parameters22.7GB model download size
New ModelsOpen weights
Inkling
Thinking Machines releases Inkling, a 975B open-weights MoE trained on 45T multimodal tokens
Mira Murati's Thinking Machines shipped Inkling, a 975B-total/41B-active Mixture-of-Experts transformer pretrained from scratch on 45 trillion tokens of text, images, audio and video, released under Apache 2.0. The ThursdAI panel called it the top US open-weights model right now — 41 on the Artificial Analysis Index — with encoder-free native reasoning over text, image and audio and a 1M-token context window. A leaner Inkling-Small (276B/12B active) was previewed alongside, and both run on the Tinker platform at a limited-time 50% discount.
975B / 41B total / active parameters45T multimodal training tokens41 Artificial Analysis Index — top US open-weights model
New ModelsOpen weights
Bonsai 27B
PrismML compresses a full 27B model to 3.9GB so it runs on a phone
PrismML released Bonsai 27B, extreme quantizations of Qwen 3.6 27B under Apache 2.0: a 1-bit build at 3.9GB keeping ~90% of full-precision quality — small enough for an iPhone 17 Pro's memory budget — and a ternary build at 5.9GB keeping ~95%. Both stay multimodal with the full 262K-token context window. Nisten demoed it live on the show running on a phone and on a 6GB GTX 1660 Ti.
3.9GB / 90% 1-bit variant size / quality retained5.9GB / 95% ternary variant size / quality retained262K context window (tokens)
New ModelsOpen weights
Robostral Navigate
Mistral releases Robostral Navigate, its first embodied-navigation model
An 8B robotics model that guides robots through natural-language task instructions using a single RGB camera, claiming state of the art on the R2R-CE benchmark. Mistral's first move into embodied AI, and one of the week's most-discussed releases on Hacker News.
8B ParametersSOTA R2R-CE benchmark
Dev ToolsOpen weights
PyTorch 2.13
PyTorch 2.13 lands FlexAttention on Apple Silicon and big memory wins
3,328 commits from 526 contributors: FlexAttention on Apple Silicon at roughly 12x over SDPA for sparse patterns, a deterministic CUDA backward path, nn.LinearCrossEntropyLoss with up to 4x peak-memory reduction, torchcomms for large-cluster training, and expanded ROCm/Arm/XPU support.
~12x FlexAttention on Apple Silicon vs SDPA3,328 Commits from 526 contributors
New ModelsOpen weights
Transcribe Arabic
Cohere open-sources Transcribe Arabic, topping the Arabic ASR leaderboard
A 2B-parameter Apache 2.0 speech-to-text model that leads the Hugging Face Arabic ASR leaderboard at 25.87 WER — about 11 points better than Whisper Large V3 — with human evaluators preferring it in roughly 96% of head-to-head tests. Handles dialect variety, code-switching and Arabic-English bilingual speech, with day-0 mlx-audio support.
25.87 WER (leaderboard #1)2B Parameters, Apache 2.096% Human preference vs Whisper
Papers & ResearchOpen weights
Antidoom
Liquid AI open-sources Antidoom, removing the reasoning doom-loop
An open method that suppresses the failure mode where reasoning models spiral into repetitive degenerate output: doom-loop rates dropped from 22.9% to 1% on Qwen3.5-4B and from 10.2% to 1.4% on an LFM2.5 checkpoint, with eval scores improving across the board.
22.9%→1% Doom-loop rate, Qwen3.5-4B
New ModelsOpen weights
Agents-A1
Shanghai AI Lab releases Agents-A1, an Apache 2.0 agentic MoE
A 35B MoE built on Qwen3.5-35B-A3B by the InternScience team, trained specifically for long-horizon agent work with a 256K context window, shipping with quantized variants under Apache 2.0.
35B MoE parameters256K Context window
Products & Apps
local.ai
Exo Labs launches local.ai to track the local-AI frontier
Announced live on ThursdAI at AI Engineer: local.ai tracks the best model for your hardware, the performance trade versus the cloud, and whether running local beats API-token pricing. Early access is live with signup codes, and the Exo CLI — 'vLLM for consumer devices, with the configs figured out for you' — ships in the coming weeks.
71% Terminal Bench 2.1, REAP-pruned GLM 5.2550B Nemotron-3 Ultra running on 4 NVIDIA Sparks
New ModelsOpen weights
LongCat-2.0
Meituan reveals LongCat-2.0, a 1.6T MoE trained entirely on Chinese ASICs
Meituan disclosed LongCat-2.0, a 1.6-trillion-parameter MoE trained entirely on Chinese ASICs without NVIDIA hardware. It scores 59.5 on SWE-bench Pro and runs at $0.038 per million tokens with free cache hits. The model had been serving anonymously as 'Owl Alpha' and ranks among OpenRouter's top models by volume — part of a surge that puts Chinese open-weight models at ~30% of global usage, up from 1.2% eleven months ago.
1.6T MoE parameters, no NVIDIA in training59.5 SWE-bench Pro$0.038 per 1M tokens, free cache hits