Thinking Machines

4 releases covered on ThursdAI · thinkingmachines.ai ↗

July 2026

Thinking Machines
New ModelsOpen weights

Inkling-Small

Thinking Machines releases Inkling-Small: 276B/12B open MoE that beats its 975B sibling on agentic coding

The efficient sibling previewed alongside Inkling ships as open weights: 276B total with 12B active, natively multimodal with an encoder-free architecture (images via hierarchical patch encoding, audio via dMel spectrograms straight into the decoder), and variable thinking effort. On-policy distillation from Inkling plus two extra weeks of agentic-coding RL let it beat the 975B teacher on SWE-Bench Verified (80.2% vs 77.6%) and ARC-AGI-2 (40.1% vs 36.5%), though factual recall regressed hard (SimpleQA 20.6% vs 43.9%). Priced at $0.30/$1.20 per million tokens, roughly 3-4x cheaper than Inkling, with day-zero SGLang, Unsloth GGUF, and Baseten support. Dropped just after the July 30 show aired.

276B / 12B total / active parameters80.2% SWE-Bench Verified, beating the 975B Inkling's 77.6%$0.30 / $1.20 per 1M tokens in/out
Thinking Machines
New ModelsOpen weights

Inkling

Thinking Machines releases Inkling, a 975B open-weights MoE trained on 45T multimodal tokens

Mira Murati's Thinking Machines shipped Inkling, a 975B-total/41B-active Mixture-of-Experts transformer pretrained from scratch on 45 trillion tokens of text, images, audio and video, released under Apache 2.0. The ThursdAI panel called it the top US open-weights model right now — 41 on the Artificial Analysis Index — with encoder-free native reasoning over text, image and audio and a 1M-token context window. A leaner Inkling-Small (276B/12B active) was previewed alongside, and both run on the Tinker platform at a limited-time 50% discount.

975B / 41B total / active parameters45T multimodal training tokens41 Artificial Analysis Index — top US open-weights model

May 2026

Thinking Machines Lab
New Models

Interaction Models

Thinking Machines Lab drops Interaction Models: real-time multimodal 276B MoE

Mira Murati's Thinking Machines Lab released Interaction Models, a 276B-parameter MoE (12B active) trained from scratch for native real-time multimodal collaboration. It supports full-duplex audio/video/text with 0.40s turn-taking latency and scores 77.8 on FD-bench v1.5. The demo can react live to events like another person entering the camera frame.

276B MoE parameters12B active parameters

December 2025

Thinking Machines Lab
Funding

Thinking Machines Lab

Thinking Machines Lab launches with a billion-dollar round

Around June, news broke that Mira Murati's Thinking Machines Lab raised its first billion-dollar round, pulling in what LDJ described as 'an absolute avalanche of top tier researchers' from OpenAI and other labs. It was one of the year's biggest talent and funding stories.