New ModelsOpen weights
Kolibri-1
Aleph Alpha Kolibri-1, a 78B German-English MoE under Apache 2.0
Aleph Alpha released Kolibri-1, a 78B mixture-of-experts model with 3.46B active parameters, trained from scratch on German and English, with 1M context, under Apache 2.0. It fits on a single H200. Wolfram's verdict from vacation: a promising specialized German tool worker.
78B total parameters3.46B active parameters1M context
New Models
Claude Haiku 5.5
Claude Haiku 5.5 at 10 cents per million input tokens
Anthropic brought Haiku back after a year: Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output (under 100K tokens), with cache reads at 1 cent, ten times cheaper than Haiku 4.5. Anthropic reports 72.4% on OSWorld versus 48.9% for GPT-6 Luna, and 39.2% on Terminal-Bench 4.0 versus 16.4%.
$0.10/M input tokens72.4% OSWorld 2.1 (vendor-reported)39.2% Terminal-Bench 4.0 (vendor-reported)
New Models
FLUX 3 Image
FLUX 3 Image: 4K, bounding-box layouts and 10 reference images
Black Forest Labs released FLUX 3 Image with 4K output, bounding-box layouts and up to 10 reference images.
4K output10 reference images
New ModelsOpen weights
Clef
Cloudflare Clef, open decision models
Cloudflare released Clef, open decision models.
New Models
GLiDE
Fastino GLiDE, a thinking decision model
Fastino released GLiDE, a decision model that thinks before it decides.
New ModelsOpen weights
EmbeddingGemma 2
EmbeddingGemma 2, an open multimodal embedding model
Google released EmbeddingGemma 2, an open multimodal embedding model, the same week Perplexity opened pplx-embed-v2.
New Models
Nano Banana 2.1
Google Nano Banana 2.1 at $0.0336 per 1K image
Google released Nano Banana 2.1 at $0.0336 per 1K image. Every infographic on the show was made with it, and Alex found it really good on high reasoning.
$0.0336 per 1K image
New ModelsOpen weights
d1-3B and d1-omni-600M
Liquid AI opens d1: d1-3B and d1-omni-600M decision models
Liquid AI opened its d1 decision models: d1 behind an API, the open d1-3B with text and vision, and d1-omni-600M, which also takes audio. d1-3B answers in 8ms on a GPU and about 50ms on a Jetson Orin Nano, tops Liquid's Decision Index under 10B parameters, and the API is a drop-in replacement for Jev.
8 ms d1-3B on a GPU~50 ms on a Jetson Orin Nano600M omni model with audio
New Models
Mistral Large 4
Mistral Large 4 "Le Chonk": 1T parameters, open weights promised for October
Mistral's Large 4, nicknamed Le Chonk, is a trillion-parameter multimodal model with about 50B active and 1M context; open weights are promised for the end of October. Artificial Analysis scores it 38, the same as GPT-6 Luna, at about $1.13 per task versus 7 cents for Luna.
1T parameters~50B active38 Artificial Analysis Intelligence Index
New ModelsOpen weights
pplx-embed-v2
Perplexity opens pplx-embed-v2 multimodal embedders
Perplexity released pplx-embed-v2, open multimodal embedding models on Hugging Face, with a write-up on multimodal embeddings beyond a single vector.
New Models
Beam
Reflection AI announces Beam, a 501B Western open-weight model
Reflection AI came out of semi-stealth with Beam: 501B parameters with 23B active, trained from scratch in the West, with Apache 2.0 weights promised this month. Reflection claims 80.9% on SWE-bench Verified, admits Kimi K3 is ahead on raw capability, and pitches 3 to 4x less inference compute than GLM 5.2.
501B total parameters23B active parameters80.9% SWE-bench Verified (claimed)
New Models
Rho-1
Reka Rho-1, a 19B research-preview omni model
Reka released Rho-1, a 19B research-preview omni model that collapses the multimodal stack into one model.
19B parameters
New Models
Griffin
Tavus Griffin: 48% of callers in Tavus's study thought it was human
Tavus released Griffin, a conversational video model; in Tavus's own study, 48% of callers thought it was human.
48% of callers thought it was human (Tavus study)