This Week in AI

30 releases from the week of Oct 8, 2026, covered live on the show. Updated every Thursday.

What happened in AI this week?

30 AI launches shipped in the week of Oct 8, 2026, led by Arena–Series funding ($200M), Kolibri-1, Claude Haiku 5.5, d1-3B and d1-omni-600M, Mistral Large 4. ThursdAI — the weekly AI news podcast hosted by Alex Volkov — covered them live on the show, with primary sources and key numbers for every entry below.

What was the biggest AI story this week?

Arena raises $200M at a $3.1B valuation, broken live on ThursdAI. Arena raised $200 million at a $3.1 billion valuation; Peter Gostev broke the news live on ThursdAI. Full context is in this week's episode segment linked below.

What open-source AI models were released this week?

5 open-weights models shipped this week: Kolibri-1 (78B total parameters), d1-3B and d1-omni-600M (8 ms d1-3B on a GPU), Clef, EmbeddingGemma 2, pplx-embed-v2. Each card below links the weights and the episode segment where we covered it.

Which companies shipped AI releases this week?

21 companies shipped AI releases in the week of Oct 8, 2026; the most active were Anthropic, OpenAI, Arena, CoreWeave, Google DeepMind. Every launch below has primary-source links and the exact episode chapter where we discussed it.

This week's verdict table: top 10 of 30 launches — who each one is for
ReleaseBest forWhy it mattersKey number
Arena–Series funding ($200M) Model evaluators Arena raises $200M at a $3.1B valuation, broken live on ThursdAI $200M raised
Kolibri-1 · Aleph Alpha Open-source builders Aleph Alpha Kolibri-1, a 78B German-English MoE under Apache 2.0 78B total parameters
Claude Haiku 5.5 · Anthropic Agent builders Claude Haiku 5.5 at 10 cents per million input tokens $0.10/M input tokens
d1-3B and d1-omni-600M · Liquid AI Local & on-device users Liquid AI opens d1: d1-3B and d1-omni-600M decision models 8 ms d1-3B on a GPU
Mistral Large 4 · Mistral AI Open-source builders Mistral Large 4 "Le Chonk": 1T parameters, open weights promised for October 1T parameters
Beam · Reflection AI Developers & coding agents Reflection AI announces Beam, a 501B Western open-weight model 501B total parameters
FLUX 3 Image · Black Forest Labs Image creators FLUX 3 Image: 4K, bounding-box layouts and 10 reference images 4K output
Alignment Index · Arena Safety & alignment researchers Arena launches the Alignment Index —
Nano Banana 2.1 · Google DeepMind Image creators Google Nano Banana 2.1 at $0.0336 per 1K image $0.0336 per 1K image
Math manuscripts · OpenAI Researchers OpenAI publishes 722 math manuscripts from an unreleased internal model 722 manuscripts

🧠 New Models 13

New ModelsOpen weights

Kolibri-1

Aleph Alpha Kolibri-1, a 78B German-English MoE under Apache 2.0

Aleph Alpha released Kolibri-1, a 78B mixture-of-experts model with 3.46B active parameters, trained from scratch on German and English, with 1M context, under Apache 2.0. It fits on a single H200. Wolfram's verdict from vacation: a promising specialized German tool worker.

78B total parameters3.46B active parameters1M context
New Models

Claude Haiku 5.5

Claude Haiku 5.5 at 10 cents per million input tokens

Anthropic brought Haiku back after a year: Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output (under 100K tokens), with cache reads at 1 cent, ten times cheaper than Haiku 4.5. Anthropic reports 72.4% on OSWorld versus 48.9% for GPT-6 Luna, and 39.2% on Terminal-Bench 4.0 versus 16.4%.

$0.10/M input tokens72.4% OSWorld 2.1 (vendor-reported)39.2% Terminal-Bench 4.0 (vendor-reported)
New ModelsOpen weights

d1-3B and d1-omni-600M

Liquid AI opens d1: d1-3B and d1-omni-600M decision models

Liquid AI opened its d1 decision models: d1 behind an API, the open d1-3B with text and vision, and d1-omni-600M, which also takes audio. d1-3B answers in 8ms on a GPU and about 50ms on a Jetson Orin Nano, tops Liquid's Decision Index under 10B parameters, and the API is a drop-in replacement for Jev.

8 ms d1-3B on a GPU~50 ms on a Jetson Orin Nano600M omni model with audio
New Models

Mistral Large 4

Mistral Large 4 "Le Chonk": 1T parameters, open weights promised for October

Mistral's Large 4, nicknamed Le Chonk, is a trillion-parameter multimodal model with about 50B active and 1M context; open weights are promised for the end of October. Artificial Analysis scores it 38, the same as GPT-6 Luna, at about $1.13 per task versus 7 cents for Luna.

1T parameters~50B active38 Artificial Analysis Intelligence Index
New Models

Beam

Reflection AI announces Beam, a 501B Western open-weight model

Reflection AI came out of semi-stealth with Beam: 501B parameters with 23B active, trained from scratch in the West, with Apache 2.0 weights promised this month. Reflection claims 80.9% on SWE-bench Verified, admits Kimi K3 is ahead on raw capability, and pitches 3 to 4x less inference compute than GLM 5.2.

501B total parameters23B active parameters80.9% SWE-bench Verified (claimed)

🚀 Products & Apps 3

Products & Apps

Serverless GPU Sandboxes

CoreWeave Serverless GPU Sandboxes are free during the preview

CoreWeave's Serverless GPU Sandboxes give you an isolated sandbox with a GPU, started from Python at forge.coreweave.com, with no salesperson in the middle. They are free during the preview; sign up through Deok's form.

Free during the preview
Products & Apps

GPT-6 in ChatGPT

GPT-6 with Intelligent UI becomes the ChatGPT default, free tier included

GPT-6 became the default model in ChatGPT for everyone, free users included, labeled simply GPT-6. It ships with Intelligent UI: answers can come back as charts, forms and small working tools instead of a wall of text. Most of ChatGPT's 1.2 billion weekly users are on the free tier.

1.2B weekly ChatGPT users

✨ Major Features & Updates 3

🔌 APIs & Platforms 2

🛠️ Dev Tools 3

📄 Papers & Research 1

Papers & Research

Math manuscripts

OpenAI publishes 722 math manuscripts from an unreleased internal model

OpenAI pushed 722 math manuscripts to GitHub, grouped into 372 families of results, from an internal model nobody outside OpenAI can use. The model was pointed at about 4,000 open problems at roughly 3 hours of ChatGPT Pro-level thinking per result; not every result is verified in Lean. Fable sized the drop at roughly 5 Navier-Stokes results, and LDJ counted solutions to 92 of a list of the 500 most important open problems in math.

722 manuscripts372 families of results92 / 500 top open problems with solutions

📊 Benchmarks & Evals 2

Benchmarks & Evals

Hermes Index

Nous Research launches the Hermes Index; Claude Opus 5.5 leads at 63.31

Nous Research's Hermes Index is an opinionated measure of agentic work run inside Hermes Agent, reporting task completion and cost per task. Claude Opus 5.5 leads with 63.31 at $4.99 per task, with GPT-6 Astra second at 56.25 and $11.61.

63.31 Claude Opus 5.5 score$4.99 Opus 5.5 per task

💰 Funding 2

Funding

Series funding ($200M)

Arena raises $200M at a $3.1B valuation, broken live on ThursdAI

Arena raised $200 million at a $3.1 billion valuation; Peter Gostev broke the news live on ThursdAI. The focus now is Agent Arena, where you work with one agent and Arena learns from how you interact with it.

$200M raised$3.1B valuation

🌀 Also Released 1

Also Released

Personal Agent Protocol

Meta and Sierra announce the Personal Agent Protocol

Meta and Sierra, Bret Taylor's company, announced an open standard for how personal agents deal with businesses, with Walmart, Shopify and Stripe on board. An agent can browse as a guest or sign in, and the business decides how it talks to the agent. OpenAI and Anthropic haven't joined yet.