StepFun AI Releases: StepAudio 3, StepAudio 2.5 & Step 3.5 Flash Base

stepfun.com ↗

ThursdAI — the weekly AI news podcast hosted by Alex Volkov — has covered 7 StepFun releases since Feb 2025, most recently StepAudio 3 on Sep 17, 2026. Highlights include StepAudio 3, Step-Video-T2V, Step-Video-TI2V, Step1X-3D. 5 of them shipped with open weights. Every entry below has primary-source links and the episode segment where we covered it live, plus key numbers where we have them.

7 releases5 open weights7 episodesFeb 2025 – Sep 2026

September 2026 1

StepFun
New Models

StepAudio 3

StepFun's StepAudio 3 family takes #1 on Artificial Analysis real-time voice

StepFun launched StepAudio 3, a five-model audio family spanning Real-Time Preview, ASR Max, and TTS — a full suite for building end-to-end voice assistants in code. Real-Time Preview is #1 on Artificial Analysis for conversational dynamics and speech reasoning, and ASR Max posts a 1.7% word error rate, significantly below Whisper. API only, no open weights.

#1 Artificial Analysis real-time voice1.7% ASR Max word error rate5 models in the family

April 2026 1

StepFun
New Models

StepAudio 2.5

StepAudio 2.5 TTS adds natural-language control of emotion and delivery

StepFun released StepAudio 2.5, a text-to-speech model that lets you steer emotion and delivery with natural-language instructions. It was covered in the show's Voice & Audio segment as the week's notable speech release.

March 2026 1

StepFun
New ModelsOpen weights

Step 3.5 Flash Base

StepFun open-sources Step 3.5 Flash Base with its training stack

StepFun released Step 3.5 Flash Base and Midtrain checkpoints, an unusually open release that includes training artifacts and the SteptronOSS training stack alongside the weights. The panel praised the Apache-2 orientation and called the continuation-pretraining flexibility a major practical unlock for builders.

February 2026 1

StepFun
New ModelsOpen weights

Step 3.5 Flash

StepFun Step 3.5 Flash: frontier reasoning claims at 11B active params

StepFun released Step 3.5 Flash, a 196B sparse MoE model with only 11B active parameters, claiming frontier-level reasoning while generating at 100-350 tokens per second. It continues the trend of sparse Chinese MoE models delivering high speed at low active parameter counts.

May 2025 1

StepFun
New ModelsOpen weights

Step1X-3D

StepFun's Step1X-3D: open two-stage framework for textured 3D assets

StepFun released Step1X-3D, an open two-stage framework for high-fidelity, controllable generation of textured 3D assets: it first synthesizes watertight geometry, then generates view-consistent textures. Trained on 2M curated meshes, the release also includes a curated dataset of 800K assets and a Hugging Face demo.

March 2025 1

February 2025 1

StepFun
New ModelsOpen weights

Step-Video-T2V

StepFun open-sources Step-Video-T2V, a SOTA 30B text-to-video model

StepFun released Step-Video-T2V (plus a T2V Turbo variant), a 30 billion parameter state-of-the-art text-to-video model under an MIT license. Results impressed especially on text integration, such as rendering 'We will open source' on a scroll as a character unfurls it, marking one of the strongest open-source video drops of the week.