StepAudio 3
StepFun's StepAudio 3 family takes #1 on Artificial Analysis real-time voice
StepFun launched StepAudio 3, a five-model audio family spanning Real-Time Preview, ASR Max, and TTS — a full suite for building end-to-end voice assistants in code. Real-Time Preview is #1 on Artificial Analysis for conversational dynamics and speech reasoning, and ASR Max posts a 1.7% word error rate, significantly below Whisper. API only, no open weights.