New Models
Bland Speech v3
Bland Speech v3 tops the Audio Realism Bench, one Elo rung below actual humans
Design Arena's blind pairwise Audio Realism benchmark puts Bland Speech v3 at 1365 Elo, above ElevenLabs, Microsoft's MAI-Voice-2, and Grok TTS, second only to real human recordings around 1500. Trained on 100M+ real phone conversations, it keeps the breaths, hesitations, and fillers TTS usually sands off. Ten seconds of audio yields an instant clone at $0.015 per thousand characters. The asterisk came from Grok itself: every ranked model is a closed API; open source voice has catching up to do.
1365 Elo, second only to humans (~1500)100M+ real conversations in training10s audio needed for an instant clone