SeedRealtime
ByteDance's SeedRealtime: a native audio-visual full-duplex LLM, free on Doubao
A single end-to-end model natively fusing audio, video, and text, replacing the cascaded ASR-VLM-TTS pipelines behind most voice agents: it listens, watches, and speaks simultaneously with turn-taking inside the model and no external VAD, cutting conversational pacing failures by 50% in human evals. It ties voices to faces in noisy rooms and speaks up proactively on scene changes, live for free in the Doubao app and its 300M+ users, ByteDance's first large-scale audio-visual full-duplex deployment.