MOSS-VL-Realtime
OpenMOSS open-sources MOSS-VL-Realtime, an 11B VLM that decides when to speak — and when to stay silent
OpenMOSS released MOSS-VL-Realtime, an open-source 11B-parameter vision-language model (~22.7GB) built for real-time streaming video that proactively speaks up or deliberately stays silent instead of only answering prompts. ThursdAI reported it as state-of-the-art on all three open proactivity benchmarks, with the base model included in the release.