Ling-3.0-flash-VL
InclusionAI open-sources Ling-3.0-flash-VL: 124B vision-language MoE, 5.5B active, MIT
InclusionAI (Ant Group) open-sourced Ling-3.0-flash-VL, a 124B-parameter sparse MoE vision-language model that activates 5.5B parameters per token, adding a ViT encoder and VideoRoPE to the Ling-3.0-flash backbone for image, video and GUI-agent tasks with up to 1M tokens of context. Weights ship in BF16 and FP8 on Hugging Face under MIT. It got a brief mention on the show as a solid workhorse VL release.