Ant Group AI Releases: Ling-3.0, Ling-3.0-Flash & LingBot-World

antgroup.com ↗

ThursdAI — the weekly AI news podcast hosted by Alex Volkov — has covered 4 Ant Group releases since Oct 2025, most recently Ling-3.0 on Aug 20, 2026. Highlights include Ling-3.0-Flash, Ming-flash-omni Preview, LingBot-World. 3 of them shipped with open weights. Every entry below has primary-source links and the episode segment where we covered it live, plus key numbers where we have them.

4 releases3 open weights4 episodesOct 2025 – Aug 2026

August 2026 1

New ModelsOpen weights

Ling-3.0

Ling-3.0: six open base checkpoints across training stages

AntLing/InclusionAI released Ling-3.0 as six open base checkpoints, including pretrained, mid-trained and WSM-merged stages for both the tiny (7.9B total / 1.3B active) and flash (124B total / 5.1B active) sizes — a rare look inside intermediate training stages.

6 open base checkpoints across training stages124B / 5.1B flash size, total / active parameters

July 2026 1

Ant Group
New Models

Ling-3.0-Flash

Ant Group's Ling-3.0-Flash: a 124B MoE that claims to match its 1T flagship on ~5B active parameters

Ant Group's inclusionAI lab released Ling-3.0-Flash, a 124B-parameter MoE activating only ~5.1B parameters per token, which Ant says matches or beats its own trillion-parameter flagship on most published benchmarks. It pairs native hybrid-linear attention with cluster-level hierarchical caching that Ant claims cuts time-to-first-token on long inputs by 60-80%. Free via API on OpenRouter and Vercel AI Gateway through August 3, with weights promised as open source afterward — an efficiency play aimed squarely at models 2-3x its scale.

124B / ~5.1B total / active parameters (MoE)60-80% claimed TTFT reduction on long inputs

February 2026 1

October 2025 1

New ModelsOpen weights

Ming-flash-omni Preview

Ming-flash-omni Preview: sparse MoE omni-modal open model

Ant Group's InclusionAI team released Ming-flash-omni Preview, a sparse mixture-of-experts omni-modal model on Hugging Face. It handles multiple input and output modalities in a single open-weights model, adding to the wave of Chinese open omni-modal releases.