LongCat-Flash-Lite-Sparse
Meituan's LongCat-Flash-Lite-Sparse: 1M native context at 3B active parameters, MIT licensed
The 'DoorDash releases a model' moment: 69B total parameters with ~3B active per token, native 1M-token context (up from 256K dense), MIT licensed. LongCat Sparse Attention lifts SWE-Bench Verified from 54.4 to 68.2 and SWE-Bench Multilingual by 21 points over the dense twin, building on DeepSeek Sparse Attention while removing its O(L²) scoring overhead. The paper reports the architecture scaling to 560B-A27B.