Qwen 3.8 Omni Flash
Qwen 3.8 Omni Flash with 1M context and Qwen 3.8 Live Translate
Alibaba's Qwen team released Qwen 3.8 Omni Flash with a 1M-token context window, alongside Qwen 3.8 Live Translate.
Alibaba's Qwen team is the highest-velocity open-weights lab on the timeline — almost every entry is a model drop, from the Qwen 2.5 and Qwen 3 language-model lines to Wan video and Z-Image generation models. ThursdAI — the weekly AI news podcast hosted by Alex Volkov — has covered 45 Alibaba (Qwen) releases since Jan 2025, most recently Qwen 3.8 Omni Flash on Sep 24, 2026. Highlights include Qwen3.8-Max, Qwen 3, Qwen3.8-27B, Qwen3.8-Flash-Next. 29 of them shipped with open weights. Every entry below has the episode segment where we covered it live, plus primary-source links and key numbers where we have them.
Qwen 3.8 Omni Flash with 1M context and Qwen 3.8 Live Translate
Alibaba's Qwen team released Qwen 3.8 Omni Flash with a 1M-token context window, alongside Qwen 3.8 Live Translate.
Qwen Image 2.1 (7B) adds native transparency
Qwen Image 2.1 is a 7B image generation model with native transparency support.
Qwen3.8-Max-0902: 2.4T API-only refresh claims #1 on Code Arena
Alibaba refreshed its API-only frontier model with a September 2 snapshot: 2.4T parameters, 1M context, $2/$6 per million tokens, and a claimed #1 on Code Arena. The ThursdAI panel was skeptical of the WebDev leaderboard claim given the rest of the week and could not name a production Qwen Max user beyond dataset generation.
Qwen3.8-Flash-Next previews the Qwen4 architecture in open weights
Alibaba released open weights for Qwen3.8-Flash-Next: 125B parameters plus a 51B N-gram embedding table with only 6B active, trained at roughly 1/9 the cost of Qwen3.7-Plus while self-reporting DeepSWE 58.7 and SWE-bench Pro 62.5. It previews the Qwen4 architecture: Qwen Sparse Attention (QSA) replaces full attention layers, and the low-bandwidth N-gram embedding table can be offloaded to slow RAM or even an SSD.
Alibaba's HappyShrimp 1.0 goes end-to-end on music generation
Yes, it's really called HappyShrimp (a shrimp-welfare meme). Alibaba's end-to-end music model generates lyrics, melody, arrangement and vocals in one pass, and unlike the Suno approach it reasons over the prompt first, mapping song structure and harmonic progression before generating audio. Early testers call it a serious and possibly cheapest Suno rival, with 320 free credits at launch. The track played on the show is extremely K-pop.
Qwen3.8-27B ties GPT-5.6 Luna and runs on a 4090
Alibaba's overnight community darling: a 27B-parameter Apache 2.0 model scoring 52 on the Artificial Analysis Intelligence Index — the same as GPT-5.6 Luna at max reasoning — and 51 on the Agentic index. It runs at ~68 tokens/sec on a 4090, ~40 on Macs via MLX, and even 11 tok/s in-browser on WebGPU kernels. The Hugging Face hub exploded with 152 fine-tunes, 650 quantizations and close to 10 million quant downloads, and Unsloth's 1-bit quants run it on 8GB of RAM at roughly 77% of BF16 quality.
Alibaba Wan-Animate-2: 14B character animation under Apache 2.0
Alibaba's Wan team released Wan-Animate-2, a 14B-parameter character animation model under Apache 2.0. It wins over 70% of blind preference comparisons.
Wan 3.0 drops into public beta live during the show, with native 30-second generation
Alibaba's Tongyi Lab pushed Wan 3.0 into public beta minutes before ThursdAI went live: native 30-second single-shot generation and Omni-Reference conditioning on text, images, audio, and video together. Given the Wan line's open-weight track record, this is the drop the open source community most wants weights for.
Qwen3.8-Max: Alibaba's 2.4T-parameter flagship, with open weights promised within a week
Alibaba's flagship MoE arrives via API: 2.4T total parameters, 95B active, 1M context, at $2/$6 per million tokens, with open weights plus a 27B sibling promised for the week of August 10, the first Max-class Qwen slated for release. It ranks #2 on EyeBench for vision behind only OpenAI's Sol, and the oh-my-cli demo ran 16 days of fully autonomous coding: 265 commits, 127 PRs, 151 issues, zero human intervention. Nisten's hands-on: best-in-class visual data labeling. Yam's counter: other frontier models pass those tests too, and the 27B is the one you'll run at home.
Alibaba previews Qwen3.8-Max at a claimed 2.4T parameters, 'second only to Fable 5'
Alibaba's Qwen team previewed Qwen3.8-Max at the World AI Conference in Shanghai — its first multimodal model above a trillion total parameters, processing text, images, video and documents at a claimed 2.4T. Alibaba shares rose as much as 5.4% in Hong Kong on the news. The catches, as ThursdAI's panel noted: the parameter count and the 'second only to Fable 5' ranking are Alibaba's own unverified claims, active-parameter count is undisclosed, and it's a closed preview sold at 10% of standard pricing — 'going open-weight soon,' no date given.
Alibaba releases Qwen 3.7-Max agentic frontier model with robotics demos
Alibaba released Qwen 3.7-Max, an agentic frontier model built for long autonomous runs, demonstrated alongside robotics demos. It continues the Qwen Max line as Alibaba's closed frontier offering aimed at agentic workloads.
Qwen3.6-27B: dense Apache-2.0 model beats Alibaba's own 400B flagship
Alibaba shipped Qwen3.6-27B, a dense 27B-parameter model under Apache 2.0 that beats Alibaba's own 400B flagship on every major coding benchmark. Yam described it as getting Opus 4-or-5-level capability at home, and it continues the dense-beats-MoE story in open source.
Qwen3.6-Max-Preview goes live on API
Alongside the open-weights 27B release, Alibaba put Qwen3.6-Max-Preview live on its API. It is the frontier closed-weights tier of the Qwen3.6 family, available API-only rather than as open weights.
Qwen 3.6-35B-A3B: Apache 2.0 MoE with 3B active hits 73.4% SWE-Verified
Alibaba Qwen open-sourced Qwen 3.6-35B-A3B under Apache 2.0 the same morning Opus 4.7 dropped: a 35B MoE with only 3B active parameters that scores 73.4% on SWE-bench Verified, rivaling models 10x its size. It is natively multimodal with 262K context extensible to 1M, and the crew called it the strongest mid-size LLM on nearly all benchmarks, putting to rest doubts about Qwen's open-source commitment after Junyang Ling's departure.
HappyHorse-1.0 takes #1 on Artificial Analysis video arena
HappyHorse-1.0, a mysterious 15B-parameter video model from Alibaba's Taotian Group, took the #1 spot on the Artificial Analysis video arena, beating Seedance 2.0, Kling 3.0, and Grok Video. Little is known about the model beyond its size and leaderboard run.
Alibaba open-sources Qwen3.5-Omni, a 397B native omni-modal model
Qwen3.5-Omni is Alibaba's natively omni-modal open model handling text, image, audio, and video, with 397B total parameters and 17B active. It extends the Qwen family's open-source momentum into unified multimodal workloads.
Alibaba ships Qwen3.6-Plus with near-Opus agentic coding and 1M context
Alibaba released Qwen3.6-Plus, an API model with agentic coding performance near Opus 4.5 and a 1M-token context window. The panel noted continued strong momentum for the Qwen family in practical coding and agent workloads.
Alibaba Wan2.7-Image unifies generation, editing, and text rendering
Alibaba's Wan team released Wan2.7-Image, a unified image model covering generation, editing, text rendering, and multi-image consistency. The panel covered it in the open ecosystem round-up alongside the Qwen updates.
Alibaba releases Qwen3.5 small models (2B, 4B, 9B) for local use
Alibaba released the Qwen3.5 small model series with 2B, 4B, and 9B variants, which the panel found highly usable on consumer hardware. The release landed alongside leadership turbulence as Junyang Lin and Binyuan Hui departed Qwen, though the panel expects Alibaba's open-source momentum to continue.
Qwen 3.5 lands: 35B/3B-active Medium outperforms the old 235B flagship
Alibaba released the Qwen 3.5 family of open-weight models, headlined by Qwen3.5-35B-A3B, a 35B model with only 3B active parameters that outperforms their previous 235B flagship. Variants include a 122B-A10B and a dense 27B, with the panel highlighting the hybrid state-space (Mamba-layer) architecture and strong practical coding and agent performance at a tiny active-parameter footprint.
Alibaba opens Qwen 3.5: 397B-param multimodal MoE with only 17B active
Alibaba released Qwen3.5-397B-A17B, billed as the first open-weight native multimodal MoE model, with 397B total parameters, just 17B active, 512 experts, and 262K native context extendable to 1M. It delivers 8.6-19x faster inference than Qwen3-Max and continues Qwen's strength in multilingual and medical tasks, scoring 52.5% on Terminal Bench, third place among open-source models. Nisten found coding still trails GLM-5.
Alibaba launches Qwen-Image-2.0 with native 2K resolution
Alibaba's Qwen team launched Qwen-Image-2.0, a 7B-parameter image generation model with native 2K resolution output and superior text rendering. Available to try on chat.qwen.ai.
Qwen3-Coder-Next hits 70.6% SWE-Bench Verified with 3B active params
Alibaba's Qwen3-Coder-Next is an 80B MoE coding agent model with only 3B active parameters that scores 70.6% on SWE-Bench Verified and 44% on the much harder SWE-Bench Pro. It was trained on 7.5T tokens with 20,000 parallel RL environments and runs under 48GB of RAM with GGUF quantization, making near-frontier agentic coding feasible on local hardware.
Tongyi Lab releases Z-Image generation model
Alibaba's Tongyi Lab released Z-Image, a new image generation model, with support landing in the open-source DiffSynth-Studio toolkit on GitHub. Covered in the AI Art segment alongside HunyuanImage 3.0.
Qwen3-TTS: open-source TTS family with 97ms latency and voice cloning
Alibaba's Qwen team released Qwen3-TTS, a full open-source text-to-speech family under Apache 2 that dropped 30 minutes before the show. It spans 5 models from 0.6B to 1.7B parameters, with 97ms latency, voice cloning from just 3 seconds of audio, voice description prompting, and 10-language support.
Qwen 3 Coder posts insane scores in the race for the coding crown
Alibaba's Qwen 3 Coder landed in July with what the crew called insane benchmark scores for an open-weights coding model. Together with Kimi K2 and GLM 4.5 it made July the peak month for Chinese open source.
Qwen launches speech-to-speech model with emotion handling
Qwen released a speech-to-speech model in March with internal emotion handling, joining the wave of voice-native models. It was part of the Qwen team's relentless 2025 release cadence across modalities.
Tongyi's Z-Image Turbo brings sub-second open image generation
Alibaba's Tongyi lab released Z-Image Turbo, a 6B-parameter open image generation model that produces images in under a second. It pushes open-source image generation toward real-time speeds at a fraction of the size of competing models.
Qwen Image Edit gains Multi-Angle LoRA for camera control
A Multi-Angle LoRA for Qwen Image Edit landed, enabling camera-control style edits that re-render a scene from new angles. Available as a Hugging Face space and on fal, it shows the fast-moving open ecosystem building on Qwen's image editing models.
Qwen3-VL adds compact 2B and 32B multimodal models
Alibaba's Qwen team extended the Qwen3-VL family with newly updated 2B and 32B checkpoints. The 2B is a generic VLM (OCR-capable) that holds up against its 4B and 8B siblings from prior weeks, while the 32B reportedly outperforms GPT-5 mini and Claude 4 Sonnet on benchmarks.
Qwen3-VL adds compact 3B and 8B open vision-language models
Alibaba's Qwen team released smaller Qwen3-VL vision-language models in 3B and 8B sizes, bringing the flagship VL capabilities down to edge- and laptop-friendly scales. Weights are open on Hugging Face as part of the Qwen3-VL collection.
Qwen3-Omni ships open-weights any-to-any audio, vision, and text
Alongside Qwen3-VL, Alibaba released Qwen3-Omni, an end-to-end omni-modal open-weights model that takes text, image, audio, and video input and can respond with streaming speech. The show treated it as direct evidence of how fast open multimodal systems are improving, with weights on Hugging Face, a GitHub repo, demos, and availability in Qwen Chat and the Model Studio API.
Qwen3-TTS-Flash multilingual text-to-speech lands via Alibaba's API
Part of the same Qwen release streak, Qwen3-TTS-Flash is a low-latency multilingual text-to-speech model with multiple voices and dialect support, offered through Alibaba Cloud Model Studio's API rather than as open weights. It fed into the episode's closing audio-demo pileup, where voice launches were treated as product proof points.
Alibaba releases Qwen3-VL open-weights vision-language flagship
Alibaba's Qwen team shipped Qwen3-VL, its new flagship open-weights vision-language family, headlining the episode's 'Qwen-mas' barrage. The panel discussed it as a practical workflow tool for visual understanding and agentic GUI tasks, not just another model card, with weights, a blog post, and a Hugging Face demo all available at launch.
Wan Animate brings open-weights character animation and replacement
Alibaba's Wan team released Wan 2.2 Animate, an open-weights model that animates a character image from a performance video, replicating motion and expressions, or swaps a character into existing footage. It landed in the episode's closing run of video releases showing multimodal product quality climbing across the board.
Tongyi DeepResearch: open-source A3B web agent rivals OpenAI Deep Research
Alibaba's Tongyi Lab open-sourced Tongyi DeepResearch, a 30B mixture-of-experts web research agent with only 3B active parameters. The lab claims parity with OpenAI's Deep Research on agentic search and report-writing tasks, and the weights are available on Hugging Face.
Alibaba's Tongyi Lab open-sources WebWatcher vision-language research agent
Alibaba's Tongyi Lab open-sourced WebWatcher, a vision-language deep research agent that sets new state-of-the-art results on agentic browsing and research tasks. The 32B model combines visual understanding with web research capabilities and is available on Hugging Face.
Alibaba launches Qwen-TTS with human-level bilingual naturalness
The Qwen team released Qwen-TTS, a bilingual Chinese/English text-to-speech model claiming human-level naturalness, available via API with a Hugging Face demo space. It was the second voice release of the week alongside Kyutai TTS.
Alibaba's Wan 2.1: open-source diffusion-transformer text-to-video suite
Alibaba, the team behind the Qwen LLMs, released Wan 2.1, a full stack of open-source diffusion-transformer text-to-video foundation models. Amid the show's discussion of video-model fatigue, this was called out as a release that cuts through the noise, with weights on Hugging Face and code on GitHub.
Qwen 2.5 Omni gets an update
Alongside the Qwen 3 launch, Alibaba updated its Qwen 2.5 Omni multimodal model line. Mentioned briefly in the open-source roundup as part of the week's Qwen ecosystem push.
Alibaba open-weights the full Qwen 3 family under Apache 2.0
Alibaba released the entire Qwen 3 stack: two MoE models (235B total/22B active and 30B/3B active) plus six dense siblings from 32B down to 0.6B, all Apache 2.0 with day-one support in LM Studio, Ollama, vLLM, MLX and llama.cpp. The headline feature is a runtime hybrid 'thinking' toggle (/think and /no_think) that trades latency for reasoning depth. Trained on ~36T tokens with 128K context and 119-language coverage, the 235B MoE rivals DeepSeek-R1, o1, o3-mini and Gemini 2.5 Pro on coding and math.
Qwen launches Omni 7B: sees, hears, reads, and talks back
Qwen released Qwen2.5-Omni-7B, an open-weights omni-modal model that perceives text, images, audio, and video, and generates both text and speech. It packs end-to-end multimodal perception and spoken output into a 7B parameter model available on Hugging Face.
Qwen releases QwQ-32B reasoning model that matches R1 on some evals
Alibaba's Qwen team released QwQ-32B, an open-weights reasoning model that matches DeepSeek R1 on several evals despite being roughly 20x smaller at 32B parameters. Qwen tech lead Junyang Lin joined the show to announce it, and the episode dubbed it Alibaba's 'R1 killer' for bringing strong reasoning to a size that runs on consumer hardware.
Alibaba launches Qwen2.5-Max flagship model with hidden video gen
Alibaba's Qwen team released Qwen2.5-Max, a large MoE flagship model available through the Qwen Chat interface and API, claiming competitive results against DeepSeek V3 and other frontier models. The chat app also quietly shipped a video generation capability powered by Alibaba's Tongyi Wanxiang.
Alibaba ships Qwen2.5-VL open vision-language model family
Alibaba's Qwen team released Qwen2.5-VL, open-weights vision-language models up to 72B that handle images, documents, video understanding, and on-screen agentic grounding. The 72B Instruct model was immediately available on Hugging Face and in Qwen Chat.
Never miss a Alibaba (Qwen) launch — we cover every release live, every Thursday.