New Models
Wan 3.0
Wan 3.0 drops into public beta live during the show, with native 30-second generation
Alibaba's Tongyi Lab pushed Wan 3.0 into public beta minutes before ThursdAI went live: native 30-second single-shot generation and Omni-Reference conditioning on text, images, audio, and video together. Given the Wan line's open-weight track record, this is the drop the open source community most wants weights for.
30s native single-shot generation4 reference modalities via Omni-Reference
New Models
SeedRealtime
ByteDance's SeedRealtime: a native audio-visual full-duplex LLM, free on Doubao
A single end-to-end model natively fusing audio, video, and text, replacing the cascaded ASR-VLM-TTS pipelines behind most voice agents: it listens, watches, and speaks simultaneously with turn-taking inside the model and no external VAD, cutting conversational pacing failures by 50% in human evals. It ties voices to faces in noisy rooms and speaks up proactively on scene changes, live for free in the Doubao app and its 300M+ users, ByteDance's first large-scale audio-visual full-duplex deployment.
50% fewer pacing failures vs cascaded stacks300M+ Doubao users getting it free
New Models
FLUX 3 Video
FLUX 3 Video: BFL's first video model generates native audio in the same pass
'Two years later, our first video model': up to 20 seconds at 24fps in 720p (1080p via upscaler), with dialogue, SFX, and ambience generated natively in the same pass and lip-sync across 14+ languages. Draft mode runs ~$0.06/s for iteration versus $0.17-0.29/s full renders, and three API modes share one endpoint (t2v, keyframe-pinned i2v, v2v continuation). BFL's internal ELO has it leading text-to-video; open weights as FLUX 3 Dev are explicitly promised.
20s @ 24fps max clip, 720p native$0.06/s draft mode vs $0.17-0.29/s full14+ lip-synced languages
New Models
Bland Speech v3
Bland Speech v3 tops the Audio Realism Bench, one Elo rung below actual humans
Design Arena's blind pairwise Audio Realism benchmark puts Bland Speech v3 at 1365 Elo, above ElevenLabs, Microsoft's MAI-Voice-2, and Grok TTS, second only to real human recordings around 1500. Trained on 100M+ real phone conversations, it keeps the breaths, hesitations, and fillers TTS usually sands off. Ten seconds of audio yields an instant clone at $0.015 per thousand characters. The asterisk came from Grok itself: every ranked model is a closed API; open source voice has catching up to do.
1365 Elo, second only to humans (~1500)100M+ real conversations in training10s audio needed for an instant clone
New ModelsOpen weights
LFM2.5-2.6B
Liquid's LFM2.5-2.6B: agentic RL trained inside real harnesses, running in 1.7GB on a phone
A 2.69B-parameter hybrid model pre-trained on ~34T tokens whose post-training ran agentic RL inside real harnesses (Hermes Agent, OpenClaw, Pi), so tool calling was learned where tool calling happens. It beats Qwen3.5-9B, three times its size, on ToolSandbox and instruction following, runs 220 tok/s on an M5 Max CPU and fits in ~1.7GB at Q4 on a phone. Liquid's own model card honestly scopes it away from agentic coding and knowledge-heavy work: this is for private, on-device agents.
2.69B parameters, 128K context77.83 ToolSandbox, above Qwen3.5-9B220 tok/s on Apple M5 Max CPU
New Models
Qwen3.8-Max
Qwen3.8-Max: Alibaba's 2.4T-parameter flagship, with open weights promised within a week
Alibaba's flagship MoE arrives via API: 2.4T total parameters, 95B active, 1M context, at $2/$6 per million tokens, with open weights plus a 27B sibling promised for the week of August 10, the first Max-class Qwen slated for release. It ranks #2 on EyeBench for vision behind only OpenAI's Sol, and the oh-my-cli demo ran 16 days of fully autonomous coding: 265 commits, 127 PRs, 151 issues, zero human intervention. Nisten's hands-on: best-in-class visual data labeling. Yam's counter: other frontier models pass those tests too, and the 27B is the one you'll run at home.
2.4T / 95B total / active parameters#2 EyeBench vision rank, behind only Sol16 days autonomous run: 265 commits, 127 PRs
New ModelsOpen weights
LongCat-Flash-Lite-Sparse
Meituan's LongCat-Flash-Lite-Sparse: 1M native context at 3B active parameters, MIT licensed
The 'DoorDash releases a model' moment: 69B total parameters with ~3B active per token, native 1M-token context (up from 256K dense), MIT licensed. LongCat Sparse Attention lifts SWE-Bench Verified from 54.4 to 68.2 and SWE-Bench Multilingual by 21 points over the dense twin, building on DeepSeek Sparse Attention while removing its O(L²) scoring overhead. The paper reports the architecture scaling to 560B-A27B.
69B / 3B total / active parameters1M native context window68.2 SWE-Bench Verified, +13.8 over dense
New ModelsOpen weights
MiniMax H3 open weights
MiniMax opens H3's weights, and the community ships LoRAs and Apple Silicon in 48 hours
H3 (Hailuo 3.0), a 33B omni-modal transformer generating up to 15 seconds at 2K with native stereo audio from unified text/image/video/audio context, landed on Hugging Face days after its announcement, per Victor Su Ortiz the first open-weight state-of-the-art omni video model. Within 48 hours the community shipped LoRA support and Apple Silicon inference, neither of which MiniMax optimized for, and X filled with recreated episodes of The Office. The panel also dug into the community license's litigation-linked restrictions on US use, the gap between downloadable and cleared.
33B open-weight omni transformer48 hrs community LoRAs + Apple Silicon support2K / 15s max resolution / clip length