New Models
MiniMax H3 Max
fal's MiniMax H3 Max generates 5-second video in under 3 seconds
fal Research debuted MiniMax H3 Max, a post-train of the open-weight MiniMax H3 that ranks #1 on image-to-video and #3 on text-to-video on Artificial Analysis while generating five-second clips in under three seconds — 2.53s in the live on-air test. Priced at $0.04/second at 768p (promo until Sept 1), with a weights release planned. The speed puts it alone on the speed-versus-quality Pareto frontier.
2.53s to generate a 5-second clip, live on the show#1 image-to-video on Artificial Analysis (#3 T2V)$0.04/s at 768p until Sept 1
New Models
Gemini Omni 1.1 Flash
Gemini Omni 1.1 Flash tops Arena text-to-video with voice-consistent scene extension
Google's new video model dropped during the show: #1 on Arena's text-to-video leaderboard and #2 on image-to-video. It analyzes up to 10 seconds of previous footage to extend scenes while keeping character identity, voice, and lighting locked; adds first/last-frame control and infinite loops; and offers 360p draft generations with built-in upscaling. Rolling out in Google AI Studio, Flow, and Gemini Enterprise with API access.
#1 Arena text-to-video (#2 image-to-video)10s of prior footage analyzed for scene extension
New ModelsOpen weights
Wan-Animate-2
Alibaba Wan-Animate-2: 14B character animation under Apache 2.0
Alibaba's Wan team released Wan-Animate-2, a 14B-parameter character animation model under Apache 2.0. It wins over 70% of blind preference comparisons.
14B parameters70%+ blind preference win rate
New ModelsOpen weights
LTX-2.5
LTX-2.5: 22B open-weights video model with multi-shot generation
Lightricks released LTX-2.5, a 22B-parameter open-weights video model with multi-shot support. It generates 10 seconds of 1080p video in 23.7 seconds on fal and needs a minimum of 16GB VRAM.
22B parameters23.7s for 10s of 1080p on fal16GB minimum VRAM
New Models
Wan 3.0
Wan 3.0 drops into public beta live during the show, with native 30-second generation
Alibaba's Tongyi Lab pushed Wan 3.0 into public beta minutes before ThursdAI went live: native 30-second single-shot generation and Omni-Reference conditioning on text, images, audio, and video together. Given the Wan line's open-weight track record, this is the drop the open source community most wants weights for.
30s native single-shot generation4 reference modalities via Omni-Reference
Products & Apps
Anywear
Decart's Anywear: real-time virtual try-on from any shopping site at 40ms a frame
A free Chrome extension: drag a garment from any shopping site onto your webcam feed and Decart's world model regenerates you wearing it, frame by frame at 40ms latency, with no retailer integration. Kfir Aberman demoed it live on ThursdAI, where Alex swapped his real jacket for a digital one on camera and bought a Dolce & Gabbana suit mid-interview, wearing it before it shipped. Aberman's frame: agentic commerce needs world models to close the loop between browsing and trying.
40ms per-frame generation latency0 retailer integrations required
New Models
FLUX 3 Video
FLUX 3 Video: BFL's first video model generates native audio in the same pass
'Two years later, our first video model': up to 20 seconds at 24fps in 720p (1080p via upscaler), with dialogue, SFX, and ambience generated natively in the same pass and lip-sync across 14+ languages. Draft mode runs ~$0.06/s for iteration versus $0.17-0.29/s full renders, and three API modes share one endpoint (t2v, keyframe-pinned i2v, v2v continuation). BFL's internal ELO has it leading text-to-video; open weights as FLUX 3 Dev are explicitly promised.
20s @ 24fps max clip, 720p native$0.06/s draft mode vs $0.17-0.29/s full14+ lip-synced languages
New ModelsOpen weights
MiniMax H3 open weights
MiniMax opens H3's weights, and the community ships LoRAs and Apple Silicon in 48 hours
H3 (Hailuo 3.0), a 33B omni-modal transformer generating up to 15 seconds at 2K with native stereo audio from unified text/image/video/audio context, landed on Hugging Face days after its announcement, per Victor Su Ortiz the first open-weight state-of-the-art omni video model. Within 48 hours the community shipped LoRA support and Apple Silicon inference, neither of which MiniMax optimized for, and X filled with recreated episodes of The Office. The panel also dug into the community license's litigation-linked restrictions on US use, the gap between downloadable and cleared.
33B open-weight omni transformer48 hrs community LoRAs + Apple Silicon support2K / 15s max resolution / clip length