Griffin
Tavus Griffin: 48% of callers in Tavus's study thought it was human
Tavus released Griffin, a conversational video model; in Tavus's own study, 48% of callers thought it was human.
Text- and image-to-video models, video editing, avatars, and animation. — 87 releases covered on the show.
Tavus Griffin: 48% of callers in Tavus's study thought it was human
Tavus released Griffin, a conversational video model; in Tavus's own study, 48% of callers thought it was human.
HeyGen Video runs on MiniMax H3 at $0.01 a second through October
HeyGen Video now runs on MiniMax H3, priced at $0.01 per second through October.
ACT 486: a research demo for talking to and reshaping a video
ACT 486 is a research demo that lets viewers talk to a video and reshape it interactively.
fal.live: viewer-steered infinite video stream built in a weekend after Twitch and Kick bans
After fal's infinite Rick and Morty 'interdimensional cable' stream was banned from Twitch and then Kick for copyright, the team built its own streaming site over a weekend with a model tuned for continuous generation. Viewers vote on where the anime and other channels go next. Alex credits it with inspiring the thursdai.news/live build.
fal H3 Max Turbo: ~97% of H3 Max quality in 1.4 seconds at $0.01 per second
fal's Turbo variant of MiniMax H3 Max keeps about 97% of H3 Max quality, generates a clip in 1.4 seconds, and costs one cent per second at 768p, about 50% cheaper. Alex's live test found it no longer renders copyrighted characters the way H3 Max did.
Runway GWM Worlds 2: real-time 720p, 24 fps world model with 48 kHz audio
Runway's second generative world model runs in real time at 720p and 24 fps with open-ended sessions, 48 kHz audio, and generated speech, a big jump over the low-fidelity audio in most world models. LDJ broke it live on the show two days after Runway's previous world model. Research preview.
World Labs Atlas: omnimodel turns 1 to 6 images into walkable 3D and bullet-time video
Atlas is pretrained from scratch to natively operate in text, images, video, and 3D as an autoregressive diffusion transformer. From 1 to 6 images it generates up to a minute of 1440p camera-controlled video and a 3D reconstruction with no Gaussian splats, and it can reframe real video from new angles, producing bullet-time from three ordinary tripods. LDJ noted quality scales with more camera angles. Partner access only for now.
fal's MiniMax H3 Max generates 5-second video in under 3 seconds
fal Research debuted MiniMax H3 Max, a post-train of the open-weight MiniMax H3 that ranks #1 on image-to-video and #3 on text-to-video on Artificial Analysis while generating five-second clips in under three seconds — 2.53s in the live on-air test. Priced at $0.04/second at 768p (promo until Sept 1), with a weights release planned. The speed puts it alone on the speed-versus-quality Pareto frontier.
Gemini Omni 1.1 Flash tops Arena text-to-video with voice-consistent scene extension
Google's new video model dropped during the show: #1 on Arena's text-to-video leaderboard and #2 on image-to-video. It analyzes up to 10 seconds of previous footage to extend scenes while keeping character identity, voice, and lighting locked; adds first/last-frame control and infinite loops; and offers 360p draft generations with built-in upscaling. Rolling out in Google AI Studio, Flow, and Gemini Enterprise with API access.
Alibaba Wan-Animate-2: 14B character animation under Apache 2.0
Alibaba's Wan team released Wan-Animate-2, a 14B-parameter character animation model under Apache 2.0. It wins over 70% of blind preference comparisons.
LTX-2.5: 22B open-weights video model with multi-shot generation
Lightricks released LTX-2.5, a 22B-parameter open-weights video model with multi-shot support. It generates 10 seconds of 1080p video in 23.7 seconds on fal and needs a minimum of 16GB VRAM.
Wan 3.0 drops into public beta live during the show, with native 30-second generation
Alibaba's Tongyi Lab pushed Wan 3.0 into public beta minutes before ThursdAI went live: native 30-second single-shot generation and Omni-Reference conditioning on text, images, audio, and video together. Given the Wan line's open-weight track record, this is the drop the open source community most wants weights for.
Decart's Anywear: real-time virtual try-on from any shopping site at 40ms a frame
A free Chrome extension: drag a garment from any shopping site onto your webcam feed and Decart's world model regenerates you wearing it, frame by frame at 40ms latency, with no retailer integration. Kfir Aberman demoed it live on ThursdAI, where Alex swapped his real jacket for a digital one on camera and bought a Dolce & Gabbana suit mid-interview, wearing it before it shipped. Aberman's frame: agentic commerce needs world models to close the loop between browsing and trying.
FLUX 3 Video: BFL's first video model generates native audio in the same pass
'Two years later, our first video model': up to 20 seconds at 24fps in 720p (1080p via upscaler), with dialogue, SFX, and ambience generated natively in the same pass and lip-sync across 14+ languages. Draft mode runs ~$0.06/s for iteration versus $0.17-0.29/s full renders, and three API modes share one endpoint (t2v, keyframe-pinned i2v, v2v continuation). BFL's internal ELO has it leading text-to-video; open weights as FLUX 3 Dev are explicitly promised.
MiniMax opens H3's weights, and the community ships LoRAs and Apple Silicon in 48 hours
H3 (Hailuo 3.0), a 33B omni-modal transformer generating up to 15 seconds at 2K with native stereo audio from unified text/image/video/audio context, landed on Hugging Face days after its announcement, per Victor Su Ortiz the first open-weight state-of-the-art omni video model. Within 48 hours the community shipped LoRA support and Apple Silicon inference, neither of which MiniMax optimized for, and X filled with recreated episodes of The Office. The panel also dug into the community license's litigation-linked restrictions on US use, the gap between downloadable and cleared.
Seedance 2.5 goes global, then lands in the US for the first time during the show
ByteDance's top-ranked video model launched globally on Dreamina with native 30-second clips, a long-video mode assembling up to 3 minutes with consistent characters, up to 50 multimodal references including 3D white models and green-screen footage, timestamp control to the second, and actual Maya and Blender plugins. US access switched on during the ThursdAI cold open, the first time Seedance has been available stateside.
FLUX 3: one model for image, 20-second video with audio, and robot action prediction
Black Forest Labs launched FLUX 3 in early access — its first model generating video, audio, and robot action-prediction from one set of weights, alongside image generation (a FLUX.3 Mimic variant was announced with it). FLUX 3 Video produces clips with native audio up to 20 seconds from text, images or footage, with continuation, keyframe transitions, multilingual dialogue and clip chaining; the same architecture is already teaching robots tasks on an Audi assembly line. BFL's own preference tests: 77% wins vs Runway Gen-4.5, 93% vs Luma Ray 3.2. Image generation and the open-weights FLUX 3 Dev come in later rollout phases.
Meta Superintelligence Labs ships Muse Image and previews Muse Video
MSL's first media-generation models: Muse Image is live in the Meta AI app, Instagram Stories (US) and WhatsApp, with agentic generation that calls web search and code execution, multi-reference composition, and Instagram social-context conditioning. Muse Video shares the same pretraining base and adds native audio, debuting at #3 on Arena text-to-video while Muse Image lands #2 on image. There is no public API, and public Instagram accounts are opted in to @-mention remixing by default.
Google DeepMind debuts OmniFlash, first of the any-to-any Omni family
OmniFlash — first of Google's any-to-any Omni family — generates videos up to 10 seconds with precise conversational multi-turn editing via the Interactions API: say 'make it daytime' and it redoes light, sky and shadows. Editing Elo 1087 at $0.10 per second of output.
xAI launches Grok Imagine Video 1.5 with faster generation and native audio
xAI launched Grok Imagine Video 1.5 with nearly 2x faster generation, native audio, and a claimed #1 leaderboard position. The episode grouped it with Gemini Omni as part of the week’s video-generation frontier.
xAI releases Grok Imagine Video 1.5 Preview with synced audio
xAI released a preview of Grok Imagine Video 1.5, an image-to-video model that generates clips with synchronized audio. It adds xAI to the week's crowded race of media-generation model updates.
Runway launches Project Luxo for solo-creator short films
Runway launched Project Luxo, claiming AI-generated video has crossed the uncanny valley for solo-creator short films. The pitch is that a single creator can now produce watchable short-form films end to end with Runway's stack.
Gemini Omni: 'create anything from anything' conversational video editor
Google DeepMind launched Gemini Omni, a multimodal 'create anything from anything' model debuting as Google's first conversational video editor. Unlike pure text-to-video systems, Omni is an iterative multi-turn editing model that combines Gemini intelligence, world knowledge, multimodal inputs and generative media, in the same way Nano Banana brought Gemini to interactive image editing. It is available in the Gemini app, Google Flow and YouTube, with API support coming soon.
Perceptron Mk1: frontier video + embodied reasoning at 1/10th the price
Perceptron released Mk1, a frontier video and embodied reasoning model priced at roughly a tenth of comparable models. It scores 88.5 on VSI-Bench and 72.4 on RefSpatialBench (versus 9.0 for GPT-5m on the latter) and is live on OpenRouter.
HeyGen HyperFrames integrates natively with Claude Design
HeyGen's HyperFrames now integrates natively with Claude Design, enabling HTML-to-MP4 motion graphics from a single CLI command. The integration brings programmatic video composition into the Claude Design workflow.
Grok Imagine update: better lip sync, sound, 30s video extensions
xAI shipped a Grok Imagine update with dramatically improved lip sync and sound. It also adds 30-second video extensions.
HappyHorse-1.0 takes #1 on Artificial Analysis video arena
HappyHorse-1.0, a mysterious 15B-parameter video model from Alibaba's Taotian Group, took the #1 spot on the Artificial Analysis video arena, beating Seedance 2.0, Kling 3.0, and Grok Video. Little is known about the model beyond its size and leaderboard run.
Seedance 2.0 launches in the US on Replicate
ByteDance's Seedance 2.0 video model became available stateside via Replicate, supporting up to 9 reference images, 3 videos, and 3 audio files per cinematic generation. Peter Gostev confirmed it sits ~80 ELO points above the next video model on Arena, a massive gap in a leaderboard where models usually cluster within 10 points.
Google launches Veo 3.1 Lite at $0.05/sec, cheapest video gen yet
Google released Veo 3.1 Lite, a lighter video generation tier priced at $0.05 per second at 720p, the cheapest video generation offering yet, with further price cuts announced for April 7. The panel framed it as a practical quality-versus-latency tradeoff tier for creator workflows.
Lightricks ships open-source LTX Video 2.3, runs on an RTX 3090
Lightricks released LTX Video 2.3, an open-source video generation model with improved motion, audio, and quality that runs on a single RTX 3090. It is available on GitHub and Hugging Face.
ByteDance Seedance 2.0 shatters video generation reality
ByteDance launched Seedance 2.0, a unified multimodal video generation model that accepts up to 9 images, 3 videos, and 3 audio clips as references and produces 15-second multi-shot clips with native stereo audio and strong character consistency (a 45-second internal test mode also exists). The panel compared the quality jump to seeing Sora for the first time. Available on the BytePlus platform.
LingBot-World: open-source world model challenges Google Genie 3
Ant Group released LingBot-World, an open-source world model that generates 10-minute playable environments at 16fps. It positions open weights as a direct challenger to Google's closed Genie 3 in interactive world generation.
Kling 3.0: 15-second multi-shot video with native audio
Kuaishou's Kling 3.0 launched as an all-in-one AI video creation engine with native multimodal generation, 15-second multi-shot sequences, built-in audio, and character consistency across scenes. Alongside Grok Imagine, it marks the week native audio and lip sync became table stakes for video models.
Grok Imagine 1.0 tops video arena with native audio and lip sync
xAI launched Grok Imagine 1.0 with 10-second 720p video generation, native audio, and lip sync, taking the #1 spot on the Artificial Analysis text-to-video arena. Generation costs roughly $0.42 per 10-second clip and an API is available.
Lucy 2.0 real-time video generation model
Lucy 2.0, a real-time video generation model, was discussed in the AI Art segment. The episode covered its real-time video capabilities.
Google DeepMind launches Project Genie 3, real-time 24fps world model
Google DeepMind's Genie 3 generates interactive, controllable 3D worlds in real time at 24 frames per second, demoed live on the show with a spaceship exploration and paint persistence on walls. It ships alongside SIMA 2, a self-improving game-playing agent built on Genie 3, and is available to Gemini Ultra subscribers in the US with a one-minute session limit.
xAI launches Grok Imagine API with video generation
xAI released the Grok Imagine API, exposing its image and video generation capabilities to developers through the xAI console. The show subtitle notes Grok Imagine ranking #1 among generation models this week.
Overworld's Waypoint-1: real-time AI world model at 60fps on consumer GPUs
Overworld released Waypoint-1, a real-time AI world model that runs at 60fps on consumer GPUs. It generates interactive environments live, bringing world-model tech out of research demos and onto hardware people actually own.
Runway 4.5 launches with image-to-video and audio
Runway launched version 4.5 of its video generation model, adding image-to-video and audio support. It was mentioned in the week's news rundown as part of a busy week for vision and video releases.
KAIST's Avatar Forcing: real-time interactive talking heads
KAIST published Avatar Forcing, a framework for real-time interactive talking-head avatars with approximately 500ms latency. The paper targets responsive, live avatar interaction rather than offline video generation.
Lightricks open-sources LTX-2 synchronized audio-video model
Lightricks open-sourced LTX-2, billed as the first truly open audio-video generation model with synchronized audio and video output, releasing full training code alongside the weights. A distilled version is available to try on Replicate.
VEO3: native audio video generation crosses the uncanny valley
Google's VEO3 stunned everyone in Q2 with video generation that included native audio, which the crew credits with crossing the uncanny valley for AI video. It was a centerpiece of Google IO 2025 and of Google's comeback year.
Sora 2 democratizes video generation and floods the internet with memes
Sora 2 opened Q4 in October by democratizing video generation, complete with a social platform, and spawned a wave of memes still circulating at year's end. The show's TL;DR credits it as part of 2025 crossing the uncanny valley for AI media.
Kling VIDEO 2.6 adds first native audio generation
Kling released VIDEO 2.6, its first video model with native audio generation, producing sound directly alongside generated footage. It was one of two Kling releases this week spanning video and image generation.
Runway Gen-4.5 takes #1 on the text-to-video leaderboard
Runway's Gen-4.5 video model climbed to the top of the text-to-video leaderboard with a 1,247 Elo rating. The result continued the weekly theme of video generation quality and multimodal consistency improving fast.
LTX Studio's Retake brings Photoshop-style object editing to video
LTX Studio launched Retake, an AI video editing tool that enables inpainting-style editing of specific objects within video frames. Wolfram called it 'the image editing moment for video' — Photoshop for video, available to try on Replicate.
Tencent releases HunyuanVideo 1.5, a lightweight open video model
Tencent released HunyuanVideo 1.5, a lightweight DiT-based open-source video generation model. It brings capable video generation to a smaller footprint, continuing the trend of open video models closing the gap with closed offerings.
Hailuo 2.3: MiniMax's cinema-grade video generation model
MiniMax's Hailuo team released version 2.3 of its video generation model, pitching cinema-grade output quality. It landed in the same week as MiniMax M2 and Speech 2.6, underlining how broadly MiniMax is shipping across text, voice, and video.
Odyssey V2: real-time interactive AI video you can steer as it generates
Odyssey ML launched V2 of its real-time interactive AI video experience, where the video stream is generated live and responds to user input. The panel grouped it with the week's evidence that video is becoming an interactive product surface rather than a render-and-wait demo.
Sora drops invite requirement and adds Character Cameos
OpenAI removed the invite requirement for the Sora app and shipped Character Cameos, letting users create reusable characters that can appear across generated videos. The update widens access to Sora as OpenAI pushes it as a consumer video product.
Decart ships real-time lip-sync API for live AI avatars
Decart AI released a real-time lip-sync API that modifies an avatar's video frames to match generated speech on the fly. Kwindla Kramer broke down the pipeline on the show: WebRTC audio capture, Whisper transcription, an LLM response, ElevenLabs voice generation, then Decart's model syncing the avatar's lips, all at sub-two-second latency, a key step toward interactive, believable AI characters.
Krea open-sources a 14B real-time video generation model
Krea AI open-sourced a 14-billion-parameter real-time video model, with weights on Hugging Face. It joins the week's clear trend of generative video racing toward live, interactive experiences rather than offline rendering.
LTX-2: native 4K audio+video generation engine from Lightricks
Lightricks announced LTX-2 as breaking news on the show: a video generation engine producing native 4K video (no upscaling) with synchronized audio, positioned as a fast, efficient open alternative to closed models like Sora. It is billed as open-source with weights coming this fall.
Reve quietly surfaces an unannounced 1080p video mode with sound
Reve's unannounced video mode was spotted this week, generating 1080p video with sound. It was covered briefly in the show's vision and video roundup with no official announcement or links yet.
Baidu's MuseStreamer pushes video generations past 20 seconds
Baidu showed off MuseStreamer, a video generation model producing clips longer than 20 seconds. It adds another Chinese lab to the long-form video generation race alongside Veo and Sora.
Veo 3.1: Google's next-gen video model launches with cinematic audio
Google DeepMind shipped Veo 3.1, the next version of its video generation model with improved quality and cinematic audio. Senior PM Jessica Gallegos joined the show to discuss how the model and its product packaging (including Flow) are evolving video generation into a real user experience story.
Sora extends generations to 15s (25s Pro) and adds storyboards
OpenAI upgraded Sora with longer generations, up to 15 seconds for standard users and 25 seconds for Pro, plus a new storyboard feature for multi-shot control. The update keeps Sora competitive as video models race on length and controllability.
Wan Animate brings open-weights character animation and replacement
Alibaba's Wan team released Wan 2.2 Animate, an open-weights model that animates a character image from a performance video, replicating motion and expressions, or swaps a character into existing footage. It landed in the episode's closing run of video releases showing multimodal product quality climbing across the board.
Kling 2.5 Turbo upgrades AI video generation quality and cost
Kuaishou's Kling AI shipped Kling 2.5 Turbo, an update to its video generation model with better motion, prompt adherence, and cinematic quality at a lower price. Together with Wan Animate it was cited on the show as proof that video model quality is being turbocharged this season.
HuMo: human-centric multimodal video generation from ByteDance/Tsinghua
ByteDance research and Tsinghua released HuMo, a human-centric video generation model that conditions on multimodal inputs (text, image, and audio) to produce videos of people. The weights are available on Hugging Face.
Luma's Ray3: a 'reasoning' video model with native HDR
Luma AI launched Ray3, a video generation model it bills as a 'reasoning' video model, with native HDR output, a fast Draft Mode, and Hi-Fi mastering. It is available in Luma's Dream Machine and feeds the episode's closing theme of a next wave of video models.
Mirage debuts as the first AI-native UGC game engine
Dynamics Lab unveiled Mirage, billed as the world's first AI-native user-generated-content game engine, with real-time photorealistic playable demos powered by world-model-style generation. Alex reacted to it live as the most visibly fun demo of the week and a preview of where interactive media is headed.
Odyssey debuts real-time interactive AI video at 30 FPS
Odyssey launched interactive video: real-time AI world exploration rendered at 30 FPS, letting you walk through generated worlds as they are created. A glimpse at world-model-driven media where the video responds to you instead of just playing back.
Tencent's HunyuanPortrait animates portraits from a single photo
Tencent's Hunyuan team published HunyuanPortrait, a model for high-fidelity portrait video generation from a single photo. It animates a still portrait into realistic talking-head video, with an accompanying paper.
Tencent releases HunyuanVideo-Avatar for audio-driven avatars
Tencent Hunyuan released HunyuanVideo-Avatar, an audio-driven full-body avatar animation model. Feed it audio and a reference image and it animates a full-body avatar in sync, pushing AI-generated humans further toward indistinguishable.
Alibaba's Wan 2.1: open-source diffusion-transformer text-to-video suite
Alibaba, the team behind the Qwen LLMs, released Wan 2.1, a full stack of open-source diffusion-transformer text-to-video foundation models. Amid the show's discussion of video-model fatigue, this was called out as a release that cuts through the noise, with weights on Hugging Face and code on GitHub.
LTX distilled model enables near real-time video generation
Lightricks shared a distilled version of its LTX video model that generates video at near real-time speeds. It was highlighted in the vision and video segment as a notable speed milestone for video generation.
Runway References brings character and scene consistency to Gen-4
Runway launched References for Gen-4 on all paid plans, letting creators supply reference images (characters, outfits, locations, even selfies) and use tags in prompts to keep those elements consistent across generations. It tackles AI video's biggest pain point, frame-to-frame identity drift, at no extra credit cost per run.
Character.AI opens early access to AvatarFX talking avatars
Character.AI announced AvatarFX, now in early access, which turns static images into speaking, emoting video avatars. It targets bringing characters to life for conversational and creative use cases.
FramePack generates 120-second videos on just 6GB of VRAM
FramePack, from ControlNet creator Lvmin Zhang (lllyasviel), is an open source next-frame prediction approach for long video generation that runs on consumer hardware. It can generate videos up to 120 seconds long on as little as 6GB of VRAM by packing input frame context into a fixed length.
Sand AI surprises with MAGI-1, a 24B streaming autoregressive video model
Sand AI released MAGI-1, a 24B autoregressive diffusion model for long-form, streaming video generation with remarkable character consistency, often the Achilles' heel of AI video. It predicts video in 24-frame chunks with causal attention between them, enabling real-time streaming generation where compute doesn't scale with length. Nisten speculated it could be a major step toward usable AI-generated movies by solving the face/character consistency problem.
ByteDance publishes Seaweed-7B video generation foundation model
ByteDance publicly presented Seaweed-7B, a 7B parameter video generation foundation model, showing competitive video quality from a comparatively small model. Details and demos were published at seaweed.video.
Veo 2 video generation hits GA in the API and Gemini App
Google made Veo 2 video generation generally available for developers and rolled it out in the Gemini App. The GA release brings Google's flagship text-to-video model out of preview and into production use.
Kling 2.0 Creative Suite launches
Kuaishou's Kling AI launched Kling 2.0 along with a broader Creative Suite, upgrading its video generation model and tooling. The release kept up the rapid pace in the closed-source video generation race during a packed vision and video week.
Test-Time Training paper one-shots minute-long videos with consistent characters
Researchers published 'One-Minute Video Generation with Test-Time Training', adding TTT layers to a pre-trained transformer to one-shot generate minute-long videos with remarkable character and scene consistency. The Tom & Jerry style demos showed the most impressive long-form AI video consistency to date.
ByteDance's OmniHuman image-to-avatar model goes public via Dreamina
ByteDance's impressive OmniHuman model, which turns a single image plus audio into a realistic talking avatar video, became publicly usable through the Dreamina (CapCut) website. The results land squarely in uncanny-valley territory, as Alex demonstrated with his own avatar thread.
Meta's MoCha generates movie-grade talking AI characters from speech and text
Meta GenAI researchers published MoCha, a model that generates stunningly realistic, movie-grade talking characters directly from speech plus text. Co-author Cong Wei joined the show to discuss the work, which points at AI actors entering Hollywood-quality territory.
Runway Gen-4 announced with major gains in video consistency
Runway announced Gen-4, its next-generation video model focused on character and world consistency across shots. Example videos showed notably coherent characters and scenes, pushing AI video further toward usable filmmaking.
StepFun releases Step-Video-TI2V image-to-video model
Chinese lab StepFun dropped Step-Video-TI2V, an open text/image-to-video generation model. Weights are on Hugging Face with code on GitHub, adding another open-weights option to the fast-moving video generation space.
OpenSora 2.0: 11B open-source video model trained for $200K
OpenSora 2.0 is an 11B parameter open-source video generation model that claims state-of-the-art results while costing only about $200,000 to train. The team claims performance approaching OpenAI's Sora on some benchmarks, underscoring how fast open-source video generation is improving.
Remade AI releases 8 open LoRA video effects for Wan 2.1
Remade AI published eight LoRA video effects for Alibaba's Wan 2.1 14B image-to-video model, including effects like squish, inflate, deflate, and cakeify. The open release shows video effects becoming trainable and customizable via LoRAs on top of open video models.
Tencent releases HunyuanVideo-I2V open image-to-video model
Tencent finally shipped the long-awaited image-to-video version of HunyuanVideo, with open weights on Hugging Face and a hosted try-it experience. It lets users animate still images using one of the strongest open video generation models.
Google's Veo 2 video model becomes available via FAL API
Google DeepMind's Veo 2 video generation model became accessible to developers through FAL's inference API. This was the first broadly available API access to Veo 2, letting builders generate high-quality video from text prompts without waiting on Google's own product surfaces.
Hao AI Lab's FastVideo makes HunyuanVideo 3x faster with no extra training
Hao AI Lab released FastVideo, a method that makes HunyuanVideo (HY-Video) three times faster with no additional training, using a technique called Sliding Tile Attention that outperforms even flash attention for this workload. Faster inference makes open-source video models far more practical, and it supports HY-Video LoRAs for fine-tuned applications.
Microsoft MUSE generates playable game worlds from a single second of video
Microsoft's MUSE can generate minutes of playable gameplay from just a single second of video frames and controller actions, preserving screen elements like health bars and percentages. It is based on the World and Human Action Model (WHAM) architecture, trained on a billion gameplay images from Xbox, with the model released on Hugging Face.
StepFun open-sources Step-Video-T2V, a SOTA 30B text-to-video model
StepFun released Step-Video-T2V (plus a T2V Turbo variant), a 30 billion parameter state-of-the-art text-to-video model under an MIT license. Results impressed especially on text integration, such as rendering 'We will open source' on a scroll as a character unfurls it, marking one of the strongest open-source video drops of the week.
Alibaba launches Qwen2.5-Max flagship model with hidden video gen
Alibaba's Qwen team released Qwen2.5-Max, a large MoE flagship model available through the Qwen Chat interface and API, claiming competitive results against DeepSeek V3 and other frontier models. The chat app also quietly shipped a video generation capability powered by Alibaba's Tongyi Wanxiang.
Follow Video Generation and everything else in AI — live every Thursday.