Alibaba's (Tongyi Wanxiang) dense 5B video generation model unifying text-to-video and image-to-video at 720p/24 fps on a single consumer GPU; Apache 2.0 licensed.
Release date
28 July 2025
Access:DownloadDeployment:💻 Local☁ Cloud
Overview
Access & deployment
Download
LocalCloud
Weights: Open source
Key parameters
📥 Input: text, image
Technical specification
Modalities
⬇ Input
textimage
⬆ Output
video
Capabilities and applications
Native model capabilities
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Text-to-video generation
Generating video sequences from a text prompt (and optionally an image), with coherent scene dynamics.
Category: video
Image-to-video
The model's ability to animate a static input image — extending it in time into a consistent video clip according to a description of motion or action.
Category: video
Multimodal understanding
Category: multimodal
