Robots Atlas>ROBOTS ATLAS
Wan2.2-TI2V-5B

Wan2.2-TI2V-5B

Wan2.2-TI2V-5B · Family: Wan
Alibaba's (Tongyi Wanxiang) dense 5B video generation model unifying text-to-video and image-to-video at 720p/24 fps on a single consumer GPU; Apache 2.0 licensed.
✓ Active✓ Public access⚖ Open sourceVideo generationMultimodal📁 Wan
Release date
28 July 2025
Access:DownloadDeployment:💻 Local☁ Cloud

Overview

Wan2.2-TI2V-5B is a video generation model developed by Alibaba's Tongyi Wanxiang lab, released on 28 July 2025 together with inference code and weights under the Apache 2.0 licence. Unlike the larger variants of the Wan2.2 family (T2V-A14B and I2V-A14B, both MoE models with 27B parameters and 14B active), TI2V-5B is dense with 5B parameters, and its name reflects the unification of two tasks: text-to-video and image-to-video in a single model.

At the model's core sits the high-compression Wan2.2-VAE with a T×H×W compression ratio of 4×16×16; an additional patchification layer raises total compression to 4×32×32. This lets the model generate 720p video at 24 frames per second (1280×704 or 704×1280) on consumer hardware — a five-second clip takes under 9 minutes without further optimization, with a minimum of 24 GB VRAM (e.g. RTX 4090). Running it requires PyTorch 2.4.0 or newer.

Beyond video content generation, the model found use in robotics: the Chinese company PsiBot adopted it as the lower-layer backbone of its Psi-R2.5 foundation model, where it generates robot action trajectories from the planning layer's output and the current observation. PsiBot's materials sometimes write the model as "Wan2.2-IT2V-5B" — a transposition of the letters in the official TI2V name.

Classification
Video generationMultimodal
Family: Wan
Access & deployment
Download
LocalCloud
Weights: Open source
Key parameters
📥 Input: text, image

Technical specification

Modalities
⬇ Input
textimage
⬆ Output
video

Capabilities and applications

Native model capabilities
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Text-to-video generation
Generating video sequences from a text prompt (and optionally an image), with coherent scene dynamics.
Category: video
Image-to-video
The model's ability to animate a static input image — extending it in time into a consistent video clip according to a description of motion or action.
Category: video
Multimodal understanding
Category: multimodal