Robots Atlas>ROBOTS ATLAS
Wan2.2

Wan2.2

Wan2.2 · Family: Wan
Alibaba’s open video-generation model (2025): the first open-source one with an MoE architecture (27B/14B). T2V and I2V, 720P@24fps, a 5B variant on RTX 4090. Apache 2.0.
✓ Active✓ Public access⚖ Open sourceVideo generationMultimodal📁 Wan
Parameters
27B (14B aktywnych, MoE); wariant TI2V-5B
parameters
Release date
28 July 2025
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud

Overview

Wan2.2 is an open video-generation model developed by the Wan team at Alibaba, released on 28 July 2025. According to the developer, it is the first open-source video-generation model to use a Mixture-of-Experts (MoE) architecture.

The MoE model combines two specialized experts: a “high-noise” expert for early denoising stages (scene layout) and a “low-noise” expert for later detail refinement. It totals 27 billion parameters, of which 14 billion are active per step. It also uses a custom Wan2.2-VAE encoder with a high compression ratio.

Wan2.2 includes variants: T2V-A14B (text-to-video, 480P/720P), I2V-A14B (image-to-video), TI2V-5B (unified text/image-to-video, 720P@24 fps), plus the later S2V-14B (speech-to-video) and Animate-14B (character animation). The TI2V-5B variant generates a 5-second 720P video in under 9 minutes on a single consumer GPU (e.g. an RTX 4090).

The model is released under the Apache 2.0 license (code and weights on GitHub and Hugging Face), permitting commercial and non-commercial use, self-hosting and fine-tuning. Wan2.2 belongs to Alibaba’s Wan (Tongyi Wanxiang) family of video models.

Classification
Video generationMultimodal
Family: Wan
Access & deployment
DownloadAPIHosted
LocalCloud
Weights: Open source
Key parameters
🧩 Parameters: 27B (14B aktywnych, MoE); wariant TI2V-5B
✓ Fine-tuning
📥 Input: text, image

Technical specification

Parameters
27B (14B aktywnych, MoE); wariant TI2V-5B
parameters
License
Apache 2.0
Hardware requirements
The TI2V-5B variant generates a 5-second 720P video in <9 min on a single consumer GPU (e.g. RTX 4090); the A14B variants require more powerful hardware.
Features:Fine-tuning
Modalities
⬇ Input
textimage
⬆ Output
video

Capabilities and applications

Native model capabilities
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Image-to-video
The model's ability to animate a static input image — extending it in time into a consistent video clip according to a description of motion or action.
Category: video

Technical architecture

Deployment and security

☁ Available on platforms