Tencent’s open-source video generation model (>13B params): text-to-video (image-to-video via HunyuanVideo-I2V), DiT + causal 3D VAE, 540p–720p, 129 frames.
Parameters
13B+
parameters
Release date
3 December 2024
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud
Overview
Access & deployment
DownloadAPIHosted
LocalCloud
Weights: Open weights
Key parameters
🧩 Parameters: 13B+
✓ Fine-tuning
📥 Input: text, image
Technical specification
Parameters
13B+
parameters
License
Tencent Hunyuan Community License Agreement
Hardware requirements
An open-weights model (Hugging Face) with code on GitHub; requires a large-VRAM GPU for local inference (FP8 quantized versions are also available to reduce requirements). Also available via Tencent Cloud.
Features:✓ Fine-tuning
Modalities
⬇ Input
textimage
⬆ Output
video
Capabilities and applications
Native model capabilities
Text-to-video generation
The model's ability to create video clips directly from a text prompt, with control over motion, framing, style and clip length.
Category: video
Image-to-video
The model's ability to animate a static input image — extending it in time into a consistent video clip according to a description of motion or action.
Category: video
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Technical architecture
Core Architecture
Training Techniques
