ByteDance (Seed) video generation model: text-to-video and image-to-video, multi-shot 1080p clips with smooth motion; version 2.5 (2026) — up to 30 s, many references, local editing.
Release date
1 June 2025
Access:APIHostedDeployment:☁ Cloud
Overview
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📥 Input: text, image
Technical specification
License
Proprietary
Hardware requirements
A hosted (cloud) model. Seedance 1.0: a 5-second 1080p video generated in about 41.4 s on an NVIDIA L20 (via a ~10× speedup from distillation). Available via API and ByteDance apps and Volcano Engine.
Modalities
⬇ Input
textimage
⬆ Output
video
Capabilities and applications
Native model capabilities
Text-to-video generation
The model's ability to create video clips directly from a text prompt, with control over motion, framing, style and clip length.
Category: video
Image-to-video
The model's ability to animate a static input image — extending it in time into a consistent video clip according to a description of motion or action.
Category: video
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Multi-shot video generation
The model's ability to generate narratively coherent sequences composed of multiple shots, preserving character identity, style and continuity across shot transitions.
Category: video
Technical architecture
Core Architecture
Training Techniques
