Robots Atlas>ROBOTS ATLAS
Seedance

Seedance

Seedance 2.5 · Family: Seedance
ByteDance (Seed) video generation model: text-to-video and image-to-video, multi-shot 1080p clips with smooth motion; version 2.5 (2026) — up to 30 s, many references, local editing.
✓ Active✓ Public accessVideo generationMultimodal📁 Seedance
Release date
1 June 2025
Access:APIHostedDeployment:☁ Cloud

Overview

Seedance is a family of video generation models developed by the Seed team at ByteDance. The first generation, Seedance 1.0, was released in June 2025; later ones are Seedance 2.0 (February 2026), Seedance 2.0 mini and Seedance 2.5 (July 2026).

The models generate video both from a text prompt (text-to-video) and from an input image (image-to-video). Seedance 1.0 produces 1080p footage with smooth, stable motion, rich detail and cinematic aesthetics, and natively supports multi-shot storytelling that preserves character identity and style across successive shots.

Technically, the models jointly learn text-to-video and image-to-video tasks, building on carefully curated data and detailed video captioning during pretraining, followed by supervised fine-tuning (SFT) and video-specific RLHF with a multi-dimensional reward mechanism. Through multi-stage distillation and system-level optimizations, Seedance 1.0 achieves roughly a 10× inference speedup — a 5-second 1080p video is generated in about 41.4 s on an NVIDIA L20.

Newer generations extend the capabilities: Seedance 2.5 generates native clips of up to 30 seconds with support for many multimodal references and offers local editing of selected scene regions without regenerating the entire shot (a beta mode extends clips up to 3 minutes).

Seedance is a proprietary model available via API and through ByteDance apps (Doubao, Jimeng/Dreamina) and the Volcano Engine platform. Its quality has been evaluated on the internal SeedVideoBench-1.0 benchmark and on the Artificial Analysis platform, among others.

Classification
Video generationMultimodal
Family: Seedance
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📥 Input: text, image

Technical specification

License
Proprietary
Hardware requirements
A hosted (cloud) model. Seedance 1.0: a 5-second 1080p video generated in about 41.4 s on an NVIDIA L20 (via a ~10× speedup from distillation). Available via API and ByteDance apps and Volcano Engine.
Modalities
⬇ Input
textimage
⬆ Output
video

Capabilities and applications

Native model capabilities
Text-to-video generation
The model's ability to create video clips directly from a text prompt, with control over motion, framing, style and clip length.
Category: video
Image-to-video
The model's ability to animate a static input image — extending it in time into a consistent video clip according to a description of motion or action.
Category: video
Video generation
The model's ability to generate video clips from a text prompt, image or another video, with control over length, resolution and visual characteristics.
Category: video
Multi-shot video generation
The model's ability to generate narratively coherent sequences composed of multiple shots, preserving character identity, style and continuity across shot transitions.
Category: video

Technical architecture