Robots Atlas>ROBOTS ATLAS
Seedream

Seedream

Seedream 4.0 · Family: Seedream
ByteDance (Seed) text-to-image model: version 4.0 (2025) unifies generation and editing, 2K/4K, native ZH/EN bilingual, ~1.4 s per 2K image.
✓ Active✓ Public accessImage generationMultimodal📁 Seedream
Release date
1 September 2025
Access:APIHostedDeployment:☁ Cloud

Overview

Seedream is a text-to-image generation model developed by the Seed team at ByteDance. Its latest generation, Seedream 4.0, was released in September 2025 and combines image generation and editing in a single, unified architecture.

Seedream 4.0 generates images at 2K and 4K resolution with flexible aspect ratios, is natively bilingual (Chinese and English, with high-quality text rendering in images) and runs more than 10× faster than Seedream 3.0 — a 2K image is produced in about 1.4 seconds.

The model uses a diffusion transformer architecture with an efficient VAE encoder. It supports multimodal tasks: knowledge-based generation, reference consistency, multi-image composition and reference-guided editing.

Seedream is available in the Doubao and Dreamina apps and via the Volcano Engine API (ByteDance), among others. It is a proprietary model; earlier generations are Seedream 2.0 and Seedream 3.0.

Classification
Image generationMultimodal
Family: Seedream
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📥 Input: text, image

Technical specification

License
Proprietary
Hardware requirements
A hosted (cloud) model; generates a 2K image in about 1.4 s. Available via ByteDance apps and the Volcano Engine API.
Modalities
⬇ Input
textimage
⬆ Output
image

Capabilities and applications

Native model capabilities
Text-to-image generation
Generating an image from a text description (prompt). The model interprets a natural-language instruction and produces a new, coherent visual from scratch — without any input image.
Category: vision
Image editing
Modifying an existing image based on a text instruction or direct annotations: removing objects, changing style, adding elements, filling in regions (inpainting/outpainting), while preserving the identity of people and scene coherence.
Category: vision
Reference-guided generation
Creating images based on previously supplied visual references — a specific person, artistic style, product, or space — preserving likeness and characteristic features.
Category: vision
Text rendering in images
Generating images containing legible, correctly spelled text — infographics, posters, menu cards, QR codes, captions in a specific graphic style. A key capability that separates new-generation models from early image generators.
Category: vision
Multi-image blending
Intelligently combining several input images into one coherent composition — e.g. placing a person from one photo into a scene from another, composing characters from different sources, or building collages that preserve style.
Category: vision

Technical architecture