ByteDance (Seed) text-to-image model: version 4.0 (2025) unifies generation and editing, 2K/4K, native ZH/EN bilingual, ~1.4 s per 2K image.
Release date
1 September 2025
Access:APIHostedDeployment:☁ Cloud
Overview
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📥 Input: text, image
Technical specification
License
Proprietary
Hardware requirements
A hosted (cloud) model; generates a 2K image in about 1.4 s. Available via ByteDance apps and the Volcano Engine API.
Modalities
⬇ Input
textimage
⬆ Output
image
Capabilities and applications
Native model capabilities
Text-to-image generation
Generating an image from a text description (prompt). The model interprets a natural-language instruction and produces a new, coherent visual from scratch — without any input image.
Category: vision
Image editing
Modifying an existing image based on a text instruction or direct annotations: removing objects, changing style, adding elements, filling in regions (inpainting/outpainting), while preserving the identity of people and scene coherence.
Category: vision
Reference-guided generation
Creating images based on previously supplied visual references — a specific person, artistic style, product, or space — preserving likeness and characteristic features.
Category: vision
Text rendering in images
Generating images containing legible, correctly spelled text — infographics, posters, menu cards, QR codes, captions in a specific graphic style. A key capability that separates new-generation models from early image generators.
Category: vision
Multi-image blending
Intelligently combining several input images into one coherent composition — e.g. placing a person from one photo into a scene from another, composing characters from different sources, or building collages that preserve style.
Category: vision
Application domains
Technical architecture
Core Architecture
Training Techniques
