Open (Apache-2.0) ~20B MMDiT image-generation foundation model from the Qwen team (Alibaba); strong at text rendering, including Chinese, and image editing.
Parameters
20B (MMDiT)
parameters
Release date
4 August 2025
Access:DownloadHostedDeployment:๐ป Localโ Cloud
Overview
Access & deployment
DownloadHosted
LocalCloud
Weights: Open source
Key parameters
๐งฉ Parameters: 20B (MMDiT)
๐ฅ Input: text, image
Technical specification
Parameters
20B (MMDiT)
parameters
License
Apache-2.0
Hardware requirements
A ~20B-class diffusion model โ local deployment requires a GPU with substantial VRAM; exact requirements depend on resolution, precision (BF16) and mode (generation vs editing).
Modalities
โฌ Input
textimage
โฌ Output
image
Capabilities and applications
Native model capabilities
Text-to-image generation
A model's ability to create images from a text description (prompt), including control of style, composition, aspect ratio and rendering text within the image.
Category: vision
Image editing
Modifying an existing image based on a text instruction or direct annotations: removing objects, changing style, adding elements, filling in regions (inpainting/outpainting), while preserving the identity of people and scene coherence.
Category: vision
Technical architecture
Core Architecture
