Robots Atlas>ROBOTS ATLAS
Qwen-Image

Qwen-Image

Qwen-Image (base)ย ยทย Family: Qwen
Open (Apache-2.0) ~20B MMDiT image-generation foundation model from the Qwen team (Alibaba); strong at text rendering, including Chinese, and image editing.
โœ“ Activeโœ“ Public accessโš– Open sourceImage generationMultimodal๐Ÿ“ Qwen
Parameters
20B (MMDiT)
parameters
Release date
4 August 2025
Access:DownloadHostedDeployment:๐Ÿ’ป Localโ˜ Cloud

Overview

Qwen-Image is the base text-to-image generation model developed by the Qwen team at Alibaba (Tongyi Lab) and released on 4 August 2025. It is a 20-billion-parameter MMDiT (Multimodal Diffusion Transformer) foundation model, published under the Apache-2.0 license together with its weights and code.

The model stands out for rendering complex in-image text โ€” including multi-line layouts and paragraph-level semantics โ€” for both alphabetic languages (e.g. English) and logographic languages (e.g. Chinese). Beyond text-to-image generation it supports consistent image editing and image-understanding tasks such as object detection, semantic segmentation, depth and edge (Canny) estimation, novel view synthesis and super-resolution.

The architecture pairs a Qwen2.5-VL text encoder with an MMDiT backbone; the technical report describes aligning the latent representations between Qwen2.5-VL and MMDiT and a progressive training strategy evolving from simple to paragraph-level text inputs. The model is distributed via Hugging Face and ModelScope.

Classification
Image generationMultimodal
Family: Qwen
Access & deployment
DownloadHosted
LocalCloud
Weights: Open source
Key parameters
๐Ÿงฉ Parameters: 20B (MMDiT)
๐Ÿ“ฅ Input: text, image

Technical specification

Parameters
20B (MMDiT)
parameters
License
Apache-2.0
Hardware requirements
A ~20B-class diffusion model โ€” local deployment requires a GPU with substantial VRAM; exact requirements depend on resolution, precision (BF16) and mode (generation vs editing).
Modalities
โฌ‡ Input
textimage
โฌ† Output
image

Capabilities and applications

Native model capabilities
Text-to-image generation
A model's ability to create images from a text description (prompt), including control of style, composition, aspect ratio and rendering text within the image.
Category: vision
Image editing
Modifying an existing image based on a text instruction or direct annotations: removing objects, changing style, adding elements, filling in regions (inpainting/outpainting), while preserving the identity of people and scene coherence.
Category: vision

Technical architecture

Core Architecture