Robots Atlas>ROBOTS ATLAS
Qwen-Image-2.0

Qwen-Image-2.0

2.0ย ยทย Family: Qwen
Next-generation foundational image generation model from Alibaba's Qwen family, announced on 2026-02-10. Unifies image generation and editing in one mode, with native 2K and professional typography rendering.
โœ“ Activeโœ“ Public accessImage generationMultimodal๐Ÿ“ Qwen
Release date
10 February 2026
Access:HostedDeployment:โ˜ Cloud

Overview

Qwen-Image-2.0 is a foundational image generation model from the Qwen family developed by Alibaba (Tongyi Lab), announced on 10 February 2026 as the next generation after Qwen-Image and Qwen-Image-2512.

Key features

  • Professional typography rendering โ€” supports instructions up to 1k tokens, enabling direct generation of professional infographics, slides (PPT), posters and comics.
  • Stronger semantic adherence โ€” native 2K resolution for finely detailed, realistic scenes (people, nature, architecture).
  • Integrated image generation and editing in a single mode, with improved text rendering.
  • Lighter model architecture โ€” smaller model size with faster inference.

Unlike the open-weights Qwen-Image and Qwen-Image-2512 releases, the weights of Qwen-Image-2.0 have not been published โ€” the model is offered as a hosted service via Qwen Chat. The Qwen-Image family is built on a diffusion transformer (MMDiT) architecture.

Classification
Image generationMultimodal
Family: Qwen
Access & deployment
Hosted
Cloud
Weights: Closed
Key parameters
๐Ÿ“ฅ Input: text, image

Technical specification

Modalities
โฌ‡ Input
textimage
โฌ† Output
image

Capabilities and applications

Native model capabilities
Text-to-image generation
A model's ability to create images from a text description (prompt), including control of style, composition, aspect ratio and rendering text within the image.
Category: vision
Image editing
Modifying an existing image based on a text instruction or direct annotations: removing objects, changing style, adding elements, filling in regions (inpainting/outpainting), while preserving the identity of people and scene coherence.
Category: vision

Technical architecture

Core Architecture