Robots Atlas>ROBOTS ATLAS
Qwen-Image-2.1

Qwen-Image-2.1

2.1 · Family: Qwen
Image generation and editing model from Alibaba's Qwen family. Lightweight Single-Stream DiT architecture (~7B) with text-to-image, reference-based editing and native RGBA transparency.
✓ Active✓ Public access⚖ Open weightsImage generationMultimodal📁 Qwen
Parameters
~7B (komponent generacji wizualnej)
parameters
Access:DownloadDeployment:💻 Local☁ Cloud

Overview

Qwen-Image-2.1 is an image generation and editing model developed by the Qwen team (Tongyi Lab) at Alibaba. It belongs to the Qwen-Image family, which began with the base Qwen-Image model (a 20B MMDiT) released in August 2025. Version 2.1 focuses on a lightweight, efficient Single-Stream DiT architecture (around 7B parameters in the visual generation component, 32 layers) with mixed-granularity attention and prefix KV cache reuse.

The model unifies text-to-image generation and image editing in a single system. It supports native transparent (RGBA) image generation, reference-based editing using up to 10 reference images, local editing driven by circles, annotations or masks, identity preservation for people and products, and text rendering within images.

Model weights are publicly available for download on Hugging Face under the Qwen Research License. Unlike the base Qwen-Image (Apache-2.0), version 2.1 is covered by a more restrictive research license.

Classification
Image generationMultimodal
Family: Qwen
Access & deployment
Download
LocalCloud
Weights: Open weights
Key parameters
🧩 Parameters: ~7B (komponent generacji wizualnej)
📥 Input: text, image

Technical specification

Parameters
~7B (komponent generacji wizualnej)
parameters
License
Qwen Research License
Hardware requirements
Requires a GPU with enough memory for a ~7B-class diffusion model; exact requirements depend on resolution and mode (generation vs editing).
Modalities
⬇ Input
textimage
⬆ Output
image

Capabilities and applications

Native model capabilities
Text-to-image generation
A model's ability to create images from a text description (prompt), including control of style, composition, aspect ratio and rendering text within the image.
Category: vision
Image editing
Modifying an existing image based on a text instruction or direct annotations: removing objects, changing style, adding elements, filling in regions (inpainting/outpainting), while preserving the identity of people and scene coherence.
Category: vision

Technical architecture

Core Architecture