Image generation model family. Qwen-Image (20B MMDiT) excels at text rendering in Chinese and English. Qwen-Image-2.0 (7B) unifies generation and editing. Qwen-Image-3.0 (July 2026) is hosted only, and the open-weight Qwen-Image-2.1 (7B, September 2026) adds native transparency and multi-reference editing.

Outputs 4

Qwen-Image

model

20B-parameter image generation foundation model built on Multimodal Diffusion Transformer (MMDiT) architecture. Exceptional text rendering accuracy for complex logographic languages.

Parameters 20B

Qwen-Image-2.0

model

Next-gen 7B image model unifying text-to-image generation and image editing. Native 2K resolution. Scores 88.32 on DPG-Bench, outperforming FLUX.1 (12B).

Parameters 7B

Qwen-Image-3.0

model

Hosted image-generation and editing model supporting prompts up to 4,500 tokens, native text rendering in 12 languages, and legible type down to roughly 10 pixels. Alibaba did not release weights or disclose scale.

Qwen-Image-2.1

model

Open-weight unified text-to-image and image-editing model with 7B parameters in its visual generation component (32 single-stream DiT layers). It generates and edits transparent (RGBA) images, extracts subjects from photographs, takes up to 10 reference images, and accepts local edits marked by circles, painted annotations or separate masks. Native 2K output (2048×2048), day-0 Diffusers (QwenImage21Pipeline) support, Qwen Research License; about 2,700 HuggingFace likes and 70,000 downloads in its first ten days.

Parameters 7B
License Qwen Research License
generationopen-weight