Image generation model family. Qwen-Image (20B MMDiT) excels at text rendering in Chinese and English. Qwen-Image-2.0 (7B) unifies generation and editing.

Outputs 3

Qwen-Image

model

20B-parameter image generation foundation model built on Multimodal Diffusion Transformer (MMDiT) architecture. Exceptional text rendering accuracy for complex logographic languages.

Parameters 20B

Qwen-Image-2.0

model

Next-gen 7B image model unifying text-to-image generation and image editing. Native 2K resolution. Scores 88.32 on DPG-Bench, outperforming FLUX.1 (12B).

Parameters 7B

Qwen-Image-3.0

model

Hosted image-generation and editing model supporting prompts up to 4,500 tokens, native text rendering in 12 languages, and legible type down to roughly 10 pixels. Alibaba did not release weights or disclose scale.

generationopen-weight