Qwen-Image
modelYour notes
Outputs 4
Qwen-Image
model20B-parameter image generation foundation model built on Multimodal Diffusion Transformer (MMDiT) architecture. Exceptional text rendering accuracy for complex logographic languages.
Qwen-Image-2.0
modelNext-gen 7B image model unifying text-to-image generation and image editing. Native 2K resolution. Scores 88.32 on DPG-Bench, outperforming FLUX.1 (12B).
Qwen-Image-3.0
modelHosted image-generation and editing model supporting prompts up to 4,500 tokens, native text rendering in 12 languages, and legible type down to roughly 10 pixels. Alibaba did not release weights or disclose scale.
Qwen-Image-2.1
modelOpen-weight unified text-to-image and image-editing model with 7B parameters in its visual generation component (32 single-stream DiT layers). It generates and edits transparent (RGBA) images, extracts subjects from photographs, takes up to 10 reference images, and accepts local edits marked by circles, painted annotations or separate masks. Native 2K output (2048×2048), day-0 Diffusers (QwenImage21Pipeline) support, Qwen Research License; about 2,700 HuggingFace likes and 70,000 downloads in its first ten days.