Qwen-Image-2.1 Open-Source Image Generation & Editing Model
Open-sourced by Alibaba Cloud's Qwen team and described as a unified text-to-image generation and editing model in the Qwen family: its visual generation component has just 7B parameters (32 Single-Stream DiT layers), balancing generation quality, inference efficiency and versatility; it supports native transparency (RGBA) generation and editing, can extract subjects from photographs, supports up to 10 reference images and local edits, and improves typography, portrait lighting and fine details, available on Hugging Face and usable with the diffusers QwenImage21Pipeline
Tool Overview
Features, steps and FAQ below
Features
- ✓ Compact and efficient: the team says Qwen-Image-2.1's visual generation component has just 7B parameters in 32 Single-Stream DiT layers and uses mixed-granularity attention and prefix KV cache reuse to deliver strong image quality at low computational cost
- ✓ Native transparency and unified creation and editing: the team says you can generate regular or transparent (RGBA) images from text, edit transparent layers and extract subjects from photographs, all in one model
- ✓ Versatile editing: the team says it supports up to 10 reference images, specifies local edits via circles, painted annotations or separate masks, and preserves identity for people and products
- ✓ Realistic textures and refined aesthetics: the team says it improves typography, portrait lighting and fine details for more visually compelling results, with supported aspect ratios including 1:1, 4:3, 3:4, 3:2, 2:3, 16:9 and 9:16 up to roughly 2048 pixels
- ✓ Open source and ecosystem: the team says the model publishes open weights on Hugging Face and ModelScope under the Qwen Research License, with a Hugging Face Space demo and local use through the diffusers QwenImage21Pipeline or memory optimization such as CPU offload
How to Use
- Download the Qwen-Image-2.1 weights from Hugging Face or ModelScope and install the dependencies such as torch, transformers, diffusers and accelerate as documented
- Load the model with the diffusers QwenImage21Pipeline, write a prompt, and set the size and number of inference steps to generate an image
- For editing, pass an input image and an editing instruction such as changing the background, or use the recommended RGBA prompt format to generate a transparent-background image
- If memory is limited, enable optimizations such as enable_model_cpu_offload, or try the Hugging Face Space demo online
FAQ
What is Qwen-Image-2.1?
The team describes Qwen-Image-2.1 as a unified text-to-image generation and image editing model in the Qwen family, with a 7B-parameter visual generation component (32 Single-Stream DiT layers) balancing generation quality, inference efficiency and versatility.
What are its key improvements?
The team lists four key improvements: compact and efficient architecture, native transparency with unified creation and editing, versatile editing with up to 10 reference images and local edits, and more realistic textures with refined typography and portrait lighting.
Can it generate images with transparent backgrounds?
Yes. The team says Qwen-Image-2.1 can generate transparent (RGBA) images from text, edit transparent layers and extract subjects from photographs, and it recommends a specific RGBA prompt format for transparent backgrounds.
How do I use the model?
The team says you can download the weights from Hugging Face and ModelScope and run text-to-image, image editing and transparent image generation with the diffusers QwenImage21Pipeline; a Hugging Face Space demo is also available, and CPU offload and similar optimizations help when memory is limited.
What about licensing and supported sizes?
The team says the model is licensed under the Qwen Research License and supports aspect ratios including 1:1, 4:3, 3:4, 3:2, 2:3, 16:9 and 9:16, at resolutions up to roughly 2048 pixels.