Qwen-Image-2.1
Qwen-Image-2.1 is an open-weight image generation and editing model developed by Alibaba’s Qwen team. Released in September 2026, the model unifies text-to-image creation and instruction-based image editing within a single architecture, allowing prompts and multiple reference images to operate seamlessly in the same workflow.
The model features a lightweight visual generation component comprising 7 billion parameters structured around 32 Single-Stream Diffusion Transformer (DiT) layers. This compact design uses mixed-granularity attention and prefix KV cache reuse to deliver strong visual fidelity and native 2K resolution outputs while maintaining high inference efficiency on consumer-grade hardware.
Key capabilities include native transparency generation, allowing users to output images with a real alpha channel (RGBA) directly from text prompts, alongside multi-reference editing that supports up to 10 input images. The architecture excels at maintaining consistency for people and products, rendering precise typography, and executing local edits using annotations or masks.
Create with Crafiq
Generate images, 3D models, video and audio in one studio.
Explore the studioHow Qwen-Image-2.1 ranks
Qwen-Image-2.1 is highlighted in the table below. Switch the metric to see how the ordering changes.