ModelsMedia Generation 🇨🇳 27.07.2026 17:05

Alibaba releases Qwen-Image: a 20B MMDiT model for text-accurate image generation

Alibaba/QwenAlibaba/Qwen
Alibaba's Qwen team has released Qwen-Image, a 20-billion-parameter MMDiT image foundation model that excels in complex text rendering (including bilingual Chinese-English), consistent image editing, and strong performance across benchmarks like GenEval and DPG. The model supports poster creation, slide generation, and precise text rendering on small regions.
Alibaba's Qwen team has released Qwen-Image, a 20-billion-parameter MMDiT image foundation model. It achieves state-of-the-art results on benchmarks such as GenEval, DPG, OneIG-Bench for generation, and GEdit, ImgEdit, GSO for editing, and LongText-Bench, ChineseWord, TextCraft for text rendering. The model excels at rendering complex multi-line text in both alphabetic languages (e.g., English) and logographic languages (e.g., Chinese), including handwritten text, posters, PPT slides, and bilingual content. It also supports consistent image editing operations like style transfer, object addition/deletion, and text editing. The model can accurately render small text, such as a paragraph on a paper occupying less than one-tenth of the image. Qwen-Image is available via Qwen Chat and on platforms like GitHub, Hugging Face, and ModelScope.
Сокращения
MMDiT = Mixture-of-Multimodal-Diffusion-Transformers — смесь мультимодальных диффузионных трансформеров
Source: Alibaba Qwen — original
Our earlier posts on this topic ↓
Fresh news