Qwen-Image: Image Generation Model with Accurate Text Rendering
Alibaba/Qwen
Alibaba released Qwen-Image, a 20-billion parameter foundational image generation model. It outperforms competitors in complex text rendering, including multi-line captions and Chinese characters, and supports precise image editing.
Alibaba (via its Qwen division) has announced Qwen-Image, a foundational image generation model with an MMDiT architecture and 20 billion parameters. Key capabilities include superior text rendering (support for multi-line layouts, paragraphs, alphabetic and logographic languages), consistent image editing while preserving semantics and realism, and strong performance on public benchmarks (GenEval, DPG, OneIG-Bench, GEdit, ImgEdit, GSO, LongText-Bench, ChineseWord, TextCraft), outperforming existing models. Demos show examples of accurate reproduction of Chinese characters on signs, scrolls, books, posters, as well as English text (including long passages in small spaces). The model supports poster and presentation creation, working with various artistic styles, and editing operations (stylization, adding/removing objects, changing text and poses).
Source: Alibaba Qwen —
original
