Qwen VLo: Unified multimodal understanding and generation model from Alibaba
Alibaba/Qwen
Alibaba releases Qwen VLo, a unified multimodal model that both understands and generates images. It supports open-ended editing, multilingual instructions, dynamic resolution, and progressive generation from top to bottom and left to right. Available as a preview on Qwen Chat.
Alibaba introduces Qwen VLo, a new unified multimodal model that can both understand and generate images. The model is an upgrade from previous QwenVL series, bridging perception and creation. It uses a progressive generation method, constructing images gradually from left to right and top to bottom. Qwen VLo supports precise content understanding, open-ended instruction-based editing (style transfer, object modification, background change), multilingual instructions (Chinese and English), and dynamic resolution with arbitrary aspect ratios. It can also perform visual perception tasks like depth map prediction, segmentation, and edge detection via editing commands. The model is in preview stage and available through Qwen Chat. Limitations include occasional inaccuracies and inconsistencies.
Source: Alibaba Qwen —
original
