Alibaba Qwen Releases QVQ-Max: Visual Reasoning Model with Evidence
Alibaba/Qwen
Alibaba Qwen has officially released QVQ-Max, the first version of its visual reasoning model. It can not only recognize images and videos but also analyze them, solving tasks from mathematics to creativity. The developers noted that the model's accuracy on the MathVision benchmark increases with the length of the thinking process.
Alibaba Qwen has unveiled QVQ-Max, the first version of its visual reasoning model, which goes beyond the preliminary QVQ-72B-Preview released in December. The model understands the content of images and videos, analyzes and reasons based on them, helping to solve problems ranging from mathematical challenges to everyday questions and creative tasks. On the MathVision benchmark, which aggregates complex multimodal mathematical problems, the model's accuracy improves as the maximum length of the reasoning process increases. QVQ-Max has three key capabilities: detailed observation (identifying objects, text, and fine details), deep reasoning (drawing conclusions from visual information and background knowledge), and flexible application (from problem-solving to creating illustrations and scenarios). Demonstration examples include assistance at work (data analysis, code writing), learning (solving problems with diagrams), and daily life (clothing and recipe recommendations). Future plans include improving recognition accuracy through verification mechanisms, developing a visual agent for multi-step tasks (such as controlling a smartphone or computer), and enhancing interaction by adding tools and visual generation capabilities.
Source: Alibaba Qwen —
original
