QVQ: To See the World with Wisdom
Alibaba/Qwen
Alibaba's Qwen team releases QVQ-72B-Preview, an open-weight multimodal reasoning model built on Qwen2-VL-72B. It achieves 70.3 on MMMU and shows strong performance on math/science benchmarks, though it has limitations like language mixing and recursive reasoning.
The Qwen team at Alibaba has released QVQ-72B-Preview, an experimental open-weight model for multimodal reasoning, built upon Qwen2-VL-72B. QVQ achieves a score of 70.3 on the MMMU benchmark and shows substantial improvements on math-related benchmarks like MathVista, MathVision, and OlympiadBench compared to its predecessor. However, the model has several limitations including language mixing, recursive reasoning, safety concerns, and potential hallucinations in multi-step visual reasoning. The team plans to integrate additional modalities into a unified model in the future.
- Сокращения
- MMMU = Massive Multi-discipline Multimodal Understanding
Source: Alibaba Qwen —
original
