Alibaba Open-Sources Qwen2.5-VL-32B: Smarter and Lighter
Alibaba/Qwen
Alibaba has open-sourced Qwen2.5-VL-32B-Instruct, a vision-language model with 32 billion parameters, under Apache 2.0. It outperforms Mistral-Small-3.1-24B and Gemma-3-27B-IT, and even surpasses the larger Qwen2-VL-72B-Instruct on multimodal reasoning benchmarks like MMMU and MathVista. The model is optimized via reinforcement learning for better human alignment and mathematical reasoning.
Alibaba's Qwen team announced the release of Qwen2.5-VL-32B-Instruct, a vision-language model with 32 billion parameters, open-sourced under Apache 2.0. The model builds on the Qwen2.5-VL series and is optimized using reinforcement learning. Key improvements include responses more aligned with human preferences, enhanced mathematical reasoning accuracy, and fine-grained image understanding and reasoning. In extensive benchmarking, Qwen2.5-VL-32B-Instruct outperforms comparable-scale SoTA models such as Mistral-Small-3.1-24B and Gemma-3-27B-IT, and even surpasses the larger Qwen2-VL-72B-Instruct on complex multimodal tasks like MMMU, MMMU-Pro, and MathVista. It also achieves top-tier performance in pure text capabilities at the same scale. The team plans to focus next on long and effective reasoning processes for highly complex visual reasoning tasks.
- Сокращения
- VL = Vision-Language — визуально-языковая модель
- SoTA = State-of-the-Art — передовые достижения
- MMMU = Massive Multi-discipline Multimodal Understanding — массовое междисциплинарное мультимодальное понимание
- MMMU-Pro = Massive Multi-discipline Multimodal Understanding Pro — MMMU-Pro (продвинутая версия)
- MathVista = Mathematical Visual Instruction Tuning Assessment — оценка математического визуального обучения
- MM-MT-Bench = Multimodal Multi-Turn Benchmark — мультимодальный многошаговый бенчмарк
Source: Alibaba Qwen —
original
