ModelsResearch 🇨🇳 28.07.2026 16:03

Alibaba Qwen Releases ProcessBench and Improved Process Reward Models for Mathematical Reasoning

Alibaba/QwenAlibaba/Qwen OpenAIOpenAI
Alibaba's Qwen team has open-sourced ProcessBench, a benchmark for detecting errors in mathematical reasoning steps, and released two new Process Reward Models (PRMs) based on Qwen2.5-Math. The models outperform existing open-source PRMs and show strong error identification capabilities.
Alibaba's Qwen team introduced ProcessBench, a benchmark with 3,400 test cases focused on competition- and Olympiad-level math problems, where human experts annotate the earliest erroneous step in step-by-step solutions. They also released two new Process Reward Models (PRMs), Qwen2.5-Math-PRM-7B and Qwen2.5-Math-PRM-72B, fine-tuned from Qwen2.5-Math-Instruct models. In Best-of-N evaluation (N=8) across seven math benchmarks, Qwen2.5-Math-PRM-7B outperformed majority voting by an average of 1.4%, and matched or exceeded other PRMs of similar scale. On ProcessBench, Qwen2.5-Math-PRM-7B surpassed GPT-4o-0806 and all open-source models, lagging only behind o1-mini. The team also noted that the Outcome Reward Model Qwen2.5-Math-RM-72B showed surprising step-error detection ability, surpassing some PRMs. The work aims to improve process supervision in LLM reasoning.
Сокращения
LLM = Large Language Model — большая языковая модель
PRM = Process Reward Model — модель вознаграждения процесса
BoN = Best-of-N — лучший из N
ORM = Outcome Reward Model — модель вознаграждения результата
Source: Alibaba Qwen — original
Our earlier posts on this topic ↓
Fresh news