ModelsResearch 🇨🇳 28.07.2026 01:05

QwQ-32B: Harnessing the Power of Reinforcement Learning

Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek OpenAIOpenAI
Alibaba Qwen unveiled the QwQ-32B model with 32 billion parameters, which achieves performance comparable to DeepSeek-R1 (671B parameters) through scalable reinforcement learning (RL). The model is open-sourced under the Apache 2.0 license and combines reasoning with agent capabilities.
Alibaba's Qwen has introduced the QwQ-32B model with 32 billion parameters, leveraging scalable reinforcement learning (RL). The model achieves performance comparable to DeepSeek-R1, which has 671 billion parameters (of which 37 billion are active). The development is based on cold start and two-stage RL: first, on mathematical and programming tasks using a correctness verifier and code execution server to validate solutions; then, on general capabilities using a general reward model and rules. QwQ-32B integrates agent capabilities, enabling the model to think critically, use tools, and adapt to feedback from the environment. The model is available as open weights on Hugging Face and ModelScope under the Apache 2.0 license, as well as via Qwen Chat and the DashScope API. In the future, Qwen plans to combine stronger base models with scalable RL computation to approach artificial general intelligence (AGI) and integrate agents with RL for long-term reasoning.
Source: Alibaba Qwen — original
Our earlier posts on this topic ↓
Fresh news