ModelsResearch 🇨🇳 28.07.2026 17:03

QwQ: Deep Reflection on the Boundaries of the Unknown

Alibaba/QwenAlibaba/Qwen
Alibaba's Qwen team has released QwQ-32B-Preview, an experimental reasoning model that excels in math and coding but has limitations like language mixing and recursive loops.
The Qwen Team at Alibaba released QwQ-32B-Preview, an experimental reasoning model designed to think deeply and question assumptions before answering. It achieves strong benchmark scores: 65.2% on GPQA, 50.0% on AIME, 90.6% on MATH-500, and 50.0% on LiveCodeBench. The model exhibits language mixing, recursive reasoning loops, and requires enhanced safety measures. It is available on GitHub, Hugging Face, and ModelScope.
Сокращения
GPQA = Graduate-Level Google-Proof Q&A Benchmark — бенчмарк вопросов-ответов на уровне аспирантуры, не поддающийся поиску Google
AIME = American Invitation Mathematics Evaluation — Американская пригласительная математическая оценка
MATH-500 = 500 test cases of the MATH benchmark — 500 тестовых случаев бенчмарка MATH
Source: Alibaba Qwen — original
Our earlier posts on this topic ↓
Fresh news