ResearchModels 🇷🇺 24.07.2026 11:01

Reasoning Models Need a New Measurement: Public Leaderboard MERA Reason Unveiled

Т-БанкТ-Банк Alibaba/QwenAlibaba/Qwen
The MERA Reason leaderboard has been launched to evaluate reasoning models in Russian. It includes four task sets: Luzitania (olympiad-level mathematics), TMath (a wide range of problems), MMReD (dense context reasoning), and an upcoming component. The goal is to objectively measure a model's ability to build chains of reasoning, not just reproduce knowledge.
Developers have launched a public leaderboard, MERA Reason, for evaluating reasoning models in Russian, consisting of four task sets. The first set, Luzitania, includes 251 complex Olympiad-level mathematical problems, filtered by difficulty using the GPT-OSS-120B model; the problems have been translated into Russian and require long chains of reasoning. The second set, TMath, contains 310 problems of varying difficulty from Russian Olympiads (All-Russian and Moscow), focusing on combinatorics, number theory, and geometry; the median reasoning length for Qwen3-32B is about 16,000 tokens. The third set, MMReD, is a synthetic benchmark for reasoning in dense contexts: it requires aggregating information across an entire sequence of steps (up to 128 steps) without using external tools. The leaderboard is part of a major update to the textual MERA and serves as a preview of the next platform release.
Source: Habr — хаб ИИ — original
Our earlier posts on this topic ↓
Fresh news