Reasoning Models Need a New Measurement: Public Leaderboard MERA Reason Unveiled
The MERA Reason leaderboard has been launched to evaluate reasoning models in Russian. It includes four task sets: Luzitania (olympiad-level mathematics), TMath (a wide range of problems), MMReD (dense context reasoning), and an upcoming component. The goal is to objectively measure a model's ability to build chains of reasoning, not just reproduce knowledge.
Т-Банк
Alibaba/Qwen
Habr — хаб ИИ24.07 · 11:01
