Research 🇷🇺

Reasoning Models Need a New Measurement: Public Leaderboard MERA Reason Unveiled

The MERA Reason leaderboard has been launched to evaluate reasoning models in Russian. It includes four task sets: Luzitania (olympiad-level mathematics), TMath (a wide range of problems), MMReD (dense context reasoning), and an upcoming component. The goal is to objectively measure a model's ability to build chains of reasoning, not just reproduce knowledge.

Т-БанкТ-Банк Alibaba/QwenAlibaba/Qwen
Habr — хаб ИИ24.07 · 11:01
AI Safety 🇷🇺

Prompt Injection Cannot Be Cured by Filters: How to Isolate Untrusted Text Using Architectural Patterns

Prompt injection is a structural defect in large language models where untrusted text is executed as an instruction. Filters and system prompts are ineffective: even the best defenses are breached in 90% of cases under adaptive attacks. Working approaches—Dual-LLM, CaMeL, and the Two Meta rule—build isolation rather than trying to distinguish good from evil.

OpenAIOpenAI AnthropicAnthropic Google DeepMindGoogle DeepMind MicrosoftMicrosoft
Habr — хаб ИИ24.07 · 11:01
Regulation 🇷🇺

Federation Council Approves Law on Experimental Legal Regimes in AI

The Federation Council has adopted a law regulating experimental legal regimes (ELRs) for the development and implementation of artificial intelligence technologies in Russia. The document establishes the legal framework for testing AI products in special zones.

GNews RU — ИИ в России24.07 · 09:03
Fresh news