Models RSS

Models 🇨🇳

Domestic Open-Source World-Class Model: What Is the Strength of Kimi K3?

Moonshot AI (月之暗面) has released its new open-source large language model, Kimi K3. The model achieves impressive results on international benchmarks: it ranks first on LiveCodeBench (programming tasks), surpassing GPT-4o and Claude 3.5 Sonnet, and also demonstrates high scores on AIME 2025 (mathematics) and MMLU-Pro (general knowledge).

Moonshot AIMoonshot AI
Moonshot Kimi (GNews)24.07 · 15:04
Research 🇷🇺

RUMBA: A New Benchmark for Evaluating Long-term Memory of Dialogue Systems in Russian

RUMBA (Russian User Memory Benchmark), the first Russian-language benchmark for evaluating long-term memory in multi-session dialogues, has been introduced. It includes 85 dialogues and 1,543 questions testing information retrieval, reasoning, and the ability to refrain from answering. The benchmark accounts for temporal context and enables diagnosis of bottlenecks in memory architecture.

СбербанкСбербанк
Habr — хаб ИИ24.07 · 15:02
Fresh news