ResearchModels 🇷🇺 24.07.2026 15:02

RUMBA: New Benchmark for Evaluating Long-Term Memory of Dialogue Systems in Russian

СбербанкСбербанк
A new Russian-language benchmark called RUMBA (Russian User Memory Benchmark) has been created to evaluate long-term memory in multi-session dialogues. It covers three supergroups: Information Extraction, Reasoning, and Abstention, with detailed taxonomy including semantic type, session scope, and temporality.
The RUMBA benchmark was developed to assess long-term memory of dialogue systems in Russian, addressing the lack of such benchmarks for the language. It consists of 85 dialogues with 1,543 questions, averaging 191 days of conversation and about 340,000 characters per dialogue. The benchmark evaluates three supergroups: Information Extraction (finding and updating facts), Reasoning (multi-step inference), and Abstention (knowing when not to answer). Each question is tagged with semantic type, session scope (single or multi-session), and temporality (temporal or atemporal), with additional temporal expression tags. Dialogues were created from scratch, with user utterances written by humans and assistant responses generated by GigaChat Max. The dataset includes Russian and English versions for cross-lingual comparison.
Сокращения
RUMBA = Russian User Memory Benchmark — Русский бенчмарк пользовательской памяти
RAG = Retrieval-Augmented Generation — генерация с дополнением через поиск
LLM = Large Language Model — большая языковая модель
Source: Habr — хаб ИИ — original
Our earlier posts on this topic ↓
Fresh news