RUMBA: A New Benchmark for Evaluating Long-term Memory of Dialogue Systems in Russian
RUMBA (Russian User Memory Benchmark), the first Russian-language benchmark for evaluating long-term memory in multi-session dialogues, has been introduced. It includes 85 dialogues and 1,543 questions testing information retrieval, reasoning, and the ability to refrain from answering. The benchmark accounts for temporal context and enables diagnosis of bottlenecks in memory architecture.
Сбербанк
Habr — хаб ИИ24.07 · 15:02

