ResearchModels 🇷🇺 27.07.2026 04:05

Finam AI Lab Updates Financial Benchmark FINESSE-Bench for LLMs

AnthropicAnthropic Moonshot AIMoonshot AI Google/DeepMindGoogle/DeepMind MiniMaxMiniMax OpenAIOpenAI MetaMeta DeepSeekDeepSeek MistralMistral СберСбер ЯндексЯндекс Т-БанкТ-Банк
Finam's Artificial Intelligence Laboratory has released an updated version of the financial benchmark FINESSE-Bench. The new version adds a technical analysis dataset CFTe-like Level 1, fixes 169 CFA-like Level 1 questions, improves metric calculation with bootstrapping, and expands the model pool to 33 in comparative tables.
Finam Artificial Intelligence Lab has released an update to the FINESSE-Bench financial benchmark, designed to evaluate the knowledge and skills of large language models (LLMs) in the financial domain. Key changes include: fixing 169 problematic questions in the CFA-like Level 1 dataset; adding a new CFTe-like Level 1 dataset with 781 questions on the fundamentals of technical analysis; transitioning to confidence interval calculation via bootstrapping for more robust model comparison; using stratified bootstrapping to aggregate results by benchmark groups; separate assessment of dataset discriminative power and saturation (it was found that most FINESSE questions fall into the intermediate zone, providing good diagnostic value). The model pool has been expanded: a total of 66 systems were tested, of which 49 have full coverage across all 11 benchmarks (8 FINESSE + 3 public), and 33 models are included in the main tables. Top results were achieved by GPT-5.5 (0.906 on exam-like, 0.876 on trading/TA), Claude Opus 4.8 (0.860 on exam-like, 0.856 on trading/TA), and others.
Source: Habr — хаб NLP — original
Our earlier posts on this topic ↓
Fresh news