Qwen3.8 Max: Alibaba's strongest model more often guesses and costs more per task
Alibaba/Qwen
Anthropic
Moonshot AI
Alibaba's Qwen3.8 Max achieves 56 points in the Artificial Analysis Intelligence Index, 10 more than Qwen3.7 Max, tying with Claude Opus 4.8. However, it is more expensive per task ($1.14 vs $0.53) and shows increased hallucination rates (40% vs 23%) on the AA-Omniscience test.
Alibaba's Qwen3.8 Max scores 56 in the Artificial Analysis Intelligence Index, 10 points more than Qwen3.7 Max (46). According to Artificial Analysis, it is on par with Claude Opus 4.8, ahead of GLM-5.2 (51), but behind Kimi K3 (57), which also operates 25% cheaper. In the GDPval-AA work-benchmark, Qwen jumps by 468 Elo to 1739, surpassing Kimi K3 (1685), with only Claude Opus 5 (1852) better. Qwen3.8 Max requires 64 instead of 14 working steps per task, and input tokens increased 15-fold because the test resends the entire conversation history at each step. The model is more thorough but slower and costlier. Although token prices fell (input $2 vs $2.50, output $6 vs $7.50, cache-hit $0.25 vs $0.50 per million), one task in the Intelligence Index costs $1.14, more than double Qwen3.7 Max's $0.53. Kimi K3 costs $0.86 with one more point, GLM-5.2 $0.57. Compared to its predecessor, Qwen3.8 Max also regresses on AA-LCR (-2 points) and AA-Omniscience (-10 points). Its hit rate stays around 31%, but the hallucination rate rises from 23% to 40%, meaning Qwen3.8 Max guesses more often instead of passing.
- Abbreviations
- AA = Artificial Analysis — Artificial Analysis (аналитическая платформа)
- GDPval = GDPval benchmark — бенчмарк GDPval (для оценки работы)
- LCR = Long Context Recall — Воспроизведение из длинного контекста
Source: The Decoder (DE) —
original
