Qwen3.8 Max: Alibaba's strongest model more often guesses and costs more per task
Alibaba's Qwen3.8 Max achieves 56 points in the Artificial Analysis Intelligence Index, 10 more than Qwen3.7 Max, tying with Claude Opus 4.8. However, it is more expensive per task ($1.14 vs $0.53) and shows increased hallucination rates (40% vs 23%) on the AA-Omniscience test.
Alibaba/Qwen
Anthropic
Moonshot AI
The Decoder (DE)06.08 · 16:01


