Qwen3.8 Max Released: Outperforms Claude Fable 5 in Agentic Index at $6 per Million Tokens
Alibaba/Qwen
Anthropic
Alibaba has released the final version of Qwen3.8 Max, a 2.4-trillion-parameter multimodal model with a million-token context. Independent tests show it surpasses Claude Fable 5 in the Agentic Index, and its API pricing is significantly lower than competitors. However, open weights are still pending.
Alibaba has officially released Qwen3.8 Max, a native multimodal mixture-of-experts model with 2.4 trillion parameters. It supports text, images, and video input, and outputs text, with a context window of up to one million tokens (983,000 input and 131,000 output tokens in thinking mode). The model features hybrid thinking enabled by default, function calling, and built-in tools, and is promoted for programming and professional agentic workflows. Independent testing by Artificial Analysis shows Qwen3.8 Max scores 58.1 in the Intelligence Index (vs. 62.1 for Claude Fable 5) and 58.4 in the Agentic Index (vs. 56.6 for Claude Fable 5), with only two Claude Opus 5 modes ranking higher. Pricing is set at $2 per million input tokens and $6 per million output tokens, with cached reads at $0.25 per million tokens — roughly five times cheaper than Claude Fable 5 on input and 8.3 times cheaper on output. However, the model is verbose: it generated about 150 million tokens on a full Intelligence Index run, about 2.3 times the median (66 million) of comparable models. Speed is about 82 output tokens per second, with first token latency around 2.8 seconds. As of the morning of August 10, open weights for Qwen3.8 Max and Qwen3.8-27B had not yet been published, though Alibaba promised they would appear the week of August 3. The large model requires about 4.8 TB in BF16 or 1.2 TB in 4-bit, making local deployment impractical for most. The article suggests Qwen3.8 Max is suitable for agentic coding, long-document analysis, multi-step processes, and as a backup to Claude, but recommends testing on one's own tasks before switching.
- Abbreviations
- API = Application Programming Interface — программный интерфейс приложения
- BF16 = Brain Floating Point 16 — 16-битный формат с плавающей точкой
- GPU = Graphics Processing Unit — графический процессор
Source: Habr — хаб ИИ —
original
