ModelsAgents 🇩🇪 03.08.2026 13:02

Alibaba releases Qwen3.8-Max with 2.4 trillion parameters, to open-source weights

Alibaba/QwenAlibaba/Qwen Moonshot AIMoonshot AI
Alibaba has unveiled Qwen3.8-Max, its most powerful language model with 2.4 trillion total parameters, designed to autonomously complete complex tasks over days. In tests, the model coded software, reproduced research results, and ran a simulated e-commerce business. It is available now, with weights to be released next week.
Alibaba's Qwen team has introduced Qwen3.8-Max, a language model scaling to 2.4 trillion total parameters with 95 billion active per request, built on the Qwen3.5 architecture. The focus is on autonomously completing complex tasks over longer periods. Three autonomous coding runs were presented as stress tests: the model built the oh-my-cli tool over 16 days, handling 265 commits, 127 pull requests, and 151 issues without human intervention; it reproduced and improved a research paper's results, exceeding its method on AIME24 by 2.7 points; and it placed ahead of 458 of 526 human teams in a multimodal challenge. Additionally, the model designed a chip, reducing logic gates from 8,298 to 678 and chip area by 81%, and simulated a year of e-commerce, quadrupling capital and outperforming GLM 5.2 by 38%. Multimodal capabilities include processing documents over 200 pages and videos over 100 hours. The model's performance is close to or above Claude Opus 4.8, Claude Fable 5, and GPT-5.6 Sol in internal benchmarks, though independent verification is pending. The weights will be released next week on Hugging Face and ModelScope. The competitor Moonshot AI released Kimi K3 with 2.8 trillion parameters and 1 million token context on July 27.
Abbreviations
GPU = Graphics Processing Unit — графический процессор
Source: The Decoder (DE) — original
Our earlier posts on this topic ↓
Fresh news