ModelsAgents 🇷🇺 27.07.2026 23:03

How Kimi K3 Works: 2.8 Trillion Parameters, Linear Attention, and Million-Token Agents

Moonshot AIMoonshot AI AnthropicAnthropic OpenAIOpenAI
Moonshot AI has released the weights of its Kimi K3 model, a multimodal model based on a Mixture of Experts architecture with 2.8 trillion parameters, of which 104 billion are activated per token. In certain programming and tool-use benchmarks, the model outperforms competitors, but on average it lags behind some of them.
Moonshot AI has released the weights of its new model Kimi K3, built on a mixture-of-experts (MoE) architecture with a total of 2.8 trillion parameters, of which 104 billion are active per token. On individual benchmarks for programming and tool use, K3 outperforms Claude Opus 4.8, GPT-5.5, GPT-5.6 Sol, and Claude Fable 5; however, the company itself acknowledges in its technical report that, on average, the model falls short of Fable 5 and GPT-5.6 Sol. According to Moonshot's technical blog, the K3 architecture was designed as a unified system: it uses a hybrid of Kimi Delta Attention and global Multi-head Latent Attention (MLA), an Attention Residuals mechanism, an MoE with 896 experts, as well as specialized training on long agent trajectories and a separate caching system. This comprehensive approach enables the model to handle contexts up to a million tokens in length and perform complex agentic tasks. However, some of Moonshot's claims have not yet been verified by independent tests and rely on the company's own data.
Source: Habr — хаб ИИ — original
Our earlier posts on this topic ↓
Fresh news