ModelsOpen Source 🇷🇺 28.07.2026 07:02

Kimi K3: Moonshot AI opens weights of a 2.8-trillion-parameter model

Moonshot AIMoonshot AI DeepSeekDeepSeek AnthropicAnthropic OpenAIOpenAI
Moonshot AI released the open weights of Kimi K3, a 2.8-trillion-parameter MoE model, along with a technical report and three internal tools. Independent benchmarks rank it first among open models with 57 points on the Artificial Analysis index, though practical deployment requires multi-GPU servers due to 1.4 TB weight size.
On July 27, Moonshot AI released the weights of Kimi K3, a 2.8-trillion-parameter MoE model with 896 experts (16 active per token, ~50B active parameters) and 1M token context. It scores 57 on the Artificial Analysis index, fourth overall and first among open models, surpassing GLM-5.2 (51) and DeepSeek V4 Pro (44). It also topped LMArena Frontend Code Arena, beating Claude Fable 5 in 76% of blind comparisons. The architecture uses Kimi Delta Attention and Attention Residuals, claimed to provide 2.5x scaling efficiency over K2 per the company's internal report. Weights occupy about 1.4 TB even in MXFP4 quantization, requiring multi-GPU server configurations. Moonshot also opened FlashKDA (CUDA kernels for attention), MoonEP (expert parallelism library), and AgentENV (sandbox environment for agent training). The license is custom, requiring brand display for products with over 100M monthly active users or $20M monthly revenue, and separate agreements for MaaS providers. Support is announced by Together AI, Fireworks, DigitalOcean, Nebius, Baseten, Modal, and vLLM. API pricing is $3 per million input tokens, $15 per million output, and $0.30 for cached input.
Сокращения
MoE = Mixture of Experts — смесь экспертов
API = Application Programming Interface — программный интерфейс
MaaS = Model as a Service — модель как услуга
GPU = Graphics Processing Unit — графический процессор
CUDA = Compute Unified Device Architecture — архитектура CUDA
Source: Habr — хаб ИИ — original
Our earlier posts on this topic ↓
Fresh news