RSS Search All 🟢 Status
Models 🇨🇳 30.07.2026 00:02

How Kimi K3 Reached the Top: 2.8 Trillion Parameters Training Journey

Moonshot AIMoonshot AI
Moonshot AI's Kimi K3 model achieves leading performance with 2.8 trillion parameters. The model uses advanced training techniques including Mixture of Experts and large-scale parallel computing.
Moonshot AI has developed Kimi K3, a large language model with 2.8 trillion parameters, achieving top performance in benchmarks. The training process involved a Mixture of Experts (MoE) architecture and massive parallel computing across thousands of GPUs. The model demonstrates significant improvements in reasoning, coding, and multilingual capabilities. The achievement highlights China's growing capabilities in large-scale AI model development.
Сокращения
MoE = Mixture of Experts
GPU = Graphics Processing Unit
Source: Moonshot Kimi (GNews) — original
Our earlier posts on this topic ↓
Fresh news