ModelsResearch 🇨🇳 05.08.2026 09:04

Understanding Kimi K3 Pretraining: A 3T Model, No Longer Magic

Moonshot AIMoonshot AI
Moonshot AI has released details on the pretraining of its Kimi K3 model, a 3 trillion parameter model. The company emphasizes a shift from 'magic' to engineering, highlighting innovative techniques in data, architecture, and infrastructure that enable efficient scaling and strong performance.
Moonshot AI has published a detailed account of the pretraining of Kimi K3, a large language model with 3 trillion parameters. The report frames the achievement not as 'magic' but as the result of systematic engineering across data, model architecture, and infrastructure. Key innovations include a refined data pipeline that improves data quality and diversity, a novel attention mechanism that enhances efficiency, and a training infrastructure optimized for stability and throughput at massive scale. These efforts enable Kimi K3 to achieve competitive performance while maintaining cost-effectiveness. The article underscores the importance of engineering rigor in advancing large-scale AI models and provides insights into the practical challenges of pretraining at the trillion-parameter scale.
Abbreviations
3T = 3 trillion — 3 триллиона
Source: Moonshot Kimi (GNews) — original
Our earlier posts on this topic ↓
Fresh news