Models RSS

Models 🇨🇳

From GPT-2 to Kimi K3: Scale Grows 22,580x in Seven Years, Main Architecture Thread Builds a "Memory Operating System"

The technical evolution of large models from GPT-2 to Kimi K3 shows how AI architecture moves from "remembering everything" to "selective memory". KV Cache solves generation efficiency, Linear Attention explores more efficient long-term memory, DeltaNet and Kimi Linear enable models to learn to update, forget, and manage information. Kimi K3 has 2.8 trillion parameters, equivalent to about 22,580 GPT-2 models.

Moonshot AIMoonshot AI
InfoQ 中国30.07 · 13:02
Fresh news