Comparison of Large Language Model Architectures: From DeepSeek V3 to OLMo 2
Sebastian Raschka published a detailed comparison of the architectures of modern open LLMs, highlighting key innovations: Multi-Head Latent Attention (MLA) and Mixture-of-Experts (MoE) in DeepSeek V3, as well as normalization features in OLMo 2. The article covers the evolution from GPT-2 to 2025 models.
DeepSeek
Allen Institute for AI











