ModelsResearch 🇺🇸 28.07.2026 00:02

Spring for Open LLMs: 10 Architectures in January–February 2026

Arcee AIArcee AI Moonshot AIMoonshot AI Alibaba/QwenAlibaba/Qwen MiniMaxMiniMax CohereCohere
A review of ten open LLMs from spring 2026, focusing on architectural similarities and differences. Key models include Trinity Large, Kimi K2.5, Step 3.5 Flash, Qwen3-Coder-Next, and others.
The author presents an overview of ten open LLMs released in January-February 2026, focusing on architectural details. Arcee AI released Trinity Large (400B MoE, 13B active) with local-global attention (3:1), QK-Norm, NoPE, gated attention, and four RMSNorm layers. Moonshot AI introduced Kimi K2.5 (1 trillion parameters) — a multimodal extension of K2 trained on 15 trillion visual and text tokens with early fusion. Step 3.5 Flash from StepFun (196B MoE, 11B active) achieved 100 tokens per second at a context length of 128k using MTP-3. Qwen3-Coder-Next (80B, 3B active) outperformed larger models in coding tasks due to a hybrid of Gated DeltaNet and Gated Attention (3:1). Also mentioned are GLM-5 from z.AI, MiniMax M2.5, Nanbeige 4.1 3B, Qwen 3.5, Ling 2.5 and Ring 2.5 from Ant Group, Tiny Aya from Cohere, as well as the Sarvam 30B and 105B models.
Source: Sebastian Raschka — original
Our earlier posts on this topic ↓
Fresh news