Ant Bailing Releases New Hybrid Reasoning Model Ling-3.0-Flash
Ant Bailing has officially released its new native hybrid reasoning model, Ling-3.0-Flash, featuring 124B total parameters with only 5.1B activated per computation. It matches or surpasses models 2-3 times its size in reasoning, instruction following, and long-context tasks, offering higher efficiency. The model is available on OpenRouter for a limited free week before open-sourcing.
On July 24, Ant Bailing launched the Ling-3.0-Flash model with 124B total parameters and 5.1B activated per step, achieving performance comparable to models 2-3 times larger while reducing computational costs. It is designed for agent applications with over 10,000 real interaction environments and improved self-correction and long-term planning. The model uses a hybrid attention architecture with a 5:1 ratio of KDA linear attention and MLA layers, and upgrades Lightning Attention to KDA for better state tracking. Expert activation per token is reduced from 1/32 to 1/64, boosting efficiency. Engineering improvements include cluster-level caching to reduce first-token latency by 60-80% and a multi-agent collaboration framework for stability.
- Сокращения
- KDA = Kimi Delta Attention — Кими Дельта Аттеншн
- MLA = Multi-Head Latent Attention — Многоголовое латентное внимание
Source: QbitAI 量子位 —
original
