Beyond Standard LLMs: Linear Attention Hybrids, Text Diffusion Models, Code World Models, and Small Recurrent Transformers
The article explores alternatives to standard autoregressive transformer-based LLMs: linear attention hybrids (MiniMax-M1, Qwen3-Next, DeepSeek V3.2, Kimi Linear), text diffusion models, code world models, and small recurrent transformers. The author notes a return to classical attention in MiniMax-M2 and the complexity of linear attention in production.
Moonshot AI
MiniMax
Alibaba/Qwen
DeepSeek
Google/DeepMind
Mistral
Meta
Hugging Face
Allen Institute for AI
xAI
OpenAI
IBM
NVIDIA
