연구모델 🇺🇸 27.07.2026 17:04

Beyond Standard LLMs: Linear Attention Hybrids, Text Diffusion Models, Code World Models, and Small Recurrent Transformers

Moonshot AIMoonshot AI MiniMaxMiniMax Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek Google/DeepMindGoogle/DeepMind MistralMistral MetaMeta Hugging FaceHugging Face Allen Institute for AIAllen Institute for AI xAIxAI OpenAIOpenAI IBMIBM NVIDIANVIDIA
The article explores alternatives to standard autoregressive transformer-based LLMs: linear attention hybrids (MiniMax-M1, Qwen3-Next, DeepSeek V3.2, Kimi Linear), text diffusion models, code world models, and small recurrent transformers. The author notes a return to classical attention in MiniMax-M2 and the complexity of linear attention in production.
저자 세바스티안 라슈카(Sebastian Raschka)는 어텐션 메커니즘을 사용하는 표준 자기회귀(autoregressive) 트랜스포머 기반 대형 언어 모델(LLM, Large Language Model)의 대안들을 분석한다. 검토된 접근법 중에는 최근 다시 부각되고 있는 선형 어텐션 하이브리드 모델들이 있다. 여기에는 MiniMax-M1(4,560억 파라미터, MoE(Mixture of Experts) 구조, lightning attention 적용), Qwen3-Next(Gated DeltaNet과 Gated Attention을 3:1 비율로 사용), DeepSeek V3.2(희소 어텐션, 준선형 복잡도), 그리고 역시 Gated DeltaNet을 사용하는 Kimi Linear가 포함된다. 다만 MiniMax-M2는 추론 및 다중 작업(multi-task) 처리에서 정확도 문제가 발생해 표준 어텐션 방식으로 회귀했다. 이 외에도 텍스트 확산(diffusion) 모델, (게임 상태 생성 등에 활용되는) 코드 기반 월드 모델(world model), 그리고 구문 분석 작업에서 GPT-4와 비교되는 15억 파라미터 모델과 같은 소형 순환 트랜스포머(recurrent transformer)도 언급된다.
출처: Sebastian Raschka — 원문
관련 게시물 ↓
새로운 뉴스