Mistral

Últimas notícias, modelos e lançamentos de IA de Mistral. ['AB-UPT', 'Connectors', 'Devstral', 'Devstral 2', 'Devstral 2 Small', 'Devstral Small 2', 'Document AI', 'Emmi AI', 'Forge', 'GyroSwin', 'Leanstral', 'Leanstral 1.5', 'Leanstral (предыдущая версия)', 'Le Chat', 'Le Chat (Vibe)', 'Magistral', 'MiMo-V2-Flash', 'Ministral 3', 'Ministral-3-3B-Instruct-2512', 'Ministral-3-8B-Instruct-2512', 'Ministral 3B', 'Mistral', 'Mistral 3', 'Mistral 7B', 'Mistral-7B', 'Mistral Large', 'Mistral Large 2', 'Mistral Large 3', 'Mistral Medium 3.1', 'Mistral Medium 3.5', 'Mistral Nemo', 'Mistral OCR 3', 'Mistral OCR 4', 'Mistral OCR4', 'Mistral Search Toolkit', 'mistral-small-2603', 'Mistral Small 3', 'Mistral Small 3.1', 'Mistral Small 3.1 24B', 'Mistral-Small-3.1-24B']

Pesquisa 🇺🇸

Além dos LLMs Padrão: Híbridos de Atenção Linear, Modelos de Difusão de Texto, Modelos de Mundo de Código e Pequenos Transformers Recorrentes

O artigo explora alternativas aos LLMs padrão baseados em transformers autorregressivos: híbridos de atenção linear (MiniMax-M1, Qwen3-Next, DeepSeek V3.2, Kimi Linear), modelos de difusão de texto, modelos de mundo de código e pequenos transformers recorrentes. O autor observa um retorno à atenção clássica no MiniMax-M2 e a complexidade da atenção linear em produção.

Moonshot AIMoonshot AI MiniMaxMiniMax Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek Google/DeepMindGoogle/DeepMind MistralMistral MetaMeta Hugging FaceHugging Face Allen Institute for AIAllen Institute for AI xAIxAI OpenAIOpenAI IBMIBM NVIDIANVIDIA
Sebastian Raschka27.07 · 17:04
Notícias frescas