Qwen2.5: A Party of Foundation Models!
Alibaba/Qwen
Meta
Mistral
OpenAI
Anthropic
DeepSeek
Microsoft
Google/DeepMind
Alibaba's Qwen team announces the open-source release of Qwen2.5 LLMs, along with specialized models for coding (Qwen2.5-Coder) and mathematics (Qwen2.5-Math). The models range from 0.5B to 72B parameters, pretrained on up to 18 trillion tokens, with significant improvements in knowledge, coding, and math capabilities. The release also includes API models Qwen-Plus and Qwen-Turbo.
Alibaba's Qwen team released Qwen2.5, a family of open-source decoder-only language models, along with specialized models Qwen2.5-Coder and Qwen2.5-Math. The models are available in sizes from 0.5B to 72B parameters, pretrained on up to 18 trillion tokens. Qwen2.5 achieves MMLU 85+, HumanEval 85+, and MATH 80+, supports up to 128K tokens context and 8K generation, and offers multilingual support for over 29 languages. Qwen2.5-Coder was trained on 5.5 trillion tokens of code data, while Qwen2.5-Math supports CoT, PoT, and TIR reasoning. The flagship API models Qwen-Plus and Qwen-Turbo are also offered. The 3B and 72B variants are not Apache 2.0 licensed.
- Сокращения
- LLM = Large Language Model — большая языковая модель
- MMLU = Massive Multitask Language Understanding
- MATH = Mathematics — математика
- CoT = Chain-of-Thought — цепочка мыслей
- PoT = Program-of-Thought — программа мыслей
- TIR = Tool-Integrated Reasoning — рассуждение с интеграцией инструментов
- API = Application Programming Interface — программный интерфейс приложения
Source: Alibaba Qwen —
original
