Zhipu.AI Open-Sources GLM: Models 8x Faster Than DeepSeek-R1 and Global Expansion Ahead of Possible IPO
Chinese AI company Zhipu.AI has open-sourced its next-generation GLM models, including the GLM-Z1 inference model that is up to eight times faster than DeepSeek-R1, achieving 200 tokens per second on consumer GPUs. The company also launched an international platform, Z.ai, and is expanding globally, potentially setting the stage for an IPO.
Zhipu.AI announced the comprehensive open-sourcing of its GLM models, including the GLM-4 series and the GLM-Z1 inference model. The GLM-Z1-32B-0414 achieves inference speeds up to eight times faster than DeepSeek-R1, delivering 200 tokens per second on consumer-grade GPUs, which is 50 times faster than human reading speed. This was achieved by optimizing GQA parameters, quantization, and speculative sampling. Zhipu also unveiled the GLM-Z1-Rumination-32B-0414 model, which can autonomously search the internet, use tools, and self-verify information. The open-source portfolio includes the GLM-4-32B-0414 enhanced for agent capabilities and smaller 9B parameter versions. All models are released under the MIT license. Zhipu launched the international platform Z.ai for global users to access these models, and its MaaS platform offers tiered API access. This open-sourcing and global expansion signal Zhipu's readiness for a potential IPO.
- Сокращения
- GLM = General Language Model — Общая языковая модель
- GQA = Grouped Query Attention — Групповое внимание запросов
- GPU = Graphics Processing Unit — Графический процессор
- MIT = Massachusetts Institute of Technology (license) — Лицензия MIT
- MaaS = Model-as-a-Service — Модель как услуга
- API = Application Programming Interface — Интерфейс прикладного программирования
Source: Synced —
original
