Zhipu.AI Opens Sources GLM: Models 8x Faster Than DeepSeek-R1 and Global Expansion Ahead of Possible IPO
Chinese company Zhipu.AI has announced the open-sourcing of its new model series GLM-4 and GLM-Z1. The inference model GLM-Z1-32B-0414 runs up to eight times faster than DeepSeek-R1, achieving 200 tokens per second on consumer GPUs. It also launched the international portal Z.ai, and all models are released under the MIT license.
Chinese company Zhipu.AI made a strategic move by announcing the full open-sourcing of its new generation models, the GLM-4 series and the inference model GLM-Z1. The flagship model GLM-Z1-32B-0414, according to the company, achieves inference speeds up to eight times faster than DeepSeek-R1 through optimizations in Grouped Query Attention (GQA) parameters, quantization, and speculative sampling. On consumer-grade GPUs, it delivers 200 tokens per second. Also introduced was the "Rumination" model GLM-Z1-Rumination-32B-0414, capable of autonomously searching the internet, using tools, and fact-checking. The open-source model set includes the GLM-4-32B-0414 with enhanced agentic abilities, including real-time generation of HTML, CSS, JavaScript, and SVG code. Zhipu also released compact 9-billion-parameter versions of GLM-4 and GLM-Z1 for resource-constrained environments. All models are distributed under the MIT license. For global access, the website Z.ai has been launched, and for enterprise customers, the MaaS platform offers tiered pricing, including the ultra-fast GLM-Z1-AirX, cost-effective GLM-Z1-Air, and free GLM-Z1-Flash. The open-sourcing and launch of the international portal are seen as signals of Zhipu.AI's readiness for a potential initial public offering (IPO).
Source: Synced —
original
