Hardware & Inference RSS

Alibaba Cloud: Zhenwu M890 Chip Adapts to Qwen3.8, Model Goes Live on Bailian for Inference

Alibaba Cloud announced that its Zhenwu M890 supernode chip has successfully adapted the Qwen3.8 model, enabling inference on the Bailian platform. The Zhenwu M890 is a new-generation AI chip from T-Head, supporting FP32 to FP4 precision and achieving 800GB/s interconnect via ICN Switch 1.0, allowing a single instance to run a 2.4-trillion-parameter model.

Alibaba/QwenAlibaba/Qwen
QbitAI 量子位24.07 · 01:03
Open Source 🇺🇸

Gigatoken: BPE Tokenizer in Rust Encodes Text at 24.53 GB/s, Up to 989x Faster Than HuggingFace Tokenizers

Gigatoken, a BPE tokenizer written in Rust by Stanford PhD student Marcel Rød, achieves text encoding speeds of 24.53 GB/s on a 144-core AMD EPYC system, outperforming HuggingFace tokenizers by 989x and OpenAI's tiktoken by 681x. The library supports 23 tokenizer families and attributes its speed to a hand-optimized pretokenizer using SWAR SIMD techniques and pretoken caching.

OpenAIOpenAI Hugging FaceHugging Face Google/DeepMindGoogle/DeepMind MetaMeta Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek Moonshot AIMoonshot AI NVIDIANVIDIA
MarkTechPost24.07 · 01:03
Fresh news