Hardware & Inference RSS

Hardware & Inference 🇨🇳

AMD Unveils Helios with World's First 2nm GPU MI455X: OpenAI, Meta, and Microsoft Among Customers

AMD has launched the Helios rack-scale system featuring the Instinct MI455X accelerator, claimed to be the first GPU with 2nm chiplets. Customers include OpenAI, Meta, Microsoft, and Oracle, with shipments expected from Q3 2026. The system uses an open architecture with UALink and Ultra Ethernet, aiming to challenge Nvidia's dominance.

OpenAIOpenAI MetaMeta MicrosoftMicrosoft OracleOracle AnthropicAnthropic Taiwan Semiconductor Manufacturing CompanyTaiwan Semiconductor Manufacturing Company
InfoQ 中国27.07 · 06:02

Google Accelerates Gemini Nano Models on Pixel with Frozen Multi-Token Prediction

Google has introduced a method to retrofit Multi-Token Prediction (MTP) onto frozen production models like Gemini Nano v3, enabling faster on-device inference on Pixel 9 and 10 devices. The MTP head attaches to the main model's final layers, leveraging its hidden states and KV cache to generate multiple tokens per inference pass without separate drafting models, achieving speedups of 50% or more and reducing memory consumption by 130MB per instance.

Google/DeepMindGoogle/DeepMind Google DeepMindGoogle DeepMind
Google Research27.07 · 05:05
Fresh news