⚡ BREAKING
Hardware & InferenceMergers & Acquisitions 🇷🇺 27.07.2026 15:06

NVIDIA Groq 3: SRAM, Disaggregation and Determinism in the New AI Platform

NVIDIANVIDIA GroqGroq
NVIDIA's $20 billion acquisition of Groq has resulted in the integration of LPU accelerators into the Vera Rubin platform, enabling disaggregated, heterogeneous AI inference. The new Groq 3 LPX rack system combines 256 LPUs for fast, deterministic token generation, while Vera Rubin NVL72 handles flexible training and inference. The platform targets low-latency interactive AI and multi-agent workloads, promising up to 35x faster performance than Grace Blackwell NVL72 at 400 TPS per user.
NVIDIA has integrated Groq's LPU (Language Processing Unit) architecture into its Vera Rubin AI platform following the $20 billion acquisition. The result is a heterogeneous system where Vera Rubin NVL72 handles general training and inference, while the Groq 3 LPX rack accelerator provides a specialized engine for low-latency token generation. LPX contains 256 Groq 3 (LP30) chips with 96 billion transistors each, connected via RealScale C2C interconnect. The LPU is designed around a deterministic execution model with 500 MB of SRAM per chip (150 TB/s bandwidth), no caches, and explicit data movement controlled by a compiler. This eliminates jitter and ensures predictable latency, crucial for interactive AI and multi-agent systems. The platform uses phase disaggregation (AFD) separating attention and FFN operations: GPU handles prefill and long-context attention, while LPU accelerates FFN/MoE decoding. NVIDIA claims up to 35x speedup over Grace Blackwell NVL72 at 400 TPS per user, and up to 10x revenue per MW for latency-sensitive workloads. The NVIDIA Dynamo software orchestrates the heterogeneous decoding, classifying requests and managing intermediate activations.
Сокращения
SRAM = Static Random-Access Memory — Статическая память с произвольным доступом
LPU = Language Processing Unit — Блок обработки языка
C2C = Chip-to-Chip — Межчиповое соединение
FFN = Feed-Forward Network — Сеть прямого распространения
MoE = Mixture of Experts — Смесь экспертов
AFD = Attention-FFN Disaggregation — Дезагрегация внимания и FFN
TPS = Tokens Per Second — Токенов в секунду
GPU = Graphics Processing Unit — Графический процессор
CPU = Central Processing Unit — Центральный процессор
Source: ServerNews — original
Our earlier posts on this topic ↓
Fresh news