NVIDIA Groq 3: SRAM, Disaggregation, and Determinism in the New AI Platform
After acquiring Groq for $20 billion, NVIDIA integrated LPU accelerators into the Vera Rubin platform, creating a heterogeneous architecture for inference. The new Groq 3 LPX rack accelerator, based on Groq 3 LP30 chips with 500 MB of SRAM and 150 TB/s memory bandwidth, ensures deterministic execution and low latency. Phase-based disaggregation of decoding (AFD) allows separating FFN computations from attention, distributing the workload between GPUs and LPUs.
NVIDIA
Groq
