⚡ BREAKING
Hardware & InferenceModels 🇷🇺 27.07.2026 17:04

NVIDIA Blackwell Ultra: 15 Pflops NVFP4, 288 GB HBM3e, no FP64

NVIDIANVIDIA Taiwan Semiconductor Manufacturing CompanyTaiwan Semiconductor Manufacturing Company
NVIDIA unveiled the Blackwell Ultra accelerator, featuring 208 billion transistors, 15 Pflops in NVFP4, 288 GB HBM3e memory, and new fifth-gen Tensor Cores. The chip omits FP64 but introduces NVFP4 format for efficient AI inference.
NVIDIA detailed the Blackwell Ultra accelerator, an enhanced version of the Blackwell architecture. The chip comprises two dies with 208 billion transistors on TSMC 4NP process, connected via NV-HBI with 10 TB/s bandwidth. It includes 160 streaming multiprocessors with 640 fifth-gen Tensor Cores delivering 15 Pflops in NVFP4 (without sparsity) and a unified L2 cache. Each SM has 128 CUDA cores, 4 Tensor Cores with second-gen Transformer Engine, 256 KB of Tensor Memory (TMEM), and special function units. The new NVFP4 format uses micro-block FP8 scaling and tensor-level FP32 scaling, achieving accuracy close to FP8 with 1.8x less memory. Blackwell Ultra boosts attention layer performance by doubling SFU throughput for softmax, reducing time to first token. Memory is 288 GB HBM3e (8 TB/s bandwidth), a 50% increase over Blackwell, enabling large models with 300+ billion parameters. External connectivity includes NVLink 5 (1.8 TB/s) and PCIe 6.0 x16. The Grace Blackwell Ultra superchip combines one Grace CPU with two Blackwell Ultras, offering up to 40 Pflops NVFP4 with sparsity and 1 TB unified memory. The GB300 NVL72 rack system integrates 36 superchips for 1.1 Exaflops FP4.
Сокращения
NV-HBI = NVIDIA High-Bandwidth Interface — высокоскоростной интерфейс NVIDIA
SM = Streaming Multiprocessor — потоковый мультипроцессор
GPC = Graphics Processing Cluster — кластер графической обработки
TMEM = Tensor Memory — тензорная память
SFU = Special Function Unit — блок специальных функций
MMA = Matrix Multiply-Accumulate — матричное умножение с накоплением
NVFP4 = NVIDIA 4-bit Floating Point — 4-битный формат с плавающей запятой NVIDIA
HBM3e = High Bandwidth Memory 3e — память с высокой пропускной способностью 3e
KV = Key-Value — ключ-значение
NVLink = NVIDIA Link — соединение NVIDIA
PCIe = Peripheral Component Interconnect Express — шина PCI Express
C2C = Chip-to-Chip — чип-к-чипу
TPS = Tokens Per Second — токенов в секунду
MW = Megawatt — мегаватт
LPDDR5X = Low Power Double Data Rate 5X — низковольтная память LPDDR5X
Source: ServerNews — original
Our earlier posts on this topic ↓
Fresh news