AI SSDs: A Storage Paradigm Shift in LLM Inference
Moonshot AI
NVIDIA
Huawei
Maxio Tech
AI infrastructure is moving beyond raw compute, with inference states becoming a first-class resource. Systems like Moonshot AI's Mooncake and NVIDIA's CMX platform are integrating storage into the real-time data path for token generation. This shift is driving the emergence of AI-specific SSDs designed for low latency, high durability, and intelligent data management.
AI infrastructure competition is shifting from raw GPU count and compute power to a more holistic approach where compute, network, memory, and storage work together per token. This is driven by long-context models, complex reasoning, and multi-agent systems, requiring efficient management of KV Cache and MoE expert weights. Two representative signals are Moonshot AI's Mooncake and NVIDIA's CMX. Mooncake is a KV Cache-centric disaggregated architecture that organizes idle CPU, DRAM, SSD, and NIC resources into a distributed cache pool, improving effective request capacity by 59% to 498% in tests. NVIDIA's CMX Context Memory Storage Platform adds a 'G3.5' layer between HBM and shared storage, providing PB-scale shared context memory within a GPU pod, potentially offering up to 5x improvement in sustained token throughput and power efficiency. These systems highlight that inference state is becoming a key resource, and storage must participate in placement, migration, and reuse. Consequently, AI SSDs are emerging, with two categories: AI-enhanced enterprise SSDs (e.g., InnoGrit Dongting-N3X, Huawei OceanDisk LC 560) and inference-participating AI SSDs. Three notable routes for the latter are Phison's aiDAPTIV (using dedicated cache SSDs and middleware), Longsys's SPU+iSA (leveraging storage-side intelligence), and Infplane-Maxio's cross-layer coordination. The competition in AI SSDs will ultimately be about the complete data path, requiring collaboration across NAND media, FTL, firmware, drivers, middleware, and inference frameworks. AI SSDs will not replace HBM, VRAM, or DRAM but will form a finer storage hierarchy, with data placed based on access temperature and migration managed by software and hardware.
- Abbreviations
- GPU = Graphics Processing Unit — графический процессор
- HBM = High Bandwidth Memory — память с высокой пропускной способностью
- KV Cache = Key-Value Cache — кэш ключ-значение
- MoE = Mixture of Experts — смесь экспертов
- DRAM = Dynamic Random Access Memory — динамическая память с произвольным доступом
- SSD = Solid State Drive — твердотельный накопитель
- NIC = Network Interface Card — сетевая интерфейсная карта
- TTFT = Time To First Token — время до первого токена
- TBT = Time Between Tokens — время между токенами
- RDMA = Remote Direct Memory Access — удаленный прямой доступ к памяти
- NVMe = Non-Volatile Memory Express — энергонезависимая память Express
- VRAM = Video Random Access Memory — видеопамять
- SLO = Service Level Objective — целевой уровень обслуживания
- NAND = Not AND (logic gate), refers to NAND flash memory — флеш-память NAND
- FTL = Flash Translation Layer — уровень трансляции флэш-памяти
- IOPS = Input/Output Operations Per Second — операций ввода/вывода в секунду
- SPU = Storage Processing Unit — процессор хранения
- iSA = Intelligent Storage Architecture — интеллектуальная архитектура хранения
- DWPD = Drive Writes Per Day — записей на диск в день
- CBRN = Chemical, Biological, Radiological, Nuclear — химическая, биологическая, радиологическая, ядерная угроза
Source: QbitAI 量子位 —
original
