Business & Market RSS

Open Source 🇺🇸

NVIDIA NeMo AutoModel Accelerates MoE Fine-Tuning 3.7x with Expert Parallelism and DeepEP

NVIDIA has introduced NeMo AutoModel, an open library that builds on Hugging Face Transformers v5, delivering 3.4-3.7x higher training throughput and 29-32% less GPU memory for MoE models via Expert Parallelism, DeepEP fused all-to-all dispatch, and TransformerEngine kernels—using the same from_pretrained() API with only an import change.

NVIDIANVIDIA Hugging FaceHugging Face Moonshot AIMoonshot AI DeepSeekDeepSeek
Hugging Face blog27.07 · 10:04
Fresh news