NVIDIA Vera Rubin: Maximum Intelligence per Dollar for Agentic AI
NVIDIA
Prime Intellect
Perplexity AI
Together AI
NVIDIA introduced the Vera Rubin platform, designed for continuous post-training of agentic AI models. The key metric becomes 'intelligence per dollar', combining token cost and training efficiency. The platform promises a fourfold reduction in GPU count compared to Blackwell.
NVIDIA는 블로그에서 에이전트 AI의 경우 포스트 트레이닝(post-training)이 더 이상 일회성 단계가 아니라 지속적인 프로세스라고 설명합니다. 생성형 모델과 달리 에이전트 모델은 계획을 세우고 도구를 사용하며 변화하는 조건에 적응해야 합니다. 포스트 트레이닝은 순방향 및 역방향 패스(forward and backward passes)를 포함하는 사이클로 이루어지며, 모델은 강화 학습(Reinforcement Learning, RL)을 통해 기술을 향상시킵니다. 핵심 지표는 '달러당 지능(intelligence per dollar)'으로, 이는 추론(inference) 시의 토큰당 비용(cost per token)과 학습 효율성을 결합한 것입니다. NVIDIA의 플랫폼 Vera Rubin(Blackwell의 후속작)은 이 시나리오에 맞게 설계되어, 4분의 1의 GPU로도 가장 큰 모델을 학습할 수 있습니다. 예시로는 5,500억 개의 파라미터(MoE 아키텍처)를 가진 Nemotron 3 Ultra 모델이 제시되었으며, SWE-bench verified에서 71.7%의 성능을 보였습니다. Vera Rubin 플랫폼은 NeMo 라이브러리(NeMo Gym, NeMo RL)와 통합되어 있으며, 수천 가지 환경의 병렬 배포에 최적화되어 있습니다. 파트너로는 Prime Intellect(Vera 프로세서에서 x86 대비 30%의 처리량 향상), Perplexity(2초 미만의 RDMA 가중치 동기화를 제공하는 RL 스택), 그리고 Together AI(AI Native Cloud 플랫폼에서 포스트 트레이닝을 서비스로 제공)가 있습니다.
출처: NVIDIA blog —
원문
