SRPO: Kuaishou's New RL Method 10x More Efficient Than GRPO for LLM Training
Researchers from Kuaishou's Kwaipilot team have introduced SRPO, a two-stage reinforcement learning framework that achieves DeepSeek-R1-Zero-level performance in math and code domains with only one-tenth of the training steps. SRPO addresses cross-domain optimization conflicts, low sample efficiency, and premature saturation through staged training and history resampling.
Synced27.07 · 18:04

