SRPO: New RL Method from Kuaishou 10x More Efficient Than GRPO for LLM Training
Researchers from Kuaishou introduced SRPO, a two-stage reinforcement learning method that achieves DeepSeek-R1-Zero level performance on math and code tasks in 1/10 of the training steps. The method addresses cross-domain conflicts and inefficient sample utilization issues inherent in standard GRPO.

