ResearchModels 🇺🇸 28.07.2026 00:03

DeepSeek hints at R2 model and introduces new inference scaling method SPCT

DeepSeekDeepSeek OpenAIOpenAI
DeepSeek has published a study proposing the Self-Principled Critique Tuning (SPCT) method to improve the scalability of general reward models during inference. At the same time, the company hinted at the upcoming release of its next model, R2, which is expected to be based on pure reinforcement learning approaches.
Deep reinforcement learning has become a key factor in improving large language models after the pretraining stage, especially during the inference phase. DeepSeek and researchers from Tsinghua University have proposed the Self-Principled Critique Tuning (SPCT) method, which includes two stages: rejection fine-tuning for cold start, enabling the model to generate principles and critique in the correct format, and rule-based online reinforcement learning for further optimization. Parallel sampling is used for efficient scaling, and a meta-reward model helps select the best reward. Experiments show that SPCT significantly improves the quality and scalability of general reward models, outperforming existing methods. At the same time, the company hinted at the development of the next generation model R2, which will likely leverage these innovations. The article also notes that the scaling paradigm for LLMs is shifting from pretraining to post-training and the inference stage, as demonstrated by OpenAI's o1 model, which generates a long internal chain of thought before responding.
Source: Synced — original
Our earlier posts on this topic ↓
Fresh news