DeepSeek hints at R2 model and introduces new inference scaling method SPCT
DeepSeek has published a study proposing the Self-Principled Critique Tuning (SPCT) method to improve the scalability of general reward models during inference. At the same time, the company hinted at the upcoming release of its next model, R2, which is expected to be based on pure reinforcement learning approaches.
DeepSeek
OpenAI


