The Art of Scaling RL Compute for LLMs
Devvrit Khatri, Lovish Madaan, Rishabh Tiwari, et al.
This paper presents the first large-scale systematic study (400,000+ GPU-hours) defining a framework for predicting RL scaling in LLMs, and proposes a best-practice recipe, ScaleRL, enabling extrapolation from small runs.