The Value of RL in Fine-Tuning
Gokul Swamy, Sanjiban Choudhury, Wen Sun, et al.
This paper investigates why two-stage RL fine-tuning outperforms offline MLE, finding that RL's value lies in searching over a reduced policy space defined by simple verifiers.