PreprintarXiv.org2025
Rethinking Reflection in Pre-Training
Essential AI Darsh J Shah, Peter Rushton, Somanshu Singla, et al.
This paper shows that language models' self-reflection ability emerges during pre-training, not just reinforcement learning, by introducing deliberate errors into chains-of-thought.
45Apr 5, 2025Reinforcement LearningReasoning
arXiv