Training Language Models to Reason Efficiently
Daman Arora, Andrea Zanette
This paper proposes using reinforcement learning to train large reasoning models to dynamically allocate inference-time compute based on task complexity, reducing inference costs while preserving accuracy.