A Deep Dive into Reasoning LLMs
Komal Kumar, Tajamul Ashraf, Omkar Thawakar, et al.
A systematic survey of post-training techniques for LLMs, covering fine-tuning, reinforcement learning, and test-time scaling, with a public repository for tracking developments.