Preprint
Large Language Models

Post Training of LLMs

Guiyao Tie, Zeli Zhao, D. Song, Fuyang Wei, Rong Zhou, Yurou Dai, Wenpeng Yin, Zhejian Yang, Jiangyue Yan, Yao Su, Zhen Dai, Yifeng Xie, Yihan Cao, Lichao Sun, Pan Zhou, Lifang He, Hechang Chen, Yu Zhang, Qingsong Wen, Tianming Liu, N. Gong, Jiliang Tang, Caiming Xiong, Heng Ji, Philip S. Yu, Jianfeng Gao
March 8, 2025arXiv.org53 citations

53

Citations

2

Influential Citations

arXiv.org

Venue

2025

Year

Abstract

The emergence of Large Language Models (LLMs) has fundamentally transformed natural language processing, making them indispensable across domains ranging from conversational systems to scientific exploration. However, their pre-trained architectures often reveal limitations in specialized contexts, including restricted reasoning capacities, ethical uncertainties, and suboptimal domain-specific performance. These challenges necessitate advanced post-training language models (PoLMs) to address these shortcomings, such as OpenAI-o1/o3 and DeepSeek-R1 (collectively known as Large Reasoning Models, or LRMs). This paper presents the first comprehensive survey of PoLMs, systematically tracing their evolution across five core paradigms: Fine-tuning, which enhances task-specific accuracy; Alignment, which ensures ethical coherence and alignment with human preferences; Reasoning, which advances multi-step inference despite challenges in reward design; Efficiency, which optimizes resource utilization amidst increasing complexity; Integration and Adaptation, which extend capabilities across diverse modalities while addressing coherence issues. Charting progress from ChatGPT's alignment strategies to DeepSeek-R1's innovative reasoning advancements, we illustrate how PoLMs leverage datasets to mitigate biases, deepen reasoning capabilities, and enhance domain adaptability. Our contributions include a pioneering synthesis of PoLM evolution, a structured taxonomy categorizing techniques and datasets, and a strategic agenda emphasizing the role of LRMs in improving reasoning proficiency and domain flexibility. As the first survey of its scope, this work consolidates recent PoLM advancements and establishes a rigorous intellectual framework for future research, fostering the development of LLMs that excel in precision, ethical robustness, and versatility across scientific and societal applications.

Analysis

Why This Paper Matters

This paper addresses a critical gap in the LLM literature: while pre-training has received extensive attention, the post-training phase—which transforms base models into useful, aligned, and reasoning-capable systems—has lacked a comprehensive synthesis. As models like OpenAI-o1 and DeepSeek-R1 demonstrate, post-training is where the magic of reasoning and alignment happens, yet practitioners often rely on fragmented knowledge. This survey provides the first unified view, making it an essential reference for researchers and engineers.

The timing is significant: with the rapid evolution of Large Reasoning Models, the field needs a structured framework to understand how techniques like RLHF, chain-of-thought, and efficient fine-tuning fit together. By charting progress from ChatGPT's alignment strategies to DeepSeek-R1's reasoning innovations, the paper offers a historical and conceptual roadmap that can inform both research directions and practical deployment decisions.

Technical Contributions

  • Five-paradigm taxonomy: The paper systematically categorizes post-training into Fine-tuning, Alignment, Reasoning, Efficiency, and Integration/Adaptation, providing a clear mental model for navigating the landscape.
  • Evolutionary synthesis: It traces the progression of techniques, showing how alignment methods evolved into reasoning-focused approaches, and how efficiency techniques enable scaling.
  • Dataset and technique categorization: The survey includes a structured taxonomy of datasets and methods, which is valuable for practitioners selecting approaches.
  • Strategic agenda for LRMs: It explicitly highlights the role of Large Reasoning Models in improving reasoning proficiency and domain flexibility, setting a research agenda.

Results

As a survey, the paper does not present new experimental metrics. Instead, its 'results' are the synthesis itself: a comprehensive mapping of post-training techniques, identification of trends (e.g., the shift from alignment to reasoning), and a framework for future work. The paper cites 53 references, indicating a broad but not exhaustive coverage. The abstract mentions challenges in reward design for reasoning and coherence issues in multimodal integration, which are key open problems.

Significance

The significance lies in establishing a common language and framework for post-training research. This can accelerate progress by helping researchers identify gaps—such as the need for better reward models for reasoning—and by providing a structured foundation for new entrants. For practitioners, the taxonomy aids in selecting appropriate post-training strategies for specific applications, from domain adaptation to safety alignment. The emphasis on Large Reasoning Models signals a shift in focus from mere task accuracy to deeper cognitive capabilities, which could shape the next generation of LLMs. By consolidating knowledge, this survey fosters a more cohesive and efficient research community, ultimately contributing to LLMs that are more precise, ethically robust, and versatile across scientific and societal applications.