ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
28
Citations
2
Influential Citations
IEEE Transactions on Pattern Analysis and Machine Intelligence
Venue
2024
Year
With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct Preference Optimization (DPO) …
Direct Preference Optimization (DPO) has emerged as a pivotal alternative to traditional Reinforcement Learning from Human Feedback (RLHF) for aligning large language models with human values. This survey arrives at a critical juncture where the field is fragmented across numerous papers, each proposing novel variants or applications. By systematically organizing datasets, theories, variants, and applications, the authors provide a much-needed map of the landscape, enabling researchers to quickly grasp the state of the art and identify gaps.
The paper's significance is amplified by the rapid adoption of DPO in both academic and industrial settings. As LLMs become more capable, ensuring they follow human instructions and reflect societal preferences is paramount. This survey not only consolidates existing knowledge but also highlights the theoretical underpinnings that differentiate DPO from RLHF, such as its implicit reward modeling and avoidance of unstable policy optimization. For practitioners, it serves as a practical guide to choosing the right DPO variant for specific tasks, from chat assistants to code generation.
The survey's key technical contributions include:
While the abstract does not provide specific quantitative metrics, the survey synthesizes findings from 28 cited papers, indicating that DPO variants often achieve comparable or superior alignment performance to RLHF with reduced computational overhead and simpler training pipelines. For instance, several variants demonstrate improved stability by avoiding explicit reward modeling and policy gradient methods. The survey also notes that DPO's performance is highly dependent on dataset quality and diversity, with some variants showing robustness to noisy labels. However, exact numbers are not available in the abstract, so readers are encouraged to consult the full text for detailed comparisons.
This survey is poised to become a standard reference for anyone working on LLM alignment. By demystifying DPO and its ecosystem, it lowers the barrier to entry for new researchers and provides a structured foundation for future innovations. The paper also underscores the shift from complex RLHF pipelines to simpler, more efficient preference optimization methods, which could democratize alignment techniques beyond large labs. As the field evolves, this survey will likely be cited as a key resource, shaping both research directions and practical deployments of aligned AI systems.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba