Preprint2024
Token-level direct preference optimization
Unknown
This paper introduces Token-level Direct Preference Optimization (TDPO), a method that extends DPO to optimize at the token level for improved alignment of LLMs with human preferences.
0Apr 1, 2024Large Language ModelsFine Tuning
arXiv