Rm-bench: Benchmarking reward models of language models with subtlety and style
Unknown
RM-Bench is a novel benchmark for evaluating reward models of language models, focusing on subtle distinctions and stylistic preferences.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
RM-Bench is a novel benchmark for evaluating reward models of language models, focusing on subtle distinctions and stylistic preferences.
Unknown
Diffusion-DPO aligns diffusion models to human preferences by directly optimizing the model on user preference data using Direct Preference Optimization.
Wenyi Xiao, Zechuan Wang, Leilei Gan, et al.
A comprehensive survey of Direct Preference Optimization covering datasets, theories, variants, and applications for aligning LLMs with human preferences.
Unknown
This paper introduces Token-level Direct Preference Optimization (TDPO), a method that extends DPO to optimize at the token level for improved alignment of LLMs with human preferences.
Shunyu Liu, Wenkai Fang, Zetian Hu, et al.
This survey systematically reviews Direct Preference Optimization (DPO), a streamlined alternative to RLHF for aligning LLMs with human preferences, covering its variants, applications, and future directions.
Unknown
This paper introduces -DPO, a dynamic preference optimization method that improves DPO by adapting the reference model during training to better align LLMs with human preferences.
Unknown
This paper applies representation engineering to identify and steer internal representations of high-level human preferences in LLMs, improving alignment.
Unknown
This paper critiques the dominant 'preferentist' approach to AI alignment and proposes a reframing of alignment targets beyond human preferences.
Elias Fernández Domingos, The Anh Han
An AI race experiment shows unsafe development is driven by competitive dynamics and fear of falling behind, not just risk preferences.
Unknown
DPO fine-tunes LLMs to align with human preferences directly via a simple classification loss, eliminating RL complexity.
Unknown
ReST iteratively aligns LLMs with human preferences by alternating between data generation and offline RL fine-tuning with increasing quality thresholds.
David A Cook, Joshua Overgaard, V Shane Pankratz, et al.
LLM-powered virtual patients can generate authentic clinician-patient dialogues, represent preferences, and provide personalized feedback at low cost.