Critique-out-loud reward models
Unknown
Introduces Critique-out-Loud (CLoud) reward models that are trained to generate critiques, improving reward modeling for LLMs.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Introduces Critique-out-Loud (CLoud) reward models that are trained to generate critiques, improving reward modeling for LLMs.
Unknown
This paper critiques the dominant 'preferentist' approach to AI alignment and proposes a reframing of alignment targets beyond human preferences.
Unknown
CriticGPT uses RLHF to train a GPT-4-based model that critiques ChatGPT code outputs, helping humans catch bugs more accurately.
Unknown
Proposes a multi-agent debate framework where multiple language models iteratively critique and refine answers to improve factual accuracy and reasoning.
Unknown
Critique Fine-Tuning trains models to critique noisy responses, improving reasoning more than imitating correct answers.
Unknown
Proposes Direct Judgement Preference Optimization to enhance LLM judges via Chain-of-Thought Critique, Standard Judgement, and Response Deduction for rating, comparison, and classification.