Preprint2024
Self-Taught Evaluators
Unknown
Proposes an iterative self-training method using only synthetic data to improve LLM-as-a-Judge without human annotations.
0Aug 1, 2024Reasoning
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Proposes an iterative self-training method using only synthetic data to improve LLM-as-a-Judge without human annotations.
Unknown
ReST iteratively aligns LLMs with human preferences by alternating between data generation and offline RL fine-tuning with increasing quality thresholds.