Safe model-based reinforcement learning with stability guarantees
Unknown
This paper introduces a safe model-based reinforcement learning approach that guarantees stability during exploration and learning.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper introduces a safe model-based reinforcement learning approach that guarantees stability during exploration and learning.
Unknown
Reflora introduces RefLoRA, a refactored low-rank adaptation method that improves the efficiency and stability of fine-tuning large models.
Unknown
This paper introduces a scalable pairwise meta-evaluator to diagnose bias and instability in LLM evaluation, providing a multifaceted view of evaluation dynamics.
Yuzhong Zhao, Yue Liu, Junpeng Liu, et al.
GMPO improves GRPO stability by replacing arithmetic mean with geometric mean of token rewards, reducing outlier sensitivity and boosting reasoning performance.
Unknown
RAFT iteratively fine-tunes generative models on top-ranked samples to align them with a reward function, improving stability and efficiency over RLHF.