Preprint
Reinforcement Learning

M-rewardbench: Evaluating reward models in multilingual settings

January 1, 2025

0

Citations

0

Influential Citations

Venue

2025

Year

Abstract

… open-source reward models on MRewardbench. We find that current reward models exhibit a … • We provide analyses and insights (§6) on how robust the current reward models are in a …

Analysis

Why This Paper Matters

Reward models are a cornerstone of Reinforcement Learning from Human Feedback (RLHF), guiding the alignment of large language models (LLMs) with human preferences. However, most reward models are trained and evaluated predominantly on English data, raising concerns about their effectiveness in multilingual settings. As LLMs are increasingly deployed globally, it is critical to understand whether reward models—and consequently the aligned models—perform equitably across languages. This paper addresses this gap by introducing M-RewardBench, a benchmark specifically designed to evaluate reward models in multilingual contexts.

The significance of this work lies in its systematic evaluation of open-source reward models on a multilingual benchmark. The findings reveal that current reward models exhibit substantial performance degradation on non-English languages, which has profound implications for the fairness and reliability of RLHF-based systems in non-English-speaking regions. By quantifying these gaps, the paper provides a clear call to action for the research community to develop more language-agnostic reward models.

Technical Contributions

  • M-RewardBench Benchmark: The paper introduces a new benchmark dataset that covers multiple languages, enabling standardized evaluation of reward models beyond English.
  • Comprehensive Evaluation: It evaluates a range of open-source reward models, providing a baseline for future comparisons.
  • Robustness Analysis: The paper includes detailed analyses of how reward models perform across different languages, identifying specific weaknesses and patterns.
  • Insights for Improvement: The findings offer actionable insights for researchers aiming to improve multilingual reward modeling.

Results

The paper reports that current open-source reward models show significantly lower performance on non-English languages compared to English. While specific metrics are not detailed in the abstract, the overall trend indicates a clear lack of multilingual robustness. The analyses highlight that performance gaps vary across languages, suggesting that some languages are more underserved than others. These results underscore the need for more diverse training data and language-aware reward model architectures.

Significance

This work has broad implications for the AI field, particularly for the deployment of LLMs in multilingual and multicultural contexts. By exposing the limitations of current reward models, it encourages the development of more inclusive alignment techniques. The benchmark itself serves as a valuable resource for the community, enabling standardized evaluation and progress tracking. Ultimately, this research contributes to the goal of creating AI systems that serve users equitably across the world's languages.