ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
… open-source reward models on MRewardbench. We find that current reward models exhibit a … • We provide analyses and insights (§6) on how robust the current reward models are in a …
Reward models are a cornerstone of Reinforcement Learning from Human Feedback (RLHF), guiding the alignment of large language models (LLMs) with human preferences. However, most reward models are trained and evaluated predominantly on English data, raising concerns about their effectiveness in multilingual settings. As LLMs are increasingly deployed globally, it is critical to understand whether reward models—and consequently the aligned models—perform equitably across languages. This paper addresses this gap by introducing M-RewardBench, a benchmark specifically designed to evaluate reward models in multilingual contexts.
The significance of this work lies in its systematic evaluation of open-source reward models on a multilingual benchmark. The findings reveal that current reward models exhibit substantial performance degradation on non-English languages, which has profound implications for the fairness and reliability of RLHF-based systems in non-English-speaking regions. By quantifying these gaps, the paper provides a clear call to action for the research community to develop more language-agnostic reward models.
The paper reports that current open-source reward models show significantly lower performance on non-English languages compared to English. While specific metrics are not detailed in the abstract, the overall trend indicates a clear lack of multilingual robustness. The analyses highlight that performance gaps vary across languages, suggesting that some languages are more underserved than others. These results underscore the need for more diverse training data and language-aware reward model architectures.
This work has broad implications for the AI field, particularly for the deployment of LLMs in multilingual and multicultural contexts. By exposing the limitations of current reward models, it encourages the development of more inclusive alignment techniques. The benchmark itself serves as a valuable resource for the community, enabling standardized evaluation and progress tracking. Ultimately, this research contributes to the goal of creating AI systems that serve users equitably across the world's languages.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba