ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2024
Year
7B & 8x7B evaluation LLMs that score high correlations with both human evaluators and proprietary LM-based judges on both direct assessment and pairwise ranking, obtained by merging Mistral models trained on Feedback Collection and Preference Collection (curated in this work.
As large language models (LLMs) become ubiquitous, evaluating their outputs reliably and affordably is a critical challenge. Proprietary evaluation models like GPT-4 are expensive, opaque, and not reproducible. Prometheus 2 addresses this gap by introducing open-source evaluation LLMs that rival proprietary judges in correlation with human judgments. This democratizes LLM evaluation, allowing practitioners to assess models without relying on closed APIs.
The paper's focus on both direct assessment (scoring a single output) and pairwise ranking (comparing two outputs) covers the two most common evaluation paradigms. By achieving high correlations in both, Prometheus 2 offers a versatile tool for model development, benchmarking, and alignment research.
The abstract states that Prometheus 2 "scores high correlations with both human evaluators and proprietary LM-based judges on both direct assessment and pairwise ranking." While specific correlation coefficients are not provided in the abstract, the claim suggests that the model's judgments align closely with human preferences and with established proprietary judges like GPT-4. This positions Prometheus 2 as a strong open-source alternative for evaluation tasks.
Prometheus 2 has the potential to become a standard tool for LLM evaluation, similar to how BLEU and ROUGE are used for text generation. By providing an open-source, high-quality evaluator, it enables more transparent and reproducible research. It also reduces the cost barrier for startups and academic labs that cannot afford proprietary APIs. The model merging technique may inspire further work on combining specialized models for multi-task capabilities.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba