ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
477
Citations
71
Influential Citations
Conference on Empirical Methods in Natural Language Processing
Venue
2023
Year
A 13B fully open source evaluation LLM trained on Feedback Collection curated using GPT-4 (in this work).
This paper addresses a critical bottleneck in AI research: the evaluation of large language models (LLMs). Traditionally, evaluation relies on expensive, proprietary models like GPT-4, which are not transparent or reproducible. Prometheus provides a fully open-source alternative, democratizing access to high-quality automated evaluation. This is significant because it allows researchers and practitioners to evaluate their models without depending on closed APIs, fostering more rigorous and reproducible research.
Moreover, the paper introduces Feedback Collection, a dataset of evaluation examples curated using GPT-4. This dataset is a valuable resource for training future evaluation models and for studying the properties of effective feedback. By open-sourcing both the model and the dataset, the authors lower the barrier to entry for LLM evaluation, which is crucial for the field's progress.
The paper reports that Prometheus achieves strong agreement with GPT-4 evaluations on several benchmarks. While exact metrics are not detailed in the abstract, the model is shown to be a viable open-source substitute for proprietary evaluation systems. This is a notable achievement given the model's relatively small size (13B parameters) compared to GPT-4.
Prometheus has broad implications for the AI field. It enables transparent and reproducible evaluation of LLMs, which is essential for scientific progress. By providing an open-source alternative, it reduces the field's dependence on proprietary models and APIs. This can accelerate research in areas like instruction following, safety, and alignment, where reliable evaluation is critical. The work also sets a precedent for creating open-source evaluation tools, potentially leading to a more collaborative and equitable AI ecosystem.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba