ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
10
Citations
1
Influential Citations
arXiv.org
Venue
2024
Year
LLMs demand significant computational resources for both pre-training and fine-tuning, requiring distributed computing capabilities due to their large model sizes \cite{…
As large language models (LLMs) become increasingly powerful and widely deployed, the need for rigorous and standardized evaluation has never been greater. This paper addresses a critical gap by providing a structured survey of the benchmarks and datasets used to assess LLM performance. With the proliferation of new benchmarks, researchers and practitioners often struggle to choose appropriate evaluation tools, leading to inconsistent and sometimes misleading comparisons. This survey offers a valuable map of the evaluation landscape, helping to demystify the options and their trade-offs.
The paper's timing is particularly relevant given the rapid evolution of LLMs and the growing concern about benchmark saturation and overfitting. By cataloging existing resources, the authors highlight the diversity of evaluation approaches and the need for more robust, dynamic, and task-specific benchmarks. This work serves as a foundational reference for anyone involved in LLM development, from academic researchers to industry engineers, and contributes to the ongoing conversation about how to measure model capabilities meaningfully.
The paper's main contribution is a comprehensive categorization of LLM evaluation benchmarks and datasets. Key innovations include:
While the paper is a survey and does not present new experimental results, it synthesizes findings from numerous studies to draw conclusions about the state of LLM evaluation. It notes that many popular benchmarks are nearing saturation, with models achieving near-human performance, which reduces their discriminative power. The survey also highlights the emergence of more complex, multi-task benchmarks that aim to better capture real-world capabilities. However, it points out that these newer benchmarks often lack standardized protocols, making cross-study comparisons difficult. The paper does not provide specific numerical metrics but instead offers a qualitative analysis of the strengths and weaknesses of various evaluation resources.
The broader impact of this survey lies in its potential to shape future evaluation practices. By providing a clear overview of available benchmarks and datasets, it helps researchers avoid duplication and encourages the adoption of more rigorous evaluation standards. It also underscores the need for continuous innovation in benchmark design to keep pace with LLM advancements. As LLMs are increasingly integrated into critical applications, the reliability of evaluation methods becomes paramount. This paper contributes to that goal by fostering a more informed and critical approach to LLM assessment, ultimately supporting the development of more trustworthy and capable AI systems.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba