ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
207
Citations
12
Influential Citations
arXiv.org
Venue
2025
Year
LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environments. This paper provides the first comprehensive survey of evaluation methods for these increasingly capable agents. We analyze the field of agent evaluation across five perspectives: (1) Core LLM capabilities needed for agentic workflows, like planning, and tool use; (2) Application-specific benchmarks such as web and SWE agents; (3) Evaluation of generalist agents; (4) Analysis of agent benchmarks'core dimensions; and (5) Evaluation frameworks and tools for agent developers. Our analysis reveals current trends, including a shift toward more realistic, challenging evaluations with continuously updated benchmarks. We also identify critical gaps that future research must address, particularly in assessing cost-efficiency, safety, and robustness, and in developing fine-grained, scalable evaluation methods.
LLM-based agents are rapidly evolving from simple chatbots to autonomous systems capable of planning, reasoning, and using tools in dynamic environments. However, evaluating these agents is a complex and fragmented field, with numerous benchmarks and frameworks that are often application-specific and lack standardization. This survey is the first to comprehensively organize and analyze the evaluation landscape, providing a crucial reference for researchers and practitioners. By synthesizing existing work across five key perspectives, it offers a structured understanding of what is being evaluated, how, and where the gaps lie.
The paper's significance is amplified by its timing. As LLM agents are increasingly deployed in real-world applications, the need for robust, reliable, and safe evaluation becomes paramount. The survey's identification of critical gaps—particularly in cost-efficiency, safety, and robustness—highlights areas where current evaluation methods are insufficient, potentially hindering the responsible deployment of these agents. This makes the survey not just an academic exercise but a practical guide for the AI community.
The survey does not present new experimental results but rather synthesizes findings from the literature. It reports that current evaluation methods are increasingly focusing on realistic and challenging scenarios, with a trend toward continuously updated benchmarks. However, it finds that most existing benchmarks are narrow and application-specific, with limited coverage of generalist capabilities. The analysis reveals that core dimensions such as reliability, safety, and cost are often overlooked. The paper does not provide specific metrics or comparisons, as its primary contribution is the qualitative analysis and categorization of the field.
The survey has significant implications for the AI community. By providing a structured overview, it enables researchers to identify gaps and opportunities in agent evaluation, potentially accelerating progress in this critical area. For practitioners, it offers a guide to selecting appropriate evaluation methods for their specific applications. The emphasis on safety, robustness, and cost-efficiency aligns with the growing concern for responsible AI deployment. This survey is likely to become a foundational reference for future work on LLM agent evaluation, shaping the development of more comprehensive and reliable evaluation methodologies.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba