Preprint
Large Language Models

Survey on Evaluation of LLM-based Agents

Asaf Yehudai, Lilach Eden, Alan Li, Guy Uziel, Yilun Zhao, Roy Bar-Haim, Arman Cohan, Michal Shmueli-Scheuer
March 20, 2025arXiv.org207 citations

207

Citations

12

Influential Citations

arXiv.org

Venue

2025

Year

Abstract

LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environments. This paper provides the first comprehensive survey of evaluation methods for these increasingly capable agents. We analyze the field of agent evaluation across five perspectives: (1) Core LLM capabilities needed for agentic workflows, like planning, and tool use; (2) Application-specific benchmarks such as web and SWE agents; (3) Evaluation of generalist agents; (4) Analysis of agent benchmarks'core dimensions; and (5) Evaluation frameworks and tools for agent developers. Our analysis reveals current trends, including a shift toward more realistic, challenging evaluations with continuously updated benchmarks. We also identify critical gaps that future research must address, particularly in assessing cost-efficiency, safety, and robustness, and in developing fine-grained, scalable evaluation methods.

Analysis

Why This Paper Matters

LLM-based agents are rapidly evolving from simple chatbots to autonomous systems capable of planning, reasoning, and using tools in dynamic environments. However, evaluating these agents is a complex and fragmented field, with numerous benchmarks and frameworks that are often application-specific and lack standardization. This survey is the first to comprehensively organize and analyze the evaluation landscape, providing a crucial reference for researchers and practitioners. By synthesizing existing work across five key perspectives, it offers a structured understanding of what is being evaluated, how, and where the gaps lie.

The paper's significance is amplified by its timing. As LLM agents are increasingly deployed in real-world applications, the need for robust, reliable, and safe evaluation becomes paramount. The survey's identification of critical gaps—particularly in cost-efficiency, safety, and robustness—highlights areas where current evaluation methods are insufficient, potentially hindering the responsible deployment of these agents. This makes the survey not just an academic exercise but a practical guide for the AI community.

Technical Contributions

  • Five-Perspective Framework: The survey organizes the field into five distinct perspectives: core LLM capabilities (e.g., planning, tool use), application-specific benchmarks (e.g., web agents, SWE agents), generalist agent evaluation, analysis of benchmark dimensions (e.g., realism, difficulty), and evaluation frameworks/tools. This taxonomy helps clarify the scope and focus of different evaluation efforts.
  • Comprehensive Categorization: It systematically categorizes a wide range of existing benchmarks and evaluation methods, providing a structured overview that is easy to navigate. This includes both well-known benchmarks and emerging ones, offering a broad view of the field.
  • Trend Analysis: The survey identifies key trends, such as the shift toward more realistic and challenging evaluations, and the use of continuously updated benchmarks to prevent saturation. This insight is valuable for understanding the direction of the field.
  • Gap Identification: It explicitly highlights underexplored areas, including cost-efficiency, safety, robustness, and the need for fine-grained, scalable evaluation methods. This provides a clear agenda for future research.

Results

The survey does not present new experimental results but rather synthesizes findings from the literature. It reports that current evaluation methods are increasingly focusing on realistic and challenging scenarios, with a trend toward continuously updated benchmarks. However, it finds that most existing benchmarks are narrow and application-specific, with limited coverage of generalist capabilities. The analysis reveals that core dimensions such as reliability, safety, and cost are often overlooked. The paper does not provide specific metrics or comparisons, as its primary contribution is the qualitative analysis and categorization of the field.

Significance

The survey has significant implications for the AI community. By providing a structured overview, it enables researchers to identify gaps and opportunities in agent evaluation, potentially accelerating progress in this critical area. For practitioners, it offers a guide to selecting appropriate evaluation methods for their specific applications. The emphasis on safety, robustness, and cost-efficiency aligns with the growing concern for responsible AI deployment. This survey is likely to become a foundational reference for future work on LLM agent evaluation, shaping the development of more comprehensive and reliable evaluation methodologies.