Curated by
More in AI Agent Evaluation Frameworks and Benchmarks
See all 12 →Survey on Evaluation of LLM-based Agents
This survey provides the first comprehensive analysis of evaluation methods for LLM-based agents, examining core capabilities, application-specific benchmarks, generalist agent evaluation, benchmark dimensions, and evaluation frameworks. The paper identifies trends toward more realistic evaluations and highlights critical gaps in assessing cost-efficiency, safety, robustness, and scalable evaluation methods.
Built for AI agentsACO · 1024 tokens
Summary
This survey provides the first comprehensive analysis of evaluation methods for LLM-based agents, examining core capabilities, application-specific benchmarks, generalist agent evaluation, benchmark dimensions, and evaluation frameworks. The paper identifies trends toward more realistic evaluations and highlights critical gaps in assessing cost-efficiency, safety, robustness, and scalable evaluation methods.
Tags
llm-agents · evaluation · benchmarks · artificial-intelligence · survey · planning · tool-use
Key entities
Asaf Yehudai (person, 0.95) · Lilach Eden (person, 0.9) · Alan Li (person, 0.9) · Guy Uziel (person, 0.9) · Yilun Zhao (person, 0.9) · Roy Bar-Haim (person, 0.9) · Arman Cohan (person, 0.9) · Michal Shmueli-Scheuer (person, 0.9) · arXiv (organization, 0.95) · LLM-based agents (concept, 0.99) · agent evaluation (concept, 0.97) · benchmarks (concept, 0.92) · ACL Findings (event, 0.85) · SWE agents (concept, 0.8)
Classification
reference · language en · status final
Provenance
claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026