Curated by

avatar

amrutha-kothapalli-307

stacklist.com/amrutha-kothapalli-307

More in AI Agent Evaluation Frameworks and Benchmarks

See all 12 →

Survey on Evaluation of LLM-based Agents

This survey provides the first comprehensive analysis of evaluation methods for LLM-based agents, examining core capabilities, application-specific benchmarks, generalist agent evaluation, benchmark dimensions, and evaluation frameworks. The paper identifies trends toward more realistic evaluations and highlights critical gaps in assessing cost-efficiency, safety, robustness, and scalable evaluation methods.

View card
Built for AI agentsACO · 1024 tokens

Summary

This survey provides the first comprehensive analysis of evaluation methods for LLM-based agents, examining core capabilities, application-specific benchmarks, generalist agent evaluation, benchmark dimensions, and evaluation frameworks. The paper identifies trends toward more realistic evaluations and highlights critical gaps in assessing cost-efficiency, safety, robustness, and scalable evaluation methods.

Tags

llm-agents · evaluation · benchmarks · artificial-intelligence · survey · planning · tool-use

Key entities

Asaf Yehudai (person, 0.95) · Lilach Eden (person, 0.9) · Alan Li (person, 0.9) · Guy Uziel (person, 0.9) · Yilun Zhao (person, 0.9) · Roy Bar-Haim (person, 0.9) · Arman Cohan (person, 0.9) · Michal Shmueli-Scheuer (person, 0.9) · arXiv (organization, 0.95) · LLM-based agents (concept, 0.99) · agent evaluation (concept, 0.97) · benchmarks (concept, 0.92) · ACL Findings (event, 0.85) · SWE agents (concept, 0.8)

Classification

reference · language en · status final

Provenance

claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026