AI Agent Evaluation Frameworks and Benchmarks

Curated byamrutha-kothapalli-30712 cardsUpdated Jul 2026
A curated set of practical frameworks, benchmark environments, metrics, and surveys for evaluating AI agents across tool use, web navigation, long-horizon autonomy, trajectory quality, and real-world task completion.
Survey on Evaluation of LLM-based Agents
arxiv.org
Survey on Evaluation of LLM-based Agents

This survey provides the first comprehensive analysis of evaluation methods for LLM-based agents, examining core capabilities, application-specific benchmarks, generalist agent evaluation, benchmark dimensions, and evaluation frameworks.

avatar
Inspect: Open-source Framework for LLM Evaluations
inspect.aisi.org.uk
Inspect: Open-source Framework for LLM Evaluations

Inspect is an open-source framework designed for evaluating large language models. It provides tools and methodologies to assess the performance and capabilities of various language models effectively.

avatar
Evaluation concepts - Docs by LangChain
docs.langchain.com
Evaluation concepts - Docs by LangChain

This page provides an overview of evaluation concepts within the LangChain framework. It covers key principles and methodologies for assessing the performance of language models and their applications.

avatar
Evals: Framework for Evaluating LLMs and Benchmarks
github.com
Evals: Framework for Evaluating LLMs and Benchmarks

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks. It provides tools and resources for assessing the performance of language models effectively.

avatar
DeepEval 5-min Quickstart | DeepEval
deepeval.com
DeepEval 5-min Quickstart | DeepEval

This quickstart takes you from installing DeepEval to your first passing eval in a few minutes. You'll create a small test case, choose a metric, and run it to evaluate LLMs effectively.

avatar
Agentic or Tool Use - Ragas
docs.ragas.io
Agentic or Tool Use - Ragas

This page provides an evaluation framework for assessing AI applications through various metrics. It outlines the available metrics related to agentic or tool use, helping developers understand and improve their AI systems.

avatar
Welcome to BrowserGym’s documentation!
browsergym.readthedocs.io
Welcome to BrowserGym’s documentation!

BrowserGym provides comprehensive documentation for users to understand and utilize its features effectively. This page serves as a guide to help users navigate through the functionalities offered by BrowserGym version 0.3.

avatar
WebArena: Benchmarks for Autonomous Web Agents
webarena.dev
WebArena: Benchmarks for Autonomous Web Agents

WebArena is a comprehensive suite of benchmarks designed to facilitate the development and evaluation of autonomous web agents. It provides tools and metrics to assess the performance and capabilities of these agents in various web environments.

avatar
$τ$-bench: A Benchmark for Tool-Agent-User Interaction
arxiv.org
$τ$-bench: A Benchmark for Tool-Agent-User Interaction

This page presents the arXiv paper 2406.12045, which introduces $τ$-bench, a benchmark designed to evaluate interactions among tools, agents, and users in real-world domains. The paper aims to provide a comprehensive framework for assessing the effectiveness and efficiency of these interactions.

avatar
GitHub - THUDM/AgentBench: A Comprehensive Benchmark
github.com
GitHub - THUDM/AgentBench: A Comprehensive Benchmark

THUDM/AgentBench is a comprehensive benchmark designed to evaluate large language models (LLMs) as agents. This resource aims to facilitate research and development in the field of artificial intelligence by providing standardized evaluation metrics and datasets.

avatar
AgentRewardBench
agent-reward-bench.github.io
AgentRewardBench

AgentRewardBench is a platform designed for evaluating and benchmarking reinforcement learning agents. It provides tools and metrics to assess agent performance in various environments.

avatar
Measuring AI Ability to Complete Long Tasks
metr.org
Measuring AI Ability to Complete Long Tasks

We propose measuring AI performance in terms of the *length* of tasks AI agents can complete. We show that this metric has been consistently exponentially increasing over the past 6 years, with a doubling time of around 7 months.

avatar