AI agent evaluation frameworks

Curated bysushmikar-51210 cardsUpdated Jul 2026
Curated frameworks, observability platforms, and benchmark harnesses for evaluating LLM applications and AI agents. Focuses on tools that help teams build datasets, run offline/online evals, trace agent behavior, score outputs, and compare coding-agent performance.
Evals: Framework for Evaluating LLMs and Benchmarks
github.com
Evals: Framework for Evaluating LLMs and Benchmarks

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks. It provides tools and resources for assessing the performance of language models effectively.

Evaluation concepts - Docs by LangChain
docs.langchain.com
Evaluation concepts - Docs by LangChain

This page provides an overview of evaluation concepts within the LangChain framework. It covers key principles and methodologies for assessing the performance of language models and their applications.

DeepEval 5-min Quickstart | DeepEval
deepeval.com
DeepEval 5-min Quickstart | DeepEval

DeepEval's 5-minute quickstart guide walks users through installing the framework, creating an LLM test case with input/output pairs, choosing a GEval metric, and running end-to-end evaluations locally.

Ragas - Evaluation Framework for AI Applications
docs.ragas.io
Ragas - Evaluation Framework for AI Applications

Ragas is an evaluation framework designed to assess the performance of AI applications. It provides tools and methodologies to ensure the effectiveness and reliability of AI systems.

Measure Performance with Evaluations - Phoenix
arize.com
Measure Performance with Evaluations - Phoenix

This page provides a comprehensive guide on how to measure the performance of machine learning models using the Evaluations feature in Phoenix. It covers the setup process, key metrics to consider, and best practices for effective evaluation.

Evaluate systematically - Braintrust
braintrust.dev
Evaluate systematically - Braintrust

Measure AI application quality, detect regressions before they reach production, and build confidence that your system is improving over time.

Inspect: Open-source Framework for LLM Evaluations
inspect.aisi.org.uk
Inspect: Open-source Framework for LLM Evaluations

Inspect is an open-source framework designed for evaluating large language models. It provides tools and methodologies to assess the performance and capabilities of various language models effectively.

Evaluation of LLM Applications - Langfuse
langfuse.com
Evaluation of LLM Applications - Langfuse

With Langfuse you can capture all your LLM evaluations in one place. You can combine a variety of different evaluation metrics like model-based evaluations (LLM-as-a-Judge), human annotations or fully custom evaluation workflows via API/SDKs. This allows you to measure quality, tonality, factual accuracy, completeness, and other dimensions of your LLM application.

TruLens: Evals and Tracing for Agents
trulens.org
TruLens: Evals and Tracing for Agents

TruLens provides tools for evaluating and tracing the performance of AI agents. It aims to enhance the understanding and reliability of AI systems through comprehensive analysis.

SWE-bench: Can Language Models Resolve GitHub Issues?
github.com
SWE-bench: Can Language Models Resolve GitHub Issues?

SWE-bench is a project that explores the capability of language models in addressing real-world issues found on GitHub. It aims to evaluate how effectively these models can assist developers in resolving software-related problems.