{"version":"1.0","type":"stack","id":"d244b35a-f040-4bb7-96ae-187b792f699b","url":"https://stacklist.com/c/technology/stack/d244b35a-f040-4bb7-96ae-187b792f699b","title":"AI agent evaluation frameworks","description":"Curated frameworks, observability platforms, and benchmark harnesses for evaluating LLM applications and AI agents. Focuses on tools that help teams build datasets, run offline/online evals, trace agent behavior, score outputs, and compare coding-agent performance.","privacy":"public","created_at":"2026-07-02T10:01:58.803Z","updated_at":"2026-07-02T10:01:58.803Z","author":{"username":"sushmikar-512","name":"Unknown","url":"https://stacklist.com/sushmikar-512","type":"person"},"category":{"id":"technology","name":"technology","url":"https://stacklist.com/c/technology"},"items":[{"id":"07068295-9490-443c-8a79-bc42f333390c","position":1,"title":"Evals: Framework for Evaluating LLMs and Benchmarks","url":"https://github.com/openai/evals","note":"Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks. It provides tools and resources for assessing the performance of language models effectively.","image":{"url":"https://ucarecdn.com/2f636057-c341-4265-a9c6-8da495313021/","alt":"Evals: Framework for Evaluating LLMs and Benchmarks","width":1200,"height":600},"direct_link":"https://stacklist.com/card/07068295-9490-443c-8a79-bc42f333390c","created_at":"2026-07-02T10:02:44.911Z","updated_at":null,"aco":{"summary":"OpenAI Evals is a framework for evaluating large language models (LLMs) and LLM-based systems, offering a registry of existing evals and the ability to create custom evaluations. The documentation covers setup, installation via pip and Git-LFS, running and writing evals, including support for advanced use cases like prompt chains, tool-using agents, and logging results to Snowflake.","tags":["openai","evals","llm-evaluation","framework","python","model-testing","custom-evals"],"key_entities":[{"name":"OpenAI","type":"organization","confidence":1},{"name":"OpenAI Evals","type":"technology","confidence":1},{"name":"Greg Brockman","type":"person","confidence":0.95},{"name":"Git-LFS","type":"technology","confidence":0.9},{"name":"Python","type":"technology","confidence":0.85},{"name":"Snowflake","type":"technology","confidence":0.8},{"name":"Weights & Biases","type":"technology","confidence":0.8},{"name":"LLM evaluation","type":"concept","confidence":0.95},{"name":"Completion Function Protocol","type":"concept","confidence":0.8}],"classification":"framework","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:48:56.388Z"},"token_counts":{"approximate":1352,"cl100k":1162},"content_hash":"sha256:ef31f15304b62eb6c223292614c401cfee761d0dde9b823e61deab6969736458","acp_version":"0.2","body_available":true,"body_tokens":1352,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"540919be-d92d-42cd-82b3-be372067c4be","position":2,"title":"Evaluation concepts - Docs by LangChain","url":"https://docs.langchain.com/langsmith/evaluation-concepts","note":"This page provides an overview of evaluation concepts within the LangChain framework. It covers key principles and methodologies for assessing the performance of language models and their applications.","image":{"url":"https://ucarecdn.com/185ef330-24e6-4340-964b-4225f3ff68e5/","alt":"Evaluation concepts - Docs by LangChain","width":2858,"height":1016},"direct_link":"https://stacklist.com/card/540919be-d92d-42cd-82b3-be372067c4be","created_at":"2026-07-02T10:02:46.665Z","updated_at":null,"aco":{"summary":"LangSmith Evaluation Concepts provides a comprehensive framework for measuring LLM application quality through offline (pre-deployment) and online (production monitoring) evaluations. The document covers evaluation lifecycle stages, core targets such as datasets and examples, and strategies for continuous improvement through iterative feedback loops.","tags":["llm-evaluation","langsmith","offline-evaluation","online-evaluation","quality-monitoring","regression-testing","rag"],"key_entities":[{"name":"LangSmith","type":"technology","confidence":0.99},{"name":"Offline Evaluation","type":"concept","confidence":0.95},{"name":"Online Evaluation","type":"concept","confidence":0.95},{"name":"RAG","type":"concept","confidence":0.85},{"name":"Evaluation Lifecycle","type":"concept","confidence":0.9},{"name":"Regression Testing","type":"concept","confidence":0.8},{"name":"Anomaly Detection","type":"concept","confidence":0.75},{"name":"LLM","type":"technology","confidence":0.95}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:48:58.280Z"},"token_counts":{"approximate":3967,"cl100k":2741},"content_hash":"sha256:20e98766884afcd0cc9bf6c3e782fe94df358be44c9a32b3bd86ee56124ef1fb","acp_version":"0.2","body_available":true,"body_tokens":3967,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"f37cde7f-d74f-49f5-917e-8eb96461aa48","position":3,"title":"DeepEval 5-min Quickstart | DeepEval","url":"https://deepeval.com/docs/getting-started","note":"DeepEval's 5-minute quickstart guide walks users through installing the framework, creating an LLM test case with input/output pairs, choosing a GEval metric, and running end-to-end evaluations locally.","image":{"url":"https://ucarecdn.com/5ce5fb36-2aab-46fc-9ee2-b4f56c83713d/","alt":"DeepEval 5-min Quickstart | DeepEval","width":3456,"height":2062},"direct_link":"https://stacklist.com/card/f37cde7f-d74f-49f5-917e-8eb96461aa48","created_at":"2026-07-02T10:02:48.681Z","updated_at":null,"aco":{"summary":"DeepEval's 5-minute quickstart guide walks users through installing the framework, creating an LLM test case with input/output pairs, choosing a GEval metric, and running end-to-end evaluations locally. The tutorial covers environment setup, single-turn and multi-turn test cases, metric thresholds, regression detection, and integration with the Confident AI cloud platform.","tags":["deepeval","llm-evaluation","quickstart","testing","ai-quality","python","confident-ai"],"key_entities":[{"name":"DeepEval","type":"technology","confidence":0.99},{"name":"Confident AI","type":"organization","confidence":0.95},{"name":"GEval","type":"concept","confidence":0.92},{"name":"LLM evaluation","type":"concept","confidence":0.95},{"name":"Python","type":"technology","confidence":0.85},{"name":"LLMTestCase","type":"concept","confidence":0.88},{"name":"test run","type":"concept","confidence":0.8}],"classification":"tutorial","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:48:58.286Z"},"token_counts":{"approximate":9170,"cl100k":8584},"content_hash":"sha256:280daed03e063c935dc0e57af6ef04ed0cb02d5a889166da362ab468fd4347d8","acp_version":"0.2","body_available":true,"body_tokens":9170,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"11bf29a4-f7f5-4d63-970e-d24eb98be3e0","position":4,"title":"Ragas - Evaluation Framework for AI Applications","url":"https://docs.ragas.io/en/stable/","note":"Ragas is an evaluation framework designed to assess the performance of AI applications. It provides tools and methodologies to ensure the effectiveness and reliability of AI systems.","image":{"url":"https://ucarecdn.com/ac2fa1d4-807d-4b93-abf2-311b95f1b21f/","alt":"Ragas - Evaluation Framework for AI Applications","width":1280,"height":800},"direct_link":"https://stacklist.com/card/11bf29a4-f7f5-4d63-970e-d24eb98be3e0","created_at":"2026-07-02T10:02:50.638Z","updated_at":null,"aco":{"summary":"Ragas is a library that replaces manual \"vibe checks\" with systematic evaluation loops for LLM applications, offering LLM-driven metrics, custom metric creation, and an experiments-first approach. It integrates with popular frameworks like LangChain and LlamaIndex, providing built-in dataset management and result tracking to enable continuous improvement of AI applications.","tags":["ragas","llm-evaluation","ai-metrics","experimentation","langchain","llamaindex","evaluation-framework"],"key_entities":[{"name":"Ragas","type":"technology","confidence":1},{"name":"LangChain","type":"technology","confidence":0.95},{"name":"LlamaIndex","type":"technology","confidence":0.95},{"name":"LLM evaluation","type":"concept","confidence":0.98},{"name":"evaluation metrics","type":"concept","confidence":0.9},{"name":"Vibrant Labs","type":"organization","confidence":0.85},{"name":"continuous improvement loop","type":"concept","confidence":0.8}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T10:02:59.091Z"},"token_counts":{"approximate":468,"cl100k":357},"content_hash":"sha256:b38a965944cfd45863a7f39c4a540e8a99652c993ee4a532e1083057f878d0f4","acp_version":"0.2","body_available":true,"body_tokens":468,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"940529c5-7a06-4039-92cf-25ff8371d9d9","position":5,"title":"Measure Performance with Evaluations - Phoenix","url":"https://arize.com/docs/phoenix/get-started/get-started-evaluations","note":"This page provides a comprehensive guide on how to measure the performance of machine learning models using the Evaluations feature in Phoenix. It covers the setup process, key metrics to consider, and best practices for effective evaluation.","image":{"url":"https://ucarecdn.com/4b30c74c-5a1d-426d-b0b2-2b32aaa76e2f/","alt":"Measure Performance with Evaluations - Phoenix","width":1280,"height":800},"direct_link":"https://stacklist.com/card/940529c5-7a06-4039-92cf-25ff8371d9d9","created_at":"2026-07-02T10:02:51.915Z","updated_at":null,"aco":{"summary":"Phoenix Evaluations is a tutorial guide for setting up and running evaluations on existing trace data to measure LLM output quality in a repeatable way. It walks through defining an LLM-as-a-judge evaluation for completeness, using a financial analysis chatbot as the example application.","tags":["phoenix","evaluations","llm-as-a-judge","tracing","model-quality","observability","ai-evaluation"],"key_entities":[{"name":"Phoenix","type":"technology","confidence":0.99},{"name":"LLM-as-a-judge","type":"concept","confidence":0.95},{"name":"evaluations","type":"concept","confidence":0.95},{"name":"tracing","type":"concept","confidence":0.85},{"name":"Financial Analysis and Research Chatbot","type":"technology","confidence":0.8},{"name":"completeness evaluation","type":"concept","confidence":0.8}],"classification":"tutorial","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T10:03:00.741Z"},"token_counts":{"approximate":2537,"cl100k":2250},"content_hash":"sha256:8a148c7e1536024478ed835e045d608d0ec99e2f2232fbcf438f9f728dc61721","acp_version":"0.2","body_available":true,"body_tokens":2537,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"fbcc14f5-9f90-4aaa-80a5-a0ce91565dc3","position":6,"title":"Evaluate systematically - Braintrust","url":"https://www.braintrust.dev/docs/evaluate","note":"Measure AI application quality, detect regressions before they reach production, and build confidence that your system is improving over time.","image":{"url":"https://ucarecdn.com/2d38c5e1-a8bd-4fa2-b024-c3ae8a86480a/","alt":"Evaluate systematically - Braintrust","width":1200,"height":630},"direct_link":"https://stacklist.com/card/fbcc14f5-9f90-4aaa-80a5-a0ce91565dc3","created_at":"2026-07-02T10:02:53.160Z","updated_at":null,"aco":{"summary":"Braintrust Evals provides a comprehensive evaluation framework for AI systems, covering the full cycle from playground iteration and offline experiments to continuous online production scoring. The platform supports datasets, tasks, and scorers as core evaluation components, with CI/CD integration to catch regressions before deployment.","tags":["ai-evaluation","braintrust","llm-as-a-judge","ci-cd","offline-evaluation","online-scoring","experiment-tracking"],"key_entities":[{"name":"Braintrust","type":"organization","confidence":0.98},{"name":"offline evaluation","type":"concept","confidence":0.95},{"name":"online evaluation","type":"concept","confidence":0.95},{"name":"LLM-as-a-judge","type":"concept","confidence":0.93},{"name":"CI/CD integration","type":"concept","confidence":0.88},{"name":"scorers and classifiers","type":"concept","confidence":0.85},{"name":"playgrounds","type":"technology","confidence":0.8}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T10:03:01.216Z"},"token_counts":{"approximate":845,"cl100k":634},"content_hash":"sha256:fbf2d785cb3e92ee35335bb1acedb2da344b698aeb7b94ec16799123c06702b6","acp_version":"0.2","body_available":true,"body_tokens":845,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"f3337bc5-d81e-4896-b326-dcfb36f2a2ab","position":7,"title":"Inspect: Open-source Framework for LLM Evaluations","url":"https://inspect.aisi.org.uk/","note":"Inspect is an open-source framework designed for evaluating large language models. It provides tools and methodologies to assess the performance and capabilities of various language models effectively.","image":{"url":"https://ucarecdn.com/ccf4369b-4387-4926-8a60-afa0c9948b06/","alt":"Inspect: Open-source Framework for LLM Evaluations","width":2400,"height":1258},"direct_link":"https://stacklist.com/card/f3337bc5-d81e-4896-b326-dcfb36f2a2ab","created_at":"2026-07-02T10:02:54.416Z","updated_at":null,"aco":{"summary":"Inspect is an AI evaluation framework developed by the UK AI Security Institute and Meridian Labs, offering composable building blocks, over 200 pre-built evaluations, and support for 20+ model providers. The framework enables coding, reasoning, knowledge, and agentic task evaluations with features including sandboxing, tool calling, multi-agent primitives, and a web-based visualization tool.","tags":["ai-evaluation","inspect","framework","llm-benchmarks","agentic-tasks","model-testing","python"],"key_entities":[{"name":"Inspect","type":"technology","confidence":1},{"name":"UK AI Security Institute","type":"organization","confidence":0.95},{"name":"Meridian Labs","type":"organization","confidence":0.9},{"name":"SimpleQA","type":"technology","confidence":0.85},{"name":"Claude Code","type":"technology","confidence":0.8},{"name":"VS Code Extension","type":"technology","confidence":0.75},{"name":"OpenAI","type":"organization","confidence":0.9},{"name":"Anthropic","type":"organization","confidence":0.9},{"name":"Google","type":"organization","confidence":0.85},{"name":"Docker","type":"technology","confidence":0.7},{"name":"Kubernetes","type":"technology","confidence":0.7},{"name":"HuggingFace","type":"technology","confidence":0.8},{"name":"AI evaluation","type":"concept","confidence":0.95},{"name":"MCP tools","type":"technology","confidence":0.7}],"classification":"framework","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:48:54.955Z"},"token_counts":{"approximate":2105,"cl100k":1781},"content_hash":"sha256:c91a7f94527cea3c006158a99487070fb4ddcca74c2e96206ea25c1fb84f9d52","acp_version":"0.2","body_available":true,"body_tokens":2105,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"a65b7ae8-5b14-4af1-9c3f-0e2b3042dfc2","position":8,"title":"Evaluation of LLM Applications - Langfuse","url":"https://langfuse.com/docs/evaluation/overview","note":"With Langfuse you can capture all your LLM evaluations in one place. You can combine a variety of different evaluation metrics like model-based evaluations (LLM-as-a-Judge), human annotations or fully custom evaluation workflows via API/SDKs. This allows you to measure quality, tonality, factual accuracy, completeness, and other dimensions of your LLM application.","image":{"url":"https://ucarecdn.com/70aedf0f-b964-4cad-9334-696a5c09f4a1/","alt":"Evaluation of LLM Applications - Langfuse","width":48,"height":48},"direct_link":"https://stacklist.com/card/a65b7ae8-5b14-4af1-9c3f-0e2b3042dfc2","created_at":"2026-07-02T10:02:56.238Z","updated_at":null,"aco":{"summary":"Langfuse Evaluation provides a comprehensive framework for checking LLM application behavior through online production trace scoring and offline experimentation with datasets, experiments, and automated or manual evaluators. The documentation covers core features including annotation queues, LLM-as-a-Judge, score analytics, CI/CD experiment integration, and code evaluators to catch regressions before deployment.","tags":["langfuse","llm-evaluation","observability","datasets","experiments","llm-as-judge","ai-engineering"],"key_entities":[{"name":"Langfuse","type":"technology","confidence":0.99},{"name":"LLM Evaluation","type":"concept","confidence":0.97},{"name":"LLM-as-a-Judge","type":"concept","confidence":0.92},{"name":"Annotation Queues","type":"concept","confidence":0.85},{"name":"Datasets","type":"concept","confidence":0.8},{"name":"Experiments","type":"concept","confidence":0.85},{"name":"CI/CD","type":"concept","confidence":0.75}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T10:03:03.313Z"},"token_counts":{"approximate":500,"cl100k":384},"content_hash":"sha256:992e0a68abb847adac380b7b9de67e6e8576ae9e654f161d893236ec245229cd","acp_version":"0.2","body_available":true,"body_tokens":500,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"4fcbdb6e-2dbd-495b-b198-fb4facc7b34c","position":9,"title":"TruLens: Evals and Tracing for Agents","url":"https://www.trulens.org/","note":"TruLens provides tools for evaluating and tracing the performance of AI agents. It aims to enhance the understanding and reliability of AI systems through comprehensive analysis.","image":{"url":"https://ucarecdn.com/8b50548d-fc5a-4ff2-aa19-bb935ddea80a/","alt":"TruLens: Evals and Tracing for Agents","width":94,"height":20},"direct_link":"https://stacklist.com/card/4fcbdb6e-2dbd-495b-b198-fb4facc7b34c","created_at":"2026-07-02T10:02:57.670Z","updated_at":null,"aco":{"summary":"TruLens is an open-source evaluation framework that helps developers objectively measure the quality and effectiveness of AI agents across metrics like groundedness, context relevance, coherence, and more. Originally created by TruEra and now shepherded by Snowflake, it supports agents, RAG, summarization, and co-pilots via a Python SDK and OpenTelemetry trace ingestion.","tags":["trulens","ai-evaluation","llm","observability","agentic-workflows","rag","opentelemetry"],"key_entities":[{"name":"TruLens","type":"technology","confidence":1},{"name":"Snowflake","type":"organization","confidence":0.98},{"name":"TruEra","type":"organization","confidence":0.95},{"name":"OpenTelemetry","type":"technology","confidence":0.92},{"name":"Retrieval Augmented Generation","type":"concept","confidence":0.9},{"name":"AI agent evaluation","type":"concept","confidence":0.93},{"name":"Groundedness","type":"concept","confidence":0.8},{"name":"Python SDK","type":"technology","confidence":0.85}],"classification":"framework","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T10:03:06.651Z"},"token_counts":{"approximate":812,"cl100k":634},"content_hash":"sha256:d17ec60676e3dbc20960ac35b40fef0e2a2bee6a4192e249fb25a2d897a09663","acp_version":"0.2","body_available":true,"body_tokens":812,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"62988ef1-05df-4fa3-9946-9da69954c2e6","position":10,"title":"SWE-bench: Can Language Models Resolve GitHub Issues?","url":"https://github.com/SWE-bench/SWE-bench","note":"SWE-bench is a project that explores the capability of language models in addressing real-world issues found on GitHub. It aims to evaluate how effectively these models can assist developers in resolving software-related problems.","image":{"url":"https://ucarecdn.com/bc0a6940-2827-44f7-9041-39e04956f8cc/","alt":"SWE-bench: Can Language Models Resolve GitHub Issues?","width":1200,"height":600},"direct_link":"https://stacklist.com/card/62988ef1-05df-4fa3-9946-9da69954c2e6","created_at":"2026-07-02T10:02:58.912Z","updated_at":null,"aco":{"summary":"SWE-bench is a benchmark for evaluating large language models on real-world software issues collected from GitHub, where models generate patches to resolve described problems. The repository includes setup instructions using Docker for reproducible evaluations, along with extensions like SWE-bench Multimodal, SWE-bench Verified, and cloud-based evaluation via Modal and sb-cli.","tags":["swe-bench","benchmark","large-language-models","github-issues","docker","software-engineering","evaluation"],"key_entities":[{"name":"SWE-bench","type":"technology","confidence":1},{"name":"Princeton NLP","type":"organization","confidence":0.95},{"name":"OpenAI","type":"organization","confidence":0.9},{"name":"Docker","type":"technology","confidence":0.95},{"name":"ICLR 2024","type":"event","confidence":0.95},{"name":"ICLR 2025","type":"event","confidence":0.9},{"name":"SWE-agent","type":"technology","confidence":0.9},{"name":"Modal","type":"technology","confidence":0.85},{"name":"SWE-bench Multimodal","type":"technology","confidence":0.9},{"name":"SWE-bench Verified","type":"technology","confidence":0.88},{"name":"OpenAI Preparedness","type":"concept","confidence":0.8},{"name":"sb-cli","type":"technology","confidence":0.85}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T10:03:10.232Z"},"token_counts":{"approximate":1728,"cl100k":1733},"content_hash":"sha256:fed9e9fdb50f0e7a660f21a09cee5ef0b3b8aa57d313a1734e08aac08d798464","acp_version":"0.2","body_available":true,"body_tokens":1728,"visibility":"public","agent_accessible":true,"status":"final"}}],"stats":{"likes_count":0,"items_count":10},"_links":{"self":"/api/public/stack/d244b35a-f040-4bb7-96ae-187b792f699b.json","html":"https://stacklist.com/c/technology/stack/d244b35a-f040-4bb7-96ae-187b792f699b","rss":"/api/feeds/stack/d244b35a-f040-4bb7-96ae-187b792f699b/rss","json_feed":"/api/feeds/stack/d244b35a-f040-4bb7-96ae-187b792f699b/json"}}