{"version":"1.0","type":"stack","id":"29e01ac8-6202-4abf-ab14-a0613026186b","url":"https://stacklist.com/c/technology/stack/29e01ac8-6202-4abf-ab14-a0613026186b","title":"AI Agent Evaluation Frameworks and Benchmarks","description":"A curated set of practical frameworks, benchmark environments, metrics, and surveys for evaluating AI agents across tool use, web navigation, long-horizon autonomy, trajectory quality, and real-world task completion.","privacy":"public","created_at":"2026-07-02T09:48:20.941Z","updated_at":"2026-07-02T09:48:20.941Z","author":{"username":"amrutha-kothapalli-307","name":"Unknown","url":"https://stacklist.com/amrutha-kothapalli-307","type":"person"},"category":{"id":"technology","name":"technology","url":"https://stacklist.com/c/technology"},"items":[{"id":"75394818-3927-400d-ab2f-56f509d6342d","position":1,"title":"Survey on Evaluation of LLM-based Agents","url":"https://arxiv.org/abs/2503.16416","note":"This survey provides the first comprehensive analysis of evaluation methods for LLM-based agents, examining core capabilities, application-specific benchmarks, generalist agent evaluation, benchmark dimensions, and evaluation frameworks.","image":{"url":"https://ucarecdn.com/22cd18df-b21b-4d8c-a68c-4f5aec4c1e4c/","alt":"Survey on Evaluation of LLM-based Agents","width":336,"height":96},"direct_link":"https://stacklist.com/card/75394818-3927-400d-ab2f-56f509d6342d","created_at":"2026-07-02T09:48:43.543Z","updated_at":null,"aco":{"summary":"This survey provides the first comprehensive analysis of evaluation methods for LLM-based agents, examining core capabilities, application-specific benchmarks, generalist agent evaluation, benchmark dimensions, and evaluation frameworks. The paper identifies trends toward more realistic evaluations and highlights critical gaps in assessing cost-efficiency, safety, robustness, and scalable evaluation methods.","tags":["llm-agents","evaluation","benchmarks","artificial-intelligence","survey","planning","tool-use"],"key_entities":[{"name":"Asaf Yehudai","type":"person","confidence":0.95},{"name":"Lilach Eden","type":"person","confidence":0.9},{"name":"Alan Li","type":"person","confidence":0.9},{"name":"Guy Uziel","type":"person","confidence":0.9},{"name":"Yilun Zhao","type":"person","confidence":0.9},{"name":"Roy Bar-Haim","type":"person","confidence":0.9},{"name":"Arman Cohan","type":"person","confidence":0.9},{"name":"Michal Shmueli-Scheuer","type":"person","confidence":0.9},{"name":"arXiv","type":"organization","confidence":0.95},{"name":"LLM-based agents","type":"concept","confidence":0.99},{"name":"agent evaluation","type":"concept","confidence":0.97},{"name":"benchmarks","type":"concept","confidence":0.92},{"name":"ACL Findings","type":"event","confidence":0.85},{"name":"SWE agents","type":"concept","confidence":0.8}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:48:53.489Z"},"token_counts":{"approximate":1024,"cl100k":983},"content_hash":"sha256:616cee134c87206809de54674ac11ee41be4a8063e883bedcdfd2589b634ab06","acp_version":"0.2","body_available":true,"body_tokens":1024,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"9ab96803-ab49-4663-a0c2-8728ded01b4e","position":2,"title":"Inspect: Open-source Framework for LLM Evaluations","url":"https://inspect.aisi.org.uk/","note":"Inspect is an open-source framework designed for evaluating large language models. It provides tools and methodologies to assess the performance and capabilities of various language models effectively.","image":{"url":"https://ucarecdn.com/ccf4369b-4387-4926-8a60-afa0c9948b06/","alt":"Inspect: Open-source Framework for LLM Evaluations","width":2400,"height":1258},"direct_link":"https://stacklist.com/card/9ab96803-ab49-4663-a0c2-8728ded01b4e","created_at":"2026-07-02T09:48:45.156Z","updated_at":null,"aco":{"summary":"Inspect is an AI evaluation framework developed by the UK AI Security Institute and Meridian Labs, offering composable building blocks, over 200 pre-built evaluations, and support for 20+ model providers. The framework enables coding, reasoning, knowledge, and agentic task evaluations with features including sandboxing, tool calling, multi-agent primitives, and a web-based visualization tool.","tags":["ai-evaluation","inspect","framework","llm-benchmarks","agentic-tasks","model-testing","python"],"key_entities":[{"name":"Inspect","type":"technology","confidence":1},{"name":"UK AI Security Institute","type":"organization","confidence":0.95},{"name":"Meridian Labs","type":"organization","confidence":0.9},{"name":"SimpleQA","type":"technology","confidence":0.85},{"name":"Claude Code","type":"technology","confidence":0.8},{"name":"VS Code Extension","type":"technology","confidence":0.75},{"name":"OpenAI","type":"organization","confidence":0.9},{"name":"Anthropic","type":"organization","confidence":0.9},{"name":"Google","type":"organization","confidence":0.85},{"name":"Docker","type":"technology","confidence":0.7},{"name":"Kubernetes","type":"technology","confidence":0.7},{"name":"HuggingFace","type":"technology","confidence":0.8},{"name":"AI evaluation","type":"concept","confidence":0.95},{"name":"MCP tools","type":"technology","confidence":0.7}],"classification":"framework","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:48:54.955Z"},"token_counts":{"approximate":2105,"cl100k":1781},"content_hash":"sha256:c91a7f94527cea3c006158a99487070fb4ddcca74c2e96206ea25c1fb84f9d52","acp_version":"0.2","body_available":true,"body_tokens":2105,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"152b6531-f45f-4cc8-8659-edea5a31bb2d","position":3,"title":"Evaluation concepts - Docs by LangChain","url":"https://docs.langchain.com/langsmith/evaluation-concepts","note":"This page provides an overview of evaluation concepts within the LangChain framework. It covers key principles and methodologies for assessing the performance of language models and their applications.","image":{"url":"https://ucarecdn.com/185ef330-24e6-4340-964b-4225f3ff68e5/","alt":"Evaluation concepts - Docs by LangChain","width":2858,"height":1016},"direct_link":"https://stacklist.com/card/152b6531-f45f-4cc8-8659-edea5a31bb2d","created_at":"2026-07-02T09:48:46.381Z","updated_at":null,"aco":{"summary":"LangSmith Evaluation Concepts provides a comprehensive framework for measuring LLM application quality through offline (pre-deployment) and online (production monitoring) evaluations. The document covers evaluation lifecycle stages, core targets such as datasets and examples, and strategies for continuous improvement through iterative feedback loops.","tags":["llm-evaluation","langsmith","offline-evaluation","online-evaluation","quality-monitoring","regression-testing","rag"],"key_entities":[{"name":"LangSmith","type":"technology","confidence":0.99},{"name":"Offline Evaluation","type":"concept","confidence":0.95},{"name":"Online Evaluation","type":"concept","confidence":0.95},{"name":"RAG","type":"concept","confidence":0.85},{"name":"Evaluation Lifecycle","type":"concept","confidence":0.9},{"name":"Regression Testing","type":"concept","confidence":0.8},{"name":"Anomaly Detection","type":"concept","confidence":0.75},{"name":"LLM","type":"technology","confidence":0.95}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:48:58.280Z"},"token_counts":{"approximate":3967,"cl100k":2741},"content_hash":"sha256:20e98766884afcd0cc9bf6c3e782fe94df358be44c9a32b3bd86ee56124ef1fb","acp_version":"0.2","body_available":true,"body_tokens":3967,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"65a259a7-1ea6-4f48-ba11-489d4d139141","position":4,"title":"Evals: Framework for Evaluating LLMs and Benchmarks","url":"https://github.com/openai/evals","note":"Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks. It provides tools and resources for assessing the performance of language models effectively.","image":{"url":"https://ucarecdn.com/2f636057-c341-4265-a9c6-8da495313021/","alt":"Evals: Framework for Evaluating LLMs and Benchmarks","width":1200,"height":600},"direct_link":"https://stacklist.com/card/65a259a7-1ea6-4f48-ba11-489d4d139141","created_at":"2026-07-02T09:48:47.710Z","updated_at":null,"aco":{"summary":"OpenAI Evals is a framework for evaluating large language models (LLMs) and LLM-based systems, offering a registry of existing evals and the ability to create custom evaluations. The documentation covers setup, installation via pip and Git-LFS, running and writing evals, including support for advanced use cases like prompt chains, tool-using agents, and logging results to Snowflake.","tags":["openai","evals","llm-evaluation","framework","python","model-testing","custom-evals"],"key_entities":[{"name":"OpenAI","type":"organization","confidence":1},{"name":"OpenAI Evals","type":"technology","confidence":1},{"name":"Greg Brockman","type":"person","confidence":0.95},{"name":"Git-LFS","type":"technology","confidence":0.9},{"name":"Python","type":"technology","confidence":0.85},{"name":"Snowflake","type":"technology","confidence":0.8},{"name":"Weights & Biases","type":"technology","confidence":0.8},{"name":"LLM evaluation","type":"concept","confidence":0.95},{"name":"Completion Function Protocol","type":"concept","confidence":0.8}],"classification":"framework","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:48:56.388Z"},"token_counts":{"approximate":1352,"cl100k":1162},"content_hash":"sha256:ef31f15304b62eb6c223292614c401cfee761d0dde9b823e61deab6969736458","acp_version":"0.2","body_available":true,"body_tokens":1352,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"f09464ad-2a5d-441a-8ada-c34551b935eb","position":5,"title":"DeepEval 5-min Quickstart | DeepEval","url":"https://deepeval.com/docs/getting-started","note":"This quickstart takes you from installing DeepEval to your first passing eval in a few minutes. You'll create a small test case, choose a metric, and run it to evaluate LLMs effectively.","image":{"url":"https://ucarecdn.com/5ce5fb36-2aab-46fc-9ee2-b4f56c83713d/","alt":"DeepEval 5-min Quickstart | DeepEval","width":3456,"height":2062},"direct_link":"https://stacklist.com/card/f09464ad-2a5d-441a-8ada-c34551b935eb","created_at":"2026-07-02T09:48:49.127Z","updated_at":null,"aco":{"summary":"DeepEval's 5-minute quickstart guide walks users through installing the framework, creating an LLM test case with input/output pairs, choosing a GEval metric, and running end-to-end evaluations locally. The tutorial covers environment setup, single-turn and multi-turn test cases, metric thresholds, regression detection, and integration with the Confident AI cloud platform.","tags":["deepeval","llm-evaluation","quickstart","testing","ai-quality","python","confident-ai"],"key_entities":[{"name":"DeepEval","type":"technology","confidence":0.99},{"name":"Confident AI","type":"organization","confidence":0.95},{"name":"GEval","type":"concept","confidence":0.92},{"name":"LLM evaluation","type":"concept","confidence":0.95},{"name":"Python","type":"technology","confidence":0.85},{"name":"LLMTestCase","type":"concept","confidence":0.88},{"name":"test run","type":"concept","confidence":0.8}],"classification":"tutorial","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:48:58.286Z"},"token_counts":{"approximate":9170,"cl100k":8584},"content_hash":"sha256:280daed03e063c935dc0e57af6ef04ed0cb02d5a889166da362ab468fd4347d8","acp_version":"0.2","body_available":true,"body_tokens":9170,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"abf80d19-21a5-40a9-919d-184da4850d25","position":6,"title":"Agentic or Tool Use - Ragas","url":"https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/agents/","note":"This page provides an evaluation framework for assessing AI applications through various metrics. It outlines the available metrics related to agentic or tool use, helping developers understand and improve their AI systems.","image":{"url":"https://ucarecdn.com/6bcaeb28-f2a6-49e3-828f-b286d94ff6f3/","alt":"Agentic or Tool Use - Ragas","width":1280,"height":800},"direct_link":"https://stacklist.com/card/abf80d19-21a5-40a9-919d-184da4850d25","created_at":"2026-07-02T09:48:50.463Z","updated_at":null,"aco":{"summary":"Ragas Agentic or Tool Use Metrics documentation describes evaluation dimensions for AI agent and tool-use workflows, focusing on the TopicAdherence metric that measures an AI system's ability to stay within predefined domains. The metric computes precision, recall, and F1 score using reference topics and user input, with a detailed Python code example demonstrating evaluation of conversational interactions.","tags":["ragas","agentic-metrics","topic-adherence","llm-evaluation","tool-use","conversational-ai","precision-recall"],"key_entities":[{"name":"Ragas","type":"technology","confidence":0.95},{"name":"TopicAdherence","type":"concept","confidence":0.95},{"name":"Agentic Metrics","type":"concept","confidence":0.85},{"name":"OpenAI","type":"technology","confidence":0.9},{"name":"gpt-4o-mini","type":"technology","confidence":0.9},{"name":"Precision-Recall-F1","type":"concept","confidence":0.85},{"name":"Albert Einstein","type":"person","confidence":0.8},{"name":"AsyncOpenAI","type":"technology","confidence":0.75}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:49:02.518Z"},"token_counts":{"approximate":5446,"cl100k":5118},"content_hash":"sha256:747a75f8a218fcce8d1ec46cfc136684c3536a74afcf3ab99d9f59555bd3132f","acp_version":"0.2","body_available":true,"body_tokens":5446,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"f3f3a0a9-9839-46cd-a891-ea28b3294379","position":7,"title":"Welcome to BrowserGym’s documentation!","url":"https://browsergym.readthedocs.io/latest/","note":"BrowserGym provides comprehensive documentation for users to understand and utilize its features effectively. This page serves as a guide to help users navigate through the functionalities offered by BrowserGym version 0.3.","image":{"url":"https://ucarecdn.com/ffea773f-30d9-4cfb-8e50-096820eccb3f/","alt":"Welcome to BrowserGym’s documentation!","width":1280,"height":800},"direct_link":"https://stacklist.com/card/f3f3a0a9-9839-46cd-a891-ea28b3294379","created_at":"2026-07-02T09:48:51.696Z","updated_at":null,"aco":{"summary":"BrowserGym is a Python library providing a gym environment for web task automation in the Chromium browser, featuring benchmarks such as MiniWob++, WebArena, and WorkArena. The documentation covers installation, example code, API details, core action and observation spaces, environments, and tutorials.","tags":["browsergym","web-automation","python","gym-environment","chromium","benchmarks","reinforcement-learning"],"key_entities":[{"name":"BrowserGym","type":"technology","confidence":0.98},{"name":"Python","type":"technology","confidence":0.95},{"name":"Chromium","type":"technology","confidence":0.93},{"name":"MiniWob++","type":"technology","confidence":0.9},{"name":"WebArena","type":"technology","confidence":0.9},{"name":"WorkArena","type":"technology","confidence":0.9},{"name":"web task automation","type":"concept","confidence":0.88},{"name":"gym environment","type":"concept","confidence":0.85}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:49:01.014Z"},"token_counts":{"approximate":97,"cl100k":70},"content_hash":"sha256:6eda1560fdd3b84e203c60ba087802c85390705f3cc4b0e90bbcfea64e25d280","acp_version":"0.2","body_available":true,"body_tokens":97,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"e6eb95cd-8fd9-4c64-b3a7-d08205fb4b01","position":8,"title":"WebArena: Benchmarks for Autonomous Web Agents","url":"https://webarena.dev/","note":"WebArena is a comprehensive suite of benchmarks designed to facilitate the development and evaluation of autonomous web agents. It provides tools and metrics to assess the performance and capabilities of these agents in various web environments.","image":{"url":"https://ucarecdn.com/a112d31f-8c39-46f6-9cee-1df10ec05511/","alt":"WebArena: Benchmarks for Autonomous Web Agents","width":1280,"height":800},"direct_link":"https://stacklist.com/card/e6eb95cd-8fd9-4c64-b3a7-d08205fb4b01","created_at":"2026-07-02T09:48:52.914Z","updated_at":null,"aco":{"summary":"WebArena-x is a suite of benchmarks for building and evaluating autonomous web agents across realistic web environments. The project includes WebArena, WebArena-Infinity, VisualWebArena, and TheAgentCompany, presented at venues such as NeurIPS 2024, ACL 2024, and ICML 2025.","tags":["web-agents","benchmarking","autonomous-agents","multimodal","web-environment","llm-agents","evaluation"],"key_entities":[{"name":"WebArena","type":"technology","confidence":0.98},{"name":"WebArena-Infinity","type":"technology","confidence":0.95},{"name":"VisualWebArena","type":"technology","confidence":0.95},{"name":"TheAgentCompany","type":"technology","confidence":0.93},{"name":"NeurIPS 2024","type":"event","confidence":0.97},{"name":"ACL 2024","type":"event","confidence":0.95},{"name":"ICML 2025","type":"event","confidence":0.95},{"name":"autonomous web agents","type":"concept","confidence":0.96},{"name":"multimodal agents","type":"concept","confidence":0.88}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:49:04.878Z"},"token_counts":{"approximate":111,"cl100k":91},"content_hash":"sha256:574a882072eddf2131082fd452470ed10b4144d0fbe8c344eeb1c817e1e7a4c3","acp_version":"0.2","body_available":true,"body_tokens":111,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"373a54aa-196d-40c3-a3ca-4fb60cb05b43","position":9,"title":"$τ$-bench: A Benchmark for Tool-Agent-User Interaction","url":"https://arxiv.org/abs/2406.12045","note":"This page presents the arXiv paper 2406.12045, which introduces $τ$-bench, a benchmark designed to evaluate interactions among tools, agents, and users in real-world domains. The paper aims to provide a comprehensive framework for assessing the effectiveness and efficiency of these interactions.","image":{"url":"https://ucarecdn.com/0b1ea719-c0fc-445d-953d-b0eabf04817c/","alt":"$τ$-bench: A Benchmark for Tool-Agent-User Interaction","width":336,"height":96},"direct_link":"https://stacklist.com/card/373a54aa-196d-40c3-a3ca-4fb60cb05b43","created_at":"2026-07-02T09:48:54.151Z","updated_at":null,"aco":{"summary":"τ-bench is a benchmark for evaluating language agent interactions with simulated human users in real-world domains, testing agents' ability to use domain-specific API tools and follow policy guidelines. The paper introduces a pass^k reliability metric and finds that even state-of-the-art models like GPT-4o succeed on less than 50% of tasks, highlighting the need for more consistent and rule-following agent methods.","tags":["benchmark","language-agents","tool-use","human-agent-interaction","evaluation","function-calling","reliability"],"key_entities":[{"name":"τ-bench","type":"concept","confidence":1},{"name":"Shunyu Yao","type":"person","confidence":0.95},{"name":"Noah Shinn","type":"person","confidence":0.9},{"name":"Pedram Razavi","type":"person","confidence":0.9},{"name":"Karthik Narasimhan","type":"person","confidence":0.95},{"name":"GPT-4o","type":"technology","confidence":0.95},{"name":"arXiv","type":"organization","confidence":0.95},{"name":"pass^k metric","type":"concept","confidence":0.9},{"name":"function-calling agents","type":"concept","confidence":0.85}],"classification":"reference","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:49:07.947Z"},"token_counts":{"approximate":988,"cl100k":947},"content_hash":"sha256:7ef81f098b387dbd454e75a834ee03961b3c1c3842b8da439bba72c34c190778","acp_version":"0.2","body_available":true,"body_tokens":988,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"6aa80af3-cff3-41d8-8f2e-6f69fa1dcc2d","position":10,"title":"GitHub - THUDM/AgentBench: A Comprehensive Benchmark","url":"https://github.com/THUDM/AgentBench","note":"THUDM/AgentBench is a comprehensive benchmark designed to evaluate large language models (LLMs) as agents. This resource aims to facilitate research and development in the field of artificial intelligence by providing standardized evaluation metrics and datasets.","image":{"url":"https://ucarecdn.com/f06f9373-3f1b-4e53-bd4c-e52adc16ce5c/","alt":"GitHub - THUDM/AgentBench: A Comprehensive Benchmark","width":1200,"height":600},"direct_link":"https://stacklist.com/card/6aa80af3-cff3-41d8-8f2e-6f69fa1dcc2d","created_at":"2026-07-02T09:48:55.382Z","updated_at":null,"aco":{"summary":"AgentBench is a comprehensive benchmark framework for evaluating LLMs as autonomous agents across diverse environments including OS interaction, databases, knowledge graphs, web shopping, and more. The latest version (AgentBench FC) introduces function-calling style prompts integrated with AgentRL, featuring fully-containerized Docker deployment for five task environments.","tags":["agentbench","llm-agents","benchmark","reinforcement-learning","function-calling","docker","evaluation"],"key_entities":[{"name":"AgentBench","type":"technology","confidence":0.99},{"name":"AgentRL","type":"technology","confidence":0.95},{"name":"VisualAgentBench","type":"technology","confidence":0.9},{"name":"Docker Compose","type":"technology","confidence":0.85},{"name":"LLM-as-Agent","type":"concept","confidence":0.95},{"name":"function-calling","type":"concept","confidence":0.9},{"name":"ALFWorld","type":"technology","confidence":0.85},{"name":"WebShop","type":"technology","confidence":0.85},{"name":"Mind2Web","type":"technology","confidence":0.8},{"name":"Redis","type":"technology","confidence":0.75}],"classification":"framework","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:49:05.463Z"},"token_counts":{"approximate":2274,"cl100k":2109},"content_hash":"sha256:83e72a6bc7ecf099156fcbc237ac6b859afa663b139b1311cd8558132c1d863b","acp_version":"0.2","body_available":true,"body_tokens":2274,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"1d05f8cd-8fb8-43dc-9225-d9a5b3a49f25","position":11,"title":"AgentRewardBench","url":"https://agent-reward-bench.github.io/","note":"AgentRewardBench is a platform designed for evaluating and benchmarking reinforcement learning agents. It provides tools and metrics to assess agent performance in various environments.","image":{"url":"https://ucarecdn.com/da0ec6de-bdac-4d9d-986c-490675e41535/","alt":"AgentRewardBench","width":1280,"height":800},"direct_link":"https://stacklist.com/card/1d05f8cd-8fb8-43dc-9225-d9a5b3a49f25","created_at":"2026-07-02T09:49:06.061Z","updated_at":null,"aco":{"summary":"AgentRewardBench is a benchmark and Python library for evaluating automatic evaluators of web agent trajectories, including LLM judges. It provides tools, environments, evaluation metrics, a Hugging Face dataset, and a leaderboard for ranking automatic evaluators across different benchmarks.","tags":["agent-evaluation","web-agents","benchmark","llm-judges","trajectories","reward-model","python-library"],"key_entities":[{"name":"AgentRewardBench","type":"technology","confidence":1},{"name":"Xing Han Lù","type":"person","confidence":0.95},{"name":"Amirhossein Kazemnejad","type":"person","confidence":0.95},{"name":"Nicholas Meade","type":"person","confidence":0.9},{"name":"Siva Reddy","type":"person","confidence":0.9},{"name":"Christopher J. Pal","type":"person","confidence":0.9},{"name":"McGill-NLP","type":"organization","confidence":0.95},{"name":"Hugging Face Hub","type":"technology","confidence":0.9},{"name":"web agent trajectories","type":"concept","confidence":0.95},{"name":"LLM judges","type":"concept","confidence":0.9}],"classification":"framework","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:49:15.134Z"},"token_counts":{"approximate":606,"cl100k":594},"content_hash":"sha256:9465a3e299b570373ee6c742dcf5e223105c8355f6747fa6a8f26fa4ec3fb1e8","acp_version":"0.2","body_available":true,"body_tokens":606,"visibility":"public","agent_accessible":true,"status":"final"}},{"id":"100ce629-805a-4644-832f-4ee4c5e60d1c","position":12,"title":"Measuring AI Ability to Complete Long Tasks","url":"https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/","note":"We propose measuring AI performance in terms of the *length* of tasks AI agents can complete. We show that this metric has been consistently exponentially increasing over the past 6 years, with a doubling time of around 7 months.","image":{"url":"https://ucarecdn.com/c499bab0-1d06-4ead-b890-40e9d9481dfe/","alt":"Measuring AI Ability to Complete Long Tasks","width":1569,"height":893},"direct_link":"https://stacklist.com/card/100ce629-805a-4644-832f-4ee4c5e60d1c","created_at":"2026-07-02T09:49:07.429Z","updated_at":null,"aco":{"summary":"METR proposes measuring AI performance by the length of tasks AI agents can autonomously complete, showing this metric has doubled approximately every 7 months over the past 6 years. Extrapolating this exponential trend predicts AI agents capable of independently completing multi-day software tasks within under a decade.","tags":["ai-agents","benchmarking","task-completion","time-horizon","capability-forecasting","exponential-growth","frontier-models"],"key_entities":[{"name":"METR","type":"organization","confidence":0.99},{"name":"Thomas Kwa","type":"person","confidence":0.95},{"name":"Ben West","type":"person","confidence":0.95},{"name":"Joel Becker","type":"person","confidence":0.95},{"name":"time horizon","type":"concept","confidence":0.97},{"name":"task-completion benchmarking","type":"concept","confidence":0.92},{"name":"exponential capability growth","type":"concept","confidence":0.9},{"name":"frontier language models","type":"technology","confidence":0.88}],"classification":"analysis","language":"en","confidence":0.85,"provenance":{"model":"claude-opus-4-6","tool":"@stacklist/mcp-server@2.0.0","confidence":0.85,"timestamp":"2026-07-02T09:49:15.129Z"},"token_counts":{"approximate":3423,"cl100k":2828},"content_hash":"sha256:a4deeb60d51f98f0ebb1246b0670c9777c9b690e1b94b28a0c6ac2a68682a3fb","acp_version":"0.2","body_available":true,"body_tokens":3423,"visibility":"public","agent_accessible":true,"status":"final"}}],"stats":{"likes_count":0,"items_count":12},"_links":{"self":"/api/public/stack/29e01ac8-6202-4abf-ab14-a0613026186b.json","html":"https://stacklist.com/c/technology/stack/29e01ac8-6202-4abf-ab14-a0613026186b","rss":"/api/feeds/stack/29e01ac8-6202-4abf-ab14-a0613026186b/rss","json_feed":"/api/feeds/stack/29e01ac8-6202-4abf-ab14-a0613026186b/json"}}