Curated by
More in AI Agent Evaluation Frameworks and Benchmarks
See all 12 →AgentRewardBench
AgentRewardBench is a benchmark and Python library for evaluating automatic evaluators of web agent trajectories, including LLM judges. It provides tools, environments, evaluation metrics, a Hugging Face dataset, and a leaderboard for ranking automatic evaluators across different benchmarks.
Built for AI agentsACO · 606 tokens
Summary
AgentRewardBench is a benchmark and Python library for evaluating automatic evaluators of web agent trajectories, including LLM judges. It provides tools, environments, evaluation metrics, a Hugging Face dataset, and a leaderboard for ranking automatic evaluators across different benchmarks.
Tags
agent-evaluation · web-agents · benchmark · llm-judges · trajectories · reward-model · python-library
Key entities
AgentRewardBench (technology, 1) · Xing Han Lù (person, 0.95) · Amirhossein Kazemnejad (person, 0.95) · Nicholas Meade (person, 0.9) · Siva Reddy (person, 0.9) · Christopher J. Pal (person, 0.9) · McGill-NLP (organization, 0.95) · Hugging Face Hub (technology, 0.9) · web agent trajectories (concept, 0.95) · LLM judges (concept, 0.9)
Classification
framework · language en · status final
Provenance
claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026