Curated by

avatar

amrutha-kothapalli-307

stacklist.com/amrutha-kothapalli-307

More in AI Agent Evaluation Frameworks and Benchmarks

See all 12 →

AgentRewardBench

AgentRewardBench is a benchmark and Python library for evaluating automatic evaluators of web agent trajectories, including LLM judges. It provides tools, environments, evaluation metrics, a Hugging Face dataset, and a leaderboard for ranking automatic evaluators across different benchmarks.

View card
Built for AI agentsACO · 606 tokens

Summary

AgentRewardBench is a benchmark and Python library for evaluating automatic evaluators of web agent trajectories, including LLM judges. It provides tools, environments, evaluation metrics, a Hugging Face dataset, and a leaderboard for ranking automatic evaluators across different benchmarks.

Tags

agent-evaluation · web-agents · benchmark · llm-judges · trajectories · reward-model · python-library

Key entities

AgentRewardBench (technology, 1) · Xing Han Lù (person, 0.95) · Amirhossein Kazemnejad (person, 0.95) · Nicholas Meade (person, 0.9) · Siva Reddy (person, 0.9) · Christopher J. Pal (person, 0.9) · McGill-NLP (organization, 0.95) · Hugging Face Hub (technology, 0.9) · web agent trajectories (concept, 0.95) · LLM judges (concept, 0.9)

Classification

framework · language en · status final

Provenance

claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026