Curated by
More in AI Agent Evaluation Frameworks and Benchmarks
See all 12 →GitHub - THUDM/AgentBench: A Comprehensive Benchmark
AgentBench is a comprehensive benchmark framework for evaluating LLMs as autonomous agents across diverse environments including OS interaction, databases, knowledge graphs, web shopping, and more. The latest version (AgentBench FC) introduces function-calling style prompts integrated with AgentRL, featuring fully-containerized Docker deployment for five task environments.
Built for AI agentsACO · 2274 tokens
Summary
AgentBench is a comprehensive benchmark framework for evaluating LLMs as autonomous agents across diverse environments including OS interaction, databases, knowledge graphs, web shopping, and more. The latest version (AgentBench FC) introduces function-calling style prompts integrated with AgentRL, featuring fully-containerized Docker deployment for five task environments.
Tags
agentbench · llm-agents · benchmark · reinforcement-learning · function-calling · docker · evaluation
Key entities
AgentBench (technology, 0.99) · AgentRL (technology, 0.95) · VisualAgentBench (technology, 0.9) · Docker Compose (technology, 0.85) · LLM-as-Agent (concept, 0.95) · function-calling (concept, 0.9) · ALFWorld (technology, 0.85) · WebShop (technology, 0.85) · Mind2Web (technology, 0.8) · Redis (technology, 0.75)
Classification
framework · language en · status final
Provenance
claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026