Curated by

avatar

amrutha-kothapalli-307

stacklist.com/amrutha-kothapalli-307

More in AI Agent Evaluation Frameworks and Benchmarks

See all 12 →

GitHub - THUDM/AgentBench: A Comprehensive Benchmark

AgentBench is a comprehensive benchmark framework for evaluating LLMs as autonomous agents across diverse environments including OS interaction, databases, knowledge graphs, web shopping, and more. The latest version (AgentBench FC) introduces function-calling style prompts integrated with AgentRL, featuring fully-containerized Docker deployment for five task environments.

View card
Built for AI agentsACO · 2274 tokens

Summary

AgentBench is a comprehensive benchmark framework for evaluating LLMs as autonomous agents across diverse environments including OS interaction, databases, knowledge graphs, web shopping, and more. The latest version (AgentBench FC) introduces function-calling style prompts integrated with AgentRL, featuring fully-containerized Docker deployment for five task environments.

Tags

agentbench · llm-agents · benchmark · reinforcement-learning · function-calling · docker · evaluation

Key entities

AgentBench (technology, 0.99) · AgentRL (technology, 0.95) · VisualAgentBench (technology, 0.9) · Docker Compose (technology, 0.85) · LLM-as-Agent (concept, 0.95) · function-calling (concept, 0.9) · ALFWorld (technology, 0.85) · WebShop (technology, 0.85) · Mind2Web (technology, 0.8) · Redis (technology, 0.75)

Classification

framework · language en · status final

Provenance

claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026