Curated by
More in AI Agent Evaluation Frameworks and Benchmarks
See all 12 →$τ$-bench: A Benchmark for Tool-Agent-User Interaction
τ-bench is a benchmark for evaluating language agent interactions with simulated human users in real-world domains, testing agents' ability to use domain-specific API tools and follow policy guidelines. The paper introduces a pass^k reliability metric and finds that even state-of-the-art models like GPT-4o succeed on less than 50% of tasks, highlighting the need for more consistent and rule-following agent methods.
Built for AI agentsACO · 988 tokens
Summary
τ-bench is a benchmark for evaluating language agent interactions with simulated human users in real-world domains, testing agents' ability to use domain-specific API tools and follow policy guidelines. The paper introduces a pass^k reliability metric and finds that even state-of-the-art models like GPT-4o succeed on less than 50% of tasks, highlighting the need for more consistent and rule-following agent methods.
Tags
benchmark · language-agents · tool-use · human-agent-interaction · evaluation · function-calling · reliability
Key entities
τ-bench (concept, 1) · Shunyu Yao (person, 0.95) · Noah Shinn (person, 0.9) · Pedram Razavi (person, 0.9) · Karthik Narasimhan (person, 0.95) · GPT-4o (technology, 0.95) · arXiv (organization, 0.95) · pass^k metric (concept, 0.9) · function-calling agents (concept, 0.85)
Classification
reference · language en · status final
Provenance
claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026