Curated by
More in AI Agent Evaluation Frameworks and Benchmarks
See all 12 →Measuring AI Ability to Complete Long Tasks
METR proposes measuring AI performance by the length of tasks AI agents can autonomously complete, showing this metric has doubled approximately every 7 months over the past 6 years. Extrapolating this exponential trend predicts AI agents capable of independently completing multi-day software tasks within under a decade.
Built for AI agentsACO · 3423 tokens
Summary
METR proposes measuring AI performance by the length of tasks AI agents can autonomously complete, showing this metric has doubled approximately every 7 months over the past 6 years. Extrapolating this exponential trend predicts AI agents capable of independently completing multi-day software tasks within under a decade.
Tags
ai-agents · benchmarking · task-completion · time-horizon · capability-forecasting · exponential-growth · frontier-models
Key entities
METR (organization, 0.99) · Thomas Kwa (person, 0.95) · Ben West (person, 0.95) · Joel Becker (person, 0.95) · time horizon (concept, 0.97) · task-completion benchmarking (concept, 0.92) · exponential capability growth (concept, 0.9) · frontier language models (technology, 0.88)
Classification
analysis · language en · status final
Provenance
claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026