Curated by
More in AI agent evaluation frameworks
See all 10 →SWE-bench: Can Language Models Resolve GitHub Issues?
SWE-bench is a benchmark for evaluating large language models on real-world software issues collected from GitHub, where models generate patches to resolve described problems. The repository includes setup instructions using Docker for reproducible evaluations, along with extensions like SWE-bench Multimodal, SWE-bench Verified, and cloud-based evaluation via Modal and sb-cli.
Built for AI agentsACO · 1728 tokens
Summary
SWE-bench is a benchmark for evaluating large language models on real-world software issues collected from GitHub, where models generate patches to resolve described problems. The repository includes setup instructions using Docker for reproducible evaluations, along with extensions like SWE-bench Multimodal, SWE-bench Verified, and cloud-based evaluation via Modal and sb-cli.
Tags
swe-bench · benchmark · large-language-models · github-issues · docker · software-engineering · evaluation
Key entities
SWE-bench (technology, 1) · Princeton NLP (organization, 0.95) · OpenAI (organization, 0.9) · Docker (technology, 0.95) · ICLR 2024 (event, 0.95) · ICLR 2025 (event, 0.9) · SWE-agent (technology, 0.9) · Modal (technology, 0.85) · SWE-bench Multimodal (technology, 0.9) · SWE-bench Verified (technology, 0.88) · OpenAI Preparedness (concept, 0.8) · sb-cli (technology, 0.85)
Classification
reference · language en · status final
Provenance
claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026