Curated by

S

sushmikar-512

stacklist.com/sushmikar-512

More in AI agent evaluation frameworks

See all 10 →

SWE-bench: Can Language Models Resolve GitHub Issues?

SWE-bench is a benchmark for evaluating large language models on real-world software issues collected from GitHub, where models generate patches to resolve described problems. The repository includes setup instructions using Docker for reproducible evaluations, along with extensions like SWE-bench Multimodal, SWE-bench Verified, and cloud-based evaluation via Modal and sb-cli.

View card
Built for AI agentsACO · 1728 tokens

Summary

SWE-bench is a benchmark for evaluating large language models on real-world software issues collected from GitHub, where models generate patches to resolve described problems. The repository includes setup instructions using Docker for reproducible evaluations, along with extensions like SWE-bench Multimodal, SWE-bench Verified, and cloud-based evaluation via Modal and sb-cli.

Tags

swe-bench · benchmark · large-language-models · github-issues · docker · software-engineering · evaluation

Key entities

SWE-bench (technology, 1) · Princeton NLP (organization, 0.95) · OpenAI (organization, 0.9) · Docker (technology, 0.95) · ICLR 2024 (event, 0.95) · ICLR 2025 (event, 0.9) · SWE-agent (technology, 0.9) · Modal (technology, 0.85) · SWE-bench Multimodal (technology, 0.9) · SWE-bench Verified (technology, 0.88) · OpenAI Preparedness (concept, 0.8) · sb-cli (technology, 0.85)

Classification

reference · language en · status final

Provenance

claude-opus-4-6 via @stacklist/mcp-server@2.0.0, confidence 0.85, 2 Jul 2026