concept
SWE-bench
Also known as SWE-bench Verified, SWE-bench Pro
SWE-bench is a benchmark that asks AI systems to resolve real software issues taken from GitHub, drawn from 2,294 problems in 12 popular Python projects.[1] The best model solved 1.96% when it launched in 2023;[2] the 2026 AI Index reports near 100% on its human-filtered Verified subset.[3]
Key facts
The 2026 AI Index calls SWE-bench Verified a key coding benchmark.[3] Instead of short puzzles, SWE-bench uses real issues and the pull requests that fixed them from popular open-source Python projects.[1] A model must change a codebase to resolve the issue, which often means working across several files and functions.[1]
How it works
The original benchmark, released in October 2023 and published at ICLR 2024, contains 2,294 problems from 12 Python repositories.[1] As of October 2026 the project maintains several leaderboards: SWE-bench Verified, a 500-task human-filtered subset, along with Lite, Multilingual and Multimodal versions.[4]
The variants
The project now runs a family of benchmarks.[4] SWE-bench Verified, created with OpenAI and released in August 2024, is a subset of 500 problems that human annotators reviewed to make sure the descriptions were clear, the tests correct and the tasks solvable.[5][4] SWE-bench Lite is a 300-task subset built for faster, cheaper runs.[6] SWE-bench Multilingual moves beyond Python and is deliberately limited to 300 tasks so that it runs quickly.[7] SWE-bench Multimodal, released in October 2024, has 617 tasks from 17 JavaScript libraries, each including at least one image; the best system at release resolved 12% of them.[8]
From 2% to near 100%
When SWE-bench launched, the best-performing model, Claude 2, solved 1.96% of issues.[2] Coding agents work in a loop, using tools and checking results such as code execution at each step.[9] The 2026 AI Index reports that performance on SWE-bench Verified rose from 60% to near 100% in a single year.[3] In April 2026 Anthropic reported 93.9% on Verified for Claude Mythos Preview, against 80.8% for Claude Opus 4.6.[10]
Saturation and harder variants
With Verified close to saturation, developers have started reporting harder tests. SWE-Bench Pro, published by Scale AI researchers in September 2025, has 1,865 problems from 41 repositories, including held-out and proprietary sets meant to resist contamination, and tasks that can take a professional engineer hours to days.[11] Anthropic gave Claude Mythos Preview 77.8% on SWE-bench Pro, against 53.4% for Opus 4.6.[10] Developers also report other agentic tests: Anthropic reported 66.4% on Terminal-Bench 4.0 for Claude Opus 5.5 in September 2026.[12]
Limits
A benchmark score is not the same as productivity in real projects. METR’s 2025 randomized trial found that experienced open-source developers using early-2025 AI tools took 19% longer to finish tasks,[13] and its 2026 follow-up was skewed by developers choosing not to take part because they did not want to work without AI.[14] METR‘s time-horizon measure asks a different question: how long a task, in human time, an agent can complete.[15]
Related benchmarks
SWE-bench sits alongside other agent tests. OSWorld measures computer use across operating systems: at its 2024 release humans succeeded on 72.36% of its 369 tasks and the best model on 12.24%,[16] and the 2026 AI Index reports agents at about 66%.[17] ARC-AGI-3 tests whether agents can learn unfamiliar games without instructions.[18]
Questions readers ask
What does SWE-bench test?
Whether an AI system can resolve real issues from GitHub by modifying a codebase, using 2,294 problems drawn from 12 popular Python repositories.[1]
What is SWE-bench Verified?
A subset of 500 human-filtered tasks, maintained by the SWE-bench project alongside Lite, Multilingual and Multimodal versions.[4]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
SWE-bench, published at ICLR 2024, contains 2,294 software engineering problems drawn from real GitHub issues and pull requests across 12 popular Python repositories. confirmedas of 2024-11-11
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · arXiv (Jimenez et al.; ICLR 2024) · 2023-10-10 · Abstract (the arXiv page renders the counts in LaTeX as $2,294$ and $12$) (retrieved 2026-10-10)
- SWE-bench Lite · SWE-bench (retrieved 2026-10-10)
- [2]
When SWE-bench was introduced in 2023, the best-performing model, Claude 2, solved 1.96% of the issues. confirmedas of 2023-10-10
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? · arXiv (Jimenez et al.; ICLR 2024) · 2023-10-10 · Abstract (retrieved 2026-10-10)
- [3]
The 2026 AI Index reports that performance on the SWE-bench Verified coding benchmark rose from 60% to near 100% in a single year. confirmedas of 2026-04-01
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [4]
As of October 2026 the SWE-bench project runs several leaderboards, including SWE-bench Verified (500 human-filtered tasks), Lite, Multilingual and Multimodal. confirmedas of 2026-10-10
- SWE-bench Verified · SWE-bench (retrieved 2026-10-10)
- SWE-bench leaderboards · SWE-bench · Site navigation (Verified, Lite, Multilingual, Multimodal leaderboards) (retrieved 2026-10-10)
- [5]
SWE-bench Verified, created with OpenAI and released in August 2024, consists of 500 problems that human annotators reviewed to confirm they were clear and solvable. confirmedas of 2026-10-10
- SWE-bench Verified · SWE-bench (retrieved 2026-10-10)
- [6]
SWE-bench Lite is a 300-task subset of the original benchmark designed for faster, cheaper evaluation. confirmedas of 2026-10-10
- SWE-bench Lite · SWE-bench (retrieved 2026-10-10)
- [7]
SWE-bench Multilingual extends the benchmark beyond Python and is deliberately limited to 300 tasks so that it runs quickly. confirmedas of 2026-10-10
- SWE-bench Multilingual · SWE-bench (retrieved 2026-10-10)
- [8]
SWE-bench Multimodal, released in October 2024, has 617 tasks from 17 JavaScript libraries, each with at least one image, and the best system at release resolved 12% of them. confirmedas of 2024-10-04
- SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? · arXiv (Yang et al.) · 2024-10-04 (retrieved 2026-10-10)
- [9]
Anthropic describes agents as typically language models using tools based on feedback from their environment in a loop, checking results such as tool outputs or code execution at each step. confirmedas of 2024-12-19
- Building effective agents · Anthropic · 2024-12-19 (retrieved 2026-10-10)
- [10]
Anthropic reported that Claude Mythos Preview scored 93.9% on SWE-bench Verified, 77.8% on SWE-bench Pro and 83.1% on CyberGym, against 80.8%, 53.4% and 66.6% for Claude Opus 4.6. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 · Benchmark table (retrieved 2026-10-10)
- [11]
SWE-Bench Pro, published by Scale AI researchers in September 2025, contains 1,865 problems from 41 repositories, including held-out and proprietary sets meant to resist training-data contamination, with tasks that may take professional engineers hours to days. confirmedas of 2025-11-14
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? · arXiv (Scale AI) · 2025-09-21 (retrieved 2026-10-10)
- [12]
Anthropic reported that Claude Opus 5.5 scored 66.4% on Terminal-Bench 4.0, 81.8% on OSWorld 2.1 and 67.7% on Humanity's Last Exam with tools. confirmedas of 2026-09-22
- Introducing Claude Opus 5.5 · Anthropic · 2026-09-22 · Benchmark table (retrieved 2026-10-10)
- [13]
In a 2025 randomized controlled trial with 16 experienced open-source developers on 246 tasks, METR found that allowing early-2025 AI tools increased task completion time by 19%, although developers had expected a 24% speedup. confirmedas of 2025-07-12
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity · arXiv (METR) · Abstract (retrieved 2026-10-10)
- [14]
METR's follow-up study with 57 developers and more than 800 tasks estimated a speedup of -18% (CI -38% to +9%) for returning developers and -4% (CI -15% to +9%) for new ones, results it said were biased by selection effects. confirmedas of 2026-02-24
- We are Changing our Developer Productivity Experiment Design · METR · 2026-02-24 (retrieved 2026-10-10)
- [15]
METR's 50%-task-completion time horizon is the length of tasks, measured by how long human experts take, that an AI model completes with 50% success. confirmedas of 2025-03-18
- Measuring AI Ability to Complete Long Tasks · arXiv (METR; NeurIPS 2025) · Abstract (retrieved 2026-10-10)
- [16]
OSWorld, a 2024 benchmark of 369 real computer tasks across operating systems, found humans succeeded on 72.36% of tasks against 12.24% for the best AI model at the time. confirmedas of 2024-04-11
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments · arXiv (Xie et al.) · 2024-04-11 · Abstract (retrieved 2026-10-10)
- [17]
The 2026 AI Index reports that AI agents went from about 12% to about 66% task success on OSWorld. confirmedas of 2026-04-01
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [18]
ARC-AGI-3 consists of game-like environments with no instructions, rules or stated goals, which an agent must explore to work out how they function. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · What ARC-AGI-3 measures (retrieved 2026-10-10)
- [19]
The MCP documentation describes MCP as an open-source standard for connecting AI applications to data sources, tools and workflows, comparing it to a USB-C port for AI applications. confirmedas of 2026-10-10
- What is the Model Context Protocol (MCP)? · Model Context Protocol project (retrieved 2026-10-10)
Revision history (2)
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"SWE-bench." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/swe-bench
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerHow frontier AI capabilities are measuredHow researchers measure what frontier AI can do: benchmarks like SWE-bench, Humanity's Last Exam and ARC-AGI, and METR time horizons.
- ExplainerHow AI agents workWhat an AI agent is and how it works: language models using tools in a loop, computer use, MCP connectors and long tasks.
- ExplainerFrontier AI in 2026: a crash courseA crash course on frontier AI in 2026: how reasoning models and AI agents work, who builds them, and where the frontier stands now.
- WikiModel Context Protocol (MCP)The Model Context Protocol is an open standard for connecting AI apps and agents to tools and data, now run by the Linux Foundation.
- WikiTest-time computeTest-time compute is the computing power an AI model spends while answering. Spending more of it is how reasoning models improve.
- ExplainerHow large language models workA plain guide to large language models: the Transformer, scaling laws and human-feedback training, at beginner and expert level.