Skip to content
ContentLora

    Tip: press / anywhere to search.

    concept

    SWE-bench

    Also known as SWE-bench Verified, SWE-bench Pro

    SWE-bench is a benchmark that asks AI systems to resolve real software issues taken from GitHub, drawn from 2,294 problems in 12 popular Python projects.[1] The best model solved 1.96% when it launched in 2023;[2] the 2026 AI Index reports near 100% on its human-filtered Verified subset.[3]

    Editor reviewedUpdated Frontier AIArtificial intelligenceComputing
    Key facts

    The 2026 AI Index calls SWE-bench Verified a key coding benchmark.[3] Instead of short puzzles, SWE-bench uses real issues and the pull requests that fixed them from popular open-source Python projects.[1] A model must change a codebase to resolve the issue, which often means working across several files and functions.[1]

    How it works

    The original benchmark, released in October 2023 and published at ICLR 2024, contains 2,294 problems from 12 Python repositories.[1] As of October 2026 the project maintains several leaderboards: SWE-bench Verified, a 500-task human-filtered subset, along with Lite, Multilingual and Multimodal versions.[4]

    The variants

    The project now runs a family of benchmarks.[4] SWE-bench Verified, created with OpenAI and released in August 2024, is a subset of 500 problems that human annotators reviewed to make sure the descriptions were clear, the tests correct and the tasks solvable.[5][4] SWE-bench Lite is a 300-task subset built for faster, cheaper runs.[6] SWE-bench Multilingual moves beyond Python and is deliberately limited to 300 tasks so that it runs quickly.[7] SWE-bench Multimodal, released in October 2024, has 617 tasks from 17 JavaScript libraries, each including at least one image; the best system at release resolved 12% of them.[8]

    From 2% to near 100%

    When SWE-bench launched, the best-performing model, Claude 2, solved 1.96% of issues.[2] Coding agents work in a loop, using tools and checking results such as code execution at each step.[9] The 2026 AI Index reports that performance on SWE-bench Verified rose from 60% to near 100% in a single year.[3] In April 2026 Anthropic reported 93.9% on Verified for Claude Mythos Preview, against 80.8% for Claude Opus 4.6.[10]

    Saturation and harder variants

    With Verified close to saturation, developers have started reporting harder tests. SWE-Bench Pro, published by Scale AI researchers in September 2025, has 1,865 problems from 41 repositories, including held-out and proprietary sets meant to resist contamination, and tasks that can take a professional engineer hours to days.[11] Anthropic gave Claude Mythos Preview 77.8% on SWE-bench Pro, against 53.4% for Opus 4.6.[10] Developers also report other agentic tests: Anthropic reported 66.4% on Terminal-Bench 4.0 for Claude Opus 5.5 in September 2026.[12]

    Limits

    A benchmark score is not the same as productivity in real projects. METR’s 2025 randomized trial found that experienced open-source developers using early-2025 AI tools took 19% longer to finish tasks,[13] and its 2026 follow-up was skewed by developers choosing not to take part because they did not want to work without AI.[14] METR‘s time-horizon measure asks a different question: how long a task, in human time, an agent can complete.[15]

    SWE-bench sits alongside other agent tests. OSWorld measures computer use across operating systems: at its 2024 release humans succeeded on 72.36% of its 369 tasks and the best model on 12.24%,[16] and the 2026 AI Index reports agents at about 66%.[17] ARC-AGI-3 tests whether agents can learn unfamiliar games without instructions.[18]

    Questions readers ask

    What does SWE-bench test?

    Whether an AI system can resolve real issues from GitHub by modifying a codebase, using 2,294 problems drawn from 12 popular Python repositories.[1]

    What is SWE-bench Verified?

    A subset of 500 human-filtered tasks, maintained by the SWE-bench project alongside Lite, Multilingual and Multimodal versions.[4]

    How fast have scores risen?

    The best model solved 1.96% of issues at launch in 2023. The 2026 AI Index reports that performance on SWE-bench Verified rose from 60% to near 100% in a single year.[2][3]

    Is SWE-bench Verified saturated?

    Close to it. With scores near 100%, developers also report harder tests such as SWE-bench Pro, where Anthropic reported 77.8% for Claude Mythos Preview.[3][10]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      SWE-bench, published at ICLR 2024, contains 2,294 software engineering problems drawn from real GitHub issues and pull requests across 12 popular Python repositories. confirmedas of 2024-11-11

    2. [2]

      When SWE-bench was introduced in 2023, the best-performing model, Claude 2, solved 1.96% of the issues. confirmedas of 2023-10-10

    3. [3]

      The 2026 AI Index reports that performance on the SWE-bench Verified coding benchmark rose from 60% to near 100% in a single year. confirmedas of 2026-04-01

    4. [4]

      As of October 2026 the SWE-bench project runs several leaderboards, including SWE-bench Verified (500 human-filtered tasks), Lite, Multilingual and Multimodal. confirmedas of 2026-10-10

    5. [5]

      SWE-bench Verified, created with OpenAI and released in August 2024, consists of 500 problems that human annotators reviewed to confirm they were clear and solvable. confirmedas of 2026-10-10

    6. [6]

      SWE-bench Lite is a 300-task subset of the original benchmark designed for faster, cheaper evaluation. confirmedas of 2026-10-10

    7. [7]

      SWE-bench Multilingual extends the benchmark beyond Python and is deliberately limited to 300 tasks so that it runs quickly. confirmedas of 2026-10-10

    8. [8]

      SWE-bench Multimodal, released in October 2024, has 617 tasks from 17 JavaScript libraries, each with at least one image, and the best system at release resolved 12% of them. confirmedas of 2024-10-04

    9. [9]

      Anthropic describes agents as typically language models using tools based on feedback from their environment in a loop, checking results such as tool outputs or code execution at each step. confirmedas of 2024-12-19

    10. [10]

      Anthropic reported that Claude Mythos Preview scored 93.9% on SWE-bench Verified, 77.8% on SWE-bench Pro and 83.1% on CyberGym, against 80.8%, 53.4% and 66.6% for Claude Opus 4.6. confirmedas of 2026-04-07

    11. [11]

      SWE-Bench Pro, published by Scale AI researchers in September 2025, contains 1,865 problems from 41 repositories, including held-out and proprietary sets meant to resist training-data contamination, with tasks that may take professional engineers hours to days. confirmedas of 2025-11-14

    12. [12]

      Anthropic reported that Claude Opus 5.5 scored 66.4% on Terminal-Bench 4.0, 81.8% on OSWorld 2.1 and 67.7% on Humanity's Last Exam with tools. confirmedas of 2026-09-22

    13. [13]

      In a 2025 randomized controlled trial with 16 experienced open-source developers on 246 tasks, METR found that allowing early-2025 AI tools increased task completion time by 19%, although developers had expected a 24% speedup. confirmedas of 2025-07-12

    14. [14]

      METR's follow-up study with 57 developers and more than 800 tasks estimated a speedup of -18% (CI -38% to +9%) for returning developers and -4% (CI -15% to +9%) for new ones, results it said were biased by selection effects. confirmedas of 2026-02-24

    15. [15]

      METR's 50%-task-completion time horizon is the length of tasks, measured by how long human experts take, that an AI model completes with 50% success. confirmedas of 2025-03-18

    16. [16]

      OSWorld, a 2024 benchmark of 369 real computer tasks across operating systems, found humans succeeded on 72.36% of tasks against 12.24% for the best AI model at the time. confirmedas of 2024-04-11

    17. [17]

      The 2026 AI Index reports that AI agents went from about 12% to about 66% task success on OSWorld. confirmedas of 2026-04-01

    18. [18]

      ARC-AGI-3 consists of game-like environments with no instructions, rules or stated goals, which an agent must explore to work out how they function. confirmedas of 2026-03-25

      • Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · What ARC-AGI-3 measures (retrieved 2026-10-10)
    19. [19]

      The MCP documentation describes MCP as an open-source standard for connecting AI applications to data sources, tools and workflows, comparing it to a USB-C port for AI applications. confirmedas of 2026-10-10

    Revision history (2)
    1. Page created.
    2. Added how SWE-bench Verified was built, the Lite, Multilingual and Multimodal variants with their sizes, and SWE-Bench Pro; re-sourced the Verified and dataset-size citations to the project's own pages.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "SWE-bench." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/swe-bench

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.