SWE-bench Pro

Scale AI's harder, contamination-resistant successor in spirit to SWE-bench Verified: 1,865 long-horizon tasks across public copyleft, held-out and private commercial codebases.

Also known as: SWE-Bench Pro

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
SubcategoryGitHub issue resolution / patch generation
Page statusactive
Metric% resolved
Directionhigher_is_better
Unit%
Dataset size1865
Dataset licenceMIT
PublisherScale AI (Scale Research Team)

What it measures

SWE-bench Pro follows the same task shape as SWE-bench — resolve a real issue in a real repository with a patch — but selects for longer, more involved changes and for source code that is unlikely to be in a model's training data. Scale AI built it because frontier models were already scoring above 70% on SWE-bench Verified, and wanted a benchmark that separates genuine software-engineering ability from memorised or near-memorised solutions on well-known open-source Python projects.

Task format

Given an issue and repository access, the system produces a patch, applied and graded inside a container against the repository's own tests, in the same fail-to-pass / pass-to-pass style as SWE-bench. Tasks are deliberately long-horizon: the paper reports resolved instances require changes averaging 107.4 lines of code across 4.1 files, more than a typical SWE-bench Verified fix.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic77.82026-04
GLM 5.1Z.ai (Zhipu AI)58.42026-04
GPT-5.4OpenAI57.72026-04
Gemini 3.1 Pro PreviewGoogle DeepMind54.22026-04
Claude Opus 4.6Anthropic53.42026-04

Data

This page as JSON · Edit on GitHub