SWE-bench Verified

A 500-task, human-screened subset of SWE-bench that OpenAI released with the SWE-bench authors to remove unfair or impossible samples; now the default SWE-bench reference.

Also known as: SWE-bench-V

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
SubcategoryGitHub issue resolution / patch generation
Page statusactive
Metric% resolved
Directionhigher_is_better
Unit%
Dataset size500
Dataset licenceMIT
PublisherOpenAI, in collaboration with the SWE-bench authors (Princeton NLP / SWE-bench Team)

What it measures

SWE-bench Verified measures the same thing as SWE-bench: whether a model can resolve a real GitHub issue by patching a Python repository, graded by the tests from the pull request that actually fixed it. The difference is curation: OpenAI ran a large human-annotation campaign to remove instances whose issue description was too vague or whose tests would reject a genuinely correct fix, so a low score is more likely to reflect a real capability gap than a broken task.

Task format

Identical to SWE-bench: given an issue description and repository access, the system outputs a patch, which is applied inside a container and graded against FAIL_TO_PASS and PASS_TO_PASS tests recovered from the original fixing pull request.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic93.92026-04
Claude Opus 4.6Anthropic80.82026-04
Gemini 3.1 Pro PreviewGoogle DeepMind80.62026-04
Claude Haiku 4.5Anthropic73.32026-04
Claude Haiku 4.5 (latest)Anthropic73.32026-04
o3-proOpenAI73.22026-04
Claude Opus 4Anthropic72.72026-04
o3OpenAI71.72026-04
o3-deep-researchOpenAI71.72026-04
GPT-5.1OpenAI71.52026-04
GPT-5.1 ChatOpenAI71.52026-04
GPT-5.1 CodexOpenAI71.52026-04
GPT-5.1 Codex MaxOpenAI71.52026-04
GPT-5.1 Codex miniOpenAI71.52026-04
o3-miniOpenAI71.22026-04
Claude Opus 4.5Anthropic70.32026-04
Claude Opus 4.5 (latest)Anthropic70.32026-04
GPT-5OpenAI69.32026-04
GPT-5 Chat (latest)OpenAI69.32026-04
GPT-5 ProOpenAI69.32026-04
GPT-5-CodexOpenAI69.32026-04
GPT-5.2OpenAI69.32026-04
GPT-5.2 ChatOpenAI69.32026-04
GPT-5.2 CodexOpenAI69.32026-04
GPT-5.2 ProOpenAI69.32026-04
GPT-5.3 Chat (latest)OpenAI69.32026-04
GPT-5.3 CodexOpenAI69.32026-04
GPT-5.3 Codex SparkOpenAI69.32026-04
GPT-5.4OpenAI69.32026-04
GPT-5.4 miniOpenAI69.32026-04
GPT-5.4 nanoOpenAI69.32026-04
GPT-5.4 ProOpenAI69.32026-04
Claude Opus 4.1Anthropic68.82026-04
Claude Opus 4.1 (latest)Anthropic68.82026-04
o4-miniOpenAI68.42026-04
o4-mini-deep-researchOpenAI68.42026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team68.22026-04
Claude Sonnet 4Anthropic65.32026-04
Claude Sonnet 4.6Anthropic65.32026-04
Grok 4xAI64.52026-04
Grok 4 FastxAI64.52026-04
Grok 4 Fast (Non-Reasoning)xAI64.52026-04
Grok 4.1 FastxAI64.52026-04
Grok 4.1 Fast (Non-Reasoning)xAI64.52026-04
Grok 4.20 (Non-Reasoning)xAI64.52026-04
Grok 4.20 (Reasoning)xAI64.52026-04
Grok 4.20 Multi-AgentxAI64.52026-04
Gemini 2.5 ProGoogle DeepMind63.82026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind63.82026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind63.82026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind63.82026-04
GPT-5 MiniOpenAI62.82026-04
GPT-5 NanoOpenAI62.82026-04
Claude Opus 4 (latest)Anthropic62.32026-04
Claude Sonnet 4.5Anthropic62.12026-04
Claude Sonnet 4.5 (latest)Anthropic62.12026-04
Devstral 2 (latest)Mistral AI61.62026-04
Devstral MediumMistral AI61.62026-04
Qwen3 235B-A22BAlibaba / Qwen Team55.82026-04
Claude Sonnet 4 (latest)Anthropic55.22026-04
GPT-4OpenAI54.62026-04
GPT-4.1OpenAI54.62026-04
GPT-4.1 nanoOpenAI54.62026-04
DeepSeek R1 0528DeepSeek53.82026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA53.82026-04
Devstral SmallMistral AI53.62026-04
Devstral Small 2 24B Instruct 2512Mistral AI53.62026-04
Devstral Small 2505Mistral AI53.62026-04
Qwen 3 235B InstructCerebras52.12026-04
DeepSeek R1DeepSeek49.22026-04
DeepSeek ReasonerDeepSeek49.22026-04
Gemini 2.5 FlashGoogle DeepMind49.22026-04
Gemini 2.5 Flash ImageGoogle DeepMind49.22026-04
Gemini 2.5 Flash Image (Preview)Google DeepMind49.22026-04
Gemini 2.5 Flash LiteGoogle DeepMind49.22026-04
Gemini 2.5 Flash Lite Preview 06-17Google DeepMind49.22026-04
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind49.22026-04
Gemini 2.5 Flash Preview 04-17Google DeepMind49.22026-04
Gemini 2.5 Flash Preview 05-20Google DeepMind49.22026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind49.22026-04
Gemini 2.5 Flash Preview TTSGoogle DeepMind49.22026-04
Claude Sonnet 3.5Anthropic49.02026-04
Claude Sonnet 3.5 v2Anthropic49.02026-04
Claude Sonnet 3.7Anthropic49.02026-04
Grok 3xAI48.52026-04
Grok 3 FastxAI48.52026-04
Grok 3 Fast LatestxAI48.52026-04
Grok 3 LatestxAI48.52026-04
DeepSeek ChatDeepSeek42.52026-04
DeepSeek V3.1DeepSeek42.52026-04
DeepSeek V3.2DeepSeek42.52026-04
DeepSeek V3.2 ExpDeepSeek42.52026-04
DeepSeek V2DeepSeek422026-04
DeepSeek V2 LiteDeepSeek422026-04
DeepSeek V2 Lite ChatDeepSeek422026-04
DeepSeek V3DeepSeek42.02026-04
DeepSeek V3 0324DeepSeek42.02026-04
o1OpenAI41.32026-04
o1-previewOpenAI41.32026-04
Claude Haiku 3.5Anthropic40.62026-04
Claude Haiku 3.5 (latest)Anthropic40.62026-04
GPT-4o miniOpenAI38.52026-04
GPT-4oOpenAI38.42026-04
GPT-4o (2024-05-13)OpenAI38.42026-04
GPT-4o (2024-08-06)OpenAI38.42026-04
GPT-4o (2024-11-20)OpenAI38.42026-04
Codestral (latest)Mistral AI35.22026-04
Mistral Large (latest)Mistral AI32.52026-04
Mistral Large 2.1Mistral AI32.52026-04
Mistral Large 3Mistral AI32.52026-04
Gemma 4 31BGoogle DeepMind30.22026-04
gemma 4 31B itGoogle DeepMind30.22026-04
gemma 4 31B it GGUFUnsloth30.22026-04
Gemma 4 31B IT NVFP4NVIDIA30.22026-04
Gemma 4 26BGoogle DeepMind28.52026-04
gemma 4 26B A4B itGoogle DeepMind28.52026-04
gemma 4 26B A4B it GGUFUnsloth28.52026-04
Llama 3.3 70B Instruct NVFP4NVIDIA25.82026-04
Llama-3.3-70B-InstructMeta25.82026-04
GPT-4.1 miniOpenAI23.62026-04

Data

This page as JSON · Edit on GitHub