A 500-task, human-screened subset of SWE-bench that OpenAI released with the SWE-bench authors to remove unfair or impossible samples; now the default SWE-bench reference.
unassessed
| Category | coding |
|---|---|
| Subcategory | GitHub issue resolution / patch generation |
| Page status | active |
| Metric | % resolved |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 500 |
| Dataset licence | MIT |
| Publisher | OpenAI, in collaboration with the SWE-bench authors (Princeton NLP / SWE-bench Team) |
SWE-bench Verified measures the same thing as SWE-bench: whether a model can resolve a real GitHub issue by patching a Python repository, graded by the tests from the pull request that actually fixed it. The difference is curation: OpenAI ran a large human-annotation campaign to remove instances whose issue description was too vague or whose tests would reject a genuinely correct fix, so a low score is more likely to reflect a real capability gap than a broken task.
Identical to SWE-bench: given an issue description and repository access, the system outputs a patch, which is applied inside a container and graded against FAIL_TO_PASS and PASS_TO_PASS tests recovered from the original fixing pull request.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 93.9 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 80.8 | 2026-04 |
| Gemini 3.1 Pro Preview | Google DeepMind | 80.6 | 2026-04 |
| Claude Haiku 4.5 | Anthropic | 73.3 | 2026-04 |
| Claude Haiku 4.5 (latest) | Anthropic | 73.3 | 2026-04 |
| o3-pro | OpenAI | 73.2 | 2026-04 |
| Claude Opus 4 | Anthropic | 72.7 | 2026-04 |
| o3 | OpenAI | 71.7 | 2026-04 |
| o3-deep-research | OpenAI | 71.7 | 2026-04 |
| GPT-5.1 | OpenAI | 71.5 | 2026-04 |
| GPT-5.1 Chat | OpenAI | 71.5 | 2026-04 |
| GPT-5.1 Codex | OpenAI | 71.5 | 2026-04 |
| GPT-5.1 Codex Max | OpenAI | 71.5 | 2026-04 |
| GPT-5.1 Codex mini | OpenAI | 71.5 | 2026-04 |
| o3-mini | OpenAI | 71.2 | 2026-04 |
| Claude Opus 4.5 | Anthropic | 70.3 | 2026-04 |
| Claude Opus 4.5 (latest) | Anthropic | 70.3 | 2026-04 |
| GPT-5 | OpenAI | 69.3 | 2026-04 |
| GPT-5 Chat (latest) | OpenAI | 69.3 | 2026-04 |
| GPT-5 Pro | OpenAI | 69.3 | 2026-04 |
| GPT-5-Codex | OpenAI | 69.3 | 2026-04 |
| GPT-5.2 | OpenAI | 69.3 | 2026-04 |
| GPT-5.2 Chat | OpenAI | 69.3 | 2026-04 |
| GPT-5.2 Codex | OpenAI | 69.3 | 2026-04 |
| GPT-5.2 Pro | OpenAI | 69.3 | 2026-04 |
| GPT-5.3 Chat (latest) | OpenAI | 69.3 | 2026-04 |
| GPT-5.3 Codex | OpenAI | 69.3 | 2026-04 |
| GPT-5.3 Codex Spark | OpenAI | 69.3 | 2026-04 |
| GPT-5.4 | OpenAI | 69.3 | 2026-04 |
| GPT-5.4 mini | OpenAI | 69.3 | 2026-04 |
| GPT-5.4 nano | OpenAI | 69.3 | 2026-04 |
| GPT-5.4 Pro | OpenAI | 69.3 | 2026-04 |
| Claude Opus 4.1 | Anthropic | 68.8 | 2026-04 |
| Claude Opus 4.1 (latest) | Anthropic | 68.8 | 2026-04 |
| o4-mini | OpenAI | 68.4 | 2026-04 |
| o4-mini-deep-research | OpenAI | 68.4 | 2026-04 |
| Qwen3-Coder 480B-A35B Instruct | Alibaba / Qwen Team | 68.2 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 65.3 | 2026-04 |
| Claude Sonnet 4.6 | Anthropic | 65.3 | 2026-04 |
| Grok 4 | xAI | 64.5 | 2026-04 |
| Grok 4 Fast | xAI | 64.5 | 2026-04 |
| Grok 4 Fast (Non-Reasoning) | xAI | 64.5 | 2026-04 |
| Grok 4.1 Fast | xAI | 64.5 | 2026-04 |
| Grok 4.1 Fast (Non-Reasoning) | xAI | 64.5 | 2026-04 |
| Grok 4.20 (Non-Reasoning) | xAI | 64.5 | 2026-04 |
| Grok 4.20 (Reasoning) | xAI | 64.5 | 2026-04 |
| Grok 4.20 Multi-Agent | xAI | 64.5 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 63.8 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 63.8 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 63.8 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 63.8 | 2026-04 |
| GPT-5 Mini | OpenAI | 62.8 | 2026-04 |
| GPT-5 Nano | OpenAI | 62.8 | 2026-04 |
| Claude Opus 4 (latest) | Anthropic | 62.3 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 62.1 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 62.1 | 2026-04 |
| Devstral 2 (latest) | Mistral AI | 61.6 | 2026-04 |
| Devstral Medium | Mistral AI | 61.6 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 55.8 | 2026-04 |
| Claude Sonnet 4 (latest) | Anthropic | 55.2 | 2026-04 |
| GPT-4 | OpenAI | 54.6 | 2026-04 |
| GPT-4.1 | OpenAI | 54.6 | 2026-04 |
| GPT-4.1 nano | OpenAI | 54.6 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 53.8 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 53.8 | 2026-04 |
| Devstral Small | Mistral AI | 53.6 | 2026-04 |
| Devstral Small 2 24B Instruct 2512 | Mistral AI | 53.6 | 2026-04 |
| Devstral Small 2505 | Mistral AI | 53.6 | 2026-04 |
| Qwen 3 235B Instruct | Cerebras | 52.1 | 2026-04 |
| DeepSeek R1 | DeepSeek | 49.2 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 49.2 | 2026-04 |
| Gemini 2.5 Flash | Google DeepMind | 49.2 | 2026-04 |
| Gemini 2.5 Flash Image | Google DeepMind | 49.2 | 2026-04 |
| Gemini 2.5 Flash Image (Preview) | Google DeepMind | 49.2 | 2026-04 |
| Gemini 2.5 Flash Lite | Google DeepMind | 49.2 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 06-17 | Google DeepMind | 49.2 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 09-25 | Google DeepMind | 49.2 | 2026-04 |
| Gemini 2.5 Flash Preview 04-17 | Google DeepMind | 49.2 | 2026-04 |
| Gemini 2.5 Flash Preview 05-20 | Google DeepMind | 49.2 | 2026-04 |
| Gemini 2.5 Flash Preview 09-25 | Google DeepMind | 49.2 | 2026-04 |
| Gemini 2.5 Flash Preview TTS | Google DeepMind | 49.2 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 49.0 | 2026-04 |
| Claude Sonnet 3.5 v2 | Anthropic | 49.0 | 2026-04 |
| Claude Sonnet 3.7 | Anthropic | 49.0 | 2026-04 |
| Grok 3 | xAI | 48.5 | 2026-04 |
| Grok 3 Fast | xAI | 48.5 | 2026-04 |
| Grok 3 Fast Latest | xAI | 48.5 | 2026-04 |
| Grok 3 Latest | xAI | 48.5 | 2026-04 |
| DeepSeek Chat | DeepSeek | 42.5 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 42.5 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 42.5 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 42.5 | 2026-04 |
| DeepSeek V2 | DeepSeek | 42 | 2026-04 |
| DeepSeek V2 Lite | DeepSeek | 42 | 2026-04 |
| DeepSeek V2 Lite Chat | DeepSeek | 42 | 2026-04 |
| DeepSeek V3 | DeepSeek | 42.0 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 42.0 | 2026-04 |
| o1 | OpenAI | 41.3 | 2026-04 |
| o1-preview | OpenAI | 41.3 | 2026-04 |
| Claude Haiku 3.5 | Anthropic | 40.6 | 2026-04 |
| Claude Haiku 3.5 (latest) | Anthropic | 40.6 | 2026-04 |
| GPT-4o mini | OpenAI | 38.5 | 2026-04 |
| GPT-4o | OpenAI | 38.4 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 38.4 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 38.4 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 38.4 | 2026-04 |
| Codestral (latest) | Mistral AI | 35.2 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 32.5 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 32.5 | 2026-04 |
| Mistral Large 3 | Mistral AI | 32.5 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 30.2 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 30.2 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 30.2 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 30.2 | 2026-04 |
| Gemma 4 26B | Google DeepMind | 28.5 | 2026-04 |
| gemma 4 26B A4B it | Google DeepMind | 28.5 | 2026-04 |
| gemma 4 26B A4B it GGUF | Unsloth | 28.5 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 25.8 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 25.8 | 2026-04 |
| GPT-4.1 mini | OpenAI | 23.6 | 2026-04 |