SWE-bench Agent

An internal model-card key for an agentic (not single-shot patch) SWE-bench score; no publisher, paper or dataset for it under this name was found.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategoryinternal repository label for an agent-harness SWE-bench score
Page statusunknown
Directionhigher_is_better
Unit%

What it measures

`swe_bench_agent` is a scoring key used by this repository's own model-card template, not a benchmark published anywhere under that name. The template comments it as "agentic SWE-bench (not just patch gen)," grouped with other agentic-environment benchmarks (tau_bench, web_arena, os_world) rather than with the SWE-bench family's other named variants. What exact dataset, instance count, harness or scaffold produces the number is not documented anywhere this research could find.

Task format

Not documented under this name. If the template comment is accurate, it denotes SWE-bench evaluated by an agent that reads, runs and edits a repository over multiple steps, as opposed to a single-shot patch generated from one prompt — but no source specific to `swe_bench_agent` describes its dataset, instance count or grading protocol.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
o3OpenAI62.82026-04
o3-deep-researchOpenAI62.82026-04
o3-miniOpenAI62.82026-04
o3-proOpenAI62.82026-04
Claude Opus 4Anthropic62.52026-04
Claude Opus 4.6Anthropic62.52026-04
Claude Opus 4.1Anthropic58.22026-04
Claude Opus 4.1 (latest)Anthropic58.22026-04
Claude Sonnet 4Anthropic55.82026-04
Claude Sonnet 4.5Anthropic55.82026-04
Claude Sonnet 4.5 (latest)Anthropic55.82026-04
Gemini 2.5 ProGoogle DeepMind55.52026-04
GPT-5OpenAI55.22026-04
GPT-5 Chat (latest)OpenAI55.22026-04
GPT-5 MiniOpenAI55.22026-04
GPT-5 NanoOpenAI55.22026-04
GPT-5 ProOpenAI55.22026-04
GPT-5-CodexOpenAI55.22026-04
GPT-5.1OpenAI55.22026-04
GPT-5.1 ChatOpenAI55.22026-04
GPT-5.1 CodexOpenAI55.22026-04
GPT-5.1 Codex MaxOpenAI55.22026-04
GPT-5.1 Codex miniOpenAI55.22026-04
GPT-5.2OpenAI55.22026-04
GPT-5.2 ChatOpenAI55.22026-04
GPT-5.2 CodexOpenAI55.22026-04
GPT-5.2 ProOpenAI55.22026-04
GPT-5.3 Chat (latest)OpenAI55.22026-04
GPT-5.3 CodexOpenAI55.22026-04
GPT-5.3 Codex SparkOpenAI55.22026-04
GPT-5.4OpenAI55.22026-04
GPT-5.4 miniOpenAI55.22026-04
GPT-5.4 nanoOpenAI55.22026-04
GPT-5.4 ProOpenAI55.22026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team52.12026-04
Grok 4xAI50.22026-04
Grok 4 FastxAI50.22026-04
Grok 4 Fast (Non-Reasoning)xAI50.22026-04
Grok 4.1 FastxAI50.22026-04
Grok 4.1 Fast (Non-Reasoning)xAI50.22026-04
Grok 4.20 (Non-Reasoning)xAI50.22026-04
Grok 4.20 (Reasoning)xAI50.22026-04
Grok 4.20 Multi-AgentxAI50.22026-04
GPT-4.1OpenAI48.52026-04
Qwen 3 235B InstructCerebras45.52026-04
Qwen3 235B-A22BAlibaba / Qwen Team45.52026-04
DeepSeek R1DeepSeek42.52026-04
DeepSeek R1 0528DeepSeek42.52026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA42.52026-04
DeepSeek ReasonerDeepSeek42.52026-04
DeepSeek ChatDeepSeek35.82026-04
DeepSeek V3DeepSeek35.82026-04
DeepSeek V3 0324DeepSeek35.82026-04
DeepSeek V3.1DeepSeek35.82026-04
DeepSeek V3.2DeepSeek35.82026-04
DeepSeek V3.2 ExpDeepSeek35.82026-04
GPT-4oOpenAI32.12026-04
GPT-4o (2024-05-13)OpenAI32.12026-04
GPT-4o (2024-08-06)OpenAI32.12026-04
GPT-4o (2024-11-20)OpenAI32.12026-04
GPT-4o miniOpenAI32.12026-04

Data

This page as JSON · Edit on GitHub