SWE-Bench ProMax

170 expert-curated multilingual refactoring tasks from post-2025 commits; gold patches average 11.4 files and 261.6 lines.

Also known as: SWE-Bench-ProMax, SWE-bench ProMax

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategorymultilingual repository-level refactoring
Page statusactive
Metricresolve rate (pass@1)
Directionhigher_is_better
Unit%
Dataset size170
PublisherSWE-Bench-ProMax authors

What it measures

SWE-Bench ProMax tests whether an agent can carry out a large, behaviour-preserving refactor in a real repository. Instances come from post-2025 GitHub commits tagged as refactoring, not from bug-fix issues. The agent must change many files so that a reviewed test suite still passes.

Task format

Given a rewritten issue description and a Dockerized pre-refactor checkout, the agent edits the tree. An instance is resolved only if every test in the suite passes. Issue text is rewritten from scratch so commit messages do not leak the gold patch.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub