170 expert-curated multilingual refactoring tasks from post-2025 commits; gold patches average 11.4 files and 261.6 lines.
unassessed
| Category | coding |
|---|---|
| Subcategory | multilingual repository-level refactoring |
| Page status | active |
| Metric | resolve rate (pass@1) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 170 |
| Publisher | SWE-Bench-ProMax authors |
SWE-Bench ProMax tests whether an agent can carry out a large, behaviour-preserving refactor in a real repository. Instances come from post-2025 GitHub commits tagged as refactoring, not from bug-fix issues. The agent must change many files so that a reviewed test suite still passes.
Given a rewritten issue description and a Dockerized pre-refactor checkout, the agent edits the tree. An instance is resolved only if every test in the suite passes. Issue text is rewritten from scratch so commit messages do not leak the gold patch.
No model card in ModelSpec reports this benchmark yet.