Multi-SWE-bench evaluates agents that modify repositories to resolve issues across Java, TypeScript, JavaScript, Go, Rust, C and C++.
unassessed
| Category | coding |
|---|---|
| Subcategory | multilingual issue resolution |
| Page status | active |
| Metric | pass rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1632 |
| Publisher | Multi-SWE-bench authors |
Whether an agent can produce a patch that resolves a real issue and passes the repository's tests.
Issue statement, repository snapshot and test environment; the agent edits the repository and is evaluated by tests.
No model card in ModelSpec reports this benchmark yet.