Multi-SWE-bench

Multi-SWE-bench evaluates agents that modify repositories to resolve issues across Java, TypeScript, JavaScript, Go, Rust, C and C++.

Also known as: Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategorymultilingual issue resolution
Page statusactive
Metricpass rate
Directionhigher_is_better
Unit%
Dataset size1632
PublisherMulti-SWE-bench authors

What it measures

Whether an agent can produce a patch that resolves a real issue and passes the repository's tests.

Task format

Issue statement, repository snapshot and test environment; the agent edits the repository and is evaluated by tests.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub