SWE-bench-java

A 91-instance, human-screened Java SWE-bench set from six popular repositories, scored by fail-to-pass unit tests.

Also known as: SWE-bench-java-verified, SWE-bench Java

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
SubcategoryJava GitHub issue resolution / patch generation
Page statusactive
Metricresolved rate
Directionhigher_is_better
Unit%
Dataset size91
Dataset licenceApache-2.0
PublisherHuawei and collaborators

What it measures

SWE-bench-java tests whether a model can resolve a real Java GitHub issue by patching the repository at the pre-fix commit. It is the Java analogue of SWE-bench Verified: issues are filtered so they compile, have fail-to-pass tests, and survive developer screening.

Task format

The system sees the issue text and a Dockerized Maven/Java checkout, then emits a patch. Tests from the fixing pull request grade the result. The authors evaluated SWE-agent rather than a one-shot RAG baseline.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub