SWE-bench-Live

A continuously updated SWE-bench-style issue-resolution set built from recent GitHub issues, with frozen Lite/Verified splits and a growing full split.

Also known as: SWE-bench Live, SWE-bench Goes Live

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategorylive GitHub issue resolution / patch generation
Page statusactive
Metric% resolved
Directionhigher_is_better
Unit%
Dataset size1888
Dataset licenceMIT
PublisherMicrosoft

What it measures

SWE-bench-Live asks an agent to resolve a real GitHub issue on a snapshot of the repository from before the fix, then grades the patch with the project's tests. Unlike the original SWE-bench pool, instances are mined automatically from issues created since 2024 and refreshed over time so evaluation is less likely to rest on patches already seen in pretraining.

Task format

The agent may use only the problem statement and the instance Docker image. It must not read FAIL_TO_PASS, PASS_TO_PASS, hints, or the test patch during rollout. A single patch is applied in the container and scored with fail-to-pass and pass-to-pass tests, following the original SWE-bench protocol.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub