SWE-bench Lite

A 300-task, single-file-edit subset of original SWE-bench, kept as a cheaper Python issue-resolution reporting split.

Also known as: SWE-bench_Lite, SWE-Bench-Lite

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
SubcategoryGitHub issue resolution / patch generation (filtered subset)
Page statusactive
Metric% resolved
Directionhigher_is_better
Unit%
Dataset size300
Dataset licenceMIT
PublisherSWE-bench Team (originally Princeton NLP)

What it measures

SWE-bench Lite measures the same skill as SWE-bench: produce a patch that resolves a real GitHub issue in a popular Python repository. The 300 test tasks are filtered to shorter, more self-contained edits so evaluation is cheaper than the full 2,294-instance set.

Task format

Identical to SWE-bench: issue text plus repository access, patch output, Docker grading with FAIL_TO_PASS and PASS_TO_PASS tests. Lite additionally drops multi-file gold patches, file create/delete, short problem statements, and several other hard cases.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub