A 300-task, single-file-edit subset of original SWE-bench, kept as a cheaper Python issue-resolution reporting split.
unassessed
| Category | coding |
|---|---|
| Subcategory | GitHub issue resolution / patch generation (filtered subset) |
| Page status | active |
| Metric | % resolved |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 300 |
| Dataset licence | MIT |
| Publisher | SWE-bench Team (originally Princeton NLP) |
SWE-bench Lite measures the same skill as SWE-bench: produce a patch that resolves a real GitHub issue in a popular Python repository. The 300 test tasks are filtered to shorter, more self-contained edits so evaluation is cheaper than the full 2,294-instance set.
Identical to SWE-bench: issue text plus repository access, patch output, Docker grading with FAIL_TO_PASS and PASS_TO_PASS tests. Lite additionally drops multi-file gold patches, file create/delete, short problem statements, and several other hard cases.
No model card in ModelSpec reports this benchmark yet.