SWE-Bench-CL

A 273-task continual-learning reformulation of SWE-bench Verified: eight chronological repository sequences with forgetting and transfer metrics.

Also known as: SWE-bench-CL

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategorycontinual learning over SWE-bench Verified sequences
Page statusproposed
Metricaverage accuracy (plus forgetting, transfer, and CL-Score)
Directionhigher_is_better
Dataset size273
Dataset licenceMIT
PublisherIndependent course project (COMS 4995)

What it measures

SWE-Bench-CL asks whether a coding agent improves, transfers, and avoids forgetting as it walks a stream of GitHub issues from one repository. The underlying issues are SWE-bench Verified tasks, reordered into curricula instead of scored as i.i.d. bugs.

Task format

Eight repository sequences (273 tasks total) are ordered first by issue creation time, then by human-estimated fix time. After each task, evaluation can re-test earlier tasks. The authors also describe a LangGraph agent with FAISS memory as a reference scaffold, not as a required runtime.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub