SWE-bench Science

119 repository-level scientific coding tasks across 98 GitHub projects and 20 domains, scored by held-out programmatic verifiers.

Also known as: SWE-bench-Science

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategoryscientific software engineering / repository-level repair
Page statusactive
Metricpass@1
Directionhigher_is_better
Unit%
Dataset size119
Dataset licenceMIT
PublisherOpenMOSS

What it measures

SWE-bench Science tests whether a coding agent can change a real scientific computing repository while preserving domain contracts such as units, file formats, numerics, and geometry. Each task starts from a fixed baseline commit and is checked in a clean environment, not by matching a gold patch.

Task format

The agent works in a pinned environment image, then a separate verifier image applies the candidate patch, rebuilds if needed, and runs held-out tests. Tasks are grouped into issue-driven, expert-exploratory, and engineering-integration paradigms. Default runs use 96 unrestricted-license tasks; 23 restricted-license tasks require an explicit opt-in.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub