119 repository-level scientific coding tasks across 98 GitHub projects and 20 domains, scored by held-out programmatic verifiers.
unassessed
| Category | coding |
|---|---|
| Subcategory | scientific software engineering / repository-level repair |
| Page status | active |
| Metric | pass@1 |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 119 |
| Dataset licence | MIT |
| Publisher | OpenMOSS |
SWE-bench Science tests whether a coding agent can change a real scientific computing repository while preserving domain contracts such as units, file formats, numerics, and geometry. Each task starts from a fixed baseline commit and is checked in a clean environment, not by matching a gold patch.
The agent works in a pinned environment image, then a separate verifier image applies the candidate patch, rebuilds if needed, and runs held-out tests. Tasks are grouped into issue-driven, expert-exploratory, and engineering-integration paradigms. Default runs use 96 unrestricted-license tasks; 23 restricted-license tasks require an explicit opt-in.
No model card in ModelSpec reports this benchmark yet.