CORE-Bench

CORE-Bench measures whether agents can reproduce scientific results from provided code and data.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorycomputational reproducibility
Page statusactive
Metrictask success rate
Directionhigher_is_better
Unit%
Dataset size270
PublisherCORE-Bench authors

What it measures

Accuracy of agents reproducing computational results from scientific papers.

Task format

270 tasks based on 90 scientific papers across computer science, social science and medicine.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub