CORE-Bench measures whether agents can reproduce scientific results from provided code and data.
unassessed
| Category | agentic |
|---|---|
| Subcategory | computational reproducibility |
| Page status | active |
| Metric | task success rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 270 |
| Publisher | CORE-Bench authors |
Accuracy of agents reproducing computational results from scientific papers.
270 tasks based on 90 scientific papers across computer science, social science and medicine.
No model card in ModelSpec reports this benchmark yet.