RE-Bench evaluates agents that modify repositories to resolve issues across Java, TypeScript, JavaScript, Go, Rust, C and C++.
unassessed
| Category | agentic |
|---|---|
| Subcategory | research engineering agents |
| Page status | active |
| Metric | environment score |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 7 |
| Publisher | RE-Bench authors |
Progress and research-engineering performance in realistic ML environments under specified time budgets.
Open-ended agent interaction with seven research-engineering environments and objective environment scores.
No model card in ModelSpec reports this benchmark yet.