RE-Bench

RE-Bench evaluates agents that modify repositories to resolve issues across Java, TypeScript, JavaScript, Go, Rust, C and C++.

Also known as: RE-Bench: A Multilingual Benchmark for Issue Resolving

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryresearch engineering agents
Page statusactive
Metricenvironment score
Directionhigher_is_better
Unit%
Dataset size7
PublisherRE-Bench authors

What it measures

Progress and research-engineering performance in realistic ML environments under specified time budgets.

Task format

Open-ended agent interaction with seven research-engineering environments and objective environment scores.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub