OJBench

OJBench evaluates competitive-level code reasoning on 232 NOI and ICPC programming problems.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategorycompetitive programming
Page statusactive
MetricPass@8
Directionhigher_is_better
Unitpercent
Dataset size232
PublisherOJBench authors

What it measures

OJBench tests whether a model can reason about and produce solutions for difficult programming-contest problems. The paper describes problems drawn from NOI and ICPC competitions.

Task format

Natural-language programming problem prompt; generate a code solution evaluated by an online-judge style checker.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub