OJBench evaluates competitive-level code reasoning on 232 NOI and ICPC programming problems.
unassessed
| Category | coding |
|---|---|
| Subcategory | competitive programming |
| Page status | active |
| Metric | Pass@8 |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 232 |
| Publisher | OJBench authors |
OJBench tests whether a model can reason about and produce solutions for difficult programming-contest problems. The paper describes problems drawn from NOI and ICPC competitions.
Natural-language programming problem prompt; generate a code solution evaluated by an online-judge style checker.
No model card in ModelSpec reports this benchmark yet.