CoCo-Bench

CoCo-Bench evaluates language models across code understanding, generation, modification and review.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategorycomprehensive code evaluation
Page statusactive
Metrictask success rate
Directionhigher_is_better
Unit%
Dataset size0
PublisherCoCo-Bench authors

What it measures

Developer-facing code capabilities across multiple programming languages and task difficulties.

Task format

Multi-language tasks across four code dimensions; total size not stated in the opened abstract.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub