CoCo-Bench evaluates language models across code understanding, generation, modification and review.
unassessed
| Category | coding |
|---|---|
| Subcategory | comprehensive code evaluation |
| Page status | active |
| Metric | task success rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 0 |
| Publisher | CoCo-Bench authors |
Developer-facing code capabilities across multiple programming languages and task difficulties.
Multi-language tasks across four code dimensions; total size not stated in the opened abstract.
No model card in ModelSpec reports this benchmark yet.