C-Eval

A 13,948-question, four-option Chinese exam benchmark across 52 subjects and four difficulty levels, whose test set was held out via a submission site until a full public release in July 2025.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryChinese multi-level multi-discipline exam suite
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size13948
Dataset licenceCC BY-NC-SA 4.0 (dataset, per the Hugging Face card and the repository's own Licenses section); the repository's evaluation code is separately released under MIT
PublisherC-Eval team (GitHub organisation hkust-nlp)

What it measures

C-Eval tests broad academic and professional knowledge and reasoning in a Chinese-language context. Each item is a four-option multiple-choice question drawn from one of 52 disciplines spanning the humanities, social sciences, STEM and other professional fields, split across four difficulty levels -- middle school, high school, college and professional -- mirroring how Chinese students and professionals actually progress through subjects. It positions itself as a Chinese-context counterpart to MMLU: exam-style questions authored in Chinese rather than translated from English content.

Task format

Four-option multiple-choice question answering in Chinese, evaluated zero-shot and five-shot. Answers are scored either by parsing a generated answer letter or, when a model does not follow the instruction format well, by taking the highest-probability option among A-D.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub