CMB (Comprehensive Medical Benchmark in Chinese)

A two-part Chinese medical benchmark: 280,839 licensing-exam questions across 28 subcategories (CMB-Exam) plus 74 multi-turn clinical cases graded on four qualitative dimensions (CMB-Clin).

Also known as: Comprehensive Medical Benchmark in Chinese

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryChinese medical licensing exams and clinical-case diagnosis
Page statusactive
Metricaccuracy (CMB-Exam); four-dimension 1-5 rating (CMB-Clin)
Directionhigher_is_better
Unit%
Dataset size280839
Dataset licenceApache-2.0 (GitHub repository and Hugging Face dataset card both state this)
PublisherThe Chinese University of Hong Kong, Shenzhen; Shenzhen Research Institute of Big Data

What it measures

CMB tests medical knowledge and clinical reasoning in a Chinese-language, China-specific medical context, deliberately built rather than translated because the authors argue medical practice, licensing structure and terminology differ by region -- the benchmark also covers traditional Chinese medicine alongside modern biomedicine. It has two parts. CMB-Exam is multiple-choice and multiple-answer questions drawn from real Chinese medical licensing and qualification exams, covering six major categories (physician, nursing, pharmacist, medical technician, professional knowledge and postgraduate-entrance exams) split into 28 subcategories. CMB-Clin is 74 complex clinical cases, each a multi-turn conversation simulating a doctor working through a patient's history, examinations and diagnosis, testing free-form clinical reasoning rather than recall.

Task format

CMB-Exam: multiple-choice and multiple-answer questions, Chinese, evaluated zero-shot and few-shot with and without chain-of-thought. CMB-Clin: multi-turn free-form clinical dialogue, graded rather than scored against a fixed key.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub