KMMLU (Korean-MMLU)

35,030 Korean four-option exam questions across 45 subjects, collected from original Korean tests rather than translated from MMLU.

Also known as: Korean-MMLU, k_mmlu

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryKorean native exam multiple choice across 45 subjects (STEM, HUMSS, applied science, other)
Page statusactive
Metricaccuracy (acc); lm-eval group kmmlu is size-weighted micro-average
Directionhigher_is_better
Unit%
Dataset size35030
Dataset licenceCC-BY-ND-4.0 (Hub card); paper text says CC-BY-ND
PublisherHAERAE-HUB (authors at Yonsei, NCSOFT, NAVER Cloud, KAIST, Carnegie Mellon, Contextual AI, EleutherAI)

What it measures

KMMLU tests expert-level Korean knowledge and reasoning with four-option questions drawn from original Korean exams, not machine-translated MMLU. Forty-five subjects span STEM, humanities and social science (HUMSS), applied science, and a residual Other bucket. Many items come from Korean licence tests, including papers that assume years of industry experience, plus items that need Korean cultural, regional, or legal knowledge. The default lm-eval group `kmmlu` is the full 45-subject test set scored as multiple-choice log-likelihood. Direct, hard, and hard-CoT groups are separate harness variants of the same project, not different ids.

Task format

Four-option multiple choice in Korean. Default yaml: output_type multiple_choice, prompt question plus A-D and "정답:", target is answer-1 (0-based index into A-D). test_split is test; fewshot_split is dev (five items per subject). The default yaml does not set num_fewshot; the paper ran every method five-shot. kmmlu_direct instead generates an option and scores exact_match.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub