35,030 Korean four-option exam questions across 45 subjects, collected from original Korean tests rather than translated from MMLU.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | Korean native exam multiple choice across 45 subjects (STEM, HUMSS, applied science, other) |
| Page status | active |
| Metric | accuracy (acc); lm-eval group kmmlu is size-weighted micro-average |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 35030 |
| Dataset licence | CC-BY-ND-4.0 (Hub card); paper text says CC-BY-ND |
| Publisher | HAERAE-HUB (authors at Yonsei, NCSOFT, NAVER Cloud, KAIST, Carnegie Mellon, Contextual AI, EleutherAI) |
KMMLU tests expert-level Korean knowledge and reasoning with four-option questions drawn from original Korean exams, not machine-translated MMLU. Forty-five subjects span STEM, humanities and social science (HUMSS), applied science, and a residual Other bucket. Many items come from Korean licence tests, including papers that assume years of industry experience, plus items that need Korean cultural, regional, or legal knowledge. The default lm-eval group `kmmlu` is the full 45-subject test set scored as multiple-choice log-likelihood. Direct, hard, and hard-CoT groups are separate harness variants of the same project, not different ids.
Four-option multiple choice in Korean. Default yaml: output_type multiple_choice, prompt question plus A-D and "정답:", target is answer-1 (0-based index into A-D). test_split is test; fewshot_split is dev (five items per subject). The default yaml does not set num_fewshot; the paper ran every method five-shot. kmmlu_direct instead generates an option and scores exact_match.
No model card in ModelSpec reports this benchmark yet.