KorMedMCQA

7,469 five-option Korean multiple-choice questions from doctor, nurse, pharmacist and dentist licensing exams (2012-2024), with a recorded human-examinee average of 77.05%.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryKorean healthcare professional licensing exam multiple-choice questions (doctor, nurse, pharmacist, dentist)
Page statusactive
Metricexact_match (accuracy on the extracted option letter), weighted by item count across the four professions for the group aggregate
Directionhigher_is_better
Unit%
Dataset size7469
Dataset licenceCC BY-NC 2.0, per the Hugging Face dataset card's licence tag; the paper's own arXiv listing separately shows a CC BY 4.0 licence for the paper text itself.
PublisherKAIST, with Ajou University School of Medicine, UNIST and Kyung Hee University College of Dentistry

What it measures

KorMedMCQA is the first Korean-language medical multiple-choice question answering benchmark, built from real Korean healthcare professional licensing examinations administered between 2012 and 2024. It covers four professions -- doctor, nurse, pharmacist and dentist -- each drawn from that profession's own official licensing exam, spanning a wide range of medical subjects within each. Every question is written in Korean and offers five answer options (A-E), matching the five-option format the source exams themselves use, so unlike this repository's four-option medqa (English, USMLE-style), guessing at random on KorMedMCQA scores 20% rather than 25%.

Task format

A Korean-language exam question with five labelled options (A-E) in; the model completes a "정답:" ("the answer is:") prompt with a single option letter, evaluated in a 5-shot setting with exemplars drawn from a fixed set of five questions per profession, originally sampled from the exams' development portion by the paper's authors.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub