Five human-written Korean NLU tasks (yes/no QA, causal alternatives, word sense, sentence completion, polarity under negation) scored as multiple-choice accuracy and macro F1.
unassessed
| Category | composite |
|---|---|
| Subcategory | five-task Korean NLU suite (BoolQ, COPA, WiC, HellaSwag, SentiNeg) |
| Page status | active |
| Metric | macro F1 (paper); lm-eval also reports accuracy (and acc_norm on HellaSwag) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 4561 |
| Dataset licence | CC-BY-SA-4.0 |
| Publisher | SK Telecom Language Super Intelligence Labs; University of Oxford (Jang) |
KoBEST is a Korean-only text suite of five multiple-choice NLU tasks. BoolQ asks whether a question is true given a paragraph. COPA picks which of two Korean alternatives is the cause or effect of a premise. WiC asks whether a word has the same sense in two sentences. HellaSwag picks the plausible next sentence from four endings. SentiNeg labels a review sentence as positive or negative, with items built around negation. Professional linguists designed the items. Not [korbench](korbench.md).
lm-eval multiple_choice on skt/kobest_v1. Group name kobest. Runnable tasks kobest_boolq, kobest_copa, kobest_hellaswag, kobest_sentineg, kobest_wic. Korean prompts; yes/no choices 아니오/예 on BoolQ and WiC; 부정/긍정 on SentiNeg.
No model card in ModelSpec reports this benchmark yet.