KoBEST (Korean Balanced Evaluation of Significant Tasks)

Five human-written Korean NLU tasks (yes/no QA, causal alternatives, word sense, sentence completion, polarity under negation) scored as multiple-choice accuracy and macro F1.

Also known as: KOBEST, KB-BoolQ, KB-COPA, KB-WiC, KB-HellaSwag, KB-SentiNeg

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategoryfive-task Korean NLU suite (BoolQ, COPA, WiC, HellaSwag, SentiNeg)
Page statusactive
Metricmacro F1 (paper); lm-eval also reports accuracy (and acc_norm on HellaSwag)
Directionhigher_is_better
Unit%
Dataset size4561
Dataset licenceCC-BY-SA-4.0
PublisherSK Telecom Language Super Intelligence Labs; University of Oxford (Jang)

What it measures

KoBEST is a Korean-only text suite of five multiple-choice NLU tasks. BoolQ asks whether a question is true given a paragraph. COPA picks which of two Korean alternatives is the cause or effect of a premise. WiC asks whether a word has the same sense in two sentences. HellaSwag picks the plausible next sentence from four endings. SentiNeg labels a review sentence as positive or negative, with items built around negation. Professional linguists designed the items. Not [korbench](korbench.md).

Task format

lm-eval multiple_choice on skt/kobest_v1. Group name kobest. Runnable tasks kobest_boolq, kobest_copa, kobest_hellaswag, kobest_sentineg, kobest_wic. Korean prompts; yes/no choices 아니오/예 on BoolQ and WiC; 부정/긍정 on SentiNeg.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub