HAE-RAE Bench (lm-eval haerae)

lm-eval group for HAE-RAE Bench: five Korean multiple-choice tasks of native vocabulary and culture; the paper's reading-comprehension slice is omitted.

Also known as: HAE-RAE Bench, HAE_RAE_BENCH, HAERAE-BENCH, HRB

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryKorean cultural and lexical multiple-choice (lm-eval five-task cut)
Page statusunknown
Metricaccuracy and length-normalized accuracy (acc, acc_norm)
Directionhigher_is_better
Unit%
Dataset size1091
Dataset licenceCC-BY-NC-ND-4.0
PublisherHAE-RAE / HAERAE-HUB

What it measures

haerae is EleutherAI lm-evaluation-harness's group over HAE-RAE Bench, a Korean multiple-choice suite built to test cultural and lexical knowledge that does not transfer easily from English. The model sees a Korean query and must pick among five lettered options. The paper's six tasks are loan words, standard nomenclature, rare words, general knowledge, history, and reading comprehension. The harness drops reading comprehension for copyright and runs the other five. This id is not [csatqa](csatqa.md) (CSAT items from the same project) and not [click](click.md) (a later Korean culture exam).

Task format

Zero-shot multiple choice in the harness (output_type multiple_choice). Prompt is the dataset `query` field. Choices are (A)–(E). Gold is `answer`. fewshot_split is test, so any few-shot run draws demonstrations from the same test items. The paper also reported 5-shot and 10-shot log-likelihood for open models, and a separate numbered-option generation protocol for GPT-3.5/GPT-4.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub