HELM wrap of human-translated MMLU clinical knowledge, college medicine, and virology items in 11 African languages, scored by exact match.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | human-translated MMLU clinical subjects in 11 African languages |
| Page status | unknown |
| Metric | exact_match |
| Direction | higher_is_better |
| Dataset size | 265 |
| Dataset licence | MIT (translation release; MMLU source also MIT) |
| Publisher | Institute for Disease Modeling, Bill & Melinda Gates Foundation, and Ghamut Corporation; HELM wrap by Stanford CRFM |
mmlu_clinical_afr is HELM's multiple-choice wrap of three MMLU health subjects translated into 11 African languages. Each item is a four-option question in the target language. Default constructor arguments are subject clinical_knowledge and lang af (Afrikaans). The run spec also accepts college_medicine and virology, and ISO codes af, zu, xh, am, bm, ig, nso, sn, st, tn, ts. It is text-only. It is not English MMLU, not MMLU-ProX, and not Global-MMLU.
Joint multiple-choice. Instruction: "The following are multiple choice questions (with answers) about {subject} in {language}." Input noun Question, output noun Answer. Adapter default max_train_instances is 5; HELM maps the 5-item dev csv to TRAIN_SPLIT, so the usual protocol is 5-shot from dev. Main metric exact_match on test.
No model card in ModelSpec reports this benchmark yet.