lm-eval group of GPT-3.5-translated ARC-Challenge questions in 31 languages on alexandrainst/m_arc, scored with acc and acc_norm.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | machine-translated ARC-Challenge multiple-choice QA in 31 lm-eval languages |
| Page status | active |
| Metric | accuracy (acc); acc_norm also reported |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 36004 |
| Dataset licence | CC-BY-NC-4.0 |
| Publisher | University of Oregon NLP (Okapi); Hugging Face mirror by Alexandra Institute |
okapi_arc_multilingual is EleutherAI's lm-evaluation-harness group over machine-translated AI2 ARC-Challenge items. Each item is a science question with labelled options and one gold letter. University of Oregon Okapi translated ARC, HellaSwag, and MMLU with ChatGPT/GPT-3.5-turbo for multilingual RLHF eval (arXiv:2307.16039). Alexandra Institute hosts the Hugging Face mirror alexandrainst/m_arc and adds Icelandic (Greynir) and Norwegian Bokmål (DeepL). English ids in the mirror are ARC-Challenge paths, not ARC-Easy. This is not English [arc](arc.md) or [arc_challenge](arc_challenge.md).
Multiple-choice log-likelihood. Template query is "Question: {instruction} \\nAnswer:" with choices from option_a..option_e when present. Gold is the index of answer in A–E. should_decontaminate true. Tag/group arc_multilingual. Per-language tasks are arc_ar, arc_bn, … arc_zh (31 names). YAML validation_split is "validation"; the Hub configs use split name "val".
No model card in ModelSpec reports this benchmark yet.