Okapi ARC multilingual (lm-eval)

lm-eval group of GPT-3.5-translated ARC-Challenge questions in 31 languages on alexandrainst/m_arc, scored with acc and acc_norm.

Also known as: arc_multilingual, okapi/arc_multilingual, m_arc

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorymachine-translated ARC-Challenge multiple-choice QA in 31 lm-eval languages
Page statusactive
Metricaccuracy (acc); acc_norm also reported
Directionhigher_is_better
Unit%
Dataset size36004
Dataset licenceCC-BY-NC-4.0
PublisherUniversity of Oregon NLP (Okapi); Hugging Face mirror by Alexandra Institute

What it measures

okapi_arc_multilingual is EleutherAI's lm-evaluation-harness group over machine-translated AI2 ARC-Challenge items. Each item is a science question with labelled options and one gold letter. University of Oregon Okapi translated ARC, HellaSwag, and MMLU with ChatGPT/GPT-3.5-turbo for multilingual RLHF eval (arXiv:2307.16039). Alexandra Institute hosts the Hugging Face mirror alexandrainst/m_arc and adds Icelandic (Greynir) and Norwegian Bokmål (DeepL). English ids in the mirror are ARC-Challenge paths, not ARC-Easy. This is not English [arc](arc.md) or [arc_challenge](arc_challenge.md).

Task format

Multiple-choice log-likelihood. Template query is "Question: {instruction} \\nAnswer:" with choices from option_a..option_e when present. Gold is the index of answer in A–E. should_decontaminate true. Tag/group arc_multilingual. Per-language tasks are arc_ar, arc_bn, … arc_zh (31 names). YAML validation_split is "validation"; the Hub configs use split name "val".

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub