Arabic EXAMS

Multiple-choice Arabic high-school exam questions across five subjects, run under two different harness names (arabic_exams, aexams) over the same 562-item set.

Also known as: aexams, AEXAMS, EXAMS (Arabic)

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryArabic high-school exam multiple-choice QA (5 subjects)
Page statusactive
Metricaccuracy (lm-evaluation-harness also reports acc_norm; HELM's main metric is exact_match)
Directionhigher_is_better
Unit%
Dataset size562

What it measures

Arabic EXAMS gives a model a high-school-level exam question written in Arabic, with four labelled answer options, and asks it to pick the correct one. The questions are drawn from real school examinations rather than written for the benchmark, and cover five subjects -- Islamic Studies, Biology, Physics, Science, and Social (Studies) -- so a score on this benchmark reflects a mix of Arabic reading comprehension and subject-matter recall rather than any single skill. It is the Arabic-language slice of EXAMS, a multilingual high-school-exam question answering dataset (Hardalov et al., 2020), later re-packaged for LLM evaluation by the AceGPT project.

Task format

Multiple-choice exam question with four labelled options (A-D), single correct answer, in Arabic; scored either by log-likelihood ranking over the options or by exact match on a generated letter, depending on harness.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub