21,805 multiple-choice questions natively sourced or authored in Greek from real exams across 45 subjects, built specifically to avoid the translation artefacts of machine-translated Greek benchmarks.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | native-sourced Greek multitask academic, professional and governmental exam knowledge |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 16857 |
| Dataset licence | MIT, per the Hugging Face dataset card; the paper states source materials were collected only from content released under open-access or educational-reuse terms, without naming a single licence for the compiled dataset |
| Publisher | Ecole Polytechnique, MBZUAI and the National Technical University of Athens, with the University of Ioannina and the University of Peloponnese |
GreekMMLU tests broad academic, professional and civic knowledge in Greek using four-option multiple-choice questions drawn from real Greek academic, professional and governmental examinations, spanning difficulty from primary school to professional licensing. Its authors built it specifically because they judged existing Greek-language evaluation material to be largely machine-translated from English, which they argue fails to capture Greek linguistic and cultural specifics -- the 45 subjects include several with no English-language equivalent to translate from at all, such as Greek driving regulations, Greek mythology, Greek traditions and Greek civil-service exam content, alongside standard STEM, humanities and social-science subjects. This page confirms, from the paper's own abstract and dataset card, that GreekMMLU is native-sourced rather than translated by any method, machine or human: every question was collected or authored directly in Greek, which changes how a score on it should be read compared with a benchmark built by translating an existing English test.
Four-option multiple-choice question answering in Greek, following Greek examination convention (answer labels mix Latin and Greek letters, A, B, Gamma, Delta). Evaluated both zero-shot and five-shot, with prompts written entirely in Greek. Open-weight models are scored by a rank-based comparison of answer-option log-likelihoods; closed API models are scored by free-form generation with the predicted label extracted by regular expression.
No model card in ModelSpec reports this benchmark yet.