CLEF 2011–2013 reading-comprehension lab: five-way questions about one document, shipped in lm-eval as English main-track sets for each year.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | CLEF multiple-choice reading comprehension over a single document |
| Page status | unknown |
| Metric | accuracy (acc and acc_norm) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 564 |
| Publisher | UNED NLP&IR Group (CLEF QA4MRE lab) |
QA4MRE tests whether a system can read one short document and pick the correct answer among five candidates. Organisers wrote the questions to require paraphrase, coreference, and sometimes facts from a background collection on the same topic (AIDS, climate change, music and society, Alzheimer's). lm-eval ships only the English main-track configs for 2011, 2012, and 2013. It does not run the Alzheimer's, entrance-exam, or modality/negation pilots, and it does not use the original c@1 metric that rewarded leaving a question unanswered.
Multiple choice. Prompt is the document plus the question; choices come from answer_options.answer_str. Target index is correct_answer_id minus one. lm-eval output_type is multiple_choice on the Hub train split of each year.
No model card in ModelSpec reports this benchmark yet.