QA4MRE

CLEF 2011–2013 reading-comprehension lab: five-way questions about one document, shipped in lm-eval as English main-track sets for each year.

Also known as: Question Answering for Machine Reading Evaluation, qa4mre_2011, qa4mre_2012, qa4mre_2013

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
SubcategoryCLEF multiple-choice reading comprehension over a single document
Page statusunknown
Metricaccuracy (acc and acc_norm)
Directionhigher_is_better
Unit%
Dataset size564
PublisherUNED NLP&IR Group (CLEF QA4MRE lab)

What it measures

QA4MRE tests whether a system can read one short document and pick the correct answer among five candidates. Organisers wrote the questions to require paraphrase, coreference, and sometimes facts from a background collection on the same topic (AIDS, climate change, music and society, Alzheimer's). lm-eval ships only the English main-track configs for 2011, 2012, and 2013. It does not run the Alzheimer's, entrance-exam, or modality/negation pilots, and it does not use the original c@1 metric that rewarded leaving a question unanswered.

Task format

Multiple choice. Prompt is the document plus the question; choices come from answer_options.answer_str. Target index is correct_answer_id minus one. lm-eval output_type is multiple_choice on the Hub train split of each year.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub