XCOPA

XCOPA translates and re-annotates the English COPA causal-reasoning test into 11 typologically diverse languages, to measure zero-shot cross-lingual transfer of commonsense reasoning.

Also known as: Cross-lingual Choice of Plausible Alternatives

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorycross-lingual causal commonsense reasoning (COPA-style)
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size6600
Dataset licenceCC BY 4.0
PublisherLanguage Technology Lab, University of Cambridge, with University of Mannheim

What it measures

XCOPA tests causal commonsense reasoning in the COPA format: a model reads a one-sentence premise, is told whether it needs the cause or the result, and must pick which of two alternative sentences is more plausible. What XCOPA adds to plain COPA is language coverage -- the same premise-and-alternatives format, translated and re-annotated by native speakers into 11 typologically diverse languages (Estonian, Haitian Creole, Indonesian, Italian, Eastern Apurímac Quechua, Swahili, Tamil, Thai, Turkish, Vietnamese and Mandarin Chinese), spanning 11 language families and several world regions. Because XCOPA ships no training data of its own, it is designed to be used as a zero-shot cross-lingual transfer test: a system is typically trained or fine-tuned on English resources (the original English COPA, sometimes alongside Social IQa) and then evaluated directly on each of the 11 target languages, so a score reflects both commonsense-reasoning ability and how well that ability transfers out of English.

Task format

Two-choice, single-turn: given a premise sentence and a prompt indicating cause or result, the model selects the more plausible of two alternative sentences, in the target language. No free-text generation or explanation is required.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub