Professionally translated TruthfulQA in Basque, Catalan, Galician and Spanish, plus English, scored with MC2 and generation metrics.
unassessed
| Category | safety |
|---|---|
| Subcategory | multilingual truthfulness / imitative falsehood (Basque, Catalan, Galician, Spanish, English) |
| Page status | active |
| Metric | MC2 accuracy (lm-eval); paper also reports LLM-as-a-Judge truthfulness |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 817 |
| Dataset licence | Apache-2.0 |
| Publisher | HiTZ Center - Ixa, University of the Basque Country (UPV/EHU), with Elhuyar, CiTIUS (Universidade de Santiago de Compostela), and Universitat Pompeu Fabra |
TruthfulQA-Multi keeps the original 817 trap questions and answers, then adds professional translations into Spanish, Catalan, Galician and Basque so the same misconceptions can be asked outside English. The questions still sit in a US/English cultural frame; translators were told not to localise them. The skill is whether a model repeats a popular falsehood in that language, not whether it knows obscure facts. The paper splits items into 288 universal-knowledge questions and 529 time- or context-dependent ones.
Per language, lm-eval ships MC1 (single best option), MC2 (probability mass on all true options), and free-form generation. Prompts are "Q: {question}\\nA:". Generation stops at "Q:", ".\\n\\n" or "!\\n\\n". The paper's preferred generation score is an LLM judge, not the harness BLEU metrics.
No model card in ModelSpec reports this benchmark yet.