TruthfulQA-Multi

Professionally translated TruthfulQA in Basque, Catalan, Galician and Spanish, plus English, scored with MC2 and generation metrics.

Also known as: truthfulqa-multi, Multilingual TruthfulQA, Truth Knows No Language

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategorymultilingual truthfulness / imitative falsehood (Basque, Catalan, Galician, Spanish, English)
Page statusactive
MetricMC2 accuracy (lm-eval); paper also reports LLM-as-a-Judge truthfulness
Directionhigher_is_better
Unit%
Dataset size817
Dataset licenceApache-2.0
PublisherHiTZ Center - Ixa, University of the Basque Country (UPV/EHU), with Elhuyar, CiTIUS (Universidade de Santiago de Compostela), and Universitat Pompeu Fabra

What it measures

TruthfulQA-Multi keeps the original 817 trap questions and answers, then adds professional translations into Spanish, Catalan, Galician and Basque so the same misconceptions can be asked outside English. The questions still sit in a US/English cultural frame; translators were told not to localise them. The skill is whether a model repeats a popular falsehood in that language, not whether it knows obscure facts. The paper splits items into 288 universal-knowledge questions and 529 time- or context-dependent ones.

Task format

Per language, lm-eval ships MC1 (single best option), MC2 (probability mass on all true options), and free-form generation. Prompts are "Q: {question}\\nA:". Generation stops at "Q:", ".\\n\\n" or "!\\n\\n". The paper's preferred generation score is an LLM judge, not the harness BLEU metrics.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub