A 49-item BIG-bench Lite task that scores whether a model assigns higher probability to a true Russian statement than to a matched misconception.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | BIG-bench Lite Russian true-vs-myth statement pairs (49 items) |
| Page status | unknown |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 49 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
misconceptions_russian gives two Russian sentences that differ mainly in whether they state a common myth or its correction. The model scores if it puts higher probability on the true sentence. Authors James Simon, Chandan Singh and Roman Sitelew translated and paraphrased Wikipedia misconceptions. The README argues Russian is common online but far less so than English, so the probe is harder than an English T/F list. Russian text only at test time.
Two-way multiple_choice_grade over the two full statements; input field is empty and choices are the statements themselves. append_choices_to_input is false. Canary GUID embedded. BIG-bench Lite keyword in the README header. 49 dummy-model queries.
No model card in ModelSpec reports this benchmark yet.