Misconceptions (Russian)

A 49-item BIG-bench Lite task that scores whether a model assigns higher probability to a true Russian statement than to a matched misconception.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
SubcategoryBIG-bench Lite Russian true-vs-myth statement pairs (49 items)
Page statusunknown
Metricmultiple_choice_grade
Directionhigher_is_better
Unit%
Dataset size49
Dataset licenceApache-2.0
PublisherGoogle (BIG-bench collaboration)

What it measures

misconceptions_russian gives two Russian sentences that differ mainly in whether they state a common myth or its correction. The model scores if it puts higher probability on the true sentence. Authors James Simon, Chandan Singh and Roman Sitelew translated and paraphrased Wikipedia misconceptions. The README argues Russian is common online but far less so than English, so the probe is harder than an English T/F list. Russian text only at test time.

Task format

Two-way multiple_choice_grade over the two full statements; input field is empty and choices are the statements themselves. append_choices_to_input is false. Canary GUID embedded. BIG-bench Lite keyword in the README header. 49 dummy-model queries.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub