Social IQa

35,364 public three-way multiple-choice questions about people's motivations and reactions in social situations; a 2019 AI2/UW benchmark with a dead leaderboard and no current model-card coverage.

Also known as: SocialIQA, Social Interaction QA

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorysocial-commonsense multiple-choice question answering
Page statussaturated
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size35364
Dataset licenceNot established: no licence tag or statement was found on the Hugging Face dataset card, and no separate licence file was confirmed in the sources checked for this page.
PublisherAllen Institute for Artificial Intelligence (AI2); Paul G. Allen School of Computer Science & Engineering, University of Washington

What it measures

Social IQa (SIQA) tests commonsense reasoning about people's motivations, reactions and mental states in everyday social interactions, rather than reasoning about the physical world the way contemporaries such as PIQA do. Each item gives a short context describing an interaction -- for example, "Jordan wanted to tell Tracy a secret, so Jordan leaned towards Tracy" -- plus a question about intent, effect or reaction, such as "Why did Jordan do this?", with three candidate answers. Contexts were seeded from event tuples in the ATOMIC commonsense knowledge graph. The authors designed a specific crowdsourcing framework to reduce a known artifact of earlier multiple-choice datasets: rather than having one worker write both a correct and an incorrect answer to the same question (which tends to give wrong answers a detectable "wrongness" in their surface style), incorrect answers were instead sourced as the correct answers to a different, related question, making them harder to rule out on style alone.

Task format

Three-way multiple-choice question answering over a short social context and question, typically zero- or few-shot, English only; no supporting passage beyond the one- or two-sentence context is supplied.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub