OpinionQA

1,498 Pew American Trends Panel multiple-choice opinion questions used to compare a model's answer distribution with 60 US demographic groups, not to score accuracy.

Also known as: OpinionsQA, opinions_qa, Opinion QA

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorysafety
Subcategoryalignment of model opinion distributions with US survey groups
Page statusactive
Metricrepresentativeness / alignment (1 minus normalized 1-Wasserstein distance)
Directionhigher_is_better
Unit0-1 alignment
Dataset size1498
PublisherStanford University and Columbia University

What it measures

OpinionQA asks a model the same multiple-choice public-opinion questions Pew Research asked US adults, covering topics such as guns, abortion, privacy, and science. There is no gold answer. The object is the distribution over choices, compared with the weighted distribution of all respondents or of a named group (for example Democrats, age 65+, or high income). A second mode prepends group context to test whether the model can be steered toward that group's views.

Task format

Multiple-choice question with ordinal options plus an optional Refused choice. Representativeness uses the question alone. Steerability prepends steer-qa, steer-bio, or steer-portray context. Scoring uses next-token log probabilities over choice letters, not generated prose.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub