1,498 Pew American Trends Panel multiple-choice opinion questions used to compare a model's answer distribution with 60 US demographic groups, not to score accuracy.
unassessed
| Category | safety |
|---|---|
| Subcategory | alignment of model opinion distributions with US survey groups |
| Page status | active |
| Metric | representativeness / alignment (1 minus normalized 1-Wasserstein distance) |
| Direction | higher_is_better |
| Unit | 0-1 alignment |
| Dataset size | 1498 |
| Publisher | Stanford University and Columbia University |
OpinionQA asks a model the same multiple-choice public-opinion questions Pew Research asked US adults, covering topics such as guns, abortion, privacy, and science. There is no gold answer. The object is the distribution over choices, compared with the weighted distribution of all respondents or of a named group (for example Democrats, age 65+, or high income). A second mode prepends group context to test whether the model can be steered toward that group's views.
Multiple-choice question with ordinal options plus an optional Refused choice. Representativeness uses the question alone. Steerability prepends steer-qa, steer-bio, or steer-portray context. Scoring uses next-token log probabilities over choice letters, not generated prose.
No model card in ModelSpec reports this benchmark yet.