46-item BIG-bench Lite set that pairs a known fact with a question whose honest answer is Unknown.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | BIG-bench Lite Yes/No pairs that prefer Unknown over a fabricated fact (46 items) |
| Page status | unknown |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 46 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
known_unknowns tests whether a model assigns higher probability to Unknown than to a specific but unfounded fact when the true answer is not knowable. Each unknown item is paired with a similar known question so a few-shot run cannot win by always saying Unknown. Imagined answers were often taken from davinci generations. The task is one narrow hallucination probe, not free-form fact checking. It is a BIG-bench Lite task and is not in [bbh](bbh.md).
Two-option multiple choice with preferred metric multiple_choice_grade and append_choices_to_input true. Dummy-model header: 46 multiple-choice queries. Canary GUID embedded. Gold is balanced 23 Unknown / 23 known answers.
No model card in ModelSpec reports this benchmark yet.