A 2,000-item BIG-bench NLI task that asks whether a nested knowledge or belief premise entails a hypothesis, scored as two-way multiple choice.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | BIG-bench nested knowledge/belief NLI (2,000 items, ten templates) |
| Page status | unknown |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2000 |
| Dataset licence | Apache-2.0 |
| Publisher | Google (BIG-bench collaboration) |
epistemic_reasoning is an English natural-language-inference probe of epistemic theory of mind. Each item pairs a premise and a hypothesis built around stacked factive verbs (knows, sees, learns, understands, recognizes, remembers) and non-factive verbs (believes, thinks, assumes, suspects). The intended skill is whether the model tracks other agents' mental states, not speaker-commitment datasets that stress negation and modals. It is not [bbh](bbh.md).
Two-option multiple choice (entailment vs non-entailment), preferred metric multiple_choice_grade. task_prefix asks the model to identify the relation between premises and hypotheses. append_choices_to_input is false. Canary GUID embedded.
No model card in ModelSpec reports this benchmark yet.