BIG-bench TimeDial asks models to select the correct answer for masked temporal spans in dialogue context.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | temporal dialogue understanding |
| Page status | active |
| Metric | multiple_choice_grade |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 2550 |
| Publisher | Google BIG-bench |
TimeDial evaluates temporal and social reasoning over dialogue. The task metadata describes choosing the correct option for a masked temporal span given the surrounding conversation.
Dialogue context with a masked temporal expression and multiple choices.
No model card in ModelSpec reports this benchmark yet.