TimeDial

BIG-bench TimeDial asks models to select the correct answer for masked temporal spans in dialogue context.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorytemporal dialogue understanding
Page statusactive
Metricmultiple_choice_grade
Directionhigher_is_better
Unitpercent
Dataset size2550
PublisherGoogle BIG-bench

What it measures

TimeDial evaluates temporal and social reasoning over dialogue. The task metadata describes choosing the correct option for a masked temporal span given the surrounding conversation.

Task format

Dialogue context with a masked temporal expression and multiple choices.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub