MC-TACO

13k English temporal-commonsense candidate answers; the paper scores question-level EM/F1, while lm-eval reports pair-level acc/F1.

Also known as: MC-TACO, MCTACO, mc-taco, Multiple Choice TemporAl COmmonsense

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorytemporal commonsense as per-candidate yes/no plausibility
Page statusunknown
Metricpaper: question-level EM and F1; lm-eval: pair-level acc and F1
Directionhigher_is_better
Unit%
Dataset size13225
PublisherUniversity of Pennsylvania / Allen Institute for AI / University of Illinois

What it measures

MC-TACO (Multiple Choice TemporAl COmmonsense) gives a MultiRC context sentence, a temporal question, and one candidate answer. The system must say whether that candidate is plausible. Five properties: duration, ordering, typical time, frequency, and stationarity. More than one candidate per question can be yes. English only. This is not a single exclusive MCQ and not a cooking-TACO code benchmark.

Task format

lm-eval multiple_choice on CogComp/mc_taco. Prompt: "{sentence} Question: {question} Answer: {answer} Plausible:". Choices no/yes. Target is the binary label. validation and test splits. should_decontaminate true. --limit shuffles pairs and can drop some options of a question.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub