ASDiv

2,305 elementary English math word problems with annotated type and grade; lm-eval runs a log-likelihood task and an 8-shot GSM8K-style CoT variant.

Also known as: ASDiv, ASDIV, Academia Sinica Diverse MWP Dataset, nlu-asdiv-dataset

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategoryelementary English math word problems with type and grade tags
Page statusunknown
Metricaccuracy (lm-eval asdiv: acc on log-likelihood; asdiv_cot_llama: exact_match)
Directionhigher_is_better
Unit%
Dataset size2305
Dataset licenceCC-BY-NC-4.0
PublisherNatural Language Understanding laboratory, Institute of Information Science, Academia Sinica

What it measures

ASDiv (Academia Sinica Diverse MWP Dataset) tests whether a solver can answer short English math word problems taught in elementary school. Each item has a story body, a question, a numeric answer that may include a unit, an annotated formula, a solution type (24 types in the authors' tag set), and a grade level. The authors built it because earlier MWP corpora were narrow in wording or in operation mix. It is one-unknown school arithmetic and related elementary types, not contest math.

Task format

Free-response English word problem. lm-eval `asdiv` concatenates body and question, then scores log-likelihood of the gold answer with the unit stripped. `asdiv_cot_llama` is 8-shot generate-until with the GSM8K CoT wording from Wei et al. 2022, exact_match after "The final answer is". Formulas are ignored in both harness tasks.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub