LMentry

Twenty-five elementary English tasks that humans solve perfectly, scored as accuracy multiplied by robustness to trivial input changes.

Also known as: LMentry, lm_entry, HELM lm_entry

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryelementary English language unit tests (25 tasks, accuracy × robustness)
Page statusactive
MetricLMentry score (mean accuracy × mean robustness); HELM reports quasi_exact_match
Directionhigher_is_better
Unit%
Dataset size110703
PublisherTel Aviv University

What it measures

LMentry checks skills an elementary-school student is expected to have: write a sentence that contains a given word, pick the longer of two words, name the first letter, say which of two numbers is bigger, pick a rhyme, and similar. Avia Efrat, Or Honovich, and Omer Levy built 25 such tasks as a compact zero-shot unit test, not as a hard reasoning contest. Each task has three prompt templates. The suite also measures brittleness to argument order, argument content, template wording, and adjacent-task flips. HELM's scenario id is lm_entry; the evaluation itself is LMentry.

Task format

Zero-shot generation (official suite) or HELM generation / multiple-choice joint. Official scoring uses about ten regexes per task. HELM generation uses quasi-exact match with a one-word instruction and max_tokens 10. HELM multiple-choice joint is offered for tasks that can be cast as options; letter/word extraction tasks cannot.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub