RACE-H (inspect_evals)

inspect_evals task for RACE-H: pick one of four answers to a high-school English-exam question about a passage.

Also known as: RACE-H, race-high

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryinspect_evals generation wrap of RACE high-school English-exam reading comprehension
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size3498
Dataset licenceNon-commercial research use only (custom RACE licence; Hugging Face tags it 'other')
PublisherCarnegie Mellon University (dataset); inspect_evals packaging by UK AI Security Institute

What it measures

race_h is the high-school slice of RACE (Lai et al., EMNLP 2017), packaged by inspect_evals. The model reads an English exam passage written for Chinese students, a question or cloze, and four options, then must emit ANSWER: <letter>. Questions are meant to need inference, not span copy. This page is the inspect_evals protocol. The dataset, splits, and licence are documented on [race](race.md).

Task format

Four-way multiple choice, English, generation. inspect_evals template requires the entire response to be 'ANSWER: $LETTER'. Solver is multiple_choice; scorer is choice(). Temperature 0. Shuffle default True. Hugging Face ehovy/race config high, split test, revision 2fec9fd81f1dc971569a9b729c43f2f0e6436637.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub