inspect_evals task for RACE-H: pick one of four answers to a high-school English-exam question about a passage.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | inspect_evals generation wrap of RACE high-school English-exam reading comprehension |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 3498 |
| Dataset licence | Non-commercial research use only (custom RACE licence; Hugging Face tags it 'other') |
| Publisher | Carnegie Mellon University (dataset); inspect_evals packaging by UK AI Security Institute |
race_h is the high-school slice of RACE (Lai et al., EMNLP 2017), packaged by inspect_evals. The model reads an English exam passage written for Chinese students, a question or cloze, and four options, then must emit ANSWER: <letter>. Questions are meant to need inference, not span copy. This page is the inspect_evals protocol. The dataset, splits, and licence are documented on [race](race.md).
Four-way multiple choice, English, generation. inspect_evals template requires the entire response to be 'ANSWER: $LETTER'. Solver is multiple_choice; scorer is choice(). Temperature 0. Shuffle default True. Hugging Face ehovy/race config high, split test, revision 2fec9fd81f1dc971569a9b729c43f2f0e6436637.
No model card in ModelSpec reports this benchmark yet.