International Corpus of English

ICE evaluates per-text perplexity across regional English corpora covering spoken and written language.

Also known as: ICE

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
SubcategoryEnglish language modeling
Page statusactive
Metricbits_per_byte
Directionlower_is_better
PublisherInternational Corpus of English

What it measures

The International Corpus of English contains written and spoken texts from regional English varieties. HELM evaluates language-model perplexity on texts after removing most corpus markup.

Task format

Language-model scoring over corpus texts, optionally filtered by region, speech/writing category, or supported gender metadata.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub