HELM evaluates language-model perplexity on a test slice of The Pile across its component domains.
unassessed
| Category | generation |
|---|---|
| Subcategory | language modeling corpus (HELM scenario; see also the lm-evaluation-harness group at pile.md) |
| Page status | active |
| Metric | bits_per_byte |
| Direction | lower_is_better |
| Unit | bits/byte |
| Dataset licence | other (per the EleutherAI/pile Hugging Face card); constituent sources keep their own licences |
| Publisher | EleutherAI |
The Pile is a large mixed-domain text corpus used for language-model evaluation. HELM scores predictive fit on test documents and supports component subsets such as ArXiv and PhilPapers.
Autoregressive next-token prediction over text documents.
No model card in ModelSpec reports this benchmark yet.