The Pile

HELM evaluates language-model perplexity on a test slice of The Pile across its component domains.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
Subcategorylanguage modeling corpus (HELM scenario; see also the lm-evaluation-harness group at pile.md)
Page statusactive
Metricbits_per_byte
Directionlower_is_better
Unitbits/byte
Dataset licenceother (per the EleutherAI/pile Hugging Face card); constituent sources keep their own licences
PublisherEleutherAI

What it measures

The Pile is a large mixed-domain text corpus used for language-model evaluation. HELM scores predictive fit on test documents and supports component subsets such as ArXiv and PhilPapers.

Task format

Autoregressive next-token prediction over text documents.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub