{
 "body": "\n## What it measures\n\nThe Pile evaluation measures autoregressive language-model fit on diverse text. HELM's `the_pile` scenario downloads the corpus's public test split and computes predictive loss on one of 22 named subsets at a time (ArXiv, Books3, Github, Wikipedia (en), and so on); it does not itself expose a single pooled score across all subsets.\n\nThe corpus spans many domains, so results on one subset (for example, PhilPapers) say little about a model's fit on another (for example, Github code). HELM's own taxonomy tags the scenario's language as \"English, code,\" reflecting that some subsets (Github) are source code rather than prose.\n\n## How it is scored\n\nHELM's scenario metadata sets `bits_per_byte` as the main metric, where lower is better, with `test` as the main split. For most of the 22 subsets, HELM subsamples the full test set using pinned index files hosted in the EleutherAI/lm_perplexity repository; three small subsets (Ubuntu IRC, BookCorpus2, PhilPapers) are used in full because they are too small to subsample. Per-subset results are more informative than any pooled value, since the scenario is parameterized by subset rather than reporting one number.\n\n## Dataset and licence\n\nThe Pile paper describes an 825GiB corpus assembled from 22 sources, with validation and test splits each about 0.1% of the data, sampled uniformly. The HELM scenario downloads the public `test.jsonl.zst` archive directly from the-eye.eu. The Hugging Face `EleutherAI/pile` dataset card lists its licence as \"other\"; the compilation code itself is separately licensed, and the constituent sources (Books3, PubMed, USPTO, and the rest) keep their own original licences, so no single licence governs the whole corpus.\n\n## Who publishes it\n\nEleutherAI introduced The Pile in a paper posted to arXiv in December 2020, by Leo Gao, Stella Biderman, and ten collaborators. Stanford CRFM HELM maintains the evaluation scenario used here. No current standalone leaderboard was established.\n\n## Lineage\n\nThe Pile is a corpus with 22 named component subsets (ArXiv, PhilPapers, and 20 others), but HELM does not assign these separate benchmark ids; they are values of the scenario's `subset` parameter, not standalone pages. The lm-evaluation-harness exposes the same underlying corpus as a differently structured task group, documented separately in this repository as [pile](pile.md) (with per-component tasks named like `pile_arxiv` and `pile_philpapers`); that page and this one describe the same corpus through two different harnesses and should not be treated as duplicates of one another.\n\n## Saturation and contamination\n\nContamination risk is high because The Pile is public and widely used for model training. Perplexity can therefore reflect memorization as well as general language modeling.\n\n## How to run it\n\nUse HELM's `the_pile` scenario, selecting one of its 22 named subsets; the scenario handles downloading and caching the test archive itself. Record which subset, the HELM and lm_perplexity index-file revisions, and the tokenizer used, since bits-per-byte depends on tokenization even though it is designed to be less sensitive to it than word perplexity. To compare against lm-evaluation-harness's `pile` group instead, see [pile](pile.md).\n\n## Reading the numbers\n\nLower bits per byte indicates better predictive compression on the selected texts. It does not establish factuality, reasoning, or quality of generated responses. Compare the same subset, tokenizer, and preprocessing, and interpret pooled scores cautiously.\n\nThe corpus\u2019s mixed provenance means that one aggregate can conceal large domain differences. Component-level reporting is essential for useful diagnosis.\n\nPerplexity is also tokenizer-dependent. A comparison should keep the tokenizer, byte accounting, document filtering, and test archive fixed, or the resulting values are not directly comparable.\n\nThe test corpus is a measurement instrument, not a guarantee of representative language use. Domain and source breakdowns help reveal where a model\u2019s predictive advantage comes from.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "the_pile",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "generation",
  "contamination": {
   "note": "The corpus is public and widely used in language-model training; exposure is likely but model-specific overlap is not measured here.",
   "risk": "high"
  },
  "dataset": {
   "languages": [
    "English"
   ],
   "license": "other (per the EleutherAI/pile Hugging Face card); constituent sources keep their own licences",
   "modalities": [
    "text",
    "code"
   ],
   "public_test_set": true,
   "size": null,
   "size_note": "The HELM scenario downloads the full public test.jsonl.zst from the-eye.eu and does not itself state a document count; the paper reports an 825GiB corpus across 22 subsets with validation and test each about 0.1% of the data. For most subsets, HELM further subsamples via pinned index files from EleutherAI/lm_perplexity; 3 small subsets (Ubuntu IRC, BookCorpus2, PhilPapers) use all instances unsampled.",
   "splits": "test",
   "url": "https://arxiv.org/abs/2101.00027"
  },
  "freshness": {
   "luna-new-001 (Codex coordinated)": null,
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna",
   "reviewed": "2026-09-08",
   "reviewed_by": "Claude Sonnet 5 independent review, luna-new-001"
  },
  "harness": {
   "bigbench": "",
   "helm": "the_pile",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": ""
  },
  "id": "the_pile",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "The Pile is a large mixed-domain text corpus used for language-model evaluation. HELM scores predictive fit on test documents and supports component subsets such as ArXiv and PhilPapers.",
  "metric": {
   "baseline_note": "No human baseline is stated in the scenario.",
   "direction": "lower_is_better",
   "human_baseline": null,
   "max_score": null,
   "name": "bits_per_byte",
   "random_baseline": null,
   "unit": "bits/byte"
  },
  "name": "The Pile",
  "page_kind": "family",
  "paper": {
   "arxiv": "2101.00027",
   "title": "The Pile: An 800GB Dataset of Diverse Text for Language Modeling",
   "url": "https://arxiv.org/abs/2101.00027",
   "year": 2020
  },
  "publisher": {
   "authors": [
    "Leo Gao",
    "Stella Biderman",
    "Sid Black",
    "Laurence Golding",
    "Travis Hoppe",
    "Charles Foster",
    "Jason Phang",
    "Horace He",
    "Anish Thite",
    "Noa Nabeshima",
    "Shawn Presser",
    "Connor Leahy"
   ],
   "org": "EleutherAI",
   "url": "https://github.com/EleutherAI/the-pile"
  },
  "released": "2020-12",
  "repo_url": "https://github.com/EleutherAI/the-pile",
  "saturation": {
   "as_of": "",
   "note": "No current leaderboard was established.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "HELM The Pile scenario",
    "url": "https://raw.githubusercontent.com/stanford-crfm/helm/main/src/helm/benchmark/scenarios/the_pile_scenario.py"
   },
   {
    "accessed": "2026-09-08",
    "title": "The Pile paper",
    "url": "https://arxiv.org/abs/2101.00027"
   },
   {
    "accessed": "2026-09-08",
    "title": "The Pile repository",
    "url": "https://github.com/EleutherAI/the-pile"
   },
   {
    "accessed": "2026-09-08",
    "title": "Hugging Face dataset card metadata for EleutherAI/pile",
    "url": "https://huggingface.co/api/datasets/EleutherAI/pile"
   }
  ],
  "status": "active",
  "subcategory": "language modeling corpus (HELM scenario; see also the lm-evaluation-harness group at pile.md)",
  "summary": "HELM evaluates language-model perplexity on a test slice of The Pile across its component domains.",
  "tags": [
   "language-modeling",
   "perplexity",
   "mixed-domain"
  ],
  "task_format": "Autoregressive next-token prediction over text documents."
 }
}