{
 "body": "## What it measures\n\nThis benchmark asks whether widely used commonsense evaluations predict performance on practical downstream reasoning tasks. The study evaluates 23 models from six model families on four established commonsense benchmarks, four reworked variants, three non-commonsense controls, and eight downstream tasks.\n\nThe downstream tasks require implicit social, pragmatic, temporal, or physical reasoning. The benchmark therefore measures predictive validity: whether a model\u2019s position on a proxy test forecasts its position on a related real-world task.\n\n## How it is scored\n\nThe authors compare model rankings, compute controlled correlations, and use leave-one-family-out cross-validation. Higher correlation means stronger predictive alignment for the tested downstream task. The protocol is about relationships among scores, so a raw benchmark accuracy should not be confused with the primary validity result. Exact item-level metrics vary by component task.\n\n## Dataset and licence\n\nThe abstract establishes the composition of four commonsense benchmarks, four revised variants, three controls, and eight downstream tasks, but it does not provide a single item count or unified licence. Dataset licences and answer visibility differ by component and are not established in this record. Reproduction requires obtaining each cited task and preserving the study\u2019s model-family splits.\n\n## Who publishes it\n\nIne Gevers and Walter Daelemans published the study as an arXiv preprint submitted in August 2026. The arXiv record is the primary source consulted. No independent leaderboard is identified; the reported outputs are validity analyses rather than a standing model-ranking board.\n\n## Lineage\n\nThis is a meta-evaluation of existing commonsense benchmarks and their reworked variants. It has no single predecessor or successor benchmark. Its controls and downstream tasks are component evaluations, not separate lineage pages in this catalogue.\n\n## Saturation and contamination\n\nThe study finds that revised benchmarks largely preserve original model rankings and do not improve downstream predictive power. It reports consistent cross-family validity only for a narrow subset of downstream tasks, with smaller or metric-specific gains elsewhere. Those findings concern transfer validity, not score saturation. Item contamination is not established by the consulted abstract, so risk is unknown.\n\n## How to run it\n\nReproduce the component benchmark protocols, evaluate the same model families, and keep family-held-out folds intact. Compute ranking comparisons, controlled correlations, and leave-one-family-out predictions as reported. Record prompt, scoring, and revision details for each component because changing one benchmark can change the validity estimate.\n\n## Reading the numbers\n\nA strong correlation means that the tested commonsense score tracks a downstream task under this study\u2019s sample and controls. It does not show that the benchmark measures broad commonsense competence. A weak correlation can reflect task mismatch rather than model failure. Read per-task and held-out-family results, plus the control comparisons, before generalizing the conclusion.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "benchmarking_the_benchmarks",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "category": "reasoning",
  "contamination": {
   "note": "The paper evaluates existing benchmarks but the consulted abstract does not establish item-level contamination.",
   "risk": "unknown"
  },
  "dataset": {
   "modalities": [
    "text"
   ],
   "public_test_set": null
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna, luna-stream-a-003 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "other": "The task implementations and benchmark protocols reported in the paper."
  },
  "id": "benchmarking_the_benchmarks",
  "measures": "This evaluation studies criterion validity rather than a single capability. It compares model rankings on established commonsense benchmarks, revised variants, non-commonsense controls, and downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning.",
  "metric": {
   "baseline_note": "The paper reports controlled correlations and leave-one-family-out validation; no universal maximum is authored here.",
   "direction": "higher_is_better",
   "name": "ranking correlation",
   "unit": "correlation"
  },
  "name": "Benchmarking the Benchmarks",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "2608.03340",
   "title": "Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks",
   "url": "https://arxiv.org/abs/2608.03340",
   "year": 2026
  },
  "publisher": {
   "authors": [
    "Ine Gevers",
    "Walter Daelemans"
   ],
   "org": "Ine Gevers and Walter Daelemans",
   "url": "https://arxiv.org/abs/2608.03340"
  },
  "released": "2026-08",
  "saturation": {
   "note": "This is a validity study; it does not define a capability ceiling.",
   "status": "unknown"
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "Benchmarking the Benchmarks paper and abstract",
    "url": "https://arxiv.org/abs/2608.03340"
   }
  ],
  "summary": "Benchmarking the Benchmarks tests whether commonsense benchmark rankings predict performance on downstream social, pragmatic, temporal, and physical reasoning tasks.",
  "tags": [
   "commonsense",
   "validity",
   "downstream-prediction"
  ],
  "task_format": "Multiple-choice or task-specific benchmark evaluations across 23 models from six model families."
 }
}