metabench

An 858-item sparse subsample of ARC, GSM8K, HellaSwag, MMLU, TruthfulQA and WinoGrande that reconstructs Open LLM Leaderboard scores from a few percent of the items.

Also known as: MetaBench, metabench-A

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategorysparse reconstruction of Open LLM Leaderboard v1 benchmarks
Page statusactive
Metricaccuracy (group mean of per-source acc; GSM8K exact match), plus optional reconstructed Open LLM Leaderboard scores
Directionhigher_is_better
Unit%
Dataset size858
Dataset licenceCC-BY-NC-SA-4.0
PublisherHuman-Centered AI, Helmholtz Munich

What it measures

metabench does not add a new skill. It keeps the most informative items from six public benchmarks used on Hugging Face's Open LLM Leaderboard v1: ARC, GSM8K, HellaSwag, MMLU, TruthfulQA and WinoGrande. The authors fit item-response models on thousands of leaderboard submissions, then keep a few hundred items per source so that latent-ability estimates can reconstruct the original six scores and their mean. English text. Five of the six sources are multiple choice; GSM8K is free-form arithmetic. A secondary 751-item set and choice-permuted copies exist for repeat evaluation.

Task format

Mix of multiple-choice log-likelihood (ARC, HellaSwag, MMLU, TruthfulQA, WinoGrande) and generate-until exact match (GSM8K). lm-evaluation-harness tasks bake in the original few-shot preprompts and set num_fewshot to 0 so those shots are not added twice.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub