tinyBenchmarks

One-hundred-item IRT-selected subsets of Open LLM Leaderboard tasks that reconstruct full-benchmark accuracy from a few percent of the original items.

Also known as: tiny Benchmarks

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
SubcategoryIRT-compressed Open LLM Leaderboard evaluation
Page statusactive
Metricgp-IRT reconstructed parent accuracy (IRT++ in the harness README)
Directionhigher_is_better
Unit%
Dataset size600
Dataset licenceMIT
PublisherUniversity of Michigan / IBM Research / Universitat Pompeu Fabra / MIT

What it measures

tinyBenchmarks does not add a new skill. It keeps about 100 examples from each of several large English benchmarks so that an item-response model can estimate the score the model would have got on the full set. The lm-evaluation-harness group covers the Open LLM Leaderboard v1 six: ARC, GSM8K, HellaSwag, MMLU, TruthfulQA and WinoGrande. The paper also releases tiny AlpacaEval 2.0 and sketches HELM Lite; those are not in the harness group. The number to trust is the reconstructed parent accuracy, not the raw 100-item hit rate.

Task format

Same formats as the parents: multiple-choice log-likelihood for ARC, HellaSwag, MMLU, TruthfulQA and WinoGrande; generate-until exact match for GSM8K. Shot counts in the group yaml follow the Open LLM Leaderboard (for example tinyArc 25-shot, tinyGSM8k 5-shot, tinyMMLU 0 extra shots with a stored prompt).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub