One-hundred-item IRT-selected subsets of Open LLM Leaderboard tasks that reconstruct full-benchmark accuracy from a few percent of the original items.
unassessed
| Category | composite |
|---|---|
| Subcategory | IRT-compressed Open LLM Leaderboard evaluation |
| Page status | active |
| Metric | gp-IRT reconstructed parent accuracy (IRT++ in the harness README) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 600 |
| Dataset licence | MIT |
| Publisher | University of Michigan / IBM Research / Universitat Pompeu Fabra / MIT |
tinyBenchmarks does not add a new skill. It keeps about 100 examples from each of several large English benchmarks so that an item-response model can estimate the score the model would have got on the full set. The lm-evaluation-harness group covers the Open LLM Leaderboard v1 six: ARC, GSM8K, HellaSwag, MMLU, TruthfulQA and WinoGrande. The paper also releases tiny AlpacaEval 2.0 and sketches HELM Lite; those are not in the harness group. The number to trust is the reconstructed parent accuracy, not the raw 100-item hit rate.
Same formats as the parents: multiple-choice log-likelihood for ARC, HellaSwag, MMLU, TruthfulQA and WinoGrande; generate-until exact match for GSM8K. Shot counts in the group yaml follow the Open LLM Leaderboard (for example tinyArc 25-shot, tinyGSM8k 5-shot, tinyMMLU 0 extra shots with a stored prompt).
No model card in ModelSpec reports this benchmark yet.