Benchmarking the Benchmarks

Benchmarking the Benchmarks tests whether commonsense benchmark rankings predict performance on downstream social, pragmatic, temporal, and physical reasoning tasks.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Metricranking correlation
Directionhigher_is_better
Unitcorrelation
PublisherIne Gevers and Walter Daelemans

What it measures

This evaluation studies criterion validity rather than a single capability. It compares model rankings on established commonsense benchmarks, revised variants, non-commonsense controls, and downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning.

Task format

Multiple-choice or task-specific benchmark evaluations across 23 models from six model families.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub