Benchmarking the Benchmarks tests whether commonsense benchmark rankings predict performance on downstream social, pragmatic, temporal, and physical reasoning tasks.
unassessed
| Category | reasoning |
|---|---|
| Metric | ranking correlation |
| Direction | higher_is_better |
| Unit | correlation |
| Publisher | Ine Gevers and Walter Daelemans |
This evaluation studies criterion validity rather than a single capability. It compares model rankings on established commonsense benchmarks, revised variants, non-commonsense controls, and downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning.
Multiple-choice or task-specific benchmark evaluations across 23 models from six model families.
No model card in ModelSpec reports this benchmark yet.