Training on Test Set

BIG-bench task designed to detect evidence that a language model was trained on benchmark data.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategorydata contamination detection
Page statusactive
Metriclog10 p-value of canary probability deviation (log10_p_dev)
Directionhigher_is_better
Dataset licenceApache-2.0
PublisherGoogle BIG-bench

What it measures

The task measures whether a model assigns anomalous conditional log-probability to a hard-coded BIG-bench canary GUID and to BIG-bench's own git commit hashes, compared with random control strings, as evidence the model was trained on BIG-bench's public repository.

Task format

Conditional log-probability scoring of fixed canary strings against random control strings; the model is not asked to generate text.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub