CL-bench

CL-bench tests whether models can learn new domain knowledge, rules and procedures from complex contexts.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorylong-context
Subcategorycontext learning
Page statusactive
Metrictask success rate
Directionhigher_is_better
Unit%
Dataset size500
PublisherCL-bench authors

What it measures

Context-grounded task solving beyond retrieval or simple in-context pattern learning.

Task format

500 complex contexts, 1,899 tasks and 31,607 verification rubrics.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub