CL-bench tests whether models can learn new domain knowledge, rules and procedures from complex contexts.
unassessed
| Category | long-context |
|---|---|
| Subcategory | context learning |
| Page status | active |
| Metric | task success rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 500 |
| Publisher | CL-bench authors |
Context-grounded task solving beyond retrieval or simple in-context pattern learning.
500 complex contexts, 1,899 tasks and 31,607 verification rubrics.
No model card in ModelSpec reports this benchmark yet.