CL-bench Life

CL-bench Life evaluates context learning over messy, fragmented real-life contexts such as conversations, archives and behavioral traces.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorylong-context
Subcategoryreal-life context learning
Page statusactive
Metrictask success rate
Directionhigher_is_better
Unit%
Dataset size405
PublisherCL-bench Life authors

What it measures

Task solving grounded in complex everyday-life contexts.

Task format

405 context-task pairs and 5,348 verification rubrics.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub