CL-bench Life evaluates context learning over messy, fragmented real-life contexts such as conversations, archives and behavioral traces.
unassessed
| Category | long-context |
|---|---|
| Subcategory | real-life context learning |
| Page status | active |
| Metric | task success rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 405 |
| Publisher | CL-bench Life authors |
Task solving grounded in complex everyday-life contexts.
405 context-task pairs and 5,348 verification rubrics.
No model card in ModelSpec reports this benchmark yet.