20 closed- and open-ended subtasks over 508 long documents (3k-200k tokens); its length-instruction-enhanced protocol curbs n-gram metrics' bias toward longer outputs.
unassessed
| Category | long-context |
|---|---|
| Subcategory | standardized long-context evaluation across 20 closed- and open-ended subtasks, 3k-200k token inputs |
| Page status | active |
| Metric | exact-match accuracy for closed-ended tasks; F1/ROUGE-L (with and without length-instruction-enhanced prompting), GPT-4/GPT-3.5 pairwise LLM-judge win rate, and 5-point human ratings for open-ended tasks |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 537 |
| Dataset licence | GPL-3.0, per both the Hugging Face dataset card and the GitHub repository. |
| Publisher | Fudan University; The University of Hong Kong; Shanghai AI Laboratory; University of Illinois Urbana-Champaign |
L-Eval tests long-context understanding across 20 subtasks split into two groups: 7 closed-ended tasks graded by exact match (TOEFL-style reading comprehension, a 16-shot long-context grade-school-math set, QuALITY multiple-choice questions, Coursera lecture-transcript questions, a topic-retrieval task, a science-fiction reasoning set, and a code-understanding set) and 13 open-ended generation tasks (financial, legal, scientific and general-domain question answering, plus summarization of government reports, patents, news, TV show transcripts, peer reviews and meeting transcripts). Inputs range from about 3,000 to 200,000 tokens across the suite, built from 508 long documents carrying more than 2,000 human-labeled query-response pairs in total. OpenCompass, the source this page's census hint names, implements 18 of these 20 subtasks -- all 13 open-ended tasks plus 5 of the 7 closed-ended ones, omitting the code-understanding (CodeU) and science-fiction (SFiction) tasks.
A long document (or documents) plus a task-specific question, instruction or exam-style prompt in; for closed-ended tasks, a short exact-match answer (a letter, a number, a word) out; for open-ended tasks, free-form generation (an answer or a summary), scored by several different metrics depending on which protocol is used.
No model card in ModelSpec reports this benchmark yet.