LawBench scores a model's Chinese legal knowledge across 20 tasks grouped into memorization, understanding and application, drawn from real legal databases, exams and court documents.
unassessed
| Category | domain |
|---|---|
| Subcategory | Chinese legal knowledge across three cognitive levels (memorization, understanding, application), 20 tasks |
| Page status | active |
| Metric | task-specific metric (accuracy, F1, rc-F1, soft-F1, F0.5, ROUGE-L, or normalized log-distance depending on task), averaged across tasks for a headline score |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 10000 |
| Dataset licence | LawBench is a mix of created and transformed datasets, and its own README asks users to follow the licence of each original source dataset (CAIL2018, CAIL2019, CAIL2021, CAIL2022, JEC_QA, LAIC2021, LEVEN and others) rather than stating one combined licence for the benchmark; the LawBench GitHub repository's own code is separately licensed Apache-2.0. |
| Publisher | Shanghai AI Laboratory, Amazon Alexa AI, Saarland University and Nanjing University |
LawBench tests a model's legal knowledge of the Chinese civil-law system across three cognitive levels: memorization (reciting law article text or answering legal-concept questions), understanding (proofreading legal documents, extracting named entities or events, identifying the point of dispute in a case, or summarizing legal news, among others) and application (predicting which article or charge applies to a set of facts, predicting a prison term, answering multiple-choice case-analysis questions in the style of China's judicial exam, or drafting a legal consultation answer). Its 20 tasks are drawn from a mix of sources: the national legal-article database, the JEC_QA judicial-exam question bank, and several years of the CAIL (Chinese AI and Law) shared-task datasets and the LAIC and LEVEN legal NLP datasets, rather than from questions written for the benchmark itself.
Task format varies by task type: single-label or multi-label classification, regression (for example, predicting a prison term in months), span extraction, or free-text generation, each read from a Chinese- language prompt built from real legal text (a statute, a case summary, a court judgment excerpt, or a consultation question).
No model card in ModelSpec reports this benchmark yet.