LawBench

LawBench scores a model's Chinese legal knowledge across 20 tasks grouped into memorization, understanding and application, drawn from real legal databases, exams and court documents.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryChinese legal knowledge across three cognitive levels (memorization, understanding, application), 20 tasks
Page statusactive
Metrictask-specific metric (accuracy, F1, rc-F1, soft-F1, F0.5, ROUGE-L, or normalized log-distance depending on task), averaged across tasks for a headline score
Directionhigher_is_better
Unit%
Dataset size10000
Dataset licenceLawBench is a mix of created and transformed datasets, and its own README asks users to follow the licence of each original source dataset (CAIL2018, CAIL2019, CAIL2021, CAIL2022, JEC_QA, LAIC2021, LEVEN and others) rather than stating one combined licence for the benchmark; the LawBench GitHub repository's own code is separately licensed Apache-2.0.
PublisherShanghai AI Laboratory, Amazon Alexa AI, Saarland University and Nanjing University

What it measures

LawBench tests a model's legal knowledge of the Chinese civil-law system across three cognitive levels: memorization (reciting law article text or answering legal-concept questions), understanding (proofreading legal documents, extracting named entities or events, identifying the point of dispute in a case, or summarizing legal news, among others) and application (predicting which article or charge applies to a set of facts, predicting a prison term, answering multiple-choice case-analysis questions in the style of China's judicial exam, or drafting a legal consultation answer). Its 20 tasks are drawn from a mix of sources: the national legal-article database, the JEC_QA judicial-exam question bank, and several years of the CAIL (Chinese AI and Law) shared-task datasets and the LAIC and LEVEN legal NLP datasets, rather than from questions written for the benchmark itself.

Task format

Task format varies by task type: single-label or multi-label classification, regression (for example, predicting a prison term in months), span extraction, or free-text generation, each read from a Chinese- language prompt built from real legal text (a statute, a case summary, a court judgment excerpt, or a consultation question).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub