LexGLUE (Legal General Language Understanding Evaluation)

English legal NLU suite of seven public datasets; HELM scores each subset as generation with classification_macro_f1 on test.

Also known as: LexGLUE, Legal GLUE

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryseven-task English legal NLU suite (ECtHR A/B, SCOTUS, EUR-LEX, LEDGAR, UNFAIR-ToS, CaseHOLD)
Page statusactive
Metricclassification_macro_f1 (HELM schema); papers also report micro-F1
Directionhigher_is_better
Unit%
Dataset size23607
Dataset licenceCC-BY-4.0
PublisherUniversity of Copenhagen and co-authors (LexGLUE); Stanford CRFM (HELM scenario)

What it measures

LexGLUE packs seven English legal datasets behind one evaluation recipe. ECtHR A predicts violated Convention articles from facts. ECtHR B predicts articles the court considered. SCOTUS maps an opinion to a Supreme Court Database issue area. EUR-LEX assigns EuroVoc labels to an EU act. LEDGAR classifies an SEC contract provision into one of 100 topics. UNFAIR-ToS tags unfair term types in a consumer ToS sentence. CaseHOLD is five-way holding identification. HELM prompts these as generation, not encoder fine-tuning.

Task format

HELM run spec lex_glue:subset=<ecthr_a|ecthr_b|scotus|eurlex|ledgar|unfair_tos|case_hold> or subset=all. Scenario loads Hugging Face config lex_glue. Generation adapter, input noun Passage, output noun Answer. MLTC tasks use comma separated labels; CaseHOLD is QA with numbered endings.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub