CaseHOLD (Case Holdings On Legal Decisions)

53,137 five-way multiple-choice questions that ask which holding statement matches a citing passage mined from US judicial opinions.

Also known as: Case Holdings On Legal Decisions

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryUS case-law holding identification
Page statusactive
Metricmacro F1 (paper); exact_match (HELM)
Directionhigher_is_better
Unit%
Dataset size53137
Dataset licenceApache-2.0
PublisherStanford RegLab

What it measures

CaseHOLD tests whether a model can pick the holding of a cited case from five candidate holding sentences. The prompt is the citing passage from a US opinion. One candidate is the holding that actually follows in that opinion; four are other holdings used as distractors. Holdings are the precedential rule of a decision, so the task is legal citation sense-making, not open legal advice.

Task format

Five-way multiple choice. The paper reports mean macro F1 over 10 folds. HELM loads Hugging Face casehold/casehold config `all`, keeps train and test (skips validation), prompts with “Give a letter answer among A, B, C, D, or E,” allows two in-context train items, and scores exact match.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub