LegalBench

A collaboratively built suite of 162 tasks, contributed by lawyers and computer scientists, testing six categories of legal reasoning in language models.

Also known as: LEGALBENCH

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorylegal reasoning
Page statusactive
Metricaccuracy, balanced accuracy, F1, or human-graded correctness/analysis (varies by task)
Directionhigher_is_better
Unit%
Dataset size162
Dataset licenceMixed, task by task: most tasks are CC BY 4.0; a smaller number are CC BY-NC 4.0 (e.g. Canada Tax Court Outcomes, Consumer Contracts QA), CC BY-SA 4.0 (Definition tasks), CC BY-NC-SA 4.0 (Learned Hands tasks), MIT (NY Judicial Ethics, Privacy Policy QA, SARA), CC BY-NC (OPP-115) or CC BY-NC 3.0 (Privacy Policy Entailment), inherited from each task's original source dataset.
PublisherStanford University, with an interdisciplinary, 40-author collaboration spanning law schools, computer science departments and legal practice at institutions including the University of Chicago, Harvard Law School, Georgetown University Law Center and others

What it measures

LegalBench tests whether a model can perform the kinds of reasoning lawyers actually use, rather than treating "legal reasoning" as one undifferentiated skill. Its 162 tasks are organised into six categories drawn from how the legal profession itself frames reasoning: issue-spotting (does a fact pattern raise a given legal question), rule-recall (state or identify the applicable rule), rule-application (explain how a rule applies to facts, with reasoning graded for correctness and analysis), rule-conclusion (state the resulting legal outcome), interpretation (parse a contract, privacy policy or statute), and rhetorical-understanding (reason about legal argument and judicial writing). Inputs are short clauses, fact patterns, questions or case excerpts; most tasks are single-turn, English-only, and text-only.

Task format

Varies by task: multiple-choice (35 tasks), open-ended generation (7), binary classification (112), and multi-class or multi-label classification (8). Few-shot prompting is standard, using 0 to 8 in-context demonstrations drawn from each task's own small training split.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Mistral Medium (latest)Mistral AI89.42026-04
Mistral Medium 3Mistral AI89.42026-04
Mistral Medium 3.1Mistral AI89.42026-04
Gemini 3.1 Pro PreviewGoogle DeepMind87.42026-04
Gemini 3 Pro PreviewGoogle DeepMind87.02026-04
Gemini 3 Flash PreviewGoogle DeepMind86.92026-04
GPT-5OpenAI86.02026-04
Claude Opus 4Anthropic75.52026-04
Claude Opus 4.6Anthropic75.52026-04
Gemini 2.5 ProGoogle DeepMind73.82026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind73.82026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind73.82026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind73.82026-04
GPT-4oOpenAI72.12026-04
GPT-4o (2024-05-13)OpenAI72.12026-04
GPT-4o (2024-08-06)OpenAI72.12026-04
GPT-4o (2024-11-20)OpenAI72.12026-04
GPT-4o miniOpenAI72.12026-04
Llama 3.1 70BMeta58.52026-04
Llama 3.1 70B InstructMeta58.52026-04

Data

This page as JSON · Edit on GitHub