A collaboratively built suite of 162 tasks, contributed by lawyers and computer scientists, testing six categories of legal reasoning in language models.
unassessed
| Category | domain |
|---|---|
| Subcategory | legal reasoning |
| Page status | active |
| Metric | accuracy, balanced accuracy, F1, or human-graded correctness/analysis (varies by task) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 162 |
| Dataset licence | Mixed, task by task: most tasks are CC BY 4.0; a smaller number are CC BY-NC 4.0 (e.g. Canada Tax Court Outcomes, Consumer Contracts QA), CC BY-SA 4.0 (Definition tasks), CC BY-NC-SA 4.0 (Learned Hands tasks), MIT (NY Judicial Ethics, Privacy Policy QA, SARA), CC BY-NC (OPP-115) or CC BY-NC 3.0 (Privacy Policy Entailment), inherited from each task's original source dataset. |
| Publisher | Stanford University, with an interdisciplinary, 40-author collaboration spanning law schools, computer science departments and legal practice at institutions including the University of Chicago, Harvard Law School, Georgetown University Law Center and others |
LegalBench tests whether a model can perform the kinds of reasoning lawyers actually use, rather than treating "legal reasoning" as one undifferentiated skill. Its 162 tasks are organised into six categories drawn from how the legal profession itself frames reasoning: issue-spotting (does a fact pattern raise a given legal question), rule-recall (state or identify the applicable rule), rule-application (explain how a rule applies to facts, with reasoning graded for correctness and analysis), rule-conclusion (state the resulting legal outcome), interpretation (parse a contract, privacy policy or statute), and rhetorical-understanding (reason about legal argument and judicial writing). Inputs are short clauses, fact patterns, questions or case excerpts; most tasks are single-turn, English-only, and text-only.
Varies by task: multiple-choice (35 tasks), open-ended generation (7), binary classification (112), and multi-class or multi-label classification (8). Few-shot prompting is standard, using 0 to 8 in-context demonstrations drawn from each task's own small training split.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Mistral Medium (latest) | Mistral AI | 89.4 | 2026-04 |
| Mistral Medium 3 | Mistral AI | 89.4 | 2026-04 |
| Mistral Medium 3.1 | Mistral AI | 89.4 | 2026-04 |
| Gemini 3.1 Pro Preview | Google DeepMind | 87.4 | 2026-04 |
| Gemini 3 Pro Preview | Google DeepMind | 87.0 | 2026-04 |
| Gemini 3 Flash Preview | Google DeepMind | 86.9 | 2026-04 |
| GPT-5 | OpenAI | 86.0 | 2026-04 |
| Claude Opus 4 | Anthropic | 75.5 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 75.5 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 73.8 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 73.8 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 73.8 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 73.8 | 2026-04 |
| GPT-4o | OpenAI | 72.1 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 72.1 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 72.1 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 72.1 | 2026-04 |
| GPT-4o mini | OpenAI | 72.1 | 2026-04 |
| Llama 3.1 70B | Meta | 58.5 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 58.5 | 2026-04 |