A HELM scenario built from AR-LSAT: 2,091 five-option multiple-choice logic-puzzle questions from real 1991-2016 LSAT analytical-reasoning ("logic games") sections, testing constraint satisfaction.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | analytical reasoning (logic games) from the Law School Admission Test |
| Page status | active |
| Metric | quasi_exact_match (HELM's exact-match family) on the selected option letter |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2091 |
| Dataset licence | MIT |
| Publisher | Sun Yat-sen University (School of Data and Computer Science); Microsoft Research |
This benchmark tests analytical reasoning, not legal knowledge: it is built entirely from the Analytical Reasoning ("logic games") section of the real Law School Admission Test, given a passage describing a set of elements and constraints -- for example assigning speakers to dates, grouping students into teams, or ordering classes in a schedule -- and asking the model to answer a question by working out which arrangement satisfies every stated condition. No legal terminology or domain knowledge is required to solve a question; the passages are logic puzzles that happen to be drawn from a law-school admissions exam rather than legal-reasoning exercises, which distinguishes this benchmark sharply from this repository's `legalbench` and `lawbench` pages, which test actual legal knowledge and argumentation. The underlying AR-LSAT dataset groups its questions into four types: grouping (in/out grouping, distribution grouping), ordering (simple, relative, complex ordering), assignment (determined, undetermined assignment) and miscellaneous.
A passage describing a constraint-satisfaction scenario, followed by a question and five lettered answer options (A-E), exactly one of which is correct. HELM's `lsat_qa` scenario presents this as standard joint multiple-choice: "Passage: ... Question: ... A. ... E. ..." with the model expected to output the correct option letter.
No model card in ModelSpec reports this benchmark yet.