Artificial Analysis's 100-question test of whether a model can reason across ~100k-token real-world document sets, not just retrieve a stated fact.
active
Recorded reasons:
| Category | long-context |
|---|---|
| Subcategory | multi-document long-context reasoning |
| Page status | active |
| Metric | pass@1 accuracy, graded by an LLM equality checker (GPT-5.6 Luna, medium, as of v1.1) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 100 |
| Dataset licence | Apache-2.0 |
| Publisher | Artificial Analysis |
AA-LCR measures long-document comprehension: whether a model can extract, reason about and synthesise information spread across long, real-world documents rather than retrieve a single stated fact. Each of its 100 questions is paired with a document set of about 100,000 tokens (cl100k_base) drawn from seven categories: company reports, industry reports, government consultations, academic papers, legal documents, marketing materials and survey reports. Questions are written so the answer cannot be read from one passage and must be assembled from information dispersed across the set. Artificial Analysis requires a 128K context window to score. It frames the task as an under-studied class of evaluation where, at introduction, humans still clearly outscored language models.
A question plus a set of long real-world documents (~100k tokens, cl100k_base) in; a free-text answer out, graded by an LLM equality checker rather than exact string match.
Each row was checked against its source by a reviewer.
| Model | Score | Evidence date | Source kind | Link |
|---|---|---|---|---|
| GPT-6 Astra (max) | 81.0% | 2026-09-04 published | independent_evaluator | source |
| GLM-5.3 (max) | 80.0% | 2026-09-04 published | independent_evaluator | source |
No model card in ModelSpec reports this benchmark yet.