AA-LCR (Artificial Analysis Long Context Reasoning)

Artificial Analysis's 100-question test of whether a model can reason across ~100k-token real-world document sets, not just retrieve a stated fact.

Also known as: AA-LCR, Artificial Analysis Long Context Reasoning

active

This benchmark is in the default catalogue: its identity, protocol, current model coverage and dated results were verified by a reviewer who opened the sources.

Recorded reasons:

Categorylong-context
Subcategorymulti-document long-context reasoning
Page statusactive
Metricpass@1 accuracy, graded by an LLM equality checker (GPT-5.6 Luna, medium, as of v1.1)
Directionhigher_is_better
Unit%
Dataset size100
Dataset licenceApache-2.0
PublisherArtificial Analysis

What it measures

AA-LCR measures long-document comprehension: whether a model can extract, reason about and synthesise information spread across long, real-world documents rather than retrieve a single stated fact. Each of its 100 questions is paired with a document set of about 100,000 tokens (cl100k_base) drawn from seven categories: company reports, industry reports, government consultations, academic papers, legal documents, marketing materials and survey reports. Questions are written so the answer cannot be read from one passage and must be assembled from information dispersed across the set. Artificial Analysis requires a 128K context window to score. It frames the task as an under-studied class of evaluation where, at introduction, humans still clearly outscored language models.

Task format

A question plus a set of long real-world documents (~100k tokens, cl100k_base) in; a free-text answer out, graded by an LLM equality checker rather than exact string match.

Verified results

Each row was checked against its source by a reviewer.

ModelScoreEvidence dateSource kindLink
GPT-6 Astra (max)81.0%2026-09-04 publishedindependent_evaluatorsource
GLM-5.3 (max)80.0%2026-09-04 publishedindependent_evaluatorsource

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub