GDP.pdf-AA

Artificial Analysis's independent GDP.pdf run: 100 professional PDF tasks, All-pass over 1,275 criteria, LiteParse text plus page images, GPT-5.6 Luna Medium judge.

Also known as: GDP.pdf AA, AA GDP.pdf, GDP.pdf (Artificial Analysis)

active

This benchmark is in the default catalogue: its identity, protocol, current model coverage and dated results were verified by a reviewer who opened the sources.

Recorded reasons:

Categorylong-context
Subcategoryindependent re-score of Surge AI GDP.pdf (professional PDF reasoning)
Page statusactive
MetricAll-pass rate
Directionhigher_is_better
Unit%
Dataset size100
PublisherArtificial Analysis

What it measures

GDP.pdf-AA is Artificial Analysis's implementation of Surge AI's GDP.pdf. Each item is one professional question plus one original PDF. The model must ground the answer in that file, including tables, charts, footnotes, legends, and later amendments. Coverage is 100 tasks in ten domains. AA delivers extracted page text to every model and page images only to models that accept image input. Answers are single-turn free text, with no tools and no browsing.

Task format

One PDF per task. AA prepares the file with LiteParse (OCR where needed). Every model gets full extracted text. Image-capable models also get ordered page images (150 DPI, reduced as far as 72 DPI under payload limits; composites of two or four pages when an endpoint caps image count). Five independent attempts. GPT-5.6 Luna Medium grades each atomic criterion without the source PDF or the contestant name.

Verified results

Each row was checked against its source by a reviewer.

ModelScoreEvidence dateSource kindLink
GPT-6 Astra (max)33.0%2026-09-04 publishedindependent_evaluatorsource
GLM-5.3 (max)12.0%2026-09-04 publishedindependent_evaluatorsource

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub