GDPval-AA

Artificial Analysis independently scores OpenAI's 220-task GDPval gold set with pairwise Elo, shown as clamp((Elo-500)/2000).

Also known as: GDPval-AA v2, AA GDPval, GDPval AA

active

This benchmark is in the default catalogue: its identity, protocol, current model coverage and dated results were verified by a reviewer who opened the sources.

Recorded reasons:

Categoryagentic
Subcategoryindependent AA re-score of OpenAI GDPval gold tasks
Page statusactive
Metricnormalized Elo ((Elo - 500) / 2000)
Directionhigher_is_better
Unit%
Dataset size220
PublisherArtificial Analysis

What it measures

GDPval-AA is Artificial Analysis's agentic run of OpenAI's public GDPval gold set. The model gets a professional task plus reference files and must write deliverable files (documents, slides, spreadsheets, diagrams, and similar). Coverage is 220 tasks across 44 occupations in nine U.S. GDP sectors. AA scores quality with blinded pairwise Elo against other models and human expert deliverables, not OpenAI's expert win rate or auto-grader. Current Index identity is GDPval-AA v2.

Task format

One Stirrup agent run per task in a fresh E2B sandbox (250-turn cap, one repeat). Tools: web fetch, Brave web search, optional view-image, bash code_exec, finish, and abandon_task. The model submits file paths; there is no live user in the loop.

Verified results

Each row was checked against its source by a reviewer.

ModelScoreEvidence dateSource kindLink
GPT-6 Astra (max)54.0 normalized Elo percent2026-09-04 publishedindependent_evaluatorsource
GLM-5.3 (max)59.0 normalized Elo percent2026-09-04 publishedindependent_evaluatorsource

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub