Artificial Analysis independently scores OpenAI's 220-task GDPval gold set with pairwise Elo, shown as clamp((Elo-500)/2000).
active
Recorded reasons:
| Category | agentic |
|---|---|
| Subcategory | independent AA re-score of OpenAI GDPval gold tasks |
| Page status | active |
| Metric | normalized Elo ((Elo - 500) / 2000) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 220 |
| Publisher | Artificial Analysis |
GDPval-AA is Artificial Analysis's agentic run of OpenAI's public GDPval gold set. The model gets a professional task plus reference files and must write deliverable files (documents, slides, spreadsheets, diagrams, and similar). Coverage is 220 tasks across 44 occupations in nine U.S. GDP sectors. AA scores quality with blinded pairwise Elo against other models and human expert deliverables, not OpenAI's expert win rate or auto-grader. Current Index identity is GDPval-AA v2.
One Stirrup agent run per task in a fresh E2B sandbox (250-turn cap, one repeat). Tools: web fetch, Brave web search, optional view-image, bash code_exec, finish, and abandon_task. The model submits file paths; there is no live user in the loop.
Each row was checked against its source by a reviewer.
| Model | Score | Evidence date | Source kind | Link |
|---|---|---|---|---|
| GPT-6 Astra (max) | 54.0 normalized Elo percent | 2026-09-04 published | independent_evaluator | source |
| GLM-5.3 (max) | 59.0 normalized Elo percent | 2026-09-04 published | independent_evaluator | source |
No model card in ModelSpec reports this benchmark yet.