GAIA

466 real-world assistant questions needing tools and browsing; 300 test answers are held out and scoring is quasi-exact match.

Also known as: General AI Assistants, GAIA: a benchmark for General AI Assistants, gaia-benchmark/GAIA

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryreal-world assistant questions requiring tools, browsing, and files
Page statusactive
Metricaccuracy (quasi-exact match after type-specific normalisation)
Directionhigher_is_better
Unit%
Dataset size466
PublisherFAIR, Meta and Hugging Face (paper affiliations); AutoGPT and Meta GenAI also listed

What it measures

GAIA asks conceptually simple assistant questions that still need reasoning, web browsing, file reading, and other tools. Answers are a number, a short string, or a comma-separated list, so scoring can be automatic. The authors wrote 466 questions: 146 Level 1 (at most one tool, few steps), 245 Level 2 (roughly 5–10 steps, mixed tools), and 75 Level 3 (long action sequences). English questions; some items attach a spreadsheet, image, audio, or other file. Humans scored 92% overall in the paper. GPT-4 with plugins scored 15% in the abstract (Table 4: 30.3 / 9.7 / 0 by level).

Task format

Zero-shot assistant prompt, optional attached file. Inspect default solver is a ReAct agent with bash, Python, and a web browser inside Docker. Default split is validation (answers present). Test split has no scores in-harness; those answers go to the public leaderboard. Tasks: gaia (2023_all), gaia_level1, gaia_level2, gaia_level3.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub