Artificial Analysis Intelligence Index

Artificial Analysis's own composite capability score, blending ten independently-run evaluations across agents, coding, general knowledge and scientific reasoning into one weighted number.

Also known as: Artificial Analysis Quality Index, AA Intelligence Index, AAII

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategorycomposite capability / intelligence score
Page statusactive
MetricIntelligence Index score
Directionhigher_is_better
Unitpoints
PublisherArtificial Analysis

What it measures

The Artificial Analysis Intelligence Index — this repository's artificial_analysis_quality_index — is Artificial Analysis's composite capability score, built, in the publisher's own words, to give "a single score for tracking progress toward artificial general intelligence across mathematics, science, coding, and reasoning." It is not one test: version 4.3 blends ten independently-run evaluations across four weighted categories — Agents (30%: AA-Briefcase, GDPval-AA v2, AutomationBench-AA), Coding (20%: Terminal-Bench v4.0, SciCode), General (30%: AA-Omniscience, GDP.pdf, AA-LCR v1.1) and Scientific Reasoning (20%: Humanity's Last Exam, CritPt) — into one weighted number per model. The publisher's live site uses "Intelligence Index" as the current name for this score; this page documents it under this repository's id, which predates that branding.

Task format

A weighted blend of ten independently-scored evaluation tasks spanning agentic file/tool tasks, terminal-based coding tasks, open-answer knowledge and reasoning questions, and long-document and physics-reasoning problems; each component evaluation keeps its own response format and scoring rule before being combined.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Opus 4.6Anthropic88.02026-04
GPT-5.1OpenAI88.02026-04
GPT-5.1 ChatOpenAI88.02026-04
GPT-5.1 CodexOpenAI88.02026-04
GPT-5.1 Codex MaxOpenAI88.02026-04
GPT-5.1 Codex miniOpenAI88.02026-04
Claude Sonnet 4.5Anthropic86.02026-04
Claude Sonnet 4.5 (latest)Anthropic86.02026-04
Claude Sonnet 4Anthropic85.02026-04
Claude Sonnet 4 (latest)Anthropic85.02026-04
Claude Sonnet 4.6Anthropic85.02026-04
Gemini 2.5 ProGoogle DeepMind85.02026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind85.02026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind85.02026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind85.02026-04
DeepSeek R1DeepSeek84.02026-04
DeepSeek R1 0528DeepSeek84.02026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA84.02026-04
DeepSeek R1 Distill Llama 70BDeepSeek84.02026-04
DeepSeek R1 Distill Llama 8BDeepSeek842026-04
DeepSeek R1 Distill Qwen 14BDeepSeek84.02026-04
DeepSeek R1 Distill Qwen 32BDeepSeek84.02026-04
DeepSeek R1 Distill Qwen 7BDeepSeek842026-04
GPT-4.1OpenAI84.02026-04
GPT-4oOpenAI82.02026-04
GPT-4o (2024-05-13)OpenAI82.02026-04
GPT-4o (2024-08-06)OpenAI82.02026-04
GPT-4o (2024-11-20)OpenAI82.02026-04
Grok 3xAI82.02026-04
Grok 3 FastxAI82.02026-04
Grok 3 Fast LatestxAI82.02026-04
Grok 3 LatestxAI82.02026-04
Qwen3 235B-A22BAlibaba / Qwen Team82.02026-04
DeepSeek V3DeepSeek80.02026-04
DeepSeek V3 0324DeepSeek80.02026-04
DeepSeek V3.1DeepSeek80.02026-04
DeepSeek V3.2DeepSeek80.02026-04
DeepSeek V3.2 ExpDeepSeek80.02026-04
Gemini 2.5 FlashGoogle DeepMind80.02026-04
Gemini 2.5 Flash ImageGoogle DeepMind80.02026-04
Gemini 2.5 Flash Image (Preview)Google DeepMind80.02026-04
Gemini 2.5 Flash LiteGoogle DeepMind80.02026-04
Gemini 2.5 Flash Lite Preview 06-17Google DeepMind80.02026-04
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind80.02026-04
Gemini 2.5 Flash Preview 04-17Google DeepMind80.02026-04
Gemini 2.5 Flash Preview 05-20Google DeepMind80.02026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind80.02026-04
Gemini 2.5 Flash Preview TTSGoogle DeepMind80.02026-04
Llama 3.1 405BMeta80.02026-04
Llama 3.1 405B FP8Meta80.02026-04
Llama 3.1 405B InstructMeta80.02026-04
Llama 3.1 405B Instruct FP8Meta80.02026-04
GPT-4 TurboOpenAI79.02026-04
Command ACohere782026-04
Command A ReasoningCohere782026-04
Gemini 1.5 ProGoogle DeepMind78.02026-04
Mistral Large (latest)Mistral AI78.02026-04
Mistral Large 2.1Mistral AI78.02026-04
Mistral Large 3Mistral AI78.02026-04
Gemini 2.0 FlashGoogle DeepMind77.02026-04
Grok 2xAI762026-04
Grok 2 (1212)xAI762026-04
Grok 2 LatestxAI762026-04
Llama 3.3 70B Instruct NVFP4NVIDIA76.02026-04
Llama-3.3-70B-InstructMeta76.02026-04
Qwen3 32BAlibaba / Qwen Team76.02026-04
Qwen3 32B AWQAlibaba / Qwen Team76.02026-04
Qwen3 32B NVFP4NVIDIA76.02026-04
Llama 3.1 70BMeta75.02026-04
Llama 3.1 70B InstructMeta75.02026-04
GPT-4.1 miniOpenAI74.02026-04
Qwen3 30B A3B Instruct 2507Alibaba / Qwen Team74.02026-04
Qwen3 30B A3B NVFP4NVIDIA74.02026-04
Qwen3 30B-A3BAlibaba / Qwen Team74.02026-04
Command R+Cohere73.02026-04
Gemma 4 31BGoogle DeepMind73.02026-04
gemma 4 31B itGoogle DeepMind73.02026-04
gemma 4 31B it GGUFUnsloth73.02026-04
Gemma 4 31B IT NVFP4NVIDIA73.02026-04
GPT-4o miniOpenAI72.02026-04
Grok 3 MinixAI72.02026-04
Grok 3 Mini FastxAI72.02026-04
Grok 3 Mini Fast LatestxAI72.02026-04
Grok 3 Mini LatestxAI72.02026-04
Gemini 1.5 FlashGoogle DeepMind71.02026-04
Gemini 1.5 Flash-8BGoogle DeepMind71.02026-04
Gemma 3 27BGoogle DeepMind70.02026-04
Qwen3 14BAlibaba / Qwen Team702026-04
Qwen3 14B AWQAlibaba / Qwen Team702026-04
Qwen3 14B NVFP4NVIDIA702026-04
Gemini 2.0 Flash LiteGoogle DeepMind68.02026-04
gemma 2 27B itGoogle DeepMind68.02026-04
Mistral Small (latest)Mistral AI68.02026-04
Mistral Small 24B Instruct 2501Mistral AI68.02026-04
Mistral Small 3.1 24B Instruct 2503Mistral AI68.02026-04
Mistral Small 3.2Mistral AI68.02026-04
Mistral Small 3.2 24B Instruct 2506Mistral AI68.02026-04
Mistral Small 4Mistral AI68.02026-04
Mistral Small 4 119B 2603Mistral AI68.02026-04
phi 4Microsoft68.02026-04
Phi 4 mini instructMicrosoft68.02026-04
Qwen3 8BAlibaba / Qwen Team662026-04
Qwen3 8B AWQAlibaba / Qwen Team662026-04
Qwen3 8B BaseAlibaba / Qwen Team662026-04
Skywork Reward V2 Qwen3 8BSkywork66undated
Gemma 3 12BGoogle DeepMind65.02026-04
GPT-4.1 nanoOpenAI65.02026-04
Mistral NemoMistral AI652026-04
Mistral Nemo Base 2407Mistral AI652026-04
Mistral Nemo Instruct 2407Mistral AI652026-04
Command RCohere642026-04
Arctic LSTM Speculator Llama 3.1 8B InstructSnowflake62undated
Hermes 3 Llama 3.1 8BNous Research62.02025-03
Hermes 3 Llama 3.1 8B GGUFNous Research62.02025-03
Llama 3.1 8BMeta622026-04
Llama 3.1 8B InstructMeta622026-04
Llama 3.1 8B InstructUnsloth622026-04
Llama 3.1 8B Instruct FP8NVIDIA622026-04
Llama 3.1 8B Instruct NVFP4NVIDIA622026-04
Meta Llama 3.1 8BNous Research62undated
Meta Llama 3.1 8B InstructNous Research62undated
Meta Llama 3.1 8B InstructUnsloth62undated
Meta Llama 3.1 8B Instruct bnb 4bitUnsloth62undated
Skywork Reward Llama 3.1 8B v0.2Skywork62undated
Skywork Reward V2 Llama 3.1 8BSkywork62undated
gemma 2 9BGoogle DeepMind60.02026-04
gemma 2 9B itGoogle DeepMind60.02026-04
GPT-3.5-turboOpenAI60.02026-04
Phi 3.5 mini instructMicrosoft58.02026-04
Muse SparkMeta52.02026-04

Data

This page as JSON · Edit on GitHub