Artificial Analysis's own composite capability score, blending ten independently-run evaluations across agents, coding, general knowledge and scientific reasoning into one weighted number.
unassessed
| Category | composite |
|---|---|
| Subcategory | composite capability / intelligence score |
| Page status | active |
| Metric | Intelligence Index score |
| Direction | higher_is_better |
| Unit | points |
| Publisher | Artificial Analysis |
The Artificial Analysis Intelligence Index — this repository's artificial_analysis_quality_index — is Artificial Analysis's composite capability score, built, in the publisher's own words, to give "a single score for tracking progress toward artificial general intelligence across mathematics, science, coding, and reasoning." It is not one test: version 4.3 blends ten independently-run evaluations across four weighted categories — Agents (30%: AA-Briefcase, GDPval-AA v2, AutomationBench-AA), Coding (20%: Terminal-Bench v4.0, SciCode), General (30%: AA-Omniscience, GDP.pdf, AA-LCR v1.1) and Scientific Reasoning (20%: Humanity's Last Exam, CritPt) — into one weighted number per model. The publisher's live site uses "Intelligence Index" as the current name for this score; this page documents it under this repository's id, which predates that branding.
A weighted blend of ten independently-scored evaluation tasks spanning agentic file/tool tasks, terminal-based coding tasks, open-answer knowledge and reasoning questions, and long-document and physics-reasoning problems; each component evaluation keeps its own response format and scoring rule before being combined.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Opus 4.6 | Anthropic | 88.0 | 2026-04 |
| GPT-5.1 | OpenAI | 88.0 | 2026-04 |
| GPT-5.1 Chat | OpenAI | 88.0 | 2026-04 |
| GPT-5.1 Codex | OpenAI | 88.0 | 2026-04 |
| GPT-5.1 Codex Max | OpenAI | 88.0 | 2026-04 |
| GPT-5.1 Codex mini | OpenAI | 88.0 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 86.0 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 86.0 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 85.0 | 2026-04 |
| Claude Sonnet 4 (latest) | Anthropic | 85.0 | 2026-04 |
| Claude Sonnet 4.6 | Anthropic | 85.0 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 85.0 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 85.0 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 85.0 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 85.0 | 2026-04 |
| DeepSeek R1 | DeepSeek | 84.0 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 84.0 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 84.0 | 2026-04 |
| DeepSeek R1 Distill Llama 70B | DeepSeek | 84.0 | 2026-04 |
| DeepSeek R1 Distill Llama 8B | DeepSeek | 84 | 2026-04 |
| DeepSeek R1 Distill Qwen 14B | DeepSeek | 84.0 | 2026-04 |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | 84.0 | 2026-04 |
| DeepSeek R1 Distill Qwen 7B | DeepSeek | 84 | 2026-04 |
| GPT-4.1 | OpenAI | 84.0 | 2026-04 |
| GPT-4o | OpenAI | 82.0 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 82.0 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 82.0 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 82.0 | 2026-04 |
| Grok 3 | xAI | 82.0 | 2026-04 |
| Grok 3 Fast | xAI | 82.0 | 2026-04 |
| Grok 3 Fast Latest | xAI | 82.0 | 2026-04 |
| Grok 3 Latest | xAI | 82.0 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 82.0 | 2026-04 |
| DeepSeek V3 | DeepSeek | 80.0 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 80.0 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 80.0 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 80.0 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 80.0 | 2026-04 |
| Gemini 2.5 Flash | Google DeepMind | 80.0 | 2026-04 |
| Gemini 2.5 Flash Image | Google DeepMind | 80.0 | 2026-04 |
| Gemini 2.5 Flash Image (Preview) | Google DeepMind | 80.0 | 2026-04 |
| Gemini 2.5 Flash Lite | Google DeepMind | 80.0 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 06-17 | Google DeepMind | 80.0 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 09-25 | Google DeepMind | 80.0 | 2026-04 |
| Gemini 2.5 Flash Preview 04-17 | Google DeepMind | 80.0 | 2026-04 |
| Gemini 2.5 Flash Preview 05-20 | Google DeepMind | 80.0 | 2026-04 |
| Gemini 2.5 Flash Preview 09-25 | Google DeepMind | 80.0 | 2026-04 |
| Gemini 2.5 Flash Preview TTS | Google DeepMind | 80.0 | 2026-04 |
| Llama 3.1 405B | Meta | 80.0 | 2026-04 |
| Llama 3.1 405B FP8 | Meta | 80.0 | 2026-04 |
| Llama 3.1 405B Instruct | Meta | 80.0 | 2026-04 |
| Llama 3.1 405B Instruct FP8 | Meta | 80.0 | 2026-04 |
| GPT-4 Turbo | OpenAI | 79.0 | 2026-04 |
| Command A | Cohere | 78 | 2026-04 |
| Command A Reasoning | Cohere | 78 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 78.0 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 78.0 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 78.0 | 2026-04 |
| Mistral Large 3 | Mistral AI | 78.0 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 77.0 | 2026-04 |
| Grok 2 | xAI | 76 | 2026-04 |
| Grok 2 (1212) | xAI | 76 | 2026-04 |
| Grok 2 Latest | xAI | 76 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 76.0 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 76.0 | 2026-04 |
| Qwen3 32B | Alibaba / Qwen Team | 76.0 | 2026-04 |
| Qwen3 32B AWQ | Alibaba / Qwen Team | 76.0 | 2026-04 |
| Qwen3 32B NVFP4 | NVIDIA | 76.0 | 2026-04 |
| Llama 3.1 70B | Meta | 75.0 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 75.0 | 2026-04 |
| GPT-4.1 mini | OpenAI | 74.0 | 2026-04 |
| Qwen3 30B A3B Instruct 2507 | Alibaba / Qwen Team | 74.0 | 2026-04 |
| Qwen3 30B A3B NVFP4 | NVIDIA | 74.0 | 2026-04 |
| Qwen3 30B-A3B | Alibaba / Qwen Team | 74.0 | 2026-04 |
| Command R+ | Cohere | 73.0 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 73.0 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 73.0 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 73.0 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 73.0 | 2026-04 |
| GPT-4o mini | OpenAI | 72.0 | 2026-04 |
| Grok 3 Mini | xAI | 72.0 | 2026-04 |
| Grok 3 Mini Fast | xAI | 72.0 | 2026-04 |
| Grok 3 Mini Fast Latest | xAI | 72.0 | 2026-04 |
| Grok 3 Mini Latest | xAI | 72.0 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 71.0 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 71.0 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 70.0 | 2026-04 |
| Qwen3 14B | Alibaba / Qwen Team | 70 | 2026-04 |
| Qwen3 14B AWQ | Alibaba / Qwen Team | 70 | 2026-04 |
| Qwen3 14B NVFP4 | NVIDIA | 70 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 68.0 | 2026-04 |
| gemma 2 27B it | Google DeepMind | 68.0 | 2026-04 |
| Mistral Small (latest) | Mistral AI | 68.0 | 2026-04 |
| Mistral Small 24B Instruct 2501 | Mistral AI | 68.0 | 2026-04 |
| Mistral Small 3.1 24B Instruct 2503 | Mistral AI | 68.0 | 2026-04 |
| Mistral Small 3.2 | Mistral AI | 68.0 | 2026-04 |
| Mistral Small 3.2 24B Instruct 2506 | Mistral AI | 68.0 | 2026-04 |
| Mistral Small 4 | Mistral AI | 68.0 | 2026-04 |
| Mistral Small 4 119B 2603 | Mistral AI | 68.0 | 2026-04 |
| phi 4 | Microsoft | 68.0 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 68.0 | 2026-04 |
| Qwen3 8B | Alibaba / Qwen Team | 66 | 2026-04 |
| Qwen3 8B AWQ | Alibaba / Qwen Team | 66 | 2026-04 |
| Qwen3 8B Base | Alibaba / Qwen Team | 66 | 2026-04 |
| Skywork Reward V2 Qwen3 8B | Skywork | 66 | undated |
| Gemma 3 12B | Google DeepMind | 65.0 | 2026-04 |
| GPT-4.1 nano | OpenAI | 65.0 | 2026-04 |
| Mistral Nemo | Mistral AI | 65 | 2026-04 |
| Mistral Nemo Base 2407 | Mistral AI | 65 | 2026-04 |
| Mistral Nemo Instruct 2407 | Mistral AI | 65 | 2026-04 |
| Command R | Cohere | 64 | 2026-04 |
| Arctic LSTM Speculator Llama 3.1 8B Instruct | Snowflake | 62 | undated |
| Hermes 3 Llama 3.1 8B | Nous Research | 62.0 | 2025-03 |
| Hermes 3 Llama 3.1 8B GGUF | Nous Research | 62.0 | 2025-03 |
| Llama 3.1 8B | Meta | 62 | 2026-04 |
| Llama 3.1 8B Instruct | Meta | 62 | 2026-04 |
| Llama 3.1 8B Instruct | Unsloth | 62 | 2026-04 |
| Llama 3.1 8B Instruct FP8 | NVIDIA | 62 | 2026-04 |
| Llama 3.1 8B Instruct NVFP4 | NVIDIA | 62 | 2026-04 |
| Meta Llama 3.1 8B | Nous Research | 62 | undated |
| Meta Llama 3.1 8B Instruct | Nous Research | 62 | undated |
| Meta Llama 3.1 8B Instruct | Unsloth | 62 | undated |
| Meta Llama 3.1 8B Instruct bnb 4bit | Unsloth | 62 | undated |
| Skywork Reward Llama 3.1 8B v0.2 | Skywork | 62 | undated |
| Skywork Reward V2 Llama 3.1 8B | Skywork | 62 | undated |
| gemma 2 9B | Google DeepMind | 60.0 | 2026-04 |
| gemma 2 9B it | Google DeepMind | 60.0 | 2026-04 |
| GPT-3.5-turbo | OpenAI | 60.0 | 2026-04 |
| Phi 3.5 mini instruct | Microsoft | 58.0 | 2026-04 |
| Muse Spark | Meta | 52.0 | 2026-04 |