{
 "body": "\n## What it measures\n\nArtificial Analysis is an independent benchmarking company, not a single benchmark: it runs its\nown standardized tests of language models and of the infrastructure that serves them, then\npublishes the results as composite indices and live leaderboards. Two axes anchor the family:\ncapability, aggregated into the Artificial Analysis Intelligence Index and a set of\nprofession-specific Capability Indices (Legal, Healthcare & Medical, Finance & Accounting, and\nothers), and inference performance \u2014 output speed, latency and price \u2014 measured by repeatedly\ncalling public model endpoints from Artificial Analysis's own accounts. A narrower Openness Index\nscores how open a model's weights, licence and documentation are (see Lineage). Coverage spans\nhundreds of models and providers and, for the Intelligence Index specifically, ten evaluation\ndatasets as of the version current in September 2026.\n\n## How it is scored\n\nCapability scoring is automated, not vote-based: each component evaluation in the Intelligence\nIndex runs independently under a fixed protocol (zero-shot, standardized prompts and temperature)\nand is scored by whatever method fits it \u2014 pass@1 on graded tasks, rubric or pairwise Elo judging\nby a panel of frontier-model judges for open-ended agentic work, or an \"equality checker\" LLM for\nfree-form answers \u2014 then combined into a weighted score across four categories: Agents, Coding,\nGeneral and Scientific Reasoning. Performance scoring instead sends live prompts to each model's\npublic API repeatedly through the day and reports the median of the trailing 72 hours for figures\nsuch as output tokens per second and time to first token. The domain-specific Capability Indices\nreuse several of the same component evaluations under different, profession-weighted blends.\n\n## Dataset and licence\n\nThere is no single Artificial Analysis dataset. The Intelligence Index alone currently blends ten\nevaluations \u2014 AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode,\nAA-Omniscience, GDP.pdf, AA-LCR v1.1, Humanity's Last Exam and CritPt \u2014 several built or licensed\nby Artificial Analysis itself as private test sets (AA-Briefcase, AA-Omniscience and\nAutomationBench-AA are marked \"Private Dataset\" on its evaluations page) and others adapted from\npublic academic benchmarks. Performance figures come from repeated, synthetic API calls rather\nthan a static dataset. No single public licence for Artificial Analysis's own aggregated scores or\nmethodology text was established from a source read for this page.\n\n## Who publishes it\n\nArtificial Analysis describes itself as \"the independent benchmarking company for AI,\" measuring\nmodels, the clouds that serve them and the chips they run on. It is headquartered in San\nFrancisco with an additional office in Melbourne. No individual founder or author names were\nestablished from a source read for this page; the company publishes under its own name and is\nreferenced by major AI labs, cloud providers and press outlets \u2014 including Microsoft, Meta,\nNVIDIA, OpenAI, Google, Bloomberg and the Financial Times \u2014 per its own about page.\n\n## Lineage\n\nArtificial Analysis's flagship capability score has moved through several major versions: v1.0\n(January 2024), v2.0 (February 2025), v3.0 (September 2025), and a v4.0 overhaul (January 2026)\nthat replaced most original components \u2014 dropping MMLU-Pro, LiveCodeBench and AIME 2025 \u2014 with\nharder, more agentic evaluations such as GDPval-AA and AA-Omniscience; v4.3, current at this\npage's research date, was announced 7 September 2026. In this repository,\n`artificial_analysis_quality_index` documents that capability score and\n`artificial_analysis_speed_index` documents the output-speed performance metric; the Openness\nIndex and the profession-specific Capability Indices are not yet paged here.\n\n## Saturation and contamination\n\nThe Intelligence Index is far from saturated on its current, harder suite: the top-scoring model\nreached 53 out of an implied 100 as of this page's research date, well clear of the ceiling.\nContamination risk is mixed and best treated as medium: several component evaluations run against\nprivate test sets built or licensed to resist public leakage, while others are established public\nacademic benchmarks that carry the usual risk of entering later training data. Performance\nbenchmarking has a different integrity concern \u2014 a provider serving faster or more accurate\nresults to Artificial Analysis's own traffic than to ordinary users \u2014 addressed with published\nIntegrity Terms and compliance checks rather than dataset controls.\n\n## How to run it\n\nArtificial Analysis's evaluations are not an open harness a third party runs independently; the\npublished methodology pages document prompts, scoring and grader models in enough detail to\nfollow the reasoning, but Artificial Analysis itself runs every official number. It uses e2b as\nits sandbox provider for agentic benchmarks and has open-sourced its own agent harness, Stirrup\n(github.com/ArtificialAnalysis/Stirrup). Performance figures come from Artificial Analysis's own\ntest accounts calling providers' public APIs, mostly with the official OpenAI Python library, from\na primary server in Google Cloud's `us-central1-a` zone.\n\n## Reading the numbers\n\nTreat an Artificial Analysis composite score as a synthesis, not a single fact: the Intelligence\nIndex blends ten differently-scored evaluations behind one weighted number, so two models with the\nsame score can have quite different strengths, and the score has moved with each methodology\nversion rather than sitting on a fixed scale over time. Read a capability score alongside a\nperformance figure such as output speed or price per task if latency or cost matters \u2014 Artificial\nAnalysis's own site pairs Intelligence Index against speed and cost for this reason \u2014 and check\nwhich index version a reported number used before comparing it to another source's figure.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "artificial_analysis",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [
   "AA"
  ],
  "category": "composite",
  "contamination": {
   "note": "Mixed by design: several component evaluations (AA-Briefcase, AA-Omniscience, AutomationBench-AA) are run against private test sets Artificial Analysis built or licensed specifically to resist public leakage, marked 'Private Dataset' on its own evaluations page, while others are established public academic benchmarks that carry the ordinary risk of entering later training data.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [],
   "license": "",
   "modalities": [
    "text",
    "image"
   ],
   "public_test_set": null,
   "size": null,
   "size_note": "No single dataset. The Intelligence Index alone blends 10 component evaluations as of the version current in September 2026; performance benchmarking uses freshly generated synthetic prompts rather than a fixed set; coverage spans hundreds of models and providers.",
   "splits": "",
   "url": "https://artificialanalysis.ai/evaluations"
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "sonnet-5 agent, batch 1b, slice J",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "No single harness. Artificial Analysis runs every official number itself, using e2b as its sandbox provider for agentic benchmarks and its own open-sourced agent harness, Stirrup (github.com/ArtificialAnalysis/Stirrup), for tool-using evaluations."
  },
  "id": "artificial_analysis",
  "last_updated": "",
  "leaderboard_url": "https://artificialanalysis.ai/",
  "lineage": {
   "family": "",
   "predecessor": "",
   "successors": [],
   "variants": [
    "artificial_analysis_quality_index",
    "artificial_analysis_speed_index"
   ]
  },
  "measures": "Artificial Analysis independently benchmarks language models and the inference providers that serve them, across two broad axes: capability, aggregated from many third-party and proprietary evaluation datasets into the Artificial Analysis Intelligence Index and a set of profession-specific Capability Indices, and inference performance (output speed, latency and price), measured by Artificial Analysis's own repeated live calls to public model endpoints. A narrower Openness Index scores how open a model's weights, licence and documentation are; it has no page in this repository yet. Unlike a human-preference platform such as Arena, Artificial Analysis's scores come from automated test suites it runs itself, graded by fixed rubrics, pass@1 checkers or LLM judges, not from public votes.\n",
  "metric": {
   "baseline_note": "No single metric spans the family: each product (Intelligence Index, a Capability Index, or a performance figure such as output speed) keeps its own scoring rule. See the subset pages.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": null,
   "name": "",
   "random_baseline": null,
   "unit": ""
  },
  "name": "Artificial Analysis",
  "page_kind": "family",
  "paper": {
   "arxiv": "",
   "title": "",
   "url": "",
   "year": null
  },
  "publisher": {
   "authors": [],
   "org": "Artificial Analysis",
   "url": "https://artificialanalysis.ai"
  },
  "released": "2024-01",
  "repo_url": "",
  "saturation": {
   "as_of": "2026-09",
   "note": "Using the Intelligence Index, the family's flagship capability score, as the representative figure: the top-scoring model reached 53 as of this page's research date, well clear of a 100-point ceiling.",
   "status": "open",
   "top_score": 53
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "About \u2014 Artificial Analysis",
    "url": "https://artificialanalysis.ai/about"
   },
   {
    "accessed": "2026-09-08",
    "title": "Evaluations overview \u2014 Artificial Analysis",
    "url": "https://artificialanalysis.ai/evaluations"
   },
   {
    "accessed": "2026-09-08",
    "title": "Artificial Analysis Intelligence Benchmarking Methodology",
    "url": "https://artificialanalysis.ai/methodology/intelligence-benchmarking"
   },
   {
    "accessed": "2026-09-08",
    "title": "Artificial Analysis Language Model API Performance Benchmarking Methodology",
    "url": "https://artificialanalysis.ai/methodology/performance-benchmarking"
   },
   {
    "accessed": "2026-09-08",
    "title": "Artificial Analysis (homepage)",
    "url": "https://artificialanalysis.ai/"
   }
  ],
  "status": "active",
  "subcategory": "independent AI benchmarking and inference-performance tracking",
  "summary": "An independent benchmarking company that runs its own model-capability and inference-performance tests and publishes them as composite indices and live leaderboards.",
  "tags": [
   "benchmarking-company",
   "composite-index",
   "inference-performance",
   "leaderboard-publisher"
  ],
  "task_format": "Varies by product: automated pass@1, rubric or Elo-judge scoring across a suite of capability evaluations for the Intelligence Index and Capability Indices; live, repeated API calls against model endpoints measuring output tokens per second, time to first token and price for performance benchmarking.\n"
 }
}