Artificial Analysis

An independent benchmarking company that runs its own model-capability and inference-performance tests and publishes them as composite indices and live leaderboards.

Also known as: AA

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategoryindependent AI benchmarking and inference-performance tracking
Page statusactive
Directionhigher_is_better
PublisherArtificial Analysis

What it measures

Artificial Analysis independently benchmarks language models and the inference providers that serve them, across two broad axes: capability, aggregated from many third-party and proprietary evaluation datasets into the Artificial Analysis Intelligence Index and a set of profession-specific Capability Indices, and inference performance (output speed, latency and price), measured by Artificial Analysis's own repeated live calls to public model endpoints. A narrower Openness Index scores how open a model's weights, licence and documentation are; it has no page in this repository yet. Unlike a human-preference platform such as Arena, Artificial Analysis's scores come from automated test suites it runs itself, graded by fixed rubrics, pass@1 checkers or LLM judges, not from public votes.

Task format

Varies by product: automated pass@1, rubric or Elo-judge scoring across a suite of capability evaluations for the Intelligence Index and Capability Indices; live, repeated API calls against model endpoints measuring output tokens per second, time to first token and price for performance benchmarking.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub