An independent benchmarking company that runs its own model-capability and inference-performance tests and publishes them as composite indices and live leaderboards.
unassessed
| Category | composite |
|---|---|
| Subcategory | independent AI benchmarking and inference-performance tracking |
| Page status | active |
| Direction | higher_is_better |
| Publisher | Artificial Analysis |
Artificial Analysis independently benchmarks language models and the inference providers that serve them, across two broad axes: capability, aggregated from many third-party and proprietary evaluation datasets into the Artificial Analysis Intelligence Index and a set of profession-specific Capability Indices, and inference performance (output speed, latency and price), measured by Artificial Analysis's own repeated live calls to public model endpoints. A narrower Openness Index scores how open a model's weights, licence and documentation are; it has no page in this repository yet. Unlike a human-preference platform such as Arena, Artificial Analysis's scores come from automated test suites it runs itself, graded by fixed rubrics, pass@1 checkers or LLM judges, not from public votes.
Varies by product: automated pass@1, rubric or Elo-judge scoring across a suite of capability evaluations for the Intelligence Index and Capability Indices; live, repeated API calls against model endpoints measuring output tokens per second, time to first token and price for performance benchmarking.
No model card in ModelSpec reports this benchmark yet.