Artificial Analysis Output Speed

Artificial Analysis's live-measured output speed for a model's API: tokens generated per second, a performance measure, not a correctness or quality score.

Also known as: Artificial Analysis Speed Index, AA Output Speed, Output Tokens per Second

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategoryinference throughput / output speed
Page statusactive
MetricOutput Speed (output tokens/second)
Directionhigher_is_better
Unittokens/second
PublisherArtificial Analysis

What it measures

This page documents Artificial Analysis's Output Speed metric — the throughput figure the publisher headlines under "Speed" on its own site — as this repository's artificial_analysis_speed_index. It is a pure performance measurement, not a capability or correctness score: Artificial Analysis sends a live prompt to a model's public API and times how many tokens per second stream back after the first token arrives. Unlike the Intelligence Index, Artificial Analysis does not publish a single blended "Speed Index" combining multiple timing metrics into one number; Output Speed is reported alongside, not merged with, separate metrics for time to first token and end-to-end response time. Workloads vary by input length (about 1,000, 10,000 or 100,000 input tokens, plus a vision workload of one megapixel image and roughly 1,000 text tokens) because both time-to-first-token and output speed itself shift with prompt length and technique such as speculative decoding.

Task format

Live API calls: Artificial Analysis sends standardized, freshly generated prompts of a fixed input-token length to a model's public endpoint and streams the response, timing token arrivals to derive tokens-per-second and latency figures.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
GPT-4.1 nanoOpenAI94.02026-04
Gemini 2.0 Flash LiteGoogle DeepMind93.02026-04
Arctic LSTM Speculator Llama 3.1 8B InstructSnowflake92undated
GPT-3.5-turboOpenAI92.02026-04
Hermes 3 Llama 3.1 8BNous Research92.02025-03
Hermes 3 Llama 3.1 8B GGUFNous Research92.02025-03
Llama 3.1 8BMeta922026-04
Llama 3.1 8B InstructMeta922026-04
Llama 3.1 8B InstructUnsloth922026-04
Llama 3.1 8B Instruct FP8NVIDIA922026-04
Llama 3.1 8B Instruct NVFP4NVIDIA922026-04
Meta Llama 3.1 8BNous Research92undated
Meta Llama 3.1 8B InstructNous Research92undated
Meta Llama 3.1 8B InstructUnsloth92undated
Meta Llama 3.1 8B Instruct bnb 4bitUnsloth92undated
Phi 3.5 mini instructMicrosoft92.02026-04
Skywork Reward Llama 3.1 8B v0.2Skywork92undated
Skywork Reward V2 Llama 3.1 8BSkywork92undated
gemma 2 9BGoogle DeepMind90.02026-04
gemma 2 9B itGoogle DeepMind90.02026-04
GPT-4o miniOpenAI90.02026-04
Gemini 2.0 FlashGoogle DeepMind89.02026-04
GPT-4.1 miniOpenAI89.02026-04
Gemma 3 12BGoogle DeepMind88.02026-04
phi 4Microsoft88.02026-04
Phi 4 mini instructMicrosoft88.02026-04
Qwen3 30B A3B Instruct 2507Alibaba / Qwen Team88.02026-04
Qwen3 30B A3B NVFP4NVIDIA88.02026-04
Qwen3 30B-A3BAlibaba / Qwen Team88.02026-04
Qwen3 8BAlibaba / Qwen Team882026-04
Qwen3 8B AWQAlibaba / Qwen Team882026-04
Qwen3 8B BaseAlibaba / Qwen Team882026-04
Skywork Reward V2 Qwen3 8BSkywork88undated
Gemini 2.5 FlashGoogle DeepMind87.02026-04
Gemini 2.5 Flash ImageGoogle DeepMind87.02026-04
Gemini 2.5 Flash Image (Preview)Google DeepMind87.02026-04
Gemini 2.5 Flash LiteGoogle DeepMind87.02026-04
Gemini 2.5 Flash Lite Preview 06-17Google DeepMind87.02026-04
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind87.02026-04
Gemini 2.5 Flash Preview 04-17Google DeepMind87.02026-04
Gemini 2.5 Flash Preview 05-20Google DeepMind87.02026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind87.02026-04
Gemini 2.5 Flash Preview TTSGoogle DeepMind87.02026-04
Gemini 1.5 FlashGoogle DeepMind86.02026-04
Gemini 1.5 Flash-8BGoogle DeepMind86.02026-04
Mistral NemoMistral AI852026-04
Mistral Nemo Base 2407Mistral AI852026-04
Mistral Nemo Instruct 2407Mistral AI852026-04
Gemma 4 31BGoogle DeepMind84.02026-04
gemma 4 31B itGoogle DeepMind84.02026-04
gemma 4 31B it GGUFUnsloth84.02026-04
Gemma 4 31B IT NVFP4NVIDIA84.02026-04
Gemma 3 27BGoogle DeepMind82.02026-04
Grok 3 MinixAI82.02026-04
Grok 3 Mini FastxAI82.02026-04
Grok 3 Mini Fast LatestxAI82.02026-04
Grok 3 Mini LatestxAI82.02026-04
Mistral Small (latest)Mistral AI82.02026-04
Mistral Small 24B Instruct 2501Mistral AI82.02026-04
Mistral Small 3.1 24B Instruct 2503Mistral AI82.02026-04
Mistral Small 3.2Mistral AI82.02026-04
Mistral Small 3.2 24B Instruct 2506Mistral AI82.02026-04
Mistral Small 4Mistral AI82.02026-04
Mistral Small 4 119B 2603Mistral AI82.02026-04
gemma 2 27B itGoogle DeepMind80.02026-04
Qwen3 14BAlibaba / Qwen Team802026-04
Qwen3 14B AWQAlibaba / Qwen Team802026-04
Qwen3 14B NVFP4NVIDIA802026-04
Command RCohere782026-04
GPT-4oOpenAI78.02026-04
GPT-4o (2024-05-13)OpenAI78.02026-04
GPT-4o (2024-08-06)OpenAI78.02026-04
GPT-4o (2024-11-20)OpenAI78.02026-04
GPT-4.1OpenAI76.02026-04
Claude Sonnet 4Anthropic73.02026-04
Claude Sonnet 4 (latest)Anthropic73.02026-04
Claude Sonnet 4.6Anthropic73.02026-04
Llama 3.3 70B Instruct NVFP4NVIDIA73.02026-04
Llama-3.3-70B-InstructMeta73.02026-04
Llama 3.1 70BMeta72.02026-04
Llama 3.1 70B InstructMeta72.02026-04
Claude Sonnet 4.5Anthropic71.02026-04
Claude Sonnet 4.5 (latest)Anthropic71.02026-04
Gemini 2.5 ProGoogle DeepMind71.02026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind71.02026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind71.02026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind71.02026-04
DeepSeek V3DeepSeek70.02026-04
DeepSeek V3 0324DeepSeek70.02026-04
DeepSeek V3.1DeepSeek70.02026-04
DeepSeek V3.2DeepSeek70.02026-04
DeepSeek V3.2 ExpDeepSeek70.02026-04
GPT-5.1OpenAI70.02026-04
GPT-5.1 ChatOpenAI70.02026-04
GPT-5.1 CodexOpenAI70.02026-04
GPT-5.1 Codex MaxOpenAI70.02026-04
GPT-5.1 Codex miniOpenAI70.02026-04
Grok 2xAI702026-04
Grok 2 (1212)xAI702026-04
Grok 2 LatestxAI702026-04
Qwen3 32BAlibaba / Qwen Team70.02026-04
Qwen3 32B AWQAlibaba / Qwen Team70.02026-04
Qwen3 32B NVFP4NVIDIA70.02026-04
Gemini 1.5 ProGoogle DeepMind68.02026-04
GPT-4 TurboOpenAI65.02026-04
Grok 3xAI65.02026-04
Grok 3 FastxAI65.02026-04
Grok 3 Fast LatestxAI65.02026-04
Grok 3 LatestxAI65.02026-04
Claude Opus 4.6Anthropic62.02026-04
Mistral Large (latest)Mistral AI60.02026-04
Mistral Large 2.1Mistral AI60.02026-04
Mistral Large 3Mistral AI60.02026-04
Command ACohere582026-04
Command A ReasoningCohere582026-04
Command R+Cohere55.02026-04
Qwen3 235B-A22BAlibaba / Qwen Team55.02026-04
Llama 3.1 405BMeta45.02026-04
Llama 3.1 405B FP8Meta45.02026-04
Llama 3.1 405B InstructMeta45.02026-04
Llama 3.1 405B Instruct FP8Meta45.02026-04
DeepSeek R1DeepSeek40.02026-04
DeepSeek R1 0528DeepSeek40.02026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA40.02026-04
DeepSeek R1 Distill Llama 70BDeepSeek40.02026-04
DeepSeek R1 Distill Llama 8BDeepSeek402026-04
DeepSeek R1 Distill Qwen 14BDeepSeek40.02026-04
DeepSeek R1 Distill Qwen 32BDeepSeek40.02026-04
DeepSeek R1 Distill Qwen 7BDeepSeek402026-04

Data

This page as JSON · Edit on GitHub