Artificial Analysis's live-measured output speed for a model's API: tokens generated per second, a performance measure, not a correctness or quality score.
unassessed
| Category | composite |
|---|---|
| Subcategory | inference throughput / output speed |
| Page status | active |
| Metric | Output Speed (output tokens/second) |
| Direction | higher_is_better |
| Unit | tokens/second |
| Publisher | Artificial Analysis |
This page documents Artificial Analysis's Output Speed metric — the throughput figure the publisher headlines under "Speed" on its own site — as this repository's artificial_analysis_speed_index. It is a pure performance measurement, not a capability or correctness score: Artificial Analysis sends a live prompt to a model's public API and times how many tokens per second stream back after the first token arrives. Unlike the Intelligence Index, Artificial Analysis does not publish a single blended "Speed Index" combining multiple timing metrics into one number; Output Speed is reported alongside, not merged with, separate metrics for time to first token and end-to-end response time. Workloads vary by input length (about 1,000, 10,000 or 100,000 input tokens, plus a vision workload of one megapixel image and roughly 1,000 text tokens) because both time-to-first-token and output speed itself shift with prompt length and technique such as speculative decoding.
Live API calls: Artificial Analysis sends standardized, freshly generated prompts of a fixed input-token length to a model's public endpoint and streams the response, timing token arrivals to derive tokens-per-second and latency figures.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| GPT-4.1 nano | OpenAI | 94.0 | 2026-04 |
| Gemini 2.0 Flash Lite | Google DeepMind | 93.0 | 2026-04 |
| Arctic LSTM Speculator Llama 3.1 8B Instruct | Snowflake | 92 | undated |
| GPT-3.5-turbo | OpenAI | 92.0 | 2026-04 |
| Hermes 3 Llama 3.1 8B | Nous Research | 92.0 | 2025-03 |
| Hermes 3 Llama 3.1 8B GGUF | Nous Research | 92.0 | 2025-03 |
| Llama 3.1 8B | Meta | 92 | 2026-04 |
| Llama 3.1 8B Instruct | Meta | 92 | 2026-04 |
| Llama 3.1 8B Instruct | Unsloth | 92 | 2026-04 |
| Llama 3.1 8B Instruct FP8 | NVIDIA | 92 | 2026-04 |
| Llama 3.1 8B Instruct NVFP4 | NVIDIA | 92 | 2026-04 |
| Meta Llama 3.1 8B | Nous Research | 92 | undated |
| Meta Llama 3.1 8B Instruct | Nous Research | 92 | undated |
| Meta Llama 3.1 8B Instruct | Unsloth | 92 | undated |
| Meta Llama 3.1 8B Instruct bnb 4bit | Unsloth | 92 | undated |
| Phi 3.5 mini instruct | Microsoft | 92.0 | 2026-04 |
| Skywork Reward Llama 3.1 8B v0.2 | Skywork | 92 | undated |
| Skywork Reward V2 Llama 3.1 8B | Skywork | 92 | undated |
| gemma 2 9B | Google DeepMind | 90.0 | 2026-04 |
| gemma 2 9B it | Google DeepMind | 90.0 | 2026-04 |
| GPT-4o mini | OpenAI | 90.0 | 2026-04 |
| Gemini 2.0 Flash | Google DeepMind | 89.0 | 2026-04 |
| GPT-4.1 mini | OpenAI | 89.0 | 2026-04 |
| Gemma 3 12B | Google DeepMind | 88.0 | 2026-04 |
| phi 4 | Microsoft | 88.0 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 88.0 | 2026-04 |
| Qwen3 30B A3B Instruct 2507 | Alibaba / Qwen Team | 88.0 | 2026-04 |
| Qwen3 30B A3B NVFP4 | NVIDIA | 88.0 | 2026-04 |
| Qwen3 30B-A3B | Alibaba / Qwen Team | 88.0 | 2026-04 |
| Qwen3 8B | Alibaba / Qwen Team | 88 | 2026-04 |
| Qwen3 8B AWQ | Alibaba / Qwen Team | 88 | 2026-04 |
| Qwen3 8B Base | Alibaba / Qwen Team | 88 | 2026-04 |
| Skywork Reward V2 Qwen3 8B | Skywork | 88 | undated |
| Gemini 2.5 Flash | Google DeepMind | 87.0 | 2026-04 |
| Gemini 2.5 Flash Image | Google DeepMind | 87.0 | 2026-04 |
| Gemini 2.5 Flash Image (Preview) | Google DeepMind | 87.0 | 2026-04 |
| Gemini 2.5 Flash Lite | Google DeepMind | 87.0 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 06-17 | Google DeepMind | 87.0 | 2026-04 |
| Gemini 2.5 Flash Lite Preview 09-25 | Google DeepMind | 87.0 | 2026-04 |
| Gemini 2.5 Flash Preview 04-17 | Google DeepMind | 87.0 | 2026-04 |
| Gemini 2.5 Flash Preview 05-20 | Google DeepMind | 87.0 | 2026-04 |
| Gemini 2.5 Flash Preview 09-25 | Google DeepMind | 87.0 | 2026-04 |
| Gemini 2.5 Flash Preview TTS | Google DeepMind | 87.0 | 2026-04 |
| Gemini 1.5 Flash | Google DeepMind | 86.0 | 2026-04 |
| Gemini 1.5 Flash-8B | Google DeepMind | 86.0 | 2026-04 |
| Mistral Nemo | Mistral AI | 85 | 2026-04 |
| Mistral Nemo Base 2407 | Mistral AI | 85 | 2026-04 |
| Mistral Nemo Instruct 2407 | Mistral AI | 85 | 2026-04 |
| Gemma 4 31B | Google DeepMind | 84.0 | 2026-04 |
| gemma 4 31B it | Google DeepMind | 84.0 | 2026-04 |
| gemma 4 31B it GGUF | Unsloth | 84.0 | 2026-04 |
| Gemma 4 31B IT NVFP4 | NVIDIA | 84.0 | 2026-04 |
| Gemma 3 27B | Google DeepMind | 82.0 | 2026-04 |
| Grok 3 Mini | xAI | 82.0 | 2026-04 |
| Grok 3 Mini Fast | xAI | 82.0 | 2026-04 |
| Grok 3 Mini Fast Latest | xAI | 82.0 | 2026-04 |
| Grok 3 Mini Latest | xAI | 82.0 | 2026-04 |
| Mistral Small (latest) | Mistral AI | 82.0 | 2026-04 |
| Mistral Small 24B Instruct 2501 | Mistral AI | 82.0 | 2026-04 |
| Mistral Small 3.1 24B Instruct 2503 | Mistral AI | 82.0 | 2026-04 |
| Mistral Small 3.2 | Mistral AI | 82.0 | 2026-04 |
| Mistral Small 3.2 24B Instruct 2506 | Mistral AI | 82.0 | 2026-04 |
| Mistral Small 4 | Mistral AI | 82.0 | 2026-04 |
| Mistral Small 4 119B 2603 | Mistral AI | 82.0 | 2026-04 |
| gemma 2 27B it | Google DeepMind | 80.0 | 2026-04 |
| Qwen3 14B | Alibaba / Qwen Team | 80 | 2026-04 |
| Qwen3 14B AWQ | Alibaba / Qwen Team | 80 | 2026-04 |
| Qwen3 14B NVFP4 | NVIDIA | 80 | 2026-04 |
| Command R | Cohere | 78 | 2026-04 |
| GPT-4o | OpenAI | 78.0 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 78.0 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 78.0 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 78.0 | 2026-04 |
| GPT-4.1 | OpenAI | 76.0 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 73.0 | 2026-04 |
| Claude Sonnet 4 (latest) | Anthropic | 73.0 | 2026-04 |
| Claude Sonnet 4.6 | Anthropic | 73.0 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 73.0 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 73.0 | 2026-04 |
| Llama 3.1 70B | Meta | 72.0 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 72.0 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 71.0 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 71.0 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 71.0 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 71.0 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 71.0 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 71.0 | 2026-04 |
| DeepSeek V3 | DeepSeek | 70.0 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 70.0 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 70.0 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 70.0 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 70.0 | 2026-04 |
| GPT-5.1 | OpenAI | 70.0 | 2026-04 |
| GPT-5.1 Chat | OpenAI | 70.0 | 2026-04 |
| GPT-5.1 Codex | OpenAI | 70.0 | 2026-04 |
| GPT-5.1 Codex Max | OpenAI | 70.0 | 2026-04 |
| GPT-5.1 Codex mini | OpenAI | 70.0 | 2026-04 |
| Grok 2 | xAI | 70 | 2026-04 |
| Grok 2 (1212) | xAI | 70 | 2026-04 |
| Grok 2 Latest | xAI | 70 | 2026-04 |
| Qwen3 32B | Alibaba / Qwen Team | 70.0 | 2026-04 |
| Qwen3 32B AWQ | Alibaba / Qwen Team | 70.0 | 2026-04 |
| Qwen3 32B NVFP4 | NVIDIA | 70.0 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 68.0 | 2026-04 |
| GPT-4 Turbo | OpenAI | 65.0 | 2026-04 |
| Grok 3 | xAI | 65.0 | 2026-04 |
| Grok 3 Fast | xAI | 65.0 | 2026-04 |
| Grok 3 Fast Latest | xAI | 65.0 | 2026-04 |
| Grok 3 Latest | xAI | 65.0 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 62.0 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 60.0 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 60.0 | 2026-04 |
| Mistral Large 3 | Mistral AI | 60.0 | 2026-04 |
| Command A | Cohere | 58 | 2026-04 |
| Command A Reasoning | Cohere | 58 | 2026-04 |
| Command R+ | Cohere | 55.0 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 55.0 | 2026-04 |
| Llama 3.1 405B | Meta | 45.0 | 2026-04 |
| Llama 3.1 405B FP8 | Meta | 45.0 | 2026-04 |
| Llama 3.1 405B Instruct | Meta | 45.0 | 2026-04 |
| Llama 3.1 405B Instruct FP8 | Meta | 45.0 | 2026-04 |
| DeepSeek R1 | DeepSeek | 40.0 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 40.0 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 40.0 | 2026-04 |
| DeepSeek R1 Distill Llama 70B | DeepSeek | 40.0 | 2026-04 |
| DeepSeek R1 Distill Llama 8B | DeepSeek | 40 | 2026-04 |
| DeepSeek R1 Distill Qwen 14B | DeepSeek | 40.0 | 2026-04 |
| DeepSeek R1 Distill Qwen 32B | DeepSeek | 40.0 | 2026-04 |
| DeepSeek R1 Distill Qwen 7B | DeepSeek | 40 | 2026-04 |