An automatic, LLM-judged win-rate test of instruction-following that is built and validated to track human preference votes.
unassessed
| Category | human-preference |
|---|---|
| Subcategory | instruction-following preference |
| Page status | active |
| Metric | length-controlled win rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 805 |
| Dataset licence | Apache-2.0 |
| Publisher | Stanford University (Tatsu Lab) |
AlpacaEval takes a fixed set of instructions, generates a response from the model under test and from a fixed reference model, and asks a strong LLM judge which response it prefers. The result is a win rate against the reference model rather than an accuracy score on a task with a right answer. It was designed as a fast, cheap stand-in for the kind of human preference voting done by Chatbot Arena, and its authors validate new versions of the metric against correlation with those human votes rather than against a fixed answer key.
Single-turn instruction in, free-text response out, judged pairwise against a reference model's response to the same instruction by an LLM annotator.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Opus 4 | Anthropic | 55.2 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 55.2 | 2026-04 |
| GPT-4 Turbo | OpenAI | 55.0 | 2026-04 |
| Claude Sonnet 3.5 | Anthropic | 52.4 | 2026-04 |
| GPT-4 | OpenAI | 52.1 | 2026-04 |
| GPT-4.1 | OpenAI | 52.1 | 2026-04 |
| GPT-4.1 mini | OpenAI | 52.1 | 2026-04 |
| GPT-4.1 nano | OpenAI | 52.1 | 2026-04 |
| Gemini 2.5 Pro | Google DeepMind | 51.5 | 2026-04 |
| Gemini 2.5 Pro Preview 05-06 | Google DeepMind | 51.5 | 2026-04 |
| Gemini 2.5 Pro Preview 06-05 | Google DeepMind | 51.5 | 2026-04 |
| Gemini 2.5 Pro Preview TTS | Google DeepMind | 51.5 | 2026-04 |
| Claude Sonnet 4 | Anthropic | 50.8 | 2026-04 |
| Claude Sonnet 4.5 | Anthropic | 50.8 | 2026-04 |
| Claude Sonnet 4.5 (latest) | Anthropic | 50.8 | 2026-04 |
| GPT-4o | OpenAI | 48.5 | 2026-04 |
| GPT-4o (2024-05-13) | OpenAI | 48.5 | 2026-04 |
| GPT-4o (2024-08-06) | OpenAI | 48.5 | 2026-04 |
| GPT-4o (2024-11-20) | OpenAI | 48.5 | 2026-04 |
| GPT-4o mini | OpenAI | 48.5 | 2026-04 |
| DeepSeek R1 | DeepSeek | 45.2 | 2026-04 |
| DeepSeek R1 0528 | DeepSeek | 45.2 | 2026-04 |
| DeepSeek R1 0528 NVFP4 v2 | NVIDIA | 45.2 | 2026-04 |
| DeepSeek Reasoner | DeepSeek | 45.2 | 2026-04 |
| Qwen 3 235B Instruct | Cerebras | 44.5 | 2026-04 |
| Qwen3 235B-A22B | Alibaba / Qwen Team | 44.5 | 2026-04 |
| DeepSeek Chat | DeepSeek | 42.8 | 2026-04 |
| DeepSeek V2 | DeepSeek | 42.8 | 2026-04 |
| DeepSeek V2 Lite | DeepSeek | 42.8 | 2026-04 |
| DeepSeek V2 Lite Chat | DeepSeek | 42.8 | 2026-04 |
| DeepSeek V3 | DeepSeek | 42.8 | 2026-04 |
| DeepSeek V3 0324 | DeepSeek | 42.8 | 2026-04 |
| DeepSeek V3.1 | DeepSeek | 42.8 | 2026-04 |
| DeepSeek V3.2 | DeepSeek | 42.8 | 2026-04 |
| DeepSeek V3.2 Exp | DeepSeek | 42.8 | 2026-04 |
| Claude Opus 3 | Anthropic | 40.5 | 2026-04 |
| Llama 3.3 70B Instruct NVFP4 | NVIDIA | 40.5 | 2026-04 |
| Llama-3.3-70B-Instruct | Meta | 40.5 | 2026-04 |
| Llama 3.1 405B Instruct | Meta | 39.3 | 2026-04 |
| Mistral Large (latest) | Mistral AI | 38.5 | 2026-04 |
| Mistral Large 2.1 | Mistral AI | 38.5 | 2026-04 |
| Mistral Large 3 | Mistral AI | 38.5 | 2026-04 |
| Qwen3 32B | Alibaba / Qwen Team | 38.2 | 2026-04 |
| Qwen3 32B AWQ | Alibaba / Qwen Team | 38.2 | 2026-04 |
| Qwen3 32B NVFP4 | NVIDIA | 38.2 | 2026-04 |
| Llama 3.1 70B Instruct | Meta | 38.1 | 2026-04 |
| Qwen2.5 72B Instruct | Alibaba / Qwen Team | 38.1 | 2026-04 |
| Command R+ | Cohere | 35.2 | 2026-04 |
| Claude Sonnet 3 | Anthropic | 34.9 | 2026-04 |
| phi 4 | Microsoft | 32.5 | 2026-04 |
| Phi 4 mini instruct | Microsoft | 32.5 | 2026-04 |
| Phi 4 multimodal instruct | Microsoft | 32.5 | 2026-04 |
| Mistral Medium (latest) | Mistral AI | 28.6 | 2026-04 |
| GPT-3.5-turbo | OpenAI | 25.4 | 2026-04 |
| Gemini 1.5 Pro | Google DeepMind | 24.4 | 2026-04 |
| Llama 3.1 8B Instruct | Meta | 22.9 | 2026-04 |