AlpacaEval

An automatic, LLM-judged win-rate test of instruction-following that is built and validated to track human preference votes.

Also known as: AlpacaEval 2.0, Length-Controlled AlpacaEval, LC AlpacaEval

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategoryinstruction-following preference
Page statusactive
Metriclength-controlled win rate
Directionhigher_is_better
Unit%
Dataset size805
Dataset licenceApache-2.0
PublisherStanford University (Tatsu Lab)

What it measures

AlpacaEval takes a fixed set of instructions, generates a response from the model under test and from a fixed reference model, and asks a strong LLM judge which response it prefers. The result is a win rate against the reference model rather than an accuracy score on a task with a right answer. It was designed as a fast, cheap stand-in for the kind of human preference voting done by Chatbot Arena, and its authors validate new versions of the metric against correlation with those human votes rather than against a fixed answer key.

Task format

Single-turn instruction in, free-text response out, judged pairwise against a reference model's response to the same instruction by an LLM annotator.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Opus 4Anthropic55.22026-04
Claude Opus 4.6Anthropic55.22026-04
GPT-4 TurboOpenAI55.02026-04
Claude Sonnet 3.5Anthropic52.42026-04
GPT-4OpenAI52.12026-04
GPT-4.1OpenAI52.12026-04
GPT-4.1 miniOpenAI52.12026-04
GPT-4.1 nanoOpenAI52.12026-04
Gemini 2.5 ProGoogle DeepMind51.52026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind51.52026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind51.52026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind51.52026-04
Claude Sonnet 4Anthropic50.82026-04
Claude Sonnet 4.5Anthropic50.82026-04
Claude Sonnet 4.5 (latest)Anthropic50.82026-04
GPT-4oOpenAI48.52026-04
GPT-4o (2024-05-13)OpenAI48.52026-04
GPT-4o (2024-08-06)OpenAI48.52026-04
GPT-4o (2024-11-20)OpenAI48.52026-04
GPT-4o miniOpenAI48.52026-04
DeepSeek R1DeepSeek45.22026-04
DeepSeek R1 0528DeepSeek45.22026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA45.22026-04
DeepSeek ReasonerDeepSeek45.22026-04
Qwen 3 235B InstructCerebras44.52026-04
Qwen3 235B-A22BAlibaba / Qwen Team44.52026-04
DeepSeek ChatDeepSeek42.82026-04
DeepSeek V2DeepSeek42.82026-04
DeepSeek V2 LiteDeepSeek42.82026-04
DeepSeek V2 Lite ChatDeepSeek42.82026-04
DeepSeek V3DeepSeek42.82026-04
DeepSeek V3 0324DeepSeek42.82026-04
DeepSeek V3.1DeepSeek42.82026-04
DeepSeek V3.2DeepSeek42.82026-04
DeepSeek V3.2 ExpDeepSeek42.82026-04
Claude Opus 3Anthropic40.52026-04
Llama 3.3 70B Instruct NVFP4NVIDIA40.52026-04
Llama-3.3-70B-InstructMeta40.52026-04
Llama 3.1 405B InstructMeta39.32026-04
Mistral Large (latest)Mistral AI38.52026-04
Mistral Large 2.1Mistral AI38.52026-04
Mistral Large 3Mistral AI38.52026-04
Qwen3 32BAlibaba / Qwen Team38.22026-04
Qwen3 32B AWQAlibaba / Qwen Team38.22026-04
Qwen3 32B NVFP4NVIDIA38.22026-04
Llama 3.1 70B InstructMeta38.12026-04
Qwen2.5 72B InstructAlibaba / Qwen Team38.12026-04
Command R+Cohere35.22026-04
Claude Sonnet 3Anthropic34.92026-04
phi 4Microsoft32.52026-04
Phi 4 mini instructMicrosoft32.52026-04
Phi 4 multimodal instructMicrosoft32.52026-04
Mistral Medium (latest)Mistral AI28.62026-04
GPT-3.5-turboOpenAI25.42026-04
Gemini 1.5 ProGoogle DeepMind24.42026-04
Llama 3.1 8B InstructMeta22.92026-04

Data

This page as JSON · Edit on GitHub