Elo-style ratings of chat models derived from live, anonymous, pairwise human-preference votes on real user prompts.
unassessed
| Category | human-preference |
|---|---|
| Subcategory | pairwise human preference |
| Page status | active |
| Metric | Elo |
| Direction | higher_is_better |
| Publisher | LMSYS (Large Model Systems Organization), UC Berkeley Sky Computing Lab (founding org); operates today as Arena (formerly LMArena) |
Arena Elo measures which chatbot response people prefer in head-to-head, blind comparisons — a live measurement of human preference on real, unscripted prompts, not accuracy against a fixed answer key. A person submits a prompt, receives responses from two anonymised models chosen at random, and votes for the one they prefer; model identities are revealed only after voting. This repository's subsets cover the overall text leaderboard plus five commonly quoted slices: coding, math, a curated "hard prompts" subset, a style-controlled ranking, and a separate vision (image+text) leaderboard.
Anonymous, randomized side-by-side chat: a user prompts two randomly paired models and votes for the preferred response, or ties/both-bad.
No model card in ModelSpec reports this benchmark yet.