Arena Elo (Chatbot Arena / LMArena)

Elo-style ratings of chat models derived from live, anonymous, pairwise human-preference votes on real user prompts.

Also known as: Chatbot Arena, LMArena, Arena

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryhuman-preference
Subcategorypairwise human preference
Page statusactive
MetricElo
Directionhigher_is_better
PublisherLMSYS (Large Model Systems Organization), UC Berkeley Sky Computing Lab (founding org); operates today as Arena (formerly LMArena)

What it measures

Arena Elo measures which chatbot response people prefer in head-to-head, blind comparisons — a live measurement of human preference on real, unscripted prompts, not accuracy against a fixed answer key. A person submits a prompt, receives responses from two anonymised models chosen at random, and votes for the one they prefer; model identities are revealed only after voting. This repository's subsets cover the overall text leaderboard plus five commonly quoted slices: coding, math, a curated "hard prompts" subset, a style-controlled ranking, and a separate vision (image+text) leaderboard.

Task format

Anonymous, randomized side-by-side chat: a user prompts two randomly paired models and votes for the preferred response, or ties/both-bad.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub