LiveBench

A monthly-refreshed suite across math, coding, reasoning, data analysis, language and instruction following, graded by automatic ground truth rather than an LLM judge, to limit contamination.

Also known as: Live Bench

unverified

This page is not in the default catalogue. Evidence required by the catalogue contract is missing or was not approved by a reviewer. That is a statement about the evidence we hold, not a claim that the benchmark is stale or illegitimate.

Recorded reasons:

Categorycomposite
Subcategorymonthly-refreshed multi-category benchmark
Page statusactive
Metrictask-specific automatic accuracy, averaged per category and overall
Directionhigher_is_better
Unit%
Dataset size1436

What it measures

LiveBench evaluates a model across six skill categories in one suite: math, coding, reasoning, data analysis, language comprehension and instruction following. Each category bundles several distinct task types rather than one narrow format — for example competition-math problems, LeetCode/AtCoder-style code generation, Zebra-puzzle and Web-of-Lies-style logic tasks, table reformatting and column-type annotation, and paraphrase/summarize/simplify instruction tasks — so a category score is itself a small suite average. The public leaderboard now also tracks a seventh category, Agentic Coding, added after the original paper.

Task format

Mostly single-turn text prompts (a small number of tasks use multiple turns) with a task-specific expected output — a number, a code solution graded by test cases, a puzzle answer, a reformatted table — that is checked automatically rather than judged by another model.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub