A monthly-refreshed suite across math, coding, reasoning, data analysis, language and instruction following, graded by automatic ground truth rather than an LLM judge, to limit contamination.
unverified
Recorded reasons:
| Category | composite |
|---|---|
| Subcategory | monthly-refreshed multi-category benchmark |
| Page status | active |
| Metric | task-specific automatic accuracy, averaged per category and overall |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1436 |
LiveBench evaluates a model across six skill categories in one suite: math, coding, reasoning, data analysis, language comprehension and instruction following. Each category bundles several distinct task types rather than one narrow format — for example competition-math problems, LeetCode/AtCoder-style code generation, Zebra-puzzle and Web-of-Lies-style logic tasks, table reformatting and column-type annotation, and paraphrase/summarize/simplify instruction tasks — so a category score is itself a small suite average. The public leaderboard now also tracks a seventh category, Agentic Coding, added after the original paper.
Mostly single-turn text prompts (a small number of tasks use multiple turns) with a task-specific expected output — a number, a code solution graded by test cases, a puzzle answer, a reformatted table — that is checked automatically rather than judged by another model.
No model card in ModelSpec reports this benchmark yet.