WritingBench

1,000 real-world writing prompts across 6 domains and 100 subdomains, each graded on five auto-generated, instance-specific criteria by an LLM or critic-model judge.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
SubcategoryLLM-judged generative writing, professional and creative domains
Page statusactive
MetricLLM-judged or critic-model score, 1-10 per criterion, averaged over five query-specific criteria
Directionhigher_is_better
Unitpoints (1-10 per criterion; the public leaderboard multiplies the mean by 10 to display it out of 100)
Dataset size1000
Dataset licenceApache-2.0
PublisherAlibaba Group

What it measures

WritingBench asks a model to complete an open-ended, real-world writing request - drafting a paper outline, a marketing brief, a legal summary, a lesson plan and similar tasks - drawn from six domains (Academic & Engineering, Finance & Business, Politics & Law, Literature & Art, Education, Advertising & Marketing) and 100 finer-grained subdomains. Prompts are often long and specific, averaging over 1,500 tokens and sometimes embedding source material the response must work from, such as a financial statement or a sample abstract, rather than a short one-line instruction. Rather than judging every response against one fixed rubric, the benchmark generates five criteria specific to each individual query - covering things like required structure, tone or factual grounding for that particular request - so that a lesson plan and a legal brief are not scored on the same axes.

Task format

Open-ended long-form writing prompt in, spanning six primary domains and 100 subdomains, with prompt lengths from tens to thousands of words and some prompts supplying source material to write from; free-text written response out, graded rather than matched against a reference answer.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub