1,000 real-world writing prompts across 6 domains and 100 subdomains, each graded on five auto-generated, instance-specific criteria by an LLM or critic-model judge.
unassessed
| Category | generation |
|---|---|
| Subcategory | LLM-judged generative writing, professional and creative domains |
| Page status | active |
| Metric | LLM-judged or critic-model score, 1-10 per criterion, averaged over five query-specific criteria |
| Direction | higher_is_better |
| Unit | points (1-10 per criterion; the public leaderboard multiplies the mean by 10 to display it out of 100) |
| Dataset size | 1000 |
| Dataset licence | Apache-2.0 |
| Publisher | Alibaba Group |
WritingBench asks a model to complete an open-ended, real-world writing request - drafting a paper outline, a marketing brief, a legal summary, a lesson plan and similar tasks - drawn from six domains (Academic & Engineering, Finance & Business, Politics & Law, Literature & Art, Education, Advertising & Marketing) and 100 finer-grained subdomains. Prompts are often long and specific, averaging over 1,500 tokens and sometimes embedding source material the response must work from, such as a financial statement or a sample abstract, rather than a short one-line instruction. Rather than judging every response against one fixed rubric, the benchmark generates five criteria specific to each individual query - covering things like required structure, tone or factual grounding for that particular request - so that a lesson plan and a legal brief are not scored on the same axes.
Open-ended long-form writing prompt in, spanning six primary domains and 100 subdomains, with prompt lengths from tens to thousands of words and some prompts supplying source material to write from; free-text written response out, graded rather than matched against a reference answer.
No model card in ModelSpec reports this benchmark yet.