HELM Arabic Enterprise task: write a Modern Standard Arabic business article from given facts and style; an LLM judge scores faithfulness, completeness and style.
unassessed
| Category | generation |
|---|---|
| Subcategory | MSA business-article generation from supplied facts and style, LLM-judge scored |
| Page status | proposed |
| Metric | arabic_content_generation_score (mean of three 1–5 rubrics, rescaled to 0–1) |
| Direction | higher_is_better |
| Unit | 0-1 scale |
| Dataset size | 222 |
| Dataset licence | CC-BY-4.0 |
| Publisher | Stanford CRFM (HELM Arabic Enterprise) |
arabic_content_generation asks a model to write a professional business article in Modern Standard Arabic from two lists: facts it must use, and style features it must follow. HELM's schema says the source material is summaries from real news articles and press releases, rewritten in a corporate style. The model must use every supplied fact, invent none, and match tone, formality and voice. It is a constrained generation task, not open-ended Arabic creative writing.
The prompt is two Markdown sections, Facts then Style. Instructions tell the model to reply with the article only, max_tokens 2000. HELM then sends the user request and the completion to an LLM judge three times, once per rubric (faithfulness, completeness, style), each scored 1–5.
No model card in ModelSpec reports this benchmark yet.