Arabic Content Generation (HELM Arabic Enterprise)

HELM Arabic Enterprise task: write a Modern Standard Arabic business article from given facts and style; an LLM judge scores faithfulness, completeness and style.

Also known as: Article Generation, arabic enterprise content_generation

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorygeneration
SubcategoryMSA business-article generation from supplied facts and style, LLM-judge scored
Page statusproposed
Metricarabic_content_generation_score (mean of three 1–5 rubrics, rescaled to 0–1)
Directionhigher_is_better
Unit0-1 scale
Dataset size222
Dataset licenceCC-BY-4.0
PublisherStanford CRFM (HELM Arabic Enterprise)

What it measures

arabic_content_generation asks a model to write a professional business article in Modern Standard Arabic from two lists: facts it must use, and style features it must follow. HELM's schema says the source material is summaries from real news articles and press releases, rewritten in a corporate style. The model must use every supplied fact, invent none, and match tone, formality and voice. It is a constrained generation task, not open-ended Arabic creative writing.

Task format

The prompt is two Markdown sections, Facts then Style. Instructions tell the model to reply with the article only, max_tokens 2000. HELM then sends the user request and the completion to an LLM judge three times, once per rubric (faithfulness, completeness, style), each scored 1–5.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub