SEA-HELM (Southeast Asian Holistic Evaluation of Language Models)

AI Singapore's Southeast Asian LLM suite across NLP classics, LLM-specifics, linguistics, culture and safety, covering Filipino, Indonesian, Tamil, Thai and Vietnamese.

Also known as: SEA-HELM, Southeast Asian Holistic Evaluation of Language Models, BHASA

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
SubcategorySoutheast Asian multilingual suite: NLP classics, LLM-specifics, linguistics, culture, and safety
Page statusactive
MetricSEA average (mean of per-language scores); task metrics vary
Directionhigher_is_better
Dataset licenceMIT
PublisherAI Singapore, National University of Singapore

What it measures

SEA-HELM scores LLMs on Southeast Asian languages as a bundle, not as one English task. The 2025 paper names five pillars: NLP Classics, LLM-specifics, SEA Linguistics (LINDSEA), SEA Culture, and Safety. At publication the languages were Filipino, Indonesian, Tamil, Thai, and Vietnamese. Items include localised QA, sentiment, NLI, translation, linguistic minimal pairs, and later culture and safety sets. Stanford HELM's seahelm_scenario.py implements an earlier BHASA-shaped slice (NLU, NLG, NLR, LINDSEA), not the full five-pillar runner.

Task format

Per-task generation or classification in the target language, with language-specific prompts. The official aisingapore/SEA-HELM runner uses --tasks seahelm and, as of 2026, eight independent runs plus bootstrap intervals. Stanford HELM run specs (tydiqa, xquad, nusax, flores, indonli, xcopa, lindsea_*, and others) are separate adapters with their own metrics.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub