HELM evaluates language-model bits per byte on 50,000 AAE-aligned and 50,000 White-aligned tweets.
unassessed
| Category | domain |
|---|---|
| Subcategory | language modeling across dialect-aligned tweet subsets |
| Page status | active |
| Metric | bits per byte |
| Direction | lower_is_better |
| Unit | bits/byte |
| Dataset size | 100000 |
| Publisher | Stanford Center for Research on Foundation Models (HELM scenario) |
TwitterAAE compares language-model performance on tweets selected for high estimated African-American English or White alignment. These are corpus-alignment labels, not identity claims about individual authors.
Autoregressive language modeling on raw tweet text, evaluated separately for aa and white subsets.
No model card in ModelSpec reports this benchmark yet.