TwitterAAE

HELM evaluates language-model bits per byte on 50,000 AAE-aligned and 50,000 White-aligned tweets.

Also known as: Twitter African-American English, Twitter AAE corpus

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorylanguage modeling across dialect-aligned tweet subsets
Page statusactive
Metricbits per byte
Directionlower_is_better
Unitbits/byte
Dataset size100000
PublisherStanford Center for Research on Foundation Models (HELM scenario)

What it measures

TwitterAAE compares language-model performance on tweets selected for high estimated African-American English or White alignment. These are corpus-alignment labels, not identity claims about individual authors.

Task format

Autoregressive language modeling on raw tweet text, evaluated separately for aa and white subsets.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub