EWoK-core-1.0 is a 4,374-item English set that tests whether a model matches a target sentence to the more plausible of two world-knowledge contexts.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | cognition-inspired English world-knowledge plausibility (11 domains) |
| Page status | unknown |
| Metric | exact_match (HELM); paper also reports LogProbs / Likert / Choice accuracy |
| Direction | higher_is_better |
| Dataset size | 4374 |
| Dataset licence | CC-BY-4.0 |
| Publisher | EWoK-core (MIT and collaborators) |
Elements of World Knowledge (EWoK) tests conceptual world modeling, not surface co-occurrence trivia. Each item gives two minimal-pair contexts and two targets. The matching context–target pairs are the plausible ones (C1 with T1, C2 with T2). Eleven domains range from social interactions (help/hinder) to spatial relations (left/right). This page documents EWoK-core-1.0, the public snapshot used in the paper and in HELM, not a later custom generation from the ewok-core/ewok pipeline.
HELM experimental run spec `ewok` uses joint multiple choice: the model sees one target as the "scenario" and two numbered contexts, and must answer "1" or "2". Adapter: ADAPT_MULTIPLE_CHOICE_JOINT, max_train_instances=2, max_tokens=2, temperature=0. The paper also reports LogProbs, Likert (1–5), and Choice paradigms that are not this HELM adapter.
No model card in ModelSpec reports this benchmark yet.