EWoK (Elements of World Knowledge)

EWoK-core-1.0 is a 4,374-item English set that tests whether a model matches a target sentence to the more plausible of two world-knowledge contexts.

Also known as: Elements of World Knowledge, EWoK-core-1.0, ewok-core

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
Subcategorycognition-inspired English world-knowledge plausibility (11 domains)
Page statusunknown
Metricexact_match (HELM); paper also reports LogProbs / Likert / Choice accuracy
Directionhigher_is_better
Dataset size4374
Dataset licenceCC-BY-4.0
PublisherEWoK-core (MIT and collaborators)

What it measures

Elements of World Knowledge (EWoK) tests conceptual world modeling, not surface co-occurrence trivia. Each item gives two minimal-pair contexts and two targets. The matching context–target pairs are the plausible ones (C1 with T1, C2 with T2). Eleven domains range from social interactions (help/hinder) to spatial relations (left/right). This page documents EWoK-core-1.0, the public snapshot used in the paper and in HELM, not a later custom generation from the ewok-core/ewok pipeline.

Task format

HELM experimental run spec `ewok` uses joint multiple choice: the model sees one target as the "scenario" and two numbered contexts, and must answer "1" or "2". Adapter: ADAPT_MULTIPLE_CHOICE_JOINT, max_train_instances=2, max_tokens=2, temperature=0. The paper also reports LogProbs, Likert (1–5), and Choice paradigms that are not this HELM adapter.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub