WorldSense is a synthetic benchmark that tests whether a model can maintain a consistent world model while controlling for dataset bias.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | reasoning |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | percent |
WorldSense tests whether a model can maintain a consistent internal world model from a set of statements, across three problem types and two difficulty grades, while controlling for the response-position and label biases that let models shortcut similar tasks.
Text input describing a small scenario (e.g. object placements or a scheduling puzzle) with three problem types: Infer (judge a statement true or false), Compl (pick the correct statement among three options), and Consist (judge whether a set of statements is possible or impossible).
No model card in ModelSpec reports this benchmark yet.