WorldSense

WorldSense is a synthetic benchmark that tests whether a model can maintain a consistent world model while controlling for dataset bias.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryreasoning
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unitpercent

What it measures

WorldSense tests whether a model can maintain a consistent internal world model from a set of statements, across three problem types and two difficulty grades, while controlling for the response-position and label biases that let models shortcut similar tasks.

Task format

Text input describing a small scenario (e.g. object placements or a scheduling puzzle) with three problem types: Infer (judge a statement true or false), Compl (pick the correct statement among three options), and Consist (judge whether a set of statements is possible or impossible).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub