GUI-CC

GUI-CC evaluates whether GUI world models preserve context across repeated interaction instead of only predicting plausible next screens.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
SubcategoryGUI world models
Metriccontextual consistency
Directionhigher_is_better
Unitscore
Dataset size700
PublisherGUI-CC authors

What it measures

GUI-CC tests contextual consistency of generated mobile user interfaces as agent environments. It includes offline rollouts along real trajectories and online probing-agent loops over model-generated UIs.

Task format

Mobile GUI trajectories and emulator-verified agent tasks.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub