GUI-CC evaluates whether GUI world models preserve context across repeated interaction instead of only predicting plausible next screens.
unassessed
| Category | agentic |
|---|---|
| Subcategory | GUI world models |
| Metric | contextual consistency |
| Direction | higher_is_better |
| Unit | score |
| Dataset size | 700 |
| Publisher | GUI-CC authors |
GUI-CC tests contextual consistency of generated mobile user interfaces as agent environments. It includes offline rollouts along real trajectories and online probing-agent loops over model-generated UIs.
Mobile GUI trajectories and emulator-verified agent tasks.
No model card in ModelSpec reports this benchmark yet.