1,581 instructions testing whether a model can point to the right UI element in authentic, high-resolution screenshots of 23 professional applications.
unassessed
| Category | agentic |
|---|---|
| Subcategory | GUI grounding |
| Page status | active |
| Metric | click accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1581 |
| Dataset licence | MIT |
| Publisher | Independent research collaboration (Hong Kong Baptist University and collaborators) |
ScreenSpot-Pro tests GUI grounding in professional software: given a natural-language instruction and a screenshot, a model must locate the precise on-screen element the instruction refers to. Unlike earlier grounding benchmarks built around everyday consumer apps and mobile screens, every image here is an authentic, expert-captured screenshot from real professional workflows in fields like CAD, scientific computing, creative software and IDEs, at native high resolution (over 1080p) rather than a cropped or downscaled view. Targets are correspondingly tiny: on average a target occupies about 0.07% of the screenshot area, versus 2.01% on the original ScreenSpot benchmark this one extends.
A screenshot plus a natural-language instruction describing an action ("open the layers panel", for example); the model outputs a single point or coordinate. Targets are additionally labelled as either text or icon.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 79.5 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 57.7 | 2026-04 |