ScreenSpot-Pro

1,581 instructions testing whether a model can point to the right UI element in authentic, high-resolution screenshots of 23 professional applications.

Also known as: ScreenSpot Pro

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
SubcategoryGUI grounding
Page statusactive
Metricclick accuracy
Directionhigher_is_better
Unit%
Dataset size1581
Dataset licenceMIT
PublisherIndependent research collaboration (Hong Kong Baptist University and collaborators)

What it measures

ScreenSpot-Pro tests GUI grounding in professional software: given a natural-language instruction and a screenshot, a model must locate the precise on-screen element the instruction refers to. Unlike earlier grounding benchmarks built around everyday consumer apps and mobile screens, every image here is an authentic, expert-captured screenshot from real professional workflows in fields like CAD, scientific computing, creative software and IDEs, at native high resolution (over 1080p) rather than a cropped or downscaled view. Targets are correspondingly tiny: on average a target occupies about 0.07% of the screenshot area, versus 2.01% on the original ScreenSpot benchmark this one extends.

Task format

A screenshot plus a natural-language instruction describing an action ("open the layers panel", for example); the model outputs a single point or coordinate. Targets are additionally labelled as either text or icon.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic79.52026-04
Claude Opus 4.6Anthropic57.72026-04

Data

This page as JSON · Edit on GitHub