ScreenSpot-Pro (with tools)

ScreenSpot-Pro scores produced by an iterative search, crop or zoom strategy instead of a single-shot prediction over the full screenshot.

Also known as: ScreenSpot-Pro agentic, ScreenSpot-Pro zoom-in

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
SubcategoryGUI grounding, tool-augmented
Page statusactive
Metricclick accuracy
Directionhigher_is_better
Unit%
Dataset size1581
Dataset licenceMIT
PublisherIndependent research collaboration (Hong Kong Baptist University and collaborators)

What it measures

This id captures ScreenSpot-Pro scores produced with some form of tool use or multi-step search rather than a single-shot grounding prediction over the full-resolution screenshot. No specific Anthropic or OpenAI system card documenting a named model's paired with-tools/without-tools ScreenSpot-Pro score was found in this research, so a single canonical tool setting is not established here. What is established, directly from the benchmark's own paper and its actively-maintained leaderboard: the field's dominant "tool" pattern on this task is an iterative zoom, crop or planner-guided search loop that progressively narrows the search region, rather than browsing or code execution. The paper's own proposed method, ScreenSeekeR, is an example: a planner model guides a cascaded search over image crops. Where a specific model's tool configuration is not documented by its source, treat that configuration as unknown rather than assuming it matches another source's setup.

Task format

Same screenshot-plus-instruction grounding task as screenspot_pro, but the model (or a wrapper around it) may take multiple steps — for example, requesting a cropped or zoomed view, or using a separate planner model to guide the search — before committing to a final point.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic92.82026-04
Claude Opus 4.6Anthropic83.12026-04

Data

This page as JSON · Edit on GitHub