ScreenSpot-Pro scores produced by an iterative search, crop or zoom strategy instead of a single-shot prediction over the full screenshot.
unassessed
| Category | agentic |
|---|---|
| Subcategory | GUI grounding, tool-augmented |
| Page status | active |
| Metric | click accuracy |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1581 |
| Dataset licence | MIT |
| Publisher | Independent research collaboration (Hong Kong Baptist University and collaborators) |
This id captures ScreenSpot-Pro scores produced with some form of tool use or multi-step search rather than a single-shot grounding prediction over the full-resolution screenshot. No specific Anthropic or OpenAI system card documenting a named model's paired with-tools/without-tools ScreenSpot-Pro score was found in this research, so a single canonical tool setting is not established here. What is established, directly from the benchmark's own paper and its actively-maintained leaderboard: the field's dominant "tool" pattern on this task is an iterative zoom, crop or planner-guided search loop that progressively narrows the search region, rather than browsing or code execution. The paper's own proposed method, ScreenSeekeR, is an example: a planner model guides a cascaded search over image crops. Where a specific model's tool configuration is not documented by its source, treat that configuration as unknown rather than assuming it matches another source's setup.
Same screenshot-plus-instruction grounding task as screenspot_pro, but the model (or a wrapper around it) may take multiple steps — for example, requesting a cropped or zoomed view, or using a separate planner model to guide the search — before committing to a final point.
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 92.8 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 83.1 | 2026-04 |