{
 "body": "\nPart of the [ScreenSpot-Pro](screenspot_pro.md) family.\n\n## What it measures\n\nThis id covers ScreenSpot-Pro scores produced with some form of iterative search, cropping or zooming\nrather than a single-shot prediction over the full screenshot. No Anthropic or OpenAI system card\npairing one named model's with-tools and without-tools ScreenSpot-Pro scores was found in this\nresearch, so treat any single tool configuration as unconfirmed unless its source states it. What is\ndirectly established, from the paper and its maintained leaderboard, is the general pattern: the\npaper's proposed ScreenSeekeR method uses a planner model to guide a cascaded search over image crops,\nand the leaderboard's strongest current entries are explicitly labelled \"zoom-in\" or agentic variants.\n\n## Reading the numbers\n\nAt launch, the paper's own comparison showed the pattern clearly: adding a search strategy lifted its\nbase grounding model from 18.9% to 48.1% accuracy with no additional training, using a planner model\n(GPT-4o) that itself scored under 1% at direct grounding \u2014 evidence a weak \"planner\" can still\nmeaningfully guide a stronger search process. The current leaderboard extends that pattern: top entries\nare consistently zoom-in or agentic variants, not single-shot predictions. Strategies differ in steps,\ncompute, and how much of the image they inspect, so do not treat two \"with tools\" scores as comparable\nwithout checking the strategy, and do not assume any lab's system card has published a matched pair\nfor this benchmark.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "screenspot_pro_tools",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Mythos Preview",
   "model_id": "anthropic/claude-mythos-preview",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 92.8,
   "source": "anthropic-system-card"
  },
  {
   "as_of": "2026-04",
   "attribution": "unverified-legacy",
   "display_name": "Claude Opus 4.6",
   "model_id": "anthropic/claude-opus-4-6",
   "provider": "anthropic",
   "provider_display": "Anthropic",
   "score": 83.1,
   "source": "lmarena.ai, provider-reports, multimodal-evals, safety-evals, preference-evals, domain-evals, anthropic-system-card-mythos"
  }
 ],
 "page": {
  "aliases": [
   "ScreenSpot-Pro agentic",
   "ScreenSpot-Pro zoom-in"
  ],
  "category": "agentic",
  "contamination": {
   "note": "Same considerations as screenspot_pro; tool use does not add a distinct leakage pathway for an image-grounding task.",
   "risk": "low"
  },
  "dataset": {
   "languages": [
    "en"
   ],
   "license": "MIT",
   "modalities": [
    "image",
    "text"
   ],
   "public_test_set": true,
   "size": 1581,
   "size_note": "Same instruction set as screenspot_pro; see that page for dataset detail.",
   "splits": "single evaluation set, no train/test split",
   "url": "https://huggingface.co/datasets/likaixin/ScreenSpot-Pro"
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "sonnet-5 agent, batch 1b, slice L",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "No dedicated harness task name confirmed; see screenspot_pro for the base task and its own evaluation scripts."
  },
  "id": "screenspot_pro_tools",
  "last_updated": "2026-08",
  "leaderboard_url": "https://gui-agent.github.io/grounding-leaderboard/",
  "lineage": {
   "family": "screenspot_pro",
   "predecessor": "",
   "successors": [],
   "variants": []
  },
  "measures": "This id captures ScreenSpot-Pro scores produced with some form of tool use or multi-step search rather than a single-shot grounding prediction over the full-resolution screenshot. No specific Anthropic or OpenAI system card documenting a named model's paired with-tools/without-tools ScreenSpot-Pro score was found in this research, so a single canonical tool setting is not established here. What is established, directly from the benchmark's own paper and its actively-maintained leaderboard: the field's dominant \"tool\" pattern on this task is an iterative zoom, crop or planner-guided search loop that progressively narrows the search region, rather than browsing or code execution. The paper's own proposed method, ScreenSeekeR, is an example: a planner model guides a cascaded search over image crops. Where a specific model's tool configuration is not documented by its source, treat that configuration as unknown rather than assuming it matches another source's setup.\n",
  "metric": {
   "baseline_note": "Directly comparable to screenspot_pro's click-in-box accuracy on the same 1,581 instructions; not comparable across different tool/search strategies, which vary in how many steps or how much compute they use per instruction.\n",
   "direction": "higher_is_better",
   "max_score": 100,
   "name": "click accuracy",
   "unit": "%"
  },
  "name": "ScreenSpot-Pro (with tools)",
  "page_kind": "subset",
  "paper": {
   "arxiv": "2504.07981",
   "title": "ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use",
   "url": "https://arxiv.org/abs/2504.07981",
   "year": 2025
  },
  "publisher": {
   "authors": [],
   "org": "Independent research collaboration (Hong Kong Baptist University and collaborators)",
   "url": "https://github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding"
  },
  "released": "2025-01",
  "repo_url": "https://github.com/likaixin2000/ScreenSpot-Pro-GUI-Grounding",
  "saturation": {
   "as_of": "2026-08",
   "note": "The project's leaderboard's current top entry (82.7%, fetched directly for this page) is an explicitly \"zoom-in\" variant of a specialized grounding model. The paper's own launch-era example of the same pattern was smaller but clearer: ScreenSeekeR lifted its base model, OS-Atlas-7B, from 18.9% to 48.1% using only a cascaded search strategy, with no additional training.\n",
   "status": "open",
   "top_score": 82.7
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use (arXiv:2504.07981)",
    "url": "https://arxiv.org/abs/2504.07981"
   },
   {
    "accessed": "2026-09-08",
    "title": "ScreenSpot-Pro Leaderboard (results data)",
    "url": "https://gui-agent.github.io/grounding-leaderboard/"
   }
  ],
  "status": "active",
  "subcategory": "GUI grounding, tool-augmented",
  "summary": "ScreenSpot-Pro scores produced by an iterative search, crop or zoom strategy instead of a single-shot prediction over the full screenshot.",
  "tags": [
   "agentic",
   "gui-grounding",
   "computer-use",
   "tool-use"
  ],
  "task_format": "Same screenshot-plus-instruction grounding task as screenspot_pro, but the model (or a wrapper around it) may take multiple steps \u2014 for example, requesting a cropped or zoomed view, or using a separate planner model to guide the search \u2014 before committing to a final point.\n"
 }
}