OSWorld

Tests whether a multimodal agent can complete open-ended tasks in a real, live desktop operating system.

Also known as: OS-World

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorycomputer-use / GUI agents
Page statusactive
Metricsuccess rate
Directionhigher_is_better
Unit%
Dataset size369
Dataset licenceApache-2.0 (repository licence, from the GitHub LICENSE file; the project website separately displays a CC BY-SA 4.0 notice for its own page content)
PublisherUniversity of Hong Kong, Salesforce Research, Carnegie Mellon University, University of Waterloo

What it measures

OSWorld gives an agent a natural-language instruction and a live view of a virtual machine, usually a screenshot and sometimes an accessibility tree, and asks it to operate real desktop and web applications to complete the task: finding files, editing documents, configuring settings, and other workflows that span multiple applications and OS file I/O rather than a single scripted action.

Task format

screenshot (optionally plus accessibility tree) as observation; agent issues GUI actions (click, type, drag, hotkey) in a loop until it signals completion or hits a step limit

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
Claude Mythos PreviewAnthropic79.62026-04
GPT-5.4OpenAI75.02026-04
Claude Opus 4.6Anthropic72.72026-04

Data

This page as JSON · Edit on GitHub