Tests whether a multimodal agent can complete open-ended tasks in a real, live desktop operating system.
unassessed
| Category | agentic |
|---|---|
| Subcategory | computer-use / GUI agents |
| Page status | active |
| Metric | success rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 369 |
| Dataset licence | Apache-2.0 (repository licence, from the GitHub LICENSE file; the project website separately displays a CC BY-SA 4.0 notice for its own page content) |
| Publisher | University of Hong Kong, Salesforce Research, Carnegie Mellon University, University of Waterloo |
OSWorld gives an agent a natural-language instruction and a live view of a virtual machine, usually a screenshot and sometimes an accessibility tree, and asks it to operate real desktop and web applications to complete the task: finding files, editing documents, configuring settings, and other workflows that span multiple applications and OS file I/O rather than a single scripted action.
screenshot (optionally plus accessibility tree) as observation; agent issues GUI actions (click, type, drag, hotkey) in a loop until it signals completion or hits a step limit
| Model | Provider | Score | Card as of |
|---|---|---|---|
| Claude Mythos Preview | Anthropic | 79.6 | 2026-04 |
| GPT-5.4 | OpenAI | 75.0 | 2026-04 |
| Claude Opus 4.6 | Anthropic | 72.7 | 2026-04 |