MAS-Bench

MAS-Bench scores Android GUI agents on 139 live-app tasks when they may call APIs, deep links, and RPA scripts instead of tapping through every screen.

Also known as: MAS-Bench, MAS-Bench-Eval

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryshortcut-augmented hybrid mobile GUI agents
Page statusactive
Metricsuccess rate (SR), with separate efficiency and cost metrics
Directionhigher_is_better
Unit%
Dataset size139
PublisherZhejiang University and vivo AI Lab, with Peking University

What it measures

The agent receives an English instruction on a real Android app, or several apps, and must finish the task. It may tap the GUI, or it may invoke a shortcut: an API, a deep link, or a short RPA script. All 139 tasks are solvable with GUI only; shortcuts are an efficiency option, not a hidden gold path. A second track asks the agent to generate its own shortcut library from exploration, then a fixed T3A baseline runs the tasks with that library. The suite covers shopping, news, mail, maps, health, and similar daily apps, with screenshots plus structured UI state.

Task format

Multi-step Android control on an emulator snapshot (MAS-Bench-AVD). 92 single-app tasks and 47 cross-app tasks. Shortcuts are retrieved by app name and injected as function signatures. Success uses a two-stage describe-and-judge pipeline (MAS-Bench-Eval).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub