MAS-Bench scores Android GUI agents on 139 live-app tasks when they may call APIs, deep links, and RPA scripts instead of tapping through every screen.
unassessed
| Category | agentic |
|---|---|
| Subcategory | shortcut-augmented hybrid mobile GUI agents |
| Page status | active |
| Metric | success rate (SR), with separate efficiency and cost metrics |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 139 |
| Publisher | Zhejiang University and vivo AI Lab, with Peking University |
The agent receives an English instruction on a real Android app, or several apps, and must finish the task. It may tap the GUI, or it may invoke a shortcut: an API, a deep link, or a short RPA script. All 139 tasks are solvable with GUI only; shortcuts are an efficiency option, not a hidden gold path. A second track asks the agent to generate its own shortcut library from exploration, then a fixed T3A baseline runs the tasks with that library. The suite covers shopping, news, mail, maps, health, and similar daily apps, with screenshots plus structured UI state.
Multi-step Android control on an emulator snapshot (MAS-Bench-AVD). 92 single-app tasks and 47 cross-app tasks. Shortcuts are retrieved by app name and injected as function signatures. Success uses a two-stage describe-and-judge pipeline (MAS-Bench-Eval).
No model card in ModelSpec reports this benchmark yet.