Scores personal-assistant tool use on 2,413 scenarios that mix a stateful email-calendar-messaging sandbox with time-evolving synthetic user context.
unassessed
| Category | agentic |
|---|---|
| Subcategory | stateful personal-assistant tool use over longitudinal synthetic context |
| Page status | unknown |
| Metric | LLM-judge task-completion rate (gpt-5.1); also rule-based milestone score |
| Direction | higher_is_better |
| Dataset size | 2413 |
| Publisher | Apple |
ASTRA-bench (Assistant Skills in Tool-use, Reasoning & Action-planning) tests whether a tool-calling agent can satisfy a user's request by reading and writing a simulated personal datastore. The agent must ground underspecified mentions in emails, calendars, messages, contacts, WhatsApp, and call logs that unfold over a protagonist's storyline, then plan multi-step tool calls against a live sandbox. Tasks are annotated on three complexity axes (referential, informational, and functional) plus two stress flags (misinformation and insufficient context). The main paper reports a zero-shot English study; an appendix machine-translates a subset of queries while keeping English system prompts.
Multi-turn tool-calling against a stateful personal-information sandbox derived from ToolSandbox. An LLM user simulator, bound by a knowledge boundary, holds the user goal. The agent may search, create, modify, or delete entities across 27 tools in six app domains until it meets success conditions or violates a minefield.
No model card in ModelSpec reports this benchmark yet.