ASTRA-bench

Scores personal-assistant tool use on 2,413 scenarios that mix a stateful email-calendar-messaging sandbox with time-evolving synthetic user context.

Also known as: Assistant Skills in Tool-use, Reasoning & Action-planning, ASTRA-bench (Apple)

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorystateful personal-assistant tool use over longitudinal synthetic context
Page statusunknown
MetricLLM-judge task-completion rate (gpt-5.1); also rule-based milestone score
Directionhigher_is_better
Dataset size2413
PublisherApple

What it measures

ASTRA-bench (Assistant Skills in Tool-use, Reasoning & Action-planning) tests whether a tool-calling agent can satisfy a user's request by reading and writing a simulated personal datastore. The agent must ground underspecified mentions in emails, calendars, messages, contacts, WhatsApp, and call logs that unfold over a protagonist's storyline, then plan multi-step tool calls against a live sandbox. Tasks are annotated on three complexity axes (referential, informational, and functional) plus two stress flags (misinformation and insufficient context). The main paper reports a zero-shot English study; an appendix machine-translates a subset of queries while keeping English system prompts.

Task format

Multi-turn tool-calling against a stateful personal-information sandbox derived from ToolSandbox. An LLM user simulator, bound by a knowledge boundary, holds the user goal. The agent may search, create, modify, or delete entities across 27 tools in six app domains until it meets success conditions or violates a minefield.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub