Artificial Analysis's private 91-task agentic benchmark of long-horizon professional knowledge work, scored as combined Elo from rubric success, analytical quality, and presentation.
active
Recorded reasons:
| Category | agentic |
|---|---|
| Subcategory | long-horizon professional knowledge work with file deliverables |
| Page status | active |
| Metric | combined Elo, Index-normalized as clamp((Elo − 500) / 2000) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 91 |
| Publisher | Artificial Analysis |
AA-Briefcase measures whether a model can complete realistic, multi-week professional knowledge-work projects as file deliverables. The scored set is 91 tasks across four private scenarios (data science, product management, banking operations, and heavy industry strategy). Each scenario is a linked weekly workflow with thousands of source files. The agent must produce artefacts such as spreadsheets, presentations, memos, and PDFs in an offline sandbox, without live user feedback. Tasks currently run independently, so a model does not carry its own prior submissions into later weeks.
One independent Stirrup/E2B run per task (up to 500 turns). The agent receives the scenario and week overviews, the task brief, and mounted source files, then submits named deliverable files. No internet. Vision models also get a view-image tool. One repeat. Official scores use the private 91-task set, not the public Lite scenario.
Each row was checked against its source by a reviewer.
| Model | Score | Evidence date | Source kind | Link |
|---|---|---|---|---|
| GPT-6 Astra (max) | 53.0 normalized Elo percent | 2026-09-04 published | independent_evaluator | source |
| GLM-5.3 (max) | 51.0 normalized Elo percent | 2026-09-04 published | independent_evaluator | source |
No model card in ModelSpec reports this benchmark yet.