TheAgentCompany has an agent complete 175 long-horizon professional tasks in a simulated software company to measure real-work automation.
unassessed
| Category | agentic |
|---|---|
| Subcategory | computer-use agents |
| Page status | active |
| Metric | task success rate / checkpoint completion |
| Direction | higher_is_better |
| Unit | percent |
| Dataset size | 175 |
| Dataset licence | MIT |
| Publisher | Carnegie Mellon University |
TheAgentCompany evaluates whether an LLM agent can act as a digital worker in a small software company, browsing internal web apps, writing and running code, and messaging simulated coworkers to complete tasks drawn from software engineering, project management, data science, HR, finance, and admin roles.
The agent operates in a self-hosted environment (GitLab, the Plane project tracker, ownCloud, and RocketChat, all pre-populated with company data) and is given a natural-language task instruction; it must take actions over many steps to complete the task.
No model card in ModelSpec reports this benchmark yet.