TheAgentCompany

TheAgentCompany has an agent complete 175 long-horizon professional tasks in a simulated software company to measure real-work automation.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorycomputer-use agents
Page statusactive
Metrictask success rate / checkpoint completion
Directionhigher_is_better
Unitpercent
Dataset size175
Dataset licenceMIT
PublisherCarnegie Mellon University

What it measures

TheAgentCompany evaluates whether an LLM agent can act as a digital worker in a small software company, browsing internal web apps, writing and running code, and messaging simulated coworkers to complete tasks drawn from software engineering, project management, data science, HR, finance, and admin roles.

Task format

The agent operates in a self-hosted environment (GitLab, the Plane project tracker, ownCloud, and RocketChat, all pre-populated with company data) and is given a natural-language task instruction; it must take actions over many steps to complete the task.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub