SWE-Lancer

OpenAI's benchmark of over 1,400 real Upwork freelance tasks on the Expensify codebase, worth $1 million in actual historical payouts, scored in dollars earned rather than percent resolved.

Also known as: SWE-Lancer Diamond

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycoding
Subcategoryfreelance software-engineering tasks and managerial decisions, priced in real dollars
Page statusactive
Metric$ earned (sum of the payout values of resolved tasks), alongside pass rate per task type
Directionhigher_is_better
Unit$
Dataset size1488
Dataset licenceMIT
PublisherOpenAI

What it measures

SWE-Lancer gives a model real freelance software-engineering work pulled from Upwork job postings against the Expensify open-source repository, with each task tagged at the dollar amount actually paid out for it historically. Two task types are covered: independent engineering (IC SWE) tasks, which range from small bug fixes worth $50 to large feature builds worth up to $32,000 and are graded by running an end-to-end test suite against the model's patch; and SWE Manager tasks, which give the model several competing technical implementation proposals for an issue and ask it to pick the one the real hiring manager chose. Because every task carries its real-world price, a model's performance converts directly into a dollar figure rather than an abstract percentage.

Task format

IC SWE: given an issue description and full repository access inside a Docker container, the model edits code and submits a patch, graded by an end-to-end Playwright test suite it cannot see during the attempt; payout is all-or-nothing per task, with no partial credit. SWE Manager: given an issue and several candidate implementation proposals originally written by competing freelancers, the model must select the proposal the real, original engineering manager actually chose.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub