SWE-Lancer
SWE-Lancer asks a blunter question than most coding evaluations: how much money could this model actually have earned? It takes 1,488 real freelance jobs posted on Upwork, together worth $1,000,000 that was genuinely paid to human engineers, and scores a model in dollars recovered rather than in percentage points. That single change makes it the most directly economic benchmark in the directory.
| What it measures | Whether frontier models can complete real paid freelance software jobs, both coding and technical-management tasks, well enough to earn the payouts. |
|---|---|
| Built by | OpenAI (Miserendino, Patwardhan et al.), 2025 |
| Format | 1,488 real Upwork freelance tasks worth $1,000,000 in payouts: 764 IC SWE tasks worth $414,775, graded by engineer-written end-to-end tests, and 724 SWE Manager tasks worth $585,225; the public Diamond split holds 502 tasks worth $500,800 |
| Scoring metric | Dollars earned (and % of tasks resolved) |
| Status | Active |
| Representative top score | Not independently confirmed; see the leaderboard |
How SWE-Lancer works
The set splits in two. 764 IC SWE tasks, worth $414,775, ask the model to do the engineering work itself: read the issue, change the code, and pass end-to-end tests that experienced engineers wrote and triple-checked. 724 SWE Manager tasks, worth $585,225, show the model several competing technical proposals for the same problem and ask it to pick the one the original hiring manager picked. A task pays out only on a full pass, so partial credit does not exist, and the headline figure is the sum of the payouts earned. OpenAI also released a public split, SWE-Lancer Diamond, holding 502 tasks worth $500,800, so results can be reproduced outside the lab.
History and current status
OpenAI published SWE-Lancer in February 2025 (arXiv 2502.12115), during the stretch when SWE-bench Verified was becoming the default coding citation and its numbers were climbing fast enough to be suspicious. The framing was deliberate. Rather than build yet another pass-rate set, the authors mapped capability onto money already spent in a real labour market, and shipped a Docker image plus the Diamond split so the result was not a closed internal number. The tasks come from the Expensify codebase, whose issues were publicly contracted on Upwork, which is what made the payout data available at all.
What the score does not tell you
Everything comes from one codebase, so the benchmark measures competence in a single large JavaScript and TypeScript product rather than software engineering in general. The prices reflect what one marketplace paid for one project, not the market value of the work. The manager tasks are effectively multiple choice, which is a far easier format than open-ended engineering and inflates the dollar total relative to what a model could really deliver. End-to-end tests resist gaming better than unit tests do, but a model that passes the tests has still only proven the tests pass. OpenAI does not maintain a live leaderboard, so published figures are pinned to the models of early 2025.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
SWE-Lancer is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.
What SWE-Lancer scores actually mean
In the paper, Claude 3.5 Sonnet earned $403k of the $1,000,000, o1 at high reasoning effort earned $380k, and GPT-4o earned $304k. Read the split and the picture changes completely. On IC SWE tasks, where the model writes the code, pass@1 was 21.1% for Claude 3.5 Sonnet, 20.3% for o1 and 8.6% for GPT-4o. On SWE Manager tasks, where it only chooses between proposals, the same models hit 47.0%, 46.3% and 38.7%. Models are roughly twice as good at judging engineering work as at doing it, and because the manager tasks carry $585,225 of the pot, a model can bank 40% of the money while solving only one coding task in five. A dollar total near $400k therefore does not mean the model did 40% of the engineering. It means it was a passable reviewer and a poor implementer.
Who reports SWE-Lancer, and how to read it
Almost nobody quotes it, and that absence is the interesting part. Labs report SWE-bench Verified in launch posts because a percentage that goes up reads well; a dollar figure showing the model left two thirds of the money on the table does not. The numbers that circulate come from third-party boards such as the llm-stats SWE-Lancer and IC-Diamond pages rather than from model cards. When a vendor tells you its agent can replace contract engineering work, SWE-Lancer is the benchmark to ask about by name, because it is the one denominated in the same unit as the claim.
When to weight SWE-Lancer in a model choice
Reach for SWE-Lancer when you are converting a capability claim into a budget, which is the point at which pass rates stop being useful. The dollar metric answers the question a buyer actually has, namely what fraction of a contract engineering spend a model could absorb, and the IC versus manager gap tells you where to put the human. What it cannot tell you is where the frontier sits today: the published results predate every model currently shipping, and no maintained leaderboard has replaced them. Treat it as a measuring instrument and a shape of result, not as a ranking, and pair it with a current agentic set such as Terminal-Bench for the capability question.
Benchmarks to read alongside this one
SWE-bench Verified
The same real-GitHub-issue resolution task as SWE-bench, restricted to a human-validated subset where the issue is solvable and the tests are not broken.
Terminal-Bench
Whether an AI agent can complete hard, realistic command-line tasks (build, configure, train, debug, secure) end to end inside a real terminal.
GDPval
Whether a model can produce the actual deliverables of skilled professional work (documents, slides, spreadsheets, diagrams) well enough to stand against an industry expert's version.
SWE-Lancer: frequently asked questions
- What is SWE-Lancer?
- SWE-Lancer is a 2025 OpenAI benchmark of 1,488 real freelance software engineering jobs from Upwork, together worth $1,000,000 in payouts that were actually paid to human engineers. Models are scored on the dollars they earn rather than on a pass rate.
- How much money have AI models earned on SWE-Lancer?
- In the paper, Claude 3.5 Sonnet earned $403k of the $1,000,000, o1 at high reasoning effort earned $380k, and GPT-4o earned $304k. On the public Diamond split, worth $500,800, the same three earned $208k, $166k and $139k.
- Why do models score so much higher on the manager tasks?
- SWE Manager tasks ask the model to pick between existing technical proposals, which is a multiple-choice problem. IC SWE tasks ask it to write code that passes engineer-written end-to-end tests. Pass rates run roughly 47% against 21%, so models judge engineering work about twice as well as they perform it.
Sources
- Miserendino, Wang, Patwardhan, Heidecke (2025). SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? arXiv preprint 2502.12115. arxiv.org/abs/2502.12115
- OpenAI (2025). Introducing the SWE-Lancer benchmark. openai.com/index/swe-lancer
- OpenAI. SWELancer-Benchmark reference implementation and Diamond split. github.com/openai/SWELancer-Benchmark
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →