Capital & Compute

WebArena

Benchmark· Agents, tool use & computer use· Checked 2026-09-12

WebArena puts an agent in front of fully functional, self-hosted copies of the kinds of sites people use daily, an online store, a forum, a GitLab instance, a CMS, plus a map and a wiki, and asks it to finish realistic multi-step tasks. Nothing is simulated and nothing is scored on whether the agent looked like it succeeded: a reward function checks the resulting state of the website.

Key facts about the WebArena benchmark
What it measuresWhether an autonomous agent can complete long-horizon, realistic web tasks (navigation, forms, multi-step workflows) in fully functional self-hosted websites.
Built byCarnegie Mellon University (Zhou, Xu et al.), 2023
Format812 long-horizon tasks across self-hosted sites: e-commerce, a social forum, GitLab, a CMS, plus a map and a wiki
Scoring metricFunctional success rate via execution-based reward checking the end state
StatusActive
Representative top score74.3% (third-party tracker) · WebTactix on DeepSeek v3.2 (a system, not a bare model) · read 2026-06
Official leaderboardleaderboard.steel.dev/leaderboards/webarena

How WebArena works

The benchmark holds 812 long-horizon tasks across those domains, phrased as natural-language instructions of the kind a person would actually issue. Evaluation is execution-based: rather than compare the agent's transcript to a reference trajectory, the harness checks functional correctness by inspecting the end state, so any route to the correct outcome counts and a plausible-looking path that changed nothing does not. Because every site is self-hosted and reset between runs, the environment is reproducible and cannot drift under the benchmark the way a live website would, and the agent cannot reach anything outside the sandbox.

History and current status

Zhou, Xu, Zhu, Neubig and colleagues at Carnegie Mellon published WebArena in July 2023 (arXiv 2307.13854), at a point when web-agent demos were circulating widely and none of them had a reproducible score. The paper's own result set the terms of the next three years: the best GPT-4-based agent finished 14.41% of tasks against 78.24% for humans. That 64-point gap became the standard citation for how far autonomous web agents had to go, and drove a wave of follow-on environments including VisualWebArena, WorkArena and WebChoreArena.

What the score does not tell you

The scores are harness scores, not model scores. A WebArena result reflects the model plus its scaffolding, prompting strategy, memory and action space, and published entries vary enormously on all four, so two numbers from different papers are rarely comparable and the leaderboards that aggregate them are comparing systems rather than models. The self-hosted sites are frozen snapshots, which removes contamination from live web drift but also removes the messiness that makes real browsing hard. Task instructions are templated in places, so a system can learn the phrasing rather than the task. And because the environment has been public since 2023, its pages and structure are in training data.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

WebArena is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.

What WebArena scores actually mean

The useful reading is the trajectory rather than any single figure. The 2023 paper baseline was 14.41% for the best GPT-4 agent, with humans at 78.24%. By 2026 the tracked leaders sit close to that human line: as aggregated by the Steel.dev leaderboard, last updated 29 June 2026, WebTactix on DeepSeek v3.2 reads 74.3% and OpAgent reads 71.6%, with several systems in the high sixties. Five points below the human rate is a very different world from sixty-four below, and it is the fastest closure of any agentic gap in this directory. The caveat is that human performance was measured once, in 2023, on the same frozen sites, so the ceiling is a 2023 measurement being chased by 2026 systems, and the last few points are as likely to be harness engineering as capability.

Who reports WebArena, and how to read it

Labs mostly do not; agent framework builders and academic groups do, through the official site and through third-party trackers such as the Steel.dev WebArena leaderboard, which is a secondary aggregation rather than an official board. Entries are named after systems, not models, for the reason above. Because submissions are self-reported and the harness is not fixed, treat a headline WebArena percentage as a claim about one team's whole stack on one day, and look for whether the run used the not-achievable hint, which materially changes the number.

When to weight WebArena in a model choice

Use WebArena when you need evidence that browser automation is viable for a workflow, because it is the most realistic reproducible environment available and the execution-based reward means a pass is a genuine state change. Read the system name rather than the model name, and if you are choosing a model rather than a framework, WebArena is the wrong instrument. Pair it with tau-bench if reliability matters more than peak success, since WebArena reports single-run success and says nothing about whether the same task passes eight times in a row, which is the property that decides whether an agent can be deployed.

2023
First released
Carnegie Mellon University (Zhou, Xu et al.)
Active
Status today
As of October 1, 2026
74.3% (third-party tracker)
Representative top score
Read 2026-06

Benchmarks to read alongside this one

WebArena: frequently asked questions

What is WebArena?
WebArena is a 2023 benchmark from Carnegie Mellon that runs autonomous agents against 812 long-horizon tasks on fully functional self-hosted websites covering e-commerce, a social forum, GitLab and a CMS, plus a map and a wiki. Success is checked by inspecting the end state of the site.
What is a good WebArena score?
The 2023 paper baseline was 14.41% for the best GPT-4 agent against 78.24% for humans. As aggregated by the Steel.dev leaderboard in mid-2026, leading systems read 74.3% and 71.6%, so a competitive score now sits in the high sixties or above.
Does WebArena measure the model or the agent framework?
The framework, mostly. A WebArena entry is a whole system: model plus scaffolding, prompting, memory and action space. Published runs vary on all of those, so two numbers from different papers are rarely comparable and entries are named after systems rather than models.

Sources

  • Zhou, Xu, Zhu, Zhou, Lo, Sridhar et al. (2023). WebArena: A Realistic Web Environment for Building Autonomous Agents. arXiv preprint 2307.13854. arxiv.org/abs/2307.13854
  • Carnegie Mellon University. WebArena environment, task set and official site. webarena.dev
  • Steel.dev. WebArena leaderboard, a third-party aggregation of self-reported system runs, last updated 29 June 2026. leaderboard.steel.dev/leaderboards/webarena

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 117 benchmarks in the directoryWhich benchmarks are worth trusting →