Capital & Compute

HellaSwag

Benchmark· Knowledge & general QA· Checked 2026-09-12

HellaSwag is a commonsense sentence-completion set from 2019, built so that humans score above 95% and the best models of that year scored below 48%. It is in this directory as history rather than as a live signal. Its arc, from a benchmark designed to be unbeatable to one nobody reports, is the saturation cycle compressed into a single entry.

Key facts about the HellaSwag benchmark
What it measuresCommonsense sentence completion: picking the plausible continuation of an everyday scenario.
Built byZellers, Holtzman, Bisk et al., 2019
FormatMultiple-choice continuations built by adversarial filtering, trivial for humans (above 95%) and hard for 2019 models (below 48%)
Scoring metricAccuracy
StatusRetired
Representative top scoreNot independently confirmed; see the leaderboard

How HellaSwag works

Each item gives a short everyday scenario, such as a woman sitting down at a piano, and four possible continuations, of which one actually happened. Scoring is plain multiple-choice accuracy. What made it hard was the construction method, Adversarial Filtering: a series of discriminator models iteratively select machine-generated wrong answers that fool the current state of the art while remaining obviously absurd to a person. The authors described targeting a Goldilocks zone of length and complexity where generated text is ridiculous to humans yet routinely misclassified by machines, and the filtering runs until that gap is as wide as it can be made.

History and current status

Zellers, Holtzman, Bisk, Farhadi and Choi published HellaSwag in May 2019 (arXiv 1905.07830), as a direct response to BERT having closed out their earlier SWAG dataset within a year. For the next several years it was a standard row on almost every model card and one of the six benchmarks in the original Hugging Face Open LLM Leaderboard. That leaderboard was archived in June 2024 and replaced by a v2 built on harder evaluations specifically because the v1 set, HellaSwag included, had saturated. Open LLM Leaderboard v2 itself stopped taking submissions in March 2025.

What the score does not tell you

Adversarial Filtering has a structural flaw that only shows up later: it tunes the wrong answers against a specific generation of models, so the difficulty is relative to 2019 systems rather than absolute, and it decays as architectures change. The dataset also carries a known share of mislabelled and ambiguous items, which puts a practical ceiling below 100% and makes the last few points meaningless. Because the answers have been public since 2019 they are in every large training corpus, so a modern score is a memorisation measurement as much as a commonsense one. None of that was a mistake in 2019; it is what happens to any static adversarial set given enough time.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.

HellaSwag is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.

What HellaSwag scores actually mean

HellaSwag has no useful band left at the top. Frontier models exceed the roughly 95% human accuracy the paper measured, and once a model passes the human rate on a human-labelled multiple-choice set, the remaining gap is mostly dataset noise rather than capability. The band that still means something is far below: a model scoring in the seventies or low eighties is genuinely weaker at everyday inference, which is why the benchmark persists in small-model evaluation. Set the 2019 numbers beside today's to see the shape of the cycle: under 48% for state of the art at release, above 95% within a few years, and retired from the flagship public leaderboard in 2024.

Who reports HellaSwag, and how to read it

Effectively nobody at the frontier. It vanished from major model cards as scores crossed the human rate, and the leaderboard that had carried it into public view was retired. Hugging Face said it stopped the v2 board because it was becoming obsolete and could encourage hill climbing in irrelevant directions, which is a plain statement of the underlying problem. If you see HellaSwag quoted in 2026 it is almost always in a small-model or fine-tune release, where it is cheap to run and produces a flattering number.

When to weight HellaSwag in a model choice

Do not use it to choose a model, and be sceptical of any 2026 release that leads with it. It has two remaining uses. It is a reasonable sanity check for very small or heavily quantised models, where the score still moves with real capability. And it is the cleanest teaching example of why a static benchmark expires: the adversarial construction that made it hard was defined against the models of one year, and the answers have been public ever since. When you are deciding how much weight to put on any benchmark in this directory, ask how HellaSwag looked in 2019 and how it looks now.

2019
First released
Zellers, Holtzman, Bisk et al.
Retired
Status today
As of October 1, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

HellaSwag: frequently asked questions

What is HellaSwag?
HellaSwag is a 2019 commonsense benchmark that asks a model to pick the plausible continuation of an everyday scenario from four options. It was built with Adversarial Filtering so that humans scored above 95% while the best models of 2019 scored below 48%.
Why is HellaSwag no longer used?
Frontier models passed the human accuracy rate, so the benchmark stopped separating them. The Hugging Face Open LLM Leaderboard, which carried it, was archived in June 2024 and replaced with harder evaluations for exactly that reason.
What is Adversarial Filtering?
It is a dataset construction method where a series of discriminator models iteratively select machine-generated wrong answers that fool the current state of the art while staying obviously wrong to a human. It makes a set hard relative to one generation of models, which is why the difficulty decays.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 117 benchmarks in the directoryWhich benchmarks are worth trusting →