Robots Atlas>ROBOTS ATLAS
Evaluation

HellaSwag

2019ActivePublished: 24 August 2026Updated: 24 August 2026Published
Key innovation
By scaling up Adversarial Filtering, it produced a commonsense NLI benchmark that is trivial for humans (>95%) yet very hard for then-SOTA models (<48%), showing benchmarks should co-evolve adversarially with models.
Category
Evaluation
Abstraction level
Pattern
Operation level
Evaluation (runtime)Post-training
Use cases
Evaluating commonsense reasoning of LLMsBenchmarking language modelsComponent of evaluation suites (lm-evaluation-harness, Open LLM Leaderboard)Measuring commonsense NLIComparing models in few-shot / zero-shot settings

How it works

Each example has a context (ctx, composed of ctx_a and ctx_b) and a list of four candidate endings; exactly one is correct (label field 0–3). The correct ending comes from the original source (ActivityNet or WikiHow), while distractors are machine-generated and selected via Adversarial Filtering: a discriminator is trained to tell human from machine text, generated examples that fool the model are kept, and the process iterates until the set is hard for models yet easy for humans (verified by annotators). A model is scored by accuracy in picking the correct ending; evaluation covers both in-domain and zero-shot categories (activities unseen in training).

Problem solved

Earlier commonsense benchmarks (e.g. SWAG) were quickly saturated by pretrained models like BERT, making it hard to measure real commonsense reasoning. HellaSwag constructs hard, adversarial examples that remain easy for humans, restoring the human–model gap as a measure of progress.

Components

Context and four endingsMultiple-choice task format

Task format: a context (ctx_a + ctx_b) and four candidate endings, one of which is correct (label 0–3).

Adversarial FilteringDistractor generation mechanism

Data-collection paradigm: an ensemble of discriminators iteratively selects machine-generated wrong answers that are hard for models yet ridiculous to humans.

Context sources (ActivityNet, WikiHow)Source data domains

Contexts come from ActivityNet Captions video descriptions and WikiHow how-to articles; they include in-domain and zero-shot categories.

Evolution

Original paper · 2019 · ACL 2019 · Rowan Zellers
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, Yejin Choi
2018
SWAG — predecessor of HellaSwag

Zellers et al. introduce SWAG, a commonsense NLI benchmark built with Adversarial Filtering.

2019
HellaSwag published (ACL 2019)
Inflection point

A harder successor to SWAG: longer contexts and examples from ActivityNet and WikiHow; humans >95%, SOTA models <48%.

2023
GPT-4 reaches 95.3% — benchmark saturation
Inflection point

GPT-4 (March 2023) reaches 95.3%, near the human level of 95.6%, effectively saturating the task.

Hyperparameters (configurable axes)

Number of candidate endingsHigh

Each example has 4 possible endings, one of which is correct.

4Fixed value in HellaSwag
Evaluation regime (zero-shot / few-shot)Medium

Commonly evaluated zero-shot or 10-shot in suites like lm-evaluation-harness.

10-shotSetting used on the Open LLM Leaderboard