Robots Atlas>ROBOTS ATLAS
Evaluation

PIQA

2020ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
A physical commonsense reasoning benchmark: the model picks the more sensible way to achieve an everyday physical goal out of two options.
Category
Evaluation
Abstraction level
Primitive
Operation level
Evaluation (runtime)
Use cases
Evaluating LLM physical commonsenseA standard benchmark in eval suitesStudying affordances and world knowledgeComparing open-weights models

How it works

The model is given a goal and two candidate procedures; the score is whether it deems the correct one more likely/appropriate (e.g. by comparing likelihoods or answering). The metric is accuracy on the validation/test set; human performance is near 95%.

Problem solved

Text-trained models poorly understand the physical world and object affordances. PIQA measures this gap in physical commonsense reasoning.

Components

Goal + two optionsExample structure

Task format: a physical goal and a pair of candidate solutions.

AccuracyScoring

Metric: share of correctly chosen options.