Evaluation
PIQA
2020ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key
innovation
A physical commonsense reasoning benchmark: the model picks the more sensible way to achieve an everyday physical goal out of two options.
Category
Evaluation
Abstraction level
Primitive
Operation level
Evaluation (runtime)
Use cases
Evaluating LLM physical commonsenseA standard benchmark in eval suitesStudying affordances and world knowledgeComparing open-weights models
How it works
The model is given a goal and two candidate procedures; the score is whether it deems the correct one more likely/appropriate (e.g. by comparing likelihoods or answering). The metric is accuracy on the validation/test set; human performance is near 95%.
Problem solved
Text-trained models poorly understand the physical world and object affordances. PIQA measures this gap in physical commonsense reasoning.
Components
Goal + two optionsExample structure
Task format: a physical goal and a pair of candidate solutions.
AccuracyScoring
Metric: share of correctly chosen options.
Evolution
Original paper · 2020 · Yonatan Bisk
PIQA: Reasoning about Physical Commonsense in Natural Language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, Yejin Choi
2020
PIQA published (Bisk et al., AAAI 2020)
Inflection point