The model picks the correct option among several for science questions; the score is accuracy on ARC-Easy and ARC-Challenge.
A QA benchmark requiring genuine reasoning and knowledge, not mere word matching, was needed.