The model solves 500 math problems; the score is the fraction of final answers matching the reference solution.
The full MATH set is large and costly to evaluate; a smaller, representative subset was needed.