For a group of rollouts of the same query, the reward is computed collectively (within-group comparisons/statistics) rather than separately per sample.
Rewards assigned to individual samples in isolation poorly discriminate response quality for the same query.