The advantage is recomputed and redistributed within a group of rollouts of the same query to more clearly separate better and worse responses.
Naive advantage estimates within a group can misallocate credit, leading to unstable or weak learning.