The entire RL budget is spent in one carefully designed policy-training pass rather than spread across many iterations.
Multiple RL rounds for LLMs are computationally expensive and unstable (policy drift, entropy collapse).