In each training episode the simulator parameters are sampled from chosen distributions; the real world becomes 'just another sample' from that distribution for the model, improving robustness and generalization without needing a faithful model of reality.
Models trained only in simulation usually perform poorly in reality due to the reality gap in the appearance and dynamics of the environment.