Tsinghua's SEED: self-evolving distillation teaches AI agents
A Tsinghua team introduced SEED — self-evolving on-policy distillation with hindsight supervision. On ALFWorld it improves success by 45.9 points over GRPO and matches it using 60% of the data.