SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
paper Your tags
Your notes
Closes the sparse-supervision gap in agentic RL: completed trajectories are converted into reusable hindsight skills, then distilled back into the policy as dense token-level on-policy signal — self-evolving supervision without an external teacher. Jianhua Tao's Tsinghua group with ZJU/CUHK/NTU/Tongji. Companion to the July 2026 agentic-distillation cluster (Microsoft ReOPD, NVIDIA Molt, OpenForgeRL).