Closes the sparse-supervision gap in agentic RL: completed trajectories are converted into reusable hindsight skills, then distilled back into the policy as dense token-level on-policy signal — self-evolving supervision without an external teacher. Jianhua Tao's Tsinghua group with ZJU/CUHK/NTU/Tongji. Companion to the July 2026 agentic-distillation cluster (Microsoft ReOPD, NVIDIA Molt, OpenForgeRL).

Paper

post-trainingagentsresearch