Introduces Direct On-Policy Distillation, a weak-to-strong post-training method that transfers the change induced by reinforcement learning rather than imitating a weaker teacher's final policy. Direct-OPD compares a small teacher before and after RL, interprets their log-probability ratio as a dense token-level implicit reward, and applies that signal on the stronger student's own on-policy states.

Across two 1.5B teacher pairs and students from 1.7B to 7B, the method improves every tested student, including students already stronger than the post-RL teacher. It raises Qwen3-1.7B from 48.3% to 58.3% on AIME 2024 in about four hours on 8 A100 GPUs, and sequentially composing two independently learned policy shifts reaches 63.8%. The paper, training code, and resulting checkpoints come from the joint SIA-Lab of Tsinghua AIR and ByteDance Seed.

Paper

Authors: Shiyuan Feng · Huan-ang Gao · Haohan Chi · Hanlin Wu · Zhilong Zhang · Zheng Jiang · Bingxiang He · Wei-Ying Ma · Ya-Qin Zhang · Hao Zhou
post-trainingreasoningdistillationreinforcement-learningweak-to-strong

Related