Asks whether on-policy distillation (OPD) is worth running before reinforcement learning even when it barely changes a model's accuracy, and finds that it is. Under shared GRPO-style RL settings, the paper compares three ways to prepare the same student: RL straight from the base model, supervised fine-tuning on a teacher's solutions followed by RL, and OPD (the student samples its own responses and the teacher supplies next-token distributions along them) followed by RL. In text, with a Qwen3-8B student and an RL-trained Qwen3-32B teacher, OPD lifts the eight-benchmark average only from 36.48 to 37.02, 0.54 points above the base model, yet after RL it finishes at 59.18, 7.28 points above direct RL and 11.70 above SFT-then-RL, and ahead of both on all eight benchmarks. SFT starts higher (37.75) but ends lower (47.48). On AIME24, OPD starts slightly below the base model (24.17 against 24.90) and reaches 57.71 after RL, against 47.40 from the base. With a Qwen3-VL-8B student and a Qwen3-VL-32B teacher the margins are smaller: 63.02 over twenty benchmarks after RL, 1.34 points above direct RL and 1.69 above SFT-then-RL.

Coverage does not explain the gain. Pre-RL Pass@k, the diagnostic behind recent RL-readiness work, fails to predict the post-RL ranking: similar or higher Pass@k does not guarantee a better result after RL. The authors point instead to alignment with the teacher's full distribution beyond top-1 agreement, which may favor good reasoning paths while keeping alternatives that RL can refine. The preferred divergence also flips: with student-generated trajectories, reverse-KL OPD is ahead before RL (60.66 against 60.17) but forward KL finishes higher (64.87 against 63.02) at the same Pass@64 (92.68% for both), while with teacher-generated trajectories reverse KL stays ahead at both stages. The recommendation is to choose distillation settings by performance after the RL that follows, not by the distilled model's own scores. The paper belongs to the stage-aware RL-readiness line of TailSFT, and it differs from Sequential Beats Joint, which also finds OPD followed by RL beats RL alone but explains the ordering through pass@k coverage. Of the 29 authors, 18 list the untracked JD.COM, including project leader Jiaqi Wang. The first author lists Fudan, SII and JD.COM, the corresponding authors are Zhongyu Wei (Fudan and SII) and Siyuan Wang (CUHK), and other authors come from Peking University, HKUST, SJTU and elsewhere. No code has been released.

Paper

Authors: Shuai Dong · Yongfu Zhu · Yuqi Xu · Weichu Xie · Liuwenpu · Ziyue Wang · Kaiwen Tuo · Congcong Wang
rlrl-scalingdistillationpost-trainingreasoningresearch

Related