A scaling study of on-policy distillation (OPD) as a way to move RL-acquired reasoning between model sizes, from Zhejiang University's OmniAI group in the ACES Lab and Kuaishou Technology (four of the eight authors; the first author and corresponding author Xuhong Zhang are at Zhejiang). The authors RL-train Qwen2.5 base models at 0.5B, 1.5B, 3B, 7B and 14B with GRPO on GSM8K and MATH, then distill between them in 25 teacher–student pairs covering weak-to-strong, same-base and strong-to-weak setups, all in verl. Borrowing the method of reward-model overoptimization studies, they track held-out accuracy (the "gold score") against d, the square root of the student's token-level reverse KL from its initialization. Every run opens with a useful-transfer regime in which the gold score rises roughly linearly in d, after which dynamics turn noisy (slower gains, saturation or regression). In every weak-to-strong pair the student's peak exceeds its teacher's own score, so a small RL-trained expert can transfer capability to a much larger model; for 7B and 14B students the peak rises with teacher size, from 72.7% to 81.9% and from 77.7% to 87.3%.

Fitted power laws in student size, teacher size and measured teacher score predict the peak and the early transfer rate. Conditioning on teacher score halves the leave-one-scale-out error of a size-only law, and the joint law extrapolates to the largest held-out student and teacher within 0.7 accuracy points (0.4 for the Delta-OPD variant). The laws say the peak improves with teacher scale only up to roughly the student's own size, and that at a matched score smaller teachers transfer better, so a teacher's score alone does not define its value as a supervisor. Delta-OPD, which rewards the change the teacher underwent during RL (the common core of Direct-OPD, W2S-OPD and OPD2), has a steeper transfer slope than vanilla OPD in 15 of 17 shared pairs, mostly in weak-to-strong settings. An off-policy SFT cold start on a 0.5B expert's outputs costs 6.6, 15.7 and 19.4 points for 3B, 7B and 14B students, and bootstrapping through a chain of intermediate teachers never beats direct OPD from the smallest expert. No code has been released.

Paper

Authors: Yuntai Bao · Qinfeng Li · Guoqing Jiang · Liwei Chen · Zhiheng Qin · Xuanping Li · Wenqi Zhang · Xuhong Zhang
distillationrlrl-scalingscaling-lawspost-trainingweak-to-strongresearch

Related