Polymath learning asks how far reinforcement learning can go with a single, carefully engineered training problem. Starting from Qwen2.5-7B-Base, the authors repeat one sample across a batch of 128 and run 140 steps of GRPO with verifiable binary rewards. A math problem selected for broad, transferable skills improves reasoning not only in mathematics but in physics, chemistry, biology, engineering, computer science, and less math-adjacent subjects.

The strongest engineered problem, Synthetic Prime, combines biology, chemistry, and physics with a wide spectrum of mathematical skills. It reaches a 30.8 average score across the paper's eight subject groups, versus 25.0 for LIMR training on more than 1,000 selected samples and 19.5 for the full MATH training set. The result motivates "sample engineering": designing unusually rich RL problems instead of only scaling dataset size. Joint work from Shanghai Jiao Tong University's GAIR lab and Alibaba's Taobao & Tmall Group; code and training data are open.

Paper

Authors: Yiyuan Li · Zhen Huang · Yanan Wu · Weixun Wang · Xuefeng Li · Yijia Luo · Wenbo Su · Bo Zheng · Pengfei Liu
reinforcement-learningreasoningdata-efficiencypost-trainingresearch