One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learning
paperYour notes
Polymath learning asks how far reinforcement learning can go with a single, carefully engineered training problem. Starting from Qwen2.5-7B-Base, the authors repeat one sample across a batch of 128 and run 140 steps of GRPO with verifiable binary rewards. A math problem selected for broad, transferable skills improves reasoning not only in mathematics but in physics, chemistry, biology, engineering, computer science, and less math-adjacent subjects.
The strongest engineered problem, Synthetic Prime, combines biology, chemistry, and physics with a wide spectrum of mathematical skills. It reaches a 30.8 average score across the paper's eight subject groups, versus 25.0 for LIMR training on more than 1,000 selected samples and 19.5 for the full MATH training set. The result motivates "sample engineering": designing unusually rich RL problems instead of only scaling dataset size. Joint work from Shanghai Jiao Tong University's GAIR lab and Alibaba's Taobao & Tmall Group; code and training data are open.