Formalizes cross-domain data synergy in pretraining mixtures (e.g. code→math gains): fits domain-to-benchmark and higher-order interaction terms that outperform standard scaling laws, and validates by training on predicted-optimal vs suboptimal mixtures. The data-mixture successor story for scaling-law work. Hamidieh (MIT CSAIL), Mackey (Microsoft Research), Alvarez-Melis (MSR + Harvard).

Paper

scalingtrainingresearch