Marin's first open scaling suite, a modernized Pythia built at Open Athena: a scaling recipe that maps a compute budget to a model configuration, a suite of checkpoints trained with it on Google's TPU Research Cloud up to 1023 FLOPs, and a scaling law fit on seven IsoFLOP optima (3e18 to 3e20 FLOPs). A preregistered forecast from that law predicted the final loss of the largest run (1e23 FLOPs, 25B parameters, 600B tokens) within 0.2%, extrapolating 300× past the fit, and held-out runs at 1e21 and 1e22 landed within 0.5% across seeds. Getting there took a token-horizon correction to the learning rate and other hyperparameters and an optimizer that removed weight decay from the search, after a first recipe diverged 30× out. A two-step regression also forecast MMLU, HumanEval and GSM8K. Checkpoints, fit coefficients, the recipe and scripts that rebuild the Nemotron-CC, StarCoderData and ProofPile 2 training mixture are all public.

scalingscaling-lawsopen-weightresearch

Related