The "open lab" David Hall and Percy Liang started at Stanford's CRFM in 2024 and, since 2025, have developed at Open Athena with Stanford: fully open models (weights, data, code and tokenizer, Apache 2.0) pretrained from scratch in JAX through the Levanter trainer, on TPUs from Google's TPU Research Cloud, with every experiment run as a public GitHub issue and pull request. Marin 8B (May 2025, 12.75T tokens) matched or beat Llama 3.1 8B on 16 of 19 base-model evaluations.

Marin 32B Base (October 2025, about 6.4T tokens) scaled the 8B recipe. It began on preemptible TPU v5p slices and finished on a reserved v4-2048; loss spikes around step 70,000 were cured mid-run by switching to the Qwen3 32B shape, which adds QK-Norm, warm-started from the run's own weights. On average it beat OLMo 2 32B Base, the previous best fully open base model, by 2 points (better on 14 of 19 tasks) and roughly matched the open-weight Gemma 3 27B PT. Its successors at Open Athena include the Delphi scaling suite, the Datakit data pipeline, the Snowball 67B-A2B MoE and a 535B-total, 23B-active MoE that began pretraining in August 2026.

Model Details

Architecture DENSE
Parameters 32B
License Apache 2.0

Variants

Name Parameters Notes
Marin 8B 8B May 2025; 12.75T tokens; Llama-style; matched or beat Llama 3.1 8B on 16 of 19 base evals
Marin 32B 32B October 2025; about 6.4T tokens; Qwen3-32B shape with QK-Norm; better than OLMo 2 32B Base on 14 of 19
open-weightpretraininginfrastructure

Related