Marin's 67B-total, roughly 2B-active mixture-of-experts model, built in Marin's own "Grug MoE" architecture and pretrained from scratch on 10T tokens on TPU v4 by the Marin team at Open Athena. It has 26 layers, a hidden size of 2,560, grouped-query attention with 20 query and 5 key-value heads, and 256 routed experts with four selected per token plus a shared expert, over a 128,256-token vocabulary; router biases from Marin's quantile-balancing scheme are folded into the exported weights.

Context was then extended to 262,144 tokens in a 1,000-step run (steps 156,000 to 157,000), and on September 9, 2026 Open Athena published five BF16 base checkpoints from that comparison on HuggingFace under the OpenMDW-1.1 license: two attention-scaling settings (QK multiplier 1.57 and 1.75) and three that upsample long documents 2×, 4× or 8×. They are base models with no instruction tuning, and serving them needs Marin's vLLM fork, which registers the GrugMoE architecture. Supervised fine-tuning at 262K context and RLVR runs on top of Snowball followed as further public experimental checkpoints in September; the team, which also built an end-to-end post-training pipeline for it, says the full release is coming soon. Snowball sits between the dense Marin 32B and the 535B-total, 23B-active MoE that Marin began pretraining in August 2026.

Model Details

Architecture MOE
Parameters 67B
Active params 2B
Experts 256 (top-4)
Context window 262,144
Training tokens 10T
Training hardware TPU v4
License OpenMDW-1.1

Variants

Name Parameters Notes
snowball-67b-a2b-base-262k-qk157 — Naive long-context extension, QK multiplier 1.57
snowball-67b-a2b-base-262k-qk175 — Raised attention scale for 262K context (QK multiplier 1.75)
snowball-67b-a2b-base-262k-qk175-skew2 / skew4 / skew8 — QK 1.75 with long-context documents upsampled 2x, 4x or 8x during the extension
moeopen-weightpretraininglong-context

Related