Open Athena's genomic language models, developed in the open on the Marin platform (blog post by Gonzalo Benegas and Eric Czech). Instead of a genomics-specific architecture, MarinDNA treats DNA as text and trains a stock decoder-only Transformer (the Qwen3 architecture, from scratch, with a seven-token nucleotide vocabulary) and puts its effort into data curation, hyperparameters and scale. Training data starts from annotation-derived coding, upstream and downstream sequences across many species, later adding ncRNA and enhancers projected from human annotations through whole-genome alignments, mixed with uniform weights so the large coding-sequence set does not dominate. Hyperparameters tuned on ~25M-parameter models predicted the best learning rate at 255M, 476M and 1B, and validation loss followed a clean scaling law through 4B.

The released 1.12B m5.1 model (19 layers, 255-base context, about 166B nucleotide tokens) slightly leads Evo 2 40B in zero-shot macro-average AUPRC on Mendelian variant-effect prediction while using ~1,980× fewer training FLOPs and scoring variants ~2,330× faster, though alignment-based and supervised models remain stronger overall and zero-shot scores on Mendelian missense variants got worse as the models grew even as linear-probe results improved. Weights are Apache 2.0 on HuggingFace, with a scaling suite from 46M to 4B parameters (August 2026), more than 150 datasets and a public leaderboard; the repository keeps the unsuccessful and inconclusive experiments alongside the ones in the write-up.

Model Details

Architecture DENSE
Parameters 1.1B
Context window 256
Training tokens 166B
License Apache 2.0

Variants

Name Parameters Notes
MarinDNA m5.1 1.1B 1,120,772,224 parameters; 19 layers, hidden 1,920; final checkpoint at step 59,158; released with the August 3, 2026 blog post
MarinDNA scaling suite v0.5 — 46M, 76M, 128M, 255M, 476M, 1B, 2B and 4B checkpoints, August 6, 2026
sciencebiologygenomicsscientific-fmscaling-lawsopen-weight

Related