"Mercury: Ultra-Fast Language Models Based on Diffusion." Introduces diffusion-based LLMs (dLLMs) that forecast multiple tokens simultaneously via iterative denoising, rather than sequential autoregressive generation. Mercury Coder Mini achieves 1,109 tokens/sec on H100.

Mercury 2 (February 2026) adds reasoning capability with AA Intelligence Index v4.3: 14 at ~929 tok/s — roughly 10x faster than comparable autoregressive models at similar quality. A genuinely novel architecture paradigm. By Khanna, Kharbanda, Li, Ermon, Grover, Kuleshov et al.

Provenance: initialization is undisclosed (closed weights, no parameter counts), but the tech report frames Mercury as extending the founders' MDLM from-scratch masked-diffusion pretraining line, "trained on the order of trillions of tokens" of web + proprietary data, and cites no AR-to-diffusion adaptation work — consistent with an in-house pretrain rather than an open-checkpoint conversion.

Model Details

Context window 128,000
AA Intelligence 14 was 12 on v4.3

Variants

Name Parameters Notes
Mercury 2.5 — Sep 8 2026: >1,100 tokens/s in production, 260K context (from 128K), $0.20/$0.75 per M tokens ($0.04/$0.15 launch promo); +10 points over Mercury 2, pitched against GPT-5.6 Luna, Gemini 3.5 Flash-Lite, Claude Haiku 4.5; tunable reasoning, native and parallel tool use, JSON mode; API, OpenRouter, Baseten; Mercury Voice (<170 ms TTFT) and Mercury Router previewed; AA Intelligence Index v4.3.2: 12, two below Mercury 2
Mercury Voice — Generally available Sep 29 2026 for enterprise customers: a diffusion LLM tuned for voice agents; 320 ms median time to first answer token (p95 750 ms) on production voice prompts; low/medium/high reasoning; 128K context, 50K output; $0.40 / $1.50 per M tokens (half price at launch)
Mercury Coder — —
Mercury 2 — Reasoning, AA Intelligence Index v4.3: 14, 929 tok/s

Paper

Citations 1
foundationalreasoningefficiency