DiffuGPT / DiffuLLaMA
paper Your tags
Your notes
The AR→diffusion adaptation recipe (ICLR 2025): shows autoregressive checkpoints (GPT-2, LLaMA-2 7B) can be continually adapted into discrete diffusion language models with modest compute (under 200B tokens), by connecting the two objectives rather than pretraining diffusion models from scratch. DiffuGPT and DiffuLLaMA scaled diffusion LMs to 7B at a time when open diffusion LLMs barely existed.
The direct methodological seed of hkunlp's Dream family (Dream 7B, Dream-Coder/DreamOn). hkunlp-led (Lingpeng Kong's group) with multi-institution co-authors.