How to Build a Diffusion Language Model
![]()
A July 2026 tutorial from Volodymyr Kuleshov's group (Kuleshov, Marianne Arriola, Yair Schiff and Guanghan Wang), adapted from talks at ICLR 2026 and MLSS 2026, builds diffusion language models up from image diffusion. It starts with masked diffusion (MDLM, "a generative BERT"), then covers block diffusion for variable length, encoder-decoder designs, remasking and uniform-state noise for error correction, distillation for speed, guidance and RL post-training. The authors write that diffusion models "became competitive with autoregressive models on quality" in 2024, though not yet scaled as far, and survey today's models, from LLaDA and Mercury to Gemma Diffusion and Nemotron Diffusion. Read the praise of Mercury's speed knowing that the page gives Kuleshov's affiliation as Cornell University and Inception, Mercury's maker. The tutorial ends on the view that diffusion may be to inference-time and post-training scaling what the transformer was to RNNs for pre-training.
Was this useful?