ML-GSAI/Diffusion-LLM-Papers logo

ML-GSAI/Diffusion-LLM-Papers

Free

Curated papers on diffusion language models — LLaDA, Dream, MMaDA, consistency sampling, fast inference; 169 stars, actively maintained (2026) ![](https://img.shields.io/github/stars/ML-GSAI/Diffusion-LLM-Papers?style=flat-square)

FreeFree tier
Type
Open Source

About ML-GSAI/Diffusion-LLM-Papers

A curated collection of papers on diffusion language models, maintained by the ML-GSAI community. The repository categorizes research into theoretical basis, foundation models (e.g., LLaDA, Dream, Seed Diffusion), multimodal models (including understanding, unified, and speech/ASR), fast sampling techniques with KV caching and advanced sampling methods, and reinforcement learning for diffusion LLMs. It aims to provide a structured reference for researchers and practitioners, with contributions accepted via pull requests.

Key Features

Categorizes papers into theoretical basis, foundation models, multimodal models, fast sampling, and reinforcement learning
Includes seminal and recent papers such as LLaDA, Dream 7B, MMaDA, Dimple, and d1
Community-driven via pull requests for contributions
Active GitHub repository with 180 stars and 10 forks
Papers listed with titles, authors, and publication dates

Pros & Cons

Pros
  • Comprehensive collection of diffusion language model papers in one place
  • Well-organized categories for easy navigation
  • Actively maintained with recent papers (up to August 2025)
  • Community-driven, allowing contributions from the research community
Cons
  • May not cover all diffusion LLM papers due to time constraints (acknowledged by maintainers)
  • Only includes papers up to the last commit date; requires manual curation
  • Requires GitHub familiarity to browse and contribute

Best For

Research reference for diffusion language modelsStaying up-to-date with the latest developments in diffusion LLMsLiterature review for academic or industry projectsEducation and learning about discrete diffusion, masked diffusion, and related techniques

FAQ

How can I contribute to this collection?
You can submit a pull request to add papers that are missing from the list.
What topics are covered in this repository?
The repository covers theoretical basis, foundation models, multimodal models (understanding, unified, speech/ASR), fast sampling via KV caching and advanced sampling, and reinforcement learning for diffusion language models.