Towards Efficient and Effective Diffusion Language Model Inference via Semantic-Aware Adaptive Denoising

Fan Li; Yu Gu (谷峪); Zhigang Wang; Fangling Leng; Zhenghao Liu (刘正皓); Ge Yu (于戈)

Towards Efficient and Effective Diffusion Language Model Inference via Semantic-Aware Adaptive Denoising

Fan Li, Yu Gu, Zhigang Wang, Fangling Leng, Zhenghao Liu, Ge Yu

Abstract

Diffusion language models (DLMs) have emerged as a powerful non-autoregressive alternative to GPT-style sequential generation, but suffer from substantial computational overhead due to their iterative parallel denoising. Existing acceleration works cannot accurately detect semantically stabilized tokens and then skip computation, leading to sub-optimal speedup in practice. This paper presents the first systematic study of convergence dynamics in DLMs. Innovative observations include the misalignment between traditionally used scalar detection criterion and the semantic convergence, and the post-peak confidence score, that wastes denoising computation and degrades inference quality. To address these limitations, we propose Ada-DLM, a semantic-aware adaptive denoising framework that encodes the trajectory of scalar confidence scores into an evolution-aware feature vector and then clusters vectors proactively and adaptively identify semantically converged tokens. Furthermore, we incorporate system-level optimizations to maximize runtime efficiency. Experiments show that Ada-DLM consistently outperforms the SOTA competitor, achieving up to 2x speedup and 19% quality improvement. That offers a practical path toward efficient high-quality DLM deployment.

Anthology ID:: 2026.acl-long.819
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 17988–18002
Language:
URL:: https://aclanthology.org/2026.acl-long.819/
DOI:
Bibkey:
Cite (ACL):: Fan Li, Yu Gu, Zhigang Wang, Fangling Leng, Zhenghao Liu, and Ge Yu. 2026. Towards Efficient and Effective Diffusion Language Model Inference via Semantic-Aware Adaptive Denoising. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 17988–18002, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Towards Efficient and Effective Diffusion Language Model Inference via Semantic-Aware Adaptive Denoising (Li et al., ACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.acl-long.819.pdf
Checklist:: 2026.acl-long.819.checklist.pdf

PDF Cite Search Checklist Fix data