Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

Xingyue Huang; Xueying Ding; Mingxuan Ju; Yozen Liu; Neil Shah; Tong Zhao

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

Xingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu, Neil Shah, Tong Zhao

Abstract

Softmax attention struggles with long contexts due to structural limitations: the strict sum-to-one constraint forces attention sinks on irrelevant tokens, while probability mass disperses as sequence lengths increase. We tackle these problems with Threshold Differential Attention (TDA), a sink-free attention mechanism that achieves ultra-sparsity and improved robustness at longer sequence lengths without the computational overhead of projection methods or the performance degradation caused by noise accumulation of standard rectified attention. TDA applies row-wise extreme-value thresholding with a length-dependent gate, retaining only exceedances. Inspired by the differential transformer, TDA also subtracts an inhibitory view to enhance expressivity. Theoretically, we prove that TDA controls the expected number of spurious survivors per row to O(1) and that consensus spurious matches across independent views vanish as context grows. Empirically, TDA produces >99 % exact zeros and eliminates attention sinks while maintaining competitive performance on standard and long-context benchmarks.

Anthology ID:: 2026.acl-long.824
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 18073–18092
Language:
URL:: https://aclanthology.org/2026.acl-long.824/
DOI:
Bibkey:
Cite (ACL):: Xingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu, Neil Shah, and Tong Zhao. 2026. Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 18073–18092, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling (Huang et al., ACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.acl-long.824.pdf
Checklist:: 2026.acl-long.824.checklist.pdf

PDF Cite Search Checklist Fix data