Speculative Decoding with a Speculative Vocabulary

Miles Williams; Young D. Kwon; Rui Li; Alexandros Kouris; Stylianos I. Venieris

Speculative Decoding with a Speculative Vocabulary

Miles Williams, Young D. Kwon, Rui Li, Alexandros Kouris, Stylianos I. Venieris

Abstract

Speculative decoding has rapidly emerged as a leading approach for accelerating language model (LM) inference, as it offers substantial speedups while yielding identical outputs. This relies upon a small draft model, tasked with predicting the outputs of the target model. State-of-the-art speculative decoding methods use a draft model comprising a single decoder layer and output embedding matrix, with the latter dominating drafting time for the latest LMs. Recent work has sought to address this output distribution bottleneck by reducing the vocabulary of the draft model. While this can improve throughput, it compromises speculation effectiveness when the target token is out-of-vocabulary. In this paper, we argue for vocabulary speculation as an alternative to a reduced vocabulary. We propose SpecVocab, an efficient and effective method that selects a vocabulary subset per decoding step. Across a variety of tasks, we show that SpecVocab can achieve a higher acceptance length than state-of-the-art speculative decoding method, EAGLE-3. Notably, this yields up to an 8.1% increase in average throughput over EAGLE-3.

Anthology ID:: 2026.findings-acl.2000
Volume:: Findings of the Association for Computational Linguistics: ACL 2026
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 40240–40254
Language:
URL:: https://aclanthology.org/2026.findings-acl.2000/
DOI:
Bibkey:
Cite (ACL):: Miles Williams, Young D. Kwon, Rui Li, Alexandros Kouris, and Stylianos I. Venieris. 2026. Speculative Decoding with a Speculative Vocabulary. In Findings of the Association for Computational Linguistics: ACL 2026, pages 40240–40254, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Speculative Decoding with a Speculative Vocabulary (Williams et al., Findings 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.findings-acl.2000.pdf
Checklist:: 2026.findings-acl.2000.checklist.pdf

PDF Cite Search Checklist Fix data