Angular Dispersion Accelerates k-Nearest Neighbors Machine Translation

Evgeniia Tokarchuk; Sergey Troshin; Vlad Niculae

doi:10.18653/v1/2025.findings-emnlp.759

Angular Dispersion Accelerates k-Nearest Neighbors Machine Translation

Evgeniia Tokarchuk, Sergey Troshin, Vlad Niculae

Abstract

Augmenting neural machine translation with external memory at decoding time, in the form of k-nearest neighbors machine translation (k-NN MT), is a well-established strategy for increasing translation performance. k-NN MT retrieves a set of tokens that occurred in the most similar contexts recorded in a prepared data store, using hidden state representations of translation contexts as vector lookup keys. One of the main disadvantages of this method is the high computational cost and memory requirements. Since an exhaustive search is not feasible in large data stores practitioners commonly use approximate k-NN lookup, yet even such algorithms are a bottleneck. In contrast to research directions seeking to accelerate k-NN MT by reducing data store size or the number of lookup calls, we pursue an orthogonal direction based on the performance properties of approximate k-NN lookup data structures. In particular, we propose encouraging angular dispersion of the neural hidden representations of contexts. We show that improving dispersion leads to better balance in the retrieval data structures, accelerating retrieval and slightly improving translations.

Anthology ID:: 2025.findings-emnlp.759
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2025
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 14120–14132
Language:
URL:: https://aclanthology.org/2025.findings-emnlp.759/
DOI:: 10.18653/v1/2025.findings-emnlp.759
Bibkey:
Cite (ACL):: Evgeniia Tokarchuk, Sergey Troshin, and Vlad Niculae. 2025. Angular Dispersion Accelerates k-Nearest Neighbors Machine Translation. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 14120–14132, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: Angular Dispersion Accelerates k-Nearest Neighbors Machine Translation (Tokarchuk et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-emnlp.759.pdf
Checklist:: 2025.findings-emnlp.759.checklist.pdf

PDF Cite Search Checklist Fix data