Attention over Heads: A Multi-Hop Attention for Neural Machine Translation

Shohei Iida; Ryuichiro Kimura; Hongyi Cui; Po-Hsuan Hung; Takehito Utsuro; Masaaki Nagata

doi:10.18653/v1/P19-2030

Attention over Heads: A Multi-Hop Attention for Neural Machine Translation

Shohei Iida, Ryuichiro Kimura, Hongyi Cui, Po-Hsuan Hung, Takehito Utsuro, Masaaki Nagata

Abstract

In this paper, we propose a multi-hop attention for the Transformer. It refines the attention for an output symbol by integrating that of each head, and consists of two hops. The first hop attention is the scaled dot-product attention which is the same attention mechanism used in the original Transformer. The second hop attention is a combination of multi-layer perceptron (MLP) attention and head gate, which efficiently increases the complexity of the model by adding dependencies between heads. We demonstrate that the translation accuracy of the proposed multi-hop attention outperforms the baseline Transformer significantly, +0.85 BLEU point for the IWSLT-2017 German-to-English task and +2.58 BLEU point for the WMT-2017 German-to-English task. We also find that the number of parameters required for a multi-hop attention is smaller than that for stacking another self-attention layer and the proposed model converges significantly faster than the original Transformer.

Anthology ID:: P19-2030
Volume:: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop
Month:: July
Year:: 2019
Address:: Florence, Italy
Editors:: Fernando Alva-Manchego, Eunsol Choi, Daniel Khashabi
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 217–222
Language:
URL:: https://aclanthology.org/P19-2030/
DOI:: 10.18653/v1/P19-2030
Bibkey:
Cite (ACL):: Shohei Iida, Ryuichiro Kimura, Hongyi Cui, Po-Hsuan Hung, Takehito Utsuro, and Masaaki Nagata. 2019. Attention over Heads: A Multi-Hop Attention for Neural Machine Translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, pages 217–222, Florence, Italy. Association for Computational Linguistics.
Cite (Informal):: Attention over Heads: A Multi-Hop Attention for Neural Machine Translation (Iida et al., ACL 2019)
Copy Citation:
PDF:: https://aclanthology.org/P19-2030.pdf

PDF Cite Search Fix data