Distiller: A Systematic Study of Model Distillation Methods in Natural Language Processing

Haoyu He; Xingjian Shi; Jonas Mueller; Sheng Zha; Mu Li; George Karypis

doi:10.18653/v1/2021.sustainlp-1.13

Distiller: A Systematic Study of Model Distillation Methods in Natural Language Processing

Haoyu He, Xingjian Shi, Jonas Mueller, Sheng Zha, Mu Li, George Karypis

Abstract

Knowledge Distillation (KD) offers a natural way to reduce the latency and memory/energy usage of massive pretrained models that have come to dominate Natural Language Processing (NLP) in recent years. While numerous sophisticated variants of KD algorithms have been proposed for NLP applications, the key factors underpinning the optimal distillation performance are often confounded and remain unclear. We aim to identify how different components in the KD pipeline affect the resulting performance and how much the optimal KD pipeline varies across different datasets/tasks, such as the data augmentation policy, the loss function, and the intermediate representation for transferring the knowledge between teacher and student. To tease apart their effects, we propose Distiller, a meta KD framework that systematically combines a broad range of techniques across different stages of the KD pipeline, which enables us to quantify each component’s contribution. Within Distiller, we unify commonly used objectives for distillation of intermediate representations under a universal mutual information (MI) objective and propose a class of MI-objective functions with better bias/variance trade-off for estimating the MI between the teacher and the student. On a diverse set of NLP datasets, the best Distiller configurations are identified via large-scale hyper-parameter optimization. Our experiments reveal the following: 1) the approach used to distill the intermediate representations is the most important factor in KD performance, 2) among different objectives for intermediate distillation, MI-performs the best, and 3) data augmentation provides a large boost for small training datasets or small student networks. Moreover, we find that different datasets/tasks prefer different KD algorithms, and thus propose a simple AutoDistiller algorithm that can recommend a good KD pipeline for a new dataset.

Anthology ID:: 2021.sustainlp-1.13
Volume:: Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing
Month:: November
Year:: 2021
Address:: Virtual
Editors:: Nafise Sadat Moosavi, Iryna Gurevych, Angela Fan, Thomas Wolf, Yufang Hou, Ana Marasović, Sujith Ravi
Venue:: sustainlp
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 119–133
Language:
URL:: https://aclanthology.org/2021.sustainlp-1.13/
DOI:: 10.18653/v1/2021.sustainlp-1.13
Bibkey:
Cite (ACL):: Haoyu He, Xingjian Shi, Jonas Mueller, Sheng Zha, Mu Li, and George Karypis. 2021. Distiller: A Systematic Study of Model Distillation Methods in Natural Language Processing. In Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing, pages 119–133, Virtual. Association for Computational Linguistics.
Cite (Informal):: Distiller: A Systematic Study of Model Distillation Methods in Natural Language Processing (He et al., sustainlp 2021)
Copy Citation:
PDF:: https://aclanthology.org/2021.sustainlp-1.13.pdf
Video:: https://aclanthology.org/2021.sustainlp-1.13.mp4

PDF Cite Search Video Fix data