Data-Efficient Paraphrase Generation to Bootstrap Intent Classification and Slot Labeling for New Features in Task-Oriented Dialog Systems

Shailza Jolly; Tobias Falke; Caglar Tirkaz; Daniil Sorokin

doi:10.18653/v1/2020.coling-industry.2

Data-Efficient Paraphrase Generation to Bootstrap Intent Classification and Slot Labeling for New Features in Task-Oriented Dialog Systems

Shailza Jolly, Tobias Falke, Caglar Tirkaz, Daniil Sorokin

Abstract

Recent progress through advanced neural models pushed the performance of task-oriented dialog systems to almost perfect accuracy on existing benchmark datasets for intent classification and slot labeling. However, in evolving real-world dialog systems, where new functionality is regularly added, a major additional challenge is the lack of annotated training data for such new functionality, as the necessary data collection efforts are laborious and time-consuming. A potential solution to reduce the effort is to augment initial seed data by paraphrasing existing utterances automatically. In this paper, we propose a new, data-efficient approach following this idea. Using an interpretation-to-text model for paraphrase generation, we are able to rely on existing dialog system training data, and, in combination with shuffling-based sampling techniques, we can obtain diverse and novel paraphrases from small amounts of seed data. In experiments on a public dataset and with a real-world dialog system, we observe improvements for both intent classification and slot labeling, demonstrating the usefulness of our approach.

Anthology ID:: 2020.coling-industry.2
Volume:: Proceedings of the 28th International Conference on Computational Linguistics: Industry Track
Month:: December
Year:: 2020
Address:: Online
Editors:: Ann Clifton, Courtney Napoles
Venue:: COLING
SIG:
Publisher:: International Committee on Computational Linguistics
Note:
Pages:: 10–20
Language:
URL:: https://aclanthology.org/2020.coling-industry.2/
DOI:: 10.18653/v1/2020.coling-industry.2
Bibkey:
Cite (ACL):: Shailza Jolly, Tobias Falke, Caglar Tirkaz, and Daniil Sorokin. 2020. Data-Efficient Paraphrase Generation to Bootstrap Intent Classification and Slot Labeling for New Features in Task-Oriented Dialog Systems. In Proceedings of the 28th International Conference on Computational Linguistics: Industry Track, pages 10–20, Online. International Committee on Computational Linguistics.
Cite (Informal):: Data-Efficient Paraphrase Generation to Bootstrap Intent Classification and Slot Labeling for New Features in Task-Oriented Dialog Systems (Jolly et al., COLING 2020)
Copy Citation:
PDF:: https://aclanthology.org/2020.coling-industry.2.pdf

PDF Cite Search Fix data