Bootstrapping a Music Voice Assistant with Weak Supervision

Sergio Oramas, Massimo Quadrana, Fabien Gouyon


Abstract
One of the first building blocks to create a voice assistant relates to the task of tagging entities or attributes in user queries. This can be particularly challenging when entities are in the tenth of millions, as is the case of e.g. music catalogs. Training slot tagging models at an industrial scale requires large quantities of accurately labeled user queries, which are often hard and costly to gather. On the other hand, voice assistants typically collect plenty of unlabeled queries that often remain unexploited. This paper presents a weakly-supervised methodology to label large amounts of voice query logs, enhanced with a manual filtering step. Our experimental evaluations show that slot tagging models trained on weakly-supervised data outperform models trained on hand-annotated or synthetic data, at a lower cost. Further, manual filtering of weakly-supervised data leads to a very significant reduction in Sentence Error Rate, while allowing us to drastically reduce human curation efforts from weeks to hours, with respect to hand-annotation of queries. The method is applied to successfully bootstrap a slot tagging system for a major music streaming service that currently serves several tens of thousands of daily voice queries.
Anthology ID:
2021.naacl-industry.7
Volume:
Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Papers
Month:
June
Year:
2021
Address:
Online
Editors:
Young-bum Kim, Yunyao Li, Owen Rambow
Venue:
NAACL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
49–55
Language:
URL:
https://aclanthology.org/2021.naacl-industry.7
DOI:
10.18653/v1/2021.naacl-industry.7
Bibkey:
Cite (ACL):
Sergio Oramas, Massimo Quadrana, and Fabien Gouyon. 2021. Bootstrapping a Music Voice Assistant with Weak Supervision. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Papers, pages 49–55, Online. Association for Computational Linguistics.
Cite (Informal):
Bootstrapping a Music Voice Assistant with Weak Supervision (Oramas et al., NAACL 2021)
Copy Citation:
PDF:
https://aclanthology.org/2021.naacl-industry.7.pdf
Video:
 https://aclanthology.org/2021.naacl-industry.7.mp4