Search Query Spell Correction with Weak Supervision in E-commerce

Vishal Kakkar, Chinmay Sharma, Madhura Pande, Surender Kumar


Abstract
Misspelled search queries in e-commerce can lead to empty or irrelevant products. Besides inadvertent typing mistakes, most spell mistakes occur because the user does not know the correct spelling, hence typing it as it is pronounced colloquially. This colloquial typing creates countless misspelling patterns for a single correct query. In this paper, we first systematically analyze and group different spell errors into error classes and then leverage the state-of-the-art Transformer model for contextual spell correction. We overcome the constraint of limited human labelled data by proposing novel synthetic data generation techniques for voluminous generation of training pairs needed by data hungry Transformers, without any human intervention. We further utilize weakly supervised data coupled with curriculum learning strategies to improve on tough spell mistakes without regressing on the easier ones. We show significant improvements from our model on human labeled data and online A/B experiments against multiple state-of-art models.
Anthology ID:
2023.acl-industry.66
Volume:
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Industry Track)
Month:
July
Year:
2023
Address:
Toronto, Canada
Editors:
Sunayana Sitaram, Beata Beigman Klebanov, Jason D Williams
Venue:
ACL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
687–694
Language:
URL:
https://aclanthology.org/2023.acl-industry.66
DOI:
10.18653/v1/2023.acl-industry.66
Bibkey:
Cite (ACL):
Vishal Kakkar, Chinmay Sharma, Madhura Pande, and Surender Kumar. 2023. Search Query Spell Correction with Weak Supervision in E-commerce. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Industry Track), pages 687–694, Toronto, Canada. Association for Computational Linguistics.
Cite (Informal):
Search Query Spell Correction with Weak Supervision in E-commerce (Kakkar et al., ACL 2023)
Copy Citation:
PDF:
https://aclanthology.org/2023.acl-industry.66.pdf