AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset

Dante Everaert; Rohit Patki; Tianqi Zheng; Christopher Potts

doi:10.18653/v1/2024.emnlp-industry.78

AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset

Dante Everaert, Rohit Patki, Tianqi Zheng, Christopher Potts

Abstract

Query Autocomplete (QAC) is a critical feature in modern search engines, facilitating user interaction by predicting search queries based on input prefixes. Despite its widespread adoption, the absence of large-scale, realistic datasets has hindered advancements in QAC system development. This paper addresses this gap by introducing AmazonQAC, a new QAC dataset sourced from Amazon Search logs, comprising 395M samples. The dataset includes actual sequences of user-typed prefixes leading to final search terms, as well as session IDs and timestamps that support modeling the context-dependent aspects of QAC. We assess Prefix Trees, semantic retrieval, and Large Language Models (LLMs) with and without finetuning. We find that finetuned LLMs perform best, particularly when incorporating contextual information. However, even our best system achieves only half of what we calculate is theoretically possible on our test data, which implies QAC is a challenging problem that is far from solved with existing systems. This contribution aims to stimulate further research on QAC systems to better serve user needs in diverse environments. We open-source this data on Hugging Face at https://huggingface.co/datasets/amazon/AmazonQAC.

Anthology ID:: 2024.emnlp-industry.78
Volume:: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track
Month:: November
Year:: 2024
Address:: Miami, Florida, US
Editors:: Franck Dernoncourt, Daniel Preoţiuc-Pietro, Anastasia Shimorina
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1046–1055
Language:
URL:: https://aclanthology.org/2024.emnlp-industry.78/
DOI:: 10.18653/v1/2024.emnlp-industry.78
Bibkey:
Cite (ACL):: Dante Everaert, Rohit Patki, Tianqi Zheng, and Christopher Potts. 2024. AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 1046–1055, Miami, Florida, US. Association for Computational Linguistics.
Cite (Informal):: AmazonQAC: A Large-Scale, Naturalistic Query Autocomplete Dataset (Everaert et al., EMNLP 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.emnlp-industry.78.pdf

PDF Cite Search Fix data