Effective and Efficient Query-aware Snippet Extraction for Web Search

Jingwei Yi; Fangzhao Wu; Chuhan Wu; Xiaolong Huang; Binxing Jiao; Guangzhong Sun; Xing Xie

doi:10.18653/v1/2022.emnlp-main.197

Effective and Efficient Query-aware Snippet Extraction for Web Search

Jingwei Yi, Fangzhao Wu, Chuhan Wu, Xiaolong Huang, Binxing Jiao, Guangzhong Sun, Xing Xie

Abstract

Query-aware webpage snippet extraction is widely used in search engines to help users better understand the content of the returned webpages before clicking. The extracted snippet is expected to summarize the webpage in the context of the input query. Existing snippet extraction methods mainly rely on handcrafted features of overlapping words, which cannot capture deep semantic relationships between the query and webpages. Another idea is to extract the sentences which are most relevant to queries as snippets with existing text matching methods. However, these methods ignore the contextual information of webpages, which may be sub-optimal. In this paper, we propose an effective query-aware webpage snippet extraction method named DeepQSE. In DeepQSE, the concatenation of title, query and each candidate sentence serves as an input of query-aware sentence encoder, aiming to capture the fine-grained relevance between the query and sentences. Then, these query-aware sentence representations are modeled jointly through a document-aware relevance encoder to capture contextual information of the webpage. Since the query and each sentence are jointly modeled in DeepQSE, its online inference may be slow. Thus, we further propose an efficient version of DeepQSE, named Efficient-DeepQSE, which can significantly improve the inference speed of DeepQSE without affecting its performance. The core idea of Efficient-DeepQSE is to decompose the query-aware snippet extraction task into two stages, i.e., a coarse-grained candidate sentence selection stage where sentence representations can be cached, and a fine-grained relevance modeling stage. Experiments on two datasets validate the effectiveness and efficiency of our methods.

Anthology ID:: 2022.emnlp-main.197
Volume:: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing
Month:: December
Year:: 2022
Address:: Abu Dhabi, United Arab Emirates
Editors:: Yoav Goldberg, Zornitsa Kozareva, Yue Zhang
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 3035–3046
Language:
URL:: https://aclanthology.org/2022.emnlp-main.197/
DOI:: 10.18653/v1/2022.emnlp-main.197
Bibkey:
Cite (ACL):: Jingwei Yi, Fangzhao Wu, Chuhan Wu, Xiaolong Huang, Binxing Jiao, Guangzhong Sun, and Xing Xie. 2022. Effective and Efficient Query-aware Snippet Extraction for Web Search. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3035–3046, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
Cite (Informal):: Effective and Efficient Query-aware Snippet Extraction for Web Search (Yi et al., EMNLP 2022)
Copy Citation:
PDF:: https://aclanthology.org/2022.emnlp-main.197.pdf

PDF Cite Search Fix data