Summarizing Community-based Question-Answer Pairs

Ting-Yao Hsu, Yoshi Suhara, Xiaolan Wang


Abstract
Community-based Question Answering (CQA), which allows users to acquire their desired information, has increasingly become an essential component of online services in various domains such as E-commerce, travel, and dining. However, an overwhelming number of CQA pairs makes it difficult for users without particular intent to find useful information spread over CQA pairs. To help users quickly digest the key information, we propose the novel CQA summarization task that aims to create a concise summary from CQA pairs. To this end, we first design a multi-stage data annotation process and create a benchmark dataset, COQASUM, based on the Amazon QA corpus. We then compare a collection of extractive and abstractive summarization methods and establish a strong baseline approach DedupLED for the CQA summarization task. Our experiment further confirms two key challenges, sentence-type transfer and deduplication removal, towards the CQA summarization task. Our data and code are publicly available.
Anthology ID:
2022.emnlp-main.250
Volume:
Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing
Month:
December
Year:
2022
Address:
Abu Dhabi, United Arab Emirates
Editors:
Yoav Goldberg, Zornitsa Kozareva, Yue Zhang
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
3798–3808
Language:
URL:
https://aclanthology.org/2022.emnlp-main.250
DOI:
10.18653/v1/2022.emnlp-main.250
Bibkey:
Cite (ACL):
Ting-Yao Hsu, Yoshi Suhara, and Xiaolan Wang. 2022. Summarizing Community-based Question-Answer Pairs. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3798–3808, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
Cite (Informal):
Summarizing Community-based Question-Answer Pairs (Hsu et al., EMNLP 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.emnlp-main.250.pdf