Moving beyond word lists: towards abstractive topic labels for human-like topics of scientific documents

Domenic Rosati


Abstract
Topic models represent groups of documents as a list of words (the topic labels). This work asks whether an alternative approach to topic labeling can be developed that is closer to a natural language description of a topic than a word list. To this end, we present an approach to generating human-like topic labels using abstractive multi-document summarization (MDS). We investigate our approach with an exploratory case study. We model topics in citation sentences in order to understand what further research needs to be done to fully operationalize MDS for topic labeling. Our case study shows that in addition to more human-like topics there are additional advantages to evaluation by using clustering and summarization measures instead of topic model measures. However, we find that there are several developments needed before we can design a well-powered study to evaluate MDS for topic modeling fully. Namely, improving cluster cohesion, improving the factuality and faithfulness of MDS, and increasing the number of documents that might be supported by MDS. We present a number of ideas on how these can be tackled and conclude with some thoughts on how topic modeling can also be used to improve MDS in general.
Anthology ID:
2022.wiesp-1.11
Volume:
Proceedings of the first Workshop on Information Extraction from Scientific Publications
Month:
November
Year:
2022
Address:
Online
Editors:
Tirthankar Ghosal, Sergi Blanco-Cuaresma, Alberto Accomazzi, Robert M. Patton, Felix Grezes, Thomas Allen
Venue:
WIESP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
91–99
Language:
URL:
https://aclanthology.org/2022.wiesp-1.11
DOI:
Bibkey:
Cite (ACL):
Domenic Rosati. 2022. Moving beyond word lists: towards abstractive topic labels for human-like topics of scientific documents. In Proceedings of the first Workshop on Information Extraction from Scientific Publications, pages 91–99, Online. Association for Computational Linguistics.
Cite (Informal):
Moving beyond word lists: towards abstractive topic labels for human-like topics of scientific documents (Rosati, WIESP 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.wiesp-1.11.pdf