Topic Modeling for Maternal Health Using Reddit

Shuang Gao, Shivani Pandya, Smisha Agarwal, João Sedoc


Abstract
This paper applies topic modeling to understand maternal health topics, concerns, and questions expressed in online communities on social networking sites. We examine Latent Dirichlet Analysis (LDA) and two state-of-the-art methods: neural topic model with knowledge distillation (KD) and Embedded Topic Model (ETM) on maternal health texts collected from Reddit. The models are evaluated on topic quality and topic inference, using both auto-evaluation metrics and human assessment. We analyze a disconnect between automatic metrics and human evaluations. While LDA performs the best overall with the auto-evaluation metrics NPMI and Coherence, Neural Topic Model with Knowledge Distillation is favorable by expert evaluation. We also create a new partially expert annotated gold-standard maternal health topic
Anthology ID:
2021.louhi-1.8
Volume:
Proceedings of the 12th International Workshop on Health Text Mining and Information Analysis
Month:
April
Year:
2021
Address:
online
Venues:
EACL | Louhi
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
69–76
Language:
URL:
https://aclanthology.org/2021.louhi-1.8
DOI:
Bibkey:
Cite (ACL):
Shuang Gao, Shivani Pandya, Smisha Agarwal, and João Sedoc. 2021. Topic Modeling for Maternal Health Using Reddit. In Proceedings of the 12th International Workshop on Health Text Mining and Information Analysis, pages 69–76, online. Association for Computational Linguistics.
Cite (Informal):
Topic Modeling for Maternal Health Using Reddit (Gao et al., Louhi 2021)
Copy Citation:
PDF:
https://aclanthology.org/2021.louhi-1.8.pdf