BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study

Atharva Mutsaddi; Anvi Jamkhande; Aryan Shirish Thakre; Yashodhara Haribhakta

BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study

Atharva Mutsaddi, Anvi Jamkhande, Aryan Shirish Thakre, Yashodhara Haribhakta

Abstract

As short text data in native languages like Hindi increasingly appear in modern media, robust methods for topic modeling on such data have gained importance. This study investigates the performance of BERTopic in modeling Hindi short texts, an area that has been under-explored in existing research. Using contextual embeddings, BERTopic can capture semantic relationships in data, making it potentially more effective than traditional models, especially for short and diverse texts. We evaluate BERTopic using 6 different document embedding models and compare its performance against 8 established topic modeling techniques, such as Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorization (NMF), Latent Semantic Indexing (LSI), Additive Regularization of Topic Models (ARTM), Probabilistic Latent Semantic Analysis (PLSA), Embedded Topic Model (ETM), Combined Topic Model (CTM), and Top2Vec. The models are assessed using coherence scores across a range of topic counts. Our results reveal that BERTopic consistently outperforms other models in capturing coherent topics from short Hindi texts.

Anthology ID:: 2025.indonlp-1.3
Volume:: Proceedings of the First Workshop on Natural Language Processing for Indo-Aryan and Dravidian Languages
Month:: January
Year:: 2025
Address:: Abu Dhabi
Editors:: Ruvan Weerasinghe, Isuri Anuradha, Deshan Sumanathilaka
Venues:: IndoNLP | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 22–32
Language:
URL:: https://aclanthology.org/2025.indonlp-1.3/
DOI:
Bibkey:
Cite (ACL):: Atharva Mutsaddi, Anvi Jamkhande, Aryan Shirish Thakre, and Yashodhara Haribhakta. 2025. BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study. In Proceedings of the First Workshop on Natural Language Processing for Indo-Aryan and Dravidian Languages, pages 22–32, Abu Dhabi. Association for Computational Linguistics.
Cite (Informal):: BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study (Mutsaddi et al., IndoNLP 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.indonlp-1.3.pdf

PDF Cite Search Fix data