HasNat@CHiPSAL 2026: Multimodal Hate Speech Detection in Low-Resource Nepali Memes Using Aligned Vision–Language Models

Alvee Hasan Chowdhury, Md. Abul Hasnat, Adnan Faisal


Abstract
Memes are widely used for communication on social media but are increasingly exploited to spread hate and harmful stereotypes. Detecting hate speech in memes is particularly challenging because meaning is conveyed jointly through images and embedded text, and the problem becomes more complex in low-resource languages such as Nepali. In this work, we participate in Subtask A of the CHiPSAL 2026 Shared Task, focusing on hate speech detection in Nepali-only memes. We benchmark three multimodal vision language backbones, ViT-B-32 (OpenCLIP), AltCLIP, and BLIP2+mT5, under controlled preprocessing and augmentation settings. Our best-performing system uses AltCLIP to extract aligned text and image representations, followed by a late-fusion classifier trained with stratified 5-fold cross-validation to address class imbalance. The proposed model achieves a macro F1-score of 0.66 on the validation set. Experimental results highlight the effectiveness of aligned vision language representations and demonstrate that preprocessing and augmentation strategies have model-dependent effects in low-resource multimodal hate speech detection.
Anthology ID:
2026.chipsal-1.23
Volume:
Proceedings of the Second workshop on Challenges in Processing South Asian Languages (CHiPSAL2026)
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Kengatharaiyer Sarveswaran, Ashwini Vaidya
Venues:
CHiPSAL | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
237–243
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-chipsal-23
DOI:
10.63317/54xeeoceu6qz
Bibkey:
Cite (ACL):
Alvee Hasan Chowdhury, Md. Abul Hasnat, and Adnan Faisal. 2026. HasNat@CHiPSAL 2026: Multimodal Hate Speech Detection in Low-Resource Nepali Memes Using Aligned Vision–Language Models. In Proceedings of the Second workshop on Challenges in Processing South Asian Languages (CHiPSAL2026), pages 237–243, Palma de Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
HasNat@CHiPSAL 2026: Multimodal Hate Speech Detection in Low-Resource Nepali Memes Using Aligned Vision–Language Models (Chowdhury et al., CHiPSAL 2026)
Copy Citation: