IndEuph-170: Benchmarking Cultural Pragmatics through Euphemism Detection in Indian English

Debamita Samajdar


Abstract
Large Language Models (LLMs) have shown remarkable proficiency in standard English benchmarks, yet their ability to navigate the sociopragmatic cues of non-Western English varieties remains underexplored. This paper introduces IndEuph-170, a novel benchmark dataset focused on Indian English (IndE) euphemisms — expressions whose roots lie in local social hierarchies, politeness norms, and cultural taboos (e.g., "setting," "loose character," "suitable boy"). IndEuph-170 comprises 170 curated IndE sentences, against which the performance of two distinct architectures was evaluated: a fine-tuned BART model and GPT-4. The findings reveal a significant "cultural gap". While GPT-4 achieves 82.5% accuracy, it struggles with authoritative and punitive nuances. BART achieves 55.3% accuracy but exhibits a high rate of false positives by over-classifying general Indianisms as euphemisms. The paper argues that current multilingual benchmarks such as MME (Fu et al., 2025) and GLUE (Wang et al., 2018) fail to capture these dialectal pragmatics, and that a culturally-aware evaluation framework for Global Englishes is necessary.
Anthology ID:
2026.wildre-1.12
Volume:
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Month:
May
Year:
2026
Address:
Palma, Mallorca, Spain
Editors:
Girish Nath Jha, Kalika Bali, Sobha L, Devendr Kumar
Venues:
WILDRE | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
93–97
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-wildre-12
DOI:
10.63317/3ecgg3aq3dbk
Bibkey:
Cite (ACL):
Debamita Samajdar. 2026. IndEuph-170: Benchmarking Cultural Pragmatics through Euphemism Detection in Indian English. In Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation, pages 93–97, Palma, Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
IndEuph-170: Benchmarking Cultural Pragmatics through Euphemism Detection in Indian English (Samajdar, WILDRE 2026)
Copy Citation: