Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties

Nhi Pham, Lachlan Pham, Adam Meyers


Abstract
The prevalence of social media presents a growing opportunity to collect and analyse examples of English varieties. Whilst usage of these varieties is often used only in spoken contexts or hard-to-access private messages, social media sites like Twitter provide a platform for users to communicate informally in a scrapeable format. Notably, Indian English (Hinglish), Singaporean English (Singlish), and African-American English (AAE) can be commonly found online. These varieties pose a challenge to existing natural language processing (NLP) tools as they often differ orthographically and syntactically from standard English for which the majority of these tools are built. NLP models trained on standard English texts produced biased outcomes for users of underrepresented varieties (Blodgett and O’Connor, 2017). Some research has aimed to overcome the inherent biases caused by unrepresentative data through techniques like data augmentation or adjusting training models. We aim to address the issue of bias at its root - the data itself. We curate a dataset of tweets from countries with high proportions of underserved English variety speakers, and propose an annotation framework of six categorical classifications along a pseudo-spectrum that measures the degree of standard English and that thereby indirectly aims to surface the manifestations of English varieties in these tweets.
Anthology ID:
2024.law-1.6
Volume:
Proceedings of The 18th Linguistic Annotation Workshop (LAW-XVIII)
Month:
March
Year:
2024
Address:
St. Julians, Malta
Editors:
Sophie Henning, Manfred Stede
Venues:
LAW | WS
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
61–70
Language:
URL:
https://aclanthology.org/2024.law-1.6
DOI:
Bibkey:
Cite (ACL):
Nhi Pham, Lachlan Pham, and Adam Meyers. 2024. Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties. In Proceedings of The 18th Linguistic Annotation Workshop (LAW-XVIII), pages 61–70, St. Julians, Malta. Association for Computational Linguistics.
Cite (Informal):
Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties (Pham et al., LAW-WS 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.law-1.6.pdf