QCon at SemEval-2023 Task 10: Data Augmentation and Model Ensembling for Detection of Online Sexism

Weston Feely, Prabhakar Gupta, Manas Ranjan Mohanty, Timothy Chon, Tuhin Kundu, Vijit Singh, Sandeep Atluri, Tanya Roosta, Viviane Ghaderi, Peter Schulam


Abstract
The web contains an abundance of user- generated content. While this content is useful for many applications, it poses many challenges due to the presence of offensive, biased, and overall toxic language. In this work, we present a system that identifies and classifies sexist content at different levels of granularity. Using transformer-based models, we explore the value of data augmentation, use of ensemble methods, and leverage in-context learning using foundation models to tackle the task. We evaluate the different components of our system both quantitatively and qualitatively. Our best systems achieve an F1 score of 0.84 for the binary classification task aiming to identify whether a given content is sexist or not and 0.64 and 0.47 for the two multi-class tasks that aim to identify the coarse and fine-grained types of sexism present in the given content respectively.
Anthology ID:
2023.semeval-1.175
Volume:
Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023)
Month:
July
Year:
2023
Address:
Toronto, Canada
Editors:
Atul Kr. Ojha, A. Seza Doğruöz, Giovanni Da San Martino, Harish Tayyar Madabushi, Ritesh Kumar, Elisa Sartori
Venue:
SemEval
SIG:
SIGLEX
Publisher:
Association for Computational Linguistics
Note:
Pages:
1260–1270
Language:
URL:
https://aclanthology.org/2023.semeval-1.175
DOI:
10.18653/v1/2023.semeval-1.175
Bibkey:
Cite (ACL):
Weston Feely, Prabhakar Gupta, Manas Ranjan Mohanty, Timothy Chon, Tuhin Kundu, Vijit Singh, Sandeep Atluri, Tanya Roosta, Viviane Ghaderi, and Peter Schulam. 2023. QCon at SemEval-2023 Task 10: Data Augmentation and Model Ensembling for Detection of Online Sexism. In Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023), pages 1260–1270, Toronto, Canada. Association for Computational Linguistics.
Cite (Informal):
QCon at SemEval-2023 Task 10: Data Augmentation and Model Ensembling for Detection of Online Sexism (Feely et al., SemEval 2023)
Copy Citation:
PDF:
https://aclanthology.org/2023.semeval-1.175.pdf