A Large-scale Comprehensive Abusiveness Detection Dataset with Multifaceted Labels from Reddit

Hoyun Song; Soo Hyun Ryu; Huije Lee; Jong C. Park

doi:10.18653/v1/2021.conll-1.43

A Large-scale Comprehensive Abusiveness Detection Dataset with Multifaceted Labels from Reddit

Hoyun Song, Soo Hyun Ryu, Huije Lee, Jong Park

Abstract

As users in online communities suffer from severe side effects of abusive language, many researchers attempted to detect abusive texts from social media, presenting several datasets for such detection. However, none of them contain both comprehensive labels and contextual information, which are essential for thoroughly detecting all kinds of abusiveness from texts, since datasets with such fine-grained features demand a significant amount of annotations, leading to much increased complexity. In this paper, we propose a Comprehensive Abusiveness Detection Dataset (CADD), collected from the English Reddit posts, with multifaceted labels and contexts. Our dataset is annotated hierarchically for an efficient annotation through crowdsourcing on a large-scale. We also empirically explore the characteristics of our dataset and provide a detailed analysis for novel insights. The results of our experiments with strong pre-trained natural language understanding models on our dataset show that our dataset gives rise to meaningful performance, assuring its practicality for abusive language detection.

Anthology ID:: 2021.conll-1.43
Volume:: Proceedings of the 25th Conference on Computational Natural Language Learning
Month:: November
Year:: 2021
Address:: Online
Editors:: Arianna Bisazza, Omri Abend
Venue:: CoNLL
SIG:: SIGNLL
Publisher:: Association for Computational Linguistics
Note:
Pages:: 552–561
Language:
URL:: https://aclanthology.org/2021.conll-1.43
DOI:: 10.18653/v1/2021.conll-1.43
Bibkey:
Cite (ACL):: Hoyun Song, Soo Hyun Ryu, Huije Lee, and Jong Park. 2021. A Large-scale Comprehensive Abusiveness Detection Dataset with Multifaceted Labels from Reddit. In Proceedings of the 25th Conference on Computational Natural Language Learning, pages 552–561, Online. Association for Computational Linguistics.
Cite (Informal):: A Large-scale Comprehensive Abusiveness Detection Dataset with Multifaceted Labels from Reddit (Song et al., CoNLL 2021)
Copy Citation:
PDF:: https://aclanthology.org/2021.conll-1.43.pdf
Video:: https://aclanthology.org/2021.conll-1.43.mp4
Code: nlpcl-lab/cadd_dataset
Data: Hate Speech

PDF Cite Search Code Video