Exploring Stylometric and Emotion-Based Features for Multilingual Cross-Domain Hate Speech Detection

Ilia Markov, Nikola Ljubešić, Darja Fišer, Walter Daelemans


Abstract
In this paper, we describe experiments designed to evaluate the impact of stylometric and emotion-based features on hate speech detection: the task of classifying textual content into hate or non-hate speech classes. Our experiments are conducted for three languages – English, Slovene, and Dutch – both in in-domain and cross-domain setups, and aim to investigate hate speech using features that model two linguistic phenomena: the writing style of hateful social media content operationalized as function word usage on the one hand, and emotion expression in hateful messages on the other hand. The results of experiments with features that model different combinations of these phenomena support our hypothesis that stylometric and emotion-based features are robust indicators of hate speech. Their contribution remains persistent with respect to domain and language variation. We show that the combination of features that model the targeted phenomena outperforms words and character n-gram features under cross-domain conditions, and provides a significant boost to deep learning models, which currently obtain the best results, when combined with them in an ensemble.
Anthology ID:
2021.wassa-1.16
Volume:
Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis
Month:
April
Year:
2021
Address:
Online
Editors:
Orphee De Clercq, Alexandra Balahur, Joao Sedoc, Valentin Barriere, Shabnam Tafreshi, Sven Buechel, Veronique Hoste
Venue:
WASSA
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
149–159
Language:
URL:
https://aclanthology.org/2021.wassa-1.16
DOI:
Bibkey:
Cite (ACL):
Ilia Markov, Nikola Ljubešić, Darja Fišer, and Walter Daelemans. 2021. Exploring Stylometric and Emotion-Based Features for Multilingual Cross-Domain Hate Speech Detection. In Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pages 149–159, Online. Association for Computational Linguistics.
Cite (Informal):
Exploring Stylometric and Emotion-Based Features for Multilingual Cross-Domain Hate Speech Detection (Markov et al., WASSA 2021)
Copy Citation:
PDF:
https://aclanthology.org/2021.wassa-1.16.pdf