Analyzing Offensive Language Dataset Insights from Training Dynamics and Human Agreement Level

Do Kyung Kim, Hyeseon Ahn, Youngwook Kim, Yo-Sub Han


Abstract
Implicit hate speech detection is challenging due to its subjectivity and context dependence, with existing models often struggling in outof-domain scenarios. We propose CONELA, a novel data refinement strategy that enhances model performance and generalization by integrating human annotation agreement with model training dynamics. By removing both easy and hard instances from the model’s perspective, while also considering whether humans agree or disagree and retaining ambiguous cases crucial for out-of-distribution generalization, CONELA consistently improves performance across multiple datasets and models. We also observe significant improvements in F1 scores and cross-domain generalization with the use of our CONELA strategy. Addressing data scarcity in smaller datasets, we introduce a weighted loss function and an ensemble strategy incorporating disagreement maximization, effectively balancing learning from limited data. Our findings demonstrate that refining datasets by integrating both model and human perspectives significantly enhances the effectiveness and generalization of implicit hate speech detection models. This approach lays a strong foundation for future research on dataset refinement and model robustness.
Anthology ID:
2025.coling-main.653
Volume:
Proceedings of the 31st International Conference on Computational Linguistics
Month:
January
Year:
2025
Address:
Abu Dhabi, UAE
Editors:
Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, Steven Schockaert
Venue:
COLING
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
9780–9792
Language:
URL:
https://aclanthology.org/2025.coling-main.653/
DOI:
Bibkey:
Cite (ACL):
Do Kyung Kim, Hyeseon Ahn, Youngwook Kim, and Yo-Sub Han. 2025. Analyzing Offensive Language Dataset Insights from Training Dynamics and Human Agreement Level. In Proceedings of the 31st International Conference on Computational Linguistics, pages 9780–9792, Abu Dhabi, UAE. Association for Computational Linguistics.
Cite (Informal):
Analyzing Offensive Language Dataset Insights from Training Dynamics and Human Agreement Level (Kim et al., COLING 2025)
Copy Citation:
PDF:
https://aclanthology.org/2025.coling-main.653.pdf