Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting

Jan Fillies; Michael Peter Hoffmann; Rebecca Reichel; Roman Salzwedel; Sven Bodemer; Adrian Paschke

doi:10.18653/v1/2025.emnlp-main.948

Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting

Jan Fillies, Michael Peter Hoffmann, Rebecca Reichel, Roman Salzwedel, Sven Bodemer, Adrian Paschke

Abstract

A lack of demographic context in existing toxic speech datasets limits our understanding of how different age groups communicate online. In collaboration with funk, a German public service content network, this research introduces the first large-scale German dataset annotated for toxicity and enriched with platform-provided age estimates. The dataset includes 3,024 human-annotated and 30,024 LLM-annotated anonymized comments from Instagram, TikTok, and YouTube. To ensure relevance, comments were consolidated using predefined toxic keywords, resulting in 16.7% labeled as problematic. The annotation pipeline combined human expertise with state-of-the-art language models, identifying key categories such as insults, disinformation, and criticism of broadcasting fees. The dataset reveals age-based differences in toxic speech patterns, with younger users favoring expressive language and older users more often engaging in disinformation and devaluation. This resource provides new opportunities for studying linguistic variation across demographics and supports the development of more equitable and age-aware content moderation systems.

Anthology ID:: 2025.emnlp-main.948
Volume:: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 18763–18779
Language:
URL:: https://aclanthology.org/2025.emnlp-main.948/
DOI:: 10.18653/v1/2025.emnlp-main.948
Bibkey:
Cite (ACL):: Jan Fillies, Michael Peter Hoffmann, Rebecca Reichel, Roman Salzwedel, Sven Bodemer, and Adrian Paschke. 2025. Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 18763–18779, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting (Fillies et al., EMNLP 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.emnlp-main.948.pdf
Checklist:: 2025.emnlp-main.948.checklist.pdf

PDF Cite Search Checklist Fix data