Joel Birrer

2024

pdf bib abs
The Lou Dataset - Exploring the Impact of Gender-Fair Language in German Text Classification
Andreas Waldis | Joel Birrer | Anne Lauscher | Iryna Gurevych
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing

Gender-fair language, an evolving linguistic variation in German, fosters inclusion by addressing all genders or using neutral forms. However, there is a notable lack of resources to assess the impact of this language shift on language models (LMs) might not been trained on examples of this variation. Addressing this gap, we present Lou, the first dataset providing high-quality reformulations for German text classification covering seven tasks, like stance detection and toxicity classification. We evaluate 16 mono- and multi-lingual LMs and find substantial label flips, reduced prediction certainty, and significantly altered attention patterns. However, existing evaluations remain valid, as LM rankings are consistent across original and reformulated instances. Our study provides initial insights into the impact of gender-fair language on classification for German. However, these findings are likely transferable to other languages, as we found consistent patterns in multi-lingual and English LMs.

Co-authors

Venues

emnlp1

Fix data