Proceedings of the the fifth edition of NLPerspectives
Shiran Dudy, Gavin Abercrombie, Valerio Basile, Elisa Leonardelli, Simona Frenda (Editors)
- Anthology ID:
- 2026.nlperspectives-1
- Month:
- May
- Year:
- 2026
- Address:
- Palma, Mallorca (Spain)
- Venues:
- NLPerspectives | WS
- Events:
- Fifteenth Language Resources and Evaluation Conference | Workshop on Perspectivist Approaches to NLP (2026) | Other Workshops and Events (2026)
- SIG:
- Publisher:
- ELRA Language Resources Association (ELRA)
- URL:
- https://aclanthology.org/2026.nlperspectives-1/
- DOI:
- 10.63317/5a3bvdkzb6f7
- PDF:
- https://aclanthology.org/2026.nlperspectives-1.pdf
Proceedings of the the fifth edition of NLPerspectives
Shiran Dudy | Gavin Abercrombie | Valerio Basile | Elisa Leonardelli | Simona Frenda
Shiran Dudy | Gavin Abercrombie | Valerio Basile | Elisa Leonardelli | Simona Frenda
What is Truth in NLP? Reflecting on Progress, Lessons, and Open Challenges as NLPerspectives turns Five
Gavin Abercrombie
Gavin Abercrombie
This paper reflects on five years of the Workshop on Perspectivist Approaches to NLP (NLPerspectives) and examines how this research community has helped to reconceptualise the notion of ground truth in human-labelled data. As NLP research has increasingly engaged with social and affective tasks, traditional assumptions about annotation reliability–centred on inter-annotator agreement and single ‘gold standard’ labels–have proven insufficient for capturing the genuine diversity of human perspectives. I review the developments that have driven the ‘Perspectivist Turn,’ assess its influence on mainstream NLP practice, and highlight the methodological challenges that arise when modelling disagreement, subjectivity, and annotator variation. In particular, I consider unresolved questions around evaluation paradigms, task formulation, population representation, community norms, and the implications of using pre-trained generative models as classifiers. By synthesising discussions from five years of workshops, keynotes, and related publications, I outline open challenges and propose directions for future work aimed at more rigorous perspectivist NLP. I argue that we should focus on centering minoritised standpoints and caution against viewing potentially harmful interpretations as equally legitimate reactions to ‘subjective’ phenomena.
The Rashomon Wikipedia: A Data-Perspectivist Analysis of Divergent Historical Narratives
Claudiu Creanga | Liviu P. Dinu | Anca Dinu
Claudiu Creanga | Liviu P. Dinu | Anca Dinu
Wikipedia aims to provide a unified, neutral record of history, yet its independent language editions often function as distinct epistemic communities, creating divergent narratives around contested events. This paper investigates cross-lingual historiographical bias by analyzing Wikipedia articles across five languages (Romanian, Hungarian, Russian, Turkish, and English) focusing on three contentious events in Romanian history: the Battle of Posada (1330), the Soviet occupation of Bessarabia (1940), and the Night Attack at Târgoviște (1462). Using human annotators and Large Language Models (LLMs) to classify citation stance and quantify narrative evolution from 2005 to 2024, we identify a phenomenon of “citation isolation”. In the case of the Battle of Posada, only 2 out of 119 citations were shared between language editions, with the Romanian edition exhibiting a 91% pro-national bias compared to the balanced Hungarian edition. Longitudinal analysis reveals that these narratives are volatile and responsive to contemporary geopolitics, evidenced by a significant shift in the Russian framing of Bessarabia in 2024. Finally, we propose a “Peace-Maker” pipeline to automate conflict reconciliation. We demonstrate that while standard prompting leads models to hallucinate consensus, “adversarial” prompting, which explicitly instructs the model to preserve and attribute disagreement, achieves near-perfect neutrality scores.
GSI:detect - A Perspectivist Approach to Gender Stereotypes Identification in Italian
Davide Testa | Sofia Brenna | Manuela Speranza | Gloria Comandini | Stefania Cavagnoli | Bernardo Magnini
Davide Testa | Sofia Brenna | Manuela Speranza | Gloria Comandini | Stefania Cavagnoli | Bernardo Magnini
The deconstruction of gender stereotypes is essential to prevent discrimination, marginalization and gender-based violence. Despite the increasing attention to this issue, research in this field often focuses on explicitly sexist or hateful communication, leaving out all the cases where stereotypes are produced unconsciously or even with apparently positive intentions. Moreover, the identification and analysis of gender stereotypes is often a very subjective task, heavily influenced by the researcher’s background, beliefs and personal sensitivity. In this context GSI:detect, a dataset for gender stereotypes identification in Italian, has been annotated following a perspectivist approach that gives value to the different points of view of four annotators. It has been designed to address (i) the lack of resources focusing on naturally occurring and non-hateful language conveying implicit or ambiguous forms of gender stereotypes, and (ii) the scarcity of datasets that can capture multiple interpretations as well as the inherent variation and disagreement in human perception. Baseline experiments with several LLMs confirm the challenging nature and value of such a linguistic resource, revealing both apparent differences and limitations in performance among the evaluated models, and raising questions about the extent to which current LLMs are suitable for detection and classification tasks in this field. Content warning: Examples taken from the GSI:detect dataset may contain sensitive or potentially distressing content.
It is increasingly recognized that humans do not always agree, and disagreement is inherent in many annotation tasks. However, not all items in a given task elicit the same level of opinion divergence. In this paper, we study the extent to which item-level annotation variation and variation structure can be captured from text features, focusing on inappropriate language detection, including offensive language, hate speech, and toxic language detection. We model annotation variation to assess whether the degree of annotation divergence can be predicted from item-level textual features. We also propose the Opposition Index, a metric that quantifies the extent of opposing stances among annotators based on their Likert ratings.
HurtLens: A Perspectivist Corpus Analysis of Hurtful Language
Samuele D’Avenia | Eliana Di Palma | Marta Marchiori Manerba | Valerio Basile
Samuele D’Avenia | Eliana Di Palma | Marta Marchiori Manerba | Valerio Basile
Offensive language detection systems often rely on majority-aggregated annotations, overlooking the diversity of perspectives that shape how different communities perceive harm. In this contribution, we introduce HurtLens, a perspectivist corpus of hurtful language leveraging four disaggregated datasets which are automatically enriched through HurtLex lemmas, a multilingual resource of offensive and derogatory terms. Using mixed-effects modeling, we investigate how annotators’ sociodemographic backgrounds, the presence of specific types of offensive language (through Hurtlex categories) and their interaction influence offensiveness ratings. Our analysis reveals that offensiveness ratings are influenced both by annotators’ sociodemographic characteristics (particularly when considering them in intersection) and by the presence of specific types of offensive language. Additionally, we identify significant interaction effects showing that different demographic groups vary in their sensitivity to texts containing particular types of offensive language.
We introduce a new metric that quantifies the extent of systematicity of the disagreement between annotators. The metric, called σ, is inspired by Structural Balance Theory and it approximates the clusterability of the annotators of a dataset. Paired with a standard metric of inter-annotator agreement such as Krippendorffs α, σ measures the amount of disagreement which stems from genuine subjective factors as opposed to the amount of disagreement caused by inner features of the annotation task. The metric is applied to over twenty datasets encoding a broad variety of annotations, showing its effectiveness in capturing the systematicity of annotator disagreement and its explanatory value.
Fine-Grained Perspectives: Modeling Explanations with Annotator-Specific Rationales
Olufunke O. Sarumi | Charles Welch | Daniel Braun
Olufunke O. Sarumi | Charles Welch | Daniel Braun
Beyond exploring disaggregated labels for modeling perspectives, annotator rationales provide fine-grained signals of individual perspectives. In this work, we propose a framework for jointly modeling annotator-specific label prediction and corresponding explanations, fine-tuned on the annotators’ provided rationales. Using a dataset with disaggregated natural language inference (NLI) annotations and annotator-provided explanations, we condition predictions on both annotator identity and demographic metadata through a representation-level User Passport mechanism. We further introduce two explainer architectures: a post-hoc prompt-based explainer and a prefixed bridge explainer that transfers annotator-conditioned classifier representations directly into a generative model. This design enables explanation generation aligned with individual annotator perspectives. Our results show that incorporating explanation modeling substantially improves predictive performance over a baseline annotator-aware classifier, with the prefixed bridge approach achieving more stable label alignment and higher semantic consistency, while the post-hoc approach yields stronger lexical similarity. These findings indicate that modeling explanations as expressions of fine-grained perspective provides a richer and more faithful representation of disagreement. The proposed approaches advance perspectivist modeling by integrating annotator-specific rationales into both predictive and generative components.
Structured Disagreement in Health-Literacy Annotation: Epistemic Stability, Conceptual Difficulty, and Agreement-Stratified Inference
Olga Kellert | Sriya Kondury | Candice Koo | Nemika Tyagi | Steffen Eikenberry
Olga Kellert | Sriya Kondury | Candice Koo | Nemika Tyagi | Steffen Eikenberry
Annotation pipelines in Natural Language Processing (NLP) commonly assume a single latent ground truth per instance and resolve disagreement through label aggregation. Perspectivist approaches challenge this view by treating disagreement as potentially informative rather than erroneous. We present a large-scale analysis of graded health-literacy annotations from 6,323 open-ended COVID-19 responses collected in Ecuador and Peru. Each response was independently labeled by multiple annotators using proportional correctness scores, allowing us to analyze the full distribution of judgments rather than aggregated labels. Variance decomposition shows that question-level conceptual difficulty accounts for substantially more variance than annotator identity, indicating that disagreement is structured by the task itself rather than driven by individual raters. Agreement-stratified analyses further reveal that key social-scientific effects, including country, education, and urban-rural differences, vary in magnitude and in some cases reverse direction depending on levels of inter-annotator agreement. These findings suggest that graded health-literacy evaluation contains both epistemically stable and unstable components, and that aggregating across them can obscure important inferential differences. We therefore argue that strong perspectivist modeling is not only conceptually justified but statistically necessary for valid inference in graded interpretive tasks.
SubData: Bridging Heterogeneous Datasets to Enable Theory-Driven Evaluation of Political and Demographic Perspectives in LLMs
Pietro Bernardelle | Leon Froehling | Stefano Civelli | Gianluca Demartini
Pietro Bernardelle | Leon Froehling | Stefano Civelli | Gianluca Demartini
As increasingly capable large language models (LLMs) emerge, researchers have begun exploring their potential for subjective tasks. While recent work demonstrates that LLMs can be aligned with diverse human perspectives, evaluating this alignment on downstream tasks (e.g., hate speech detection) remains challenging due to the use of inconsistent datasets across studies. To address this issue, in this resource paper we propose a two-step framework: we (1) introduce SubData, an open-source Python library designed for standardizing heterogeneous datasets to evaluate LLMs perspective alignment; and (2) present a theory-driven approach leveraging this library to test how differently-aligned LLMs (e.g., aligned with different political viewpoints) classify content targeting specific demographics. SubData’s flexible mapping and taxonomy enable customization for diverse research needs, distinguishing it from existing resources. We illustrate its usage with an example application and invite contributions to extend our initial release into a multi-construct benchmark suite for evaluating LLMs perspective alignment on natural language processing tasks.
ChatGPT, why can’t anyone afford a house? On the Effects of LLM pre-annotation on Annotator Subjectivity
Emilie Francis | Céline Leuzinger | Ricardo Muñoz Sánchez | Lee D. Gauthier
Emilie Francis | Céline Leuzinger | Ricardo Muñoz Sánchez | Lee D. Gauthier
Large language models (LLMs) have often been proposed as substitutes for human annotators in a variety of tasks. At the same time, there has been increased focus on the role that human subjectivity and perspective plays in data annotation. To avoid eliminating the human role in annotation entirely, the use of LLMs for pre-annotation has been suggested as an alternative approach. In this paper, we explore to which degree this approach affects subjectivity of social media annotation in English. We focus on comments regarding the current status of the housing market and label them for concern level, factors affecting housing affordability, and aspects that authors claim either exacerbate or improve the situation. To investigate this, we design an experiment involving two rounds of annotation: the first, a dataset annotated by humans only; and the second, a dataset with LLM pre-annotations curated by the same human annotators. We observe that the second setting leads to much higher agreement, as well as significant changes in label distribution and co-occurrence. Similar shifts do not appear in the LLM labels. Our findings show that use of LLMs in the annotation process leads to convergence in annotations and, thus, to an erosion of human subjectivity.
An Overview of Current Practices and Recommendations for Working with Stereotypes in NLP
Alessandra Teresa Cignarella | Matteo Pellegrini
Alessandra Teresa Cignarella | Matteo Pellegrini
This article presents a discussion on the main challenges and considerations involved in addressing stereotypes within Natural Language Processing (NLP), and proposes a set of guidelines and recommendations for their treatment in research and resource development. On the one hand, the growing interest in fairness, bias mitigation, and inclusivity has led to an increasing number of studies and datasets dealing with stereotypes; on the other hand, their conceptualization and operationalization remain highly heterogeneous across works. The aim of this article is therefore twofold: (1) to provide a concise yet comprehensive overview of existing annotation schemes highlighting their key features and offering a comparative analysis and (2) to propose a set of tentative guidelines and recommendations to foster clarity when working with stereotypes in NLP. Furthermore, as a case study, we conduct an annotation exercise of a subset of texts from the QUEEREOTYPES dataset, containing stereotypes targeting LGBTQIA+ people, using all labels proposed in prior work to assess their clarity, overlap, and practical usefulness.
Modeling Perspectives in NLP: Parameter-Efficient Perspective Conditioning for Span Extraction and Summarization
Harikrishnan Gurushankar Saisudha | Sabine Bergler
Harikrishnan Gurushankar Saisudha | Sabine Bergler
Understanding text through multiple perspectives is essential in domains such as healthcare community question answering, where answers frequently contain heterogeneous viewpoints, including experiences, suggestions, causes, follow-up questions, and informational claims. We present a unified perspective-conditioned framework for both span identification and perspective-aware summarization on the PerAnsSumm dataset. Our approach introduces explicit perspective signals into transformer models using two parameter-efficient mechanisms: prefix-conditioned representations and perspective-aware attention layers. We first employ multi-label perspective classification to identify relevant viewpoints, which serve as conditioning signals for downstream tasks. For span identification, we model perspective-specific extraction as a conditioned binary sequence labeling problem. For summarization, we guide generation using perspective-enriched encoder representations. Experiments demonstrate that explicit perspective conditioning substantially improves span detection performance while achieving competitive summarization quality. Notably, perspective-aware attention achieves strong results using only a small fraction of the trainable parameters required by full fine-tuning. Our findings highlight the importance of structured viewpoint modeling and show that explicit perspective control enables efficient and interpretable multi-perspective text understanding.
A Pilot Study Investigating Stakeholder Subjectivity in Collaborative Dialog Analysis
Ananya Ganesh | Martha Palmer | Katharina von der Wense
Ananya Ganesh | Martha Palmer | Katharina von der Wense
Qualitative research in education relies on“ground truth" codes or labels generated by having a trained or expert coder code observations in data such as student dialog. Although rigorous validity checks are a part of the coding process, there is limited research investigating how and to what extent, this notion of the ground truth is influenced by inherent task subjectivity. This paper presents a pilot study of task subjectivity centered around the phenomenon of verbal off-task behavior. The context for this study is real-world small-group collaborative conversations among three to five students in a middle-school science classroom. To investigate how stakeholders such as teachers and students show subjectivity in approaching this task, we recruit five teachers from the Prolific online platform, and five students from local middle and high schools as annotators of off-task speech. We show that teachers, students, and expert coders differ in their perception of off-task speech, with some of these differences being systematic. Drawing upon recent research in machine learning and natural language processing, we then outline the potential benefits of collecting and modeling a range of codes that explicitly represent the subjective perspectives of a diverse set of coders.