Quantifying and Predicting Disagreement in Graded Human Ratings

Leixin Zhang, Çağrı Çöltekin


Abstract
It is increasingly recognized that humans do not always agree, and disagreement is inherent in many annotation tasks. However, not all items in a given task elicit the same level of opinion divergence. In this paper, we study the extent to which item-level annotation variation and variation structure can be captured from text features, focusing on inappropriate language detection, including offensive language, hate speech, and toxic language detection. We model annotation variation to assess whether the degree of annotation divergence can be predicted from item-level textual features. We also propose the Opposition Index, a metric that quantifies the extent of opposing stances among annotators based on their Likert ratings.
Anthology ID:
2026.nlperspectives-1.4
Volume:
Proceedings of the the fifth edition of NLPerspectives
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Shiran Dudy, Gavin Abercrombie, Valerio Basile, Elisa Leonardelli, Simona Frenda
Venues:
NLPerspectives | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
33–43
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-nlperspectives-04
DOI:
10.63317/4qy8nuowzhpy
Bibkey:
Cite (ACL):
Leixin Zhang and Çağrı Çöltekin. 2026. Quantifying and Predicting Disagreement in Graded Human Ratings. In Proceedings of the the fifth edition of NLPerspectives, pages 33–43, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Quantifying and Predicting Disagreement in Graded Human Ratings (Zhang & Çöltekin, NLPerspectives 2026)
Copy Citation: