MASIVE: Open-Ended Affective State Identification in English and Spanish

Nicholas Deas; Elsbeth Turcan; Ivan Ernesto Perez Mejia; Kathleen McKeown

doi:10.18653/v1/2024.emnlp-main.1139

MASIVE: Open-Ended Affective State Identification in English and Spanish

Nicholas Deas, Elsbeth Turcan, Ivan Ernesto Perez Mejia, Kathleen McKeown

Abstract

In the field of emotion analysis, much NLP research focuses on identifying a limited number of discrete emotion categories, often applied across languages. These basic sets, however, are rarely designed with textual data in mind, and culture, language, and dialect can influence how particular emotions are interpreted. In this work, we broaden our scope to a practically unbounded set of affective states, which includes any terms that humans use to describe their experiences of feeling. We collect and publish MASIVE, a dataset of Reddit posts in English and Spanish containing over 1,000 unique affective states each. We then define the new problem of affective state identification for language generation models framed as a masked span prediction task. On this task, we find that smaller finetuned multilingual models outperform much larger LLMs, even on region-specific Spanish affective states. Additionally, we show that pretraining on MASIVE improves model performance on existing emotion benchmarks. Finally, through machine translation experiments, we find that native speaker-written data is vital to good performance on this task.

Anthology ID:: 2024.emnlp-main.1139
Volume:: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2024
Address:: Miami, Florida, USA
Editors:: Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 20467–20485
Language:
URL:: https://aclanthology.org/2024.emnlp-main.1139/
DOI:: 10.18653/v1/2024.emnlp-main.1139
Bibkey:
Cite (ACL):: Nicholas Deas, Elsbeth Turcan, Ivan Ernesto Perez Mejia, and Kathleen McKeown. 2024. MASIVE: Open-Ended Affective State Identification in English and Spanish. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 20467–20485, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):: MASIVE: Open-Ended Affective State Identification in English and Spanish (Deas et al., EMNLP 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.emnlp-main.1139.pdf

PDF Cite Search Fix data