Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection

Erik Arakelyan; Arnav Arora; Isabelle Augenstein

doi:10.18653/v1/2023.acl-long.752

Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection

Erik Arakelyan, Arnav Arora, Isabelle Augenstein

Abstract

The task of Stance Detection is concerned with identifying the attitudes expressed by an author towards a target of interest. This task spans a variety of domains ranging from social media opinion identification to detecting the stance for a legal claim. However, the framing of the task varies within these domains in terms of the data collection protocol, the label dictionary and the number of available annotations. Furthermore, these stance annotations are significantly imbalanced on a per-topic and inter-topic basis. These make multi-domain stance detection challenging, requiring standardization and domain adaptation. To overcome this challenge, we propose Topic Efficient StancE Detection (TESTED), consisting of a topic-guided diversity sampling technique used for creating a multi-domain data efficient training set and a contrastive objective that is used for fine-tuning a stance classifier using the produced set. We evaluate the method on an existing benchmark of 16 datasets with in-domain, i.e. all topics seen and out-of-domain, i.e. unseen topics, experiments. The results show that the method outperforms the state-of-the-art with an average of 3.5 F1 points increase in-domain and is more generalizable with an averaged 10.2 F1 on out-of-domain evaluation while using <10% of the training data. We show that our sampling technique mitigates both inter- and per-topic class imbalances. Finally, our analysis demonstrates that the contrastive learning objective allows the model for a more pronounced segmentation of samples with varying labels.

Anthology ID:: 2023.acl-long.752
Volume:: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2023
Address:: Toronto, Canada
Editors:: Anna Rogers, Jordan Boyd-Graber, Naoaki Okazaki
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 13448–13464
Language:
URL:: https://aclanthology.org/2023.acl-long.752/
DOI:: 10.18653/v1/2023.acl-long.752
Bibkey:
Cite (ACL):: Erik Arakelyan, Arnav Arora, and Isabelle Augenstein. 2023. Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13448–13464, Toronto, Canada. Association for Computational Linguistics.
Cite (Informal):: Topic-Guided Sampling For Data-Efficient Multi-Domain Stance Detection (Arakelyan et al., ACL 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.acl-long.752.pdf
Video:: https://aclanthology.org/2023.acl-long.752.mp4

PDF Cite Search Video Fix data