TRoTR: A Framework for Evaluating the Re-contextualization of Text Reuse

Francesco Periti; Pierluigi Cassotti; Stefano Montanelli; Nina Tahmasebi; Dominik Schlechtweg

doi:10.18653/v1/2024.emnlp-main.774

TRoTR: A Framework for Evaluating the Re-contextualization of Text Reuse

Francesco Periti, Pierluigi Cassotti, Stefano Montanelli, Nina Tahmasebi, Dominik Schlechtweg

Abstract

Current approaches for detecting text reuse do not focus on recontextualization, i.e., how the new context(s) of a reused text differs from its original context(s). In this paper, we propose a novel framework called TRoTR that relies on the notion of topic relatedness for evaluating the diachronic change of context in which text is reused. TRoTR includes two NLP tasks: TRiC and TRaC. TRiC is designed to evaluate the topic relatedness between a pair of recontextualizations. TRaC is designed to evaluate the overall topic variation within a set of recontextualizations. We also provide a curated TRoTR benchmark of biblical text reuse, human-annotated with topic relatedness. The benchmark exhibits an inter-annotator agreement of .811. We evaluate multiple, established SBERT models on the TRoTR tasks and find that they exhibit greater sensitivity to textual similarity than topic relatedness. Our experiments show that fine-tuning these models can mitigate such a kind of sensitivity.

Anthology ID:: 2024.emnlp-main.774
Volume:: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2024
Address:: Miami, Florida, USA
Editors:: Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 13972–13990
Language:
URL:: https://aclanthology.org/2024.emnlp-main.774/
DOI:: 10.18653/v1/2024.emnlp-main.774
Bibkey:
Cite (ACL):: Francesco Periti, Pierluigi Cassotti, Stefano Montanelli, Nina Tahmasebi, and Dominik Schlechtweg. 2024. TRoTR: A Framework for Evaluating the Re-contextualization of Text Reuse. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 13972–13990, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):: TRoTR: A Framework for Evaluating the Re-contextualization of Text Reuse (Periti et al., EMNLP 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.emnlp-main.774.pdf
Software:: 2024.emnlp-main.774.software.zip
Data:: 2024.emnlp-main.774.data.zip

PDF Cite Search Software Data Fix data