Simulating Complex Immediate Textual Variation with Large Language Models

Fernando Aguilar-Canto; Alberto Espinosa-Juarez; Hiram Calvo

Simulating Complex Immediate Textual Variation with Large Language Models

Fernando Aguilar-Canto, Alberto Espinosa-Juarez, Hiram Calvo

Abstract

Immediate Textual Variation (ITV) is defined as the process of introducing changes during text transmission from one node to another. One-step variation can be useful for testing specific philological hypotheses. In this paper, we propose using Large Language Models (LLMs) as text-modifying agents. We analyze three scenarios: (1) simple variations (omissions), (2) paraphrasing, and (3) paraphrasing with bias injection (polarity). We generate simulated news items using a predefined scheme. We hypothesize that central tendency measures—such as the mean and median vectors in the feature space of sentence transformers—can effectively approximate the original text representation. Our findings indicate that the median vector is a more accurate estimator of the original vector than most alternatives. However, in cases involving substantial rephrasing, the agent that produces the least semantic drift provides the best estimation, aligning with the principles of Bédierian textual criticism.

Anthology ID:: 2025.lm4dh-1.2
Volume:: Proceedings of the First on Natural Language Processing and Language Models for Digital Humanities
Month:: September
Year:: 2025
Address:: Varna, Bulgaria
Editors:: Isuri Nanomi Arachchige, Francesca Frontini, Ruslan Mitkov, Paul Rayson
Venues:: LM4DH | WS
SIG:
Publisher:: INCOMA Ltd., Shoumen, Bulgaria
Note:
Pages:: 25–31
Language:
URL:: https://aclanthology.org/2025.lm4dh-1.2/
DOI:
Bibkey:
Cite (ACL):: Fernando Aguilar-Canto, Alberto Espinosa-Juarez, and Hiram Calvo. 2025. Simulating Complex Immediate Textual Variation with Large Language Models. In Proceedings of the First on Natural Language Processing and Language Models for Digital Humanities, pages 25–31, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Cite (Informal):: Simulating Complex Immediate Textual Variation with Large Language Models (Aguilar-Canto et al., LM4DH 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.lm4dh-1.2.pdf
Optionalsupplementarymaterial:: 2025.lm4dh-1.2.OptionalSupplementaryMaterial.zip

PDF Cite Search Optionalsupplementarymaterial Fix data