Simonetta Montemagni
Other people with similar names: Simonetta Montemagni
Unverified author pages with similar names: Simonetta Montemagni
2026
When Lexicographic Quotations Become a Corpus: To Deduplicate or Not to Deduplicate?
Manuel Favaro | Elisa Guadagnini | Eva Sassolini | Marco Biffi | Simonetta Montemagni
Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026
Manuel Favaro | Elisa Guadagnini | Eva Sassolini | Marco Biffi | Simonetta Montemagni
Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026
Historical dictionaries are increasingly reused as sources for diachronic language corpora. In this context, lexicographic quotations represent a valuable yet challenging type of data, as they are both editorially curated and diachronically representative. A major issue in their computational reuse is the presence of duplicate and near-duplicate quotations. This paper addresses quotation deduplication in corpora derived from lexicographic resources. We introduce QRD (Quotation Reuse Detection), a multi-stage pipeline designed to identify, compare, and cluster quotations based on graded similarity rather than binary matching. The approach combines string-based similarity measures, iterative threshold analysis, and clustering, enabling both quantitative and qualitative investigation of quotation reuse. Our results show that deduplication in this context cannot be reduced to the automatic elimination of redundant data. The variability observed in the quotations - ranging from OCR-related noise to substantial editorial variation - reflects both technical and structural factors and calls for a more nuanced approach. QRD supports the identification of OCR-related errors and reveals patterns of textual reuse underlying the compilation of the dictionary. We argue that quotation deduplication should be conceived primarily as a task of identification and clustering. This perspective reframes deduplication from a data-cleaning operation into an analytical methodology for historically and editorially curated textual resources.
From Print to Digital and beyond: The Retrodigitization of a Historical Dictionary of Italian as a Hybrid Lexical Resource
Marco Biffi | Sebastiana Cucurullo | Manuel Favaro | Elisa Guadagnini | Simonetta Montemagni | Eva Sassolini
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Marco Biffi | Sebastiana Cucurullo | Manuel Favaro | Elisa Guadagnini | Simonetta Montemagni | Eva Sassolini
Proceedings of the Fifteenth Language Resources and Evaluation Conference
This paper presents the retrodigitization project of the Grande Dizionario della Lingua Italiana (GDLI), the largest and most comprehensive historical dictionary of the Italian language. The GDLI’s 23,000 pages — originally designed for human consultation — constitute an exceptional repository of linguistic and cultural-historical information, while posing significant challenges to large-scale digitization and data structuring. The project, still ongoing, will result in the development of a set of interoperable and interlinked resources: (i) a TEI-XML edition of the dictionary text, encoding its complex lexicographic and citation structure; (ii) an annotated corpus of the quoted examples, enabling linguistic and historical research across centuries; and (iii) a database of cited authors and works. Together, these components form a hybrid lexical resource that establishes the foundations for innovative and advanced modes of accessing and exploring the rich and multifaceted content of this historical dictionary.