Annotating Personal Information in Swedish Texts with SPARV

Maria Irena Szawerna, David Alfter, Elena Volodina


Abstract
Digital Humanities (DH) research, among many others, relies on data, a subset of which comes in the form of language data that contains personal information (PI). Working with and sharing such data has ethical and legal implications. The process of removing (anonymization) or replacing (pseudonymization) of personal information in texts may be used to address these issues, and often begins with a PI detection and labeling stage. We present a new tool for personal information detection and labeling for Swedish, SBX-PI-DETECTION (henceforth SBX-PI), alongside a visualization interface, (IM)PERSONAL DATA, which allows for the comparison of outputs from different tools. A valuable feature of SBX-PI is that it enables the users to run the annotation locally. It is also integrated into the text annotation pipeline SPARV, allowing for other types of annotation to be performed simultaneously and contributing to the privacy by design requirement set by the GDPR. A novel feature of (IM)PERSONAL DATA is that it allows researchers to assess the extent of detected PI in a text and how much of it will be manipulated once anonymization or pseudonymization are applied. The tools are primarily aimed at researchers within Digital Humanities and Natural Language Processing and are linked to CLARIN’s Virtual Language Observatory.
Anthology ID:
2025.lm4dh-1.15
Volume:
Proceedings of the First on Natural Language Processing and Language Models for Digital Humanities
Month:
September
Year:
2025
Address:
Varna, Bulgaria
Editors:
Isuri Nanomi Arachchige, Francesca Frontini, Ruslan Mitkov, Paul Rayson
Venues:
LM4DH | WS
SIG:
Publisher:
INCOMA Ltd., Shoumen, Bulgaria
Note:
Pages:
155–163
Language:
URL:
https://aclanthology.org/2025.lm4dh-1.15/
DOI:
Bibkey:
Cite (ACL):
Maria Irena Szawerna, David Alfter, and Elena Volodina. 2025. Annotating Personal Information in Swedish Texts with SPARV. In Proceedings of the First on Natural Language Processing and Language Models for Digital Humanities, pages 155–163, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Cite (Informal):
Annotating Personal Information in Swedish Texts with SPARV (Szawerna et al., LM4DH 2025)
Copy Citation:
PDF:
https://aclanthology.org/2025.lm4dh-1.15.pdf