Workshop on Ethical and Legal Issues in Human Language Technologies and Multilingual De-Identification of Sensitive Data In Language Resources (2026)
up
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Ingo Siegert | Maria Irena Szawerna | Khalid Choukri | Simon Dobnik | Paweł Kamocki | Therese Lindström Tiedemann | Pierre Lison | Ricardo Muñoz Sánchez | Ildikó Pilán | Lisa Södergård | Kossay Talmoudi | Elena Volodina | Xuan-Son Vu
Ingo Siegert | Maria Irena Szawerna | Khalid Choukri | Simon Dobnik | Paweł Kamocki | Therese Lindström Tiedemann | Pierre Lison | Ricardo Muñoz Sánchez | Ildikó Pilán | Lisa Södergård | Kossay Talmoudi | Elena Volodina | Xuan-Son Vu
Transparency as Architecture: Structural Compliance Gaps in EU AI Act Article 50 II
Vera Schmitt | Niklas Kruse | Premtim Sahitaj | Julius Schöning
Vera Schmitt | Niklas Kruse | Premtim Sahitaj | Julius Schöning
Art. 50 II of the EU Artificial Intelligence Act mandates dual transparency for AI-generated content: outputs must be labeled in both human-understandable and machine-readable form for automated verification. This requirement, entering into force in August 2026, collides with fundamental constraints of current generative AI systems. Using synthetic data generation and automated fact-checking as diagnostic use cases, we show that compliance cannot be reduced to post-hoc labeling. In fact-checking pipelines, provenance tracking is not feasible under iterative editorial workflows and non-deterministic LLM outputs; moreover, the assistive-function exemption does not apply, as such systems actively assign truth values rather than supporting editorial presentation. In synthetic data generation, persistent dual-mode marking is paradoxical: watermarks surviving human inspection risk being learned as spurious features during training, while marks suited for machine verification are fragile under standard data processing. Across both domains, three structural gaps obstruct compliance: (a) absent cross-platform marking formats for interleaved human-AI outputs; (b) misalignment between the regulation’s ’reliability’ criterion and probabilistic model behavior; and (c) missing guidance for adapting disclosures to heterogeneous user expertise. Closing these gaps requires transparency to be treated as an architectural design requirement, demanding interdisciplinary research across legal semantics, AI engineering, and human-centered design.
Towards Robust Evaluation for Privacy QA Systems
Anna Leschanowsky | Zahra Kolagar | Erion Çano | Ivan Habernal | Dara Hallinan | Emanuël Habets | Birgit Popp
Anna Leschanowsky | Zahra Kolagar | Erion Çano | Ivan Habernal | Dara Hallinan | Emanuël Habets | Birgit Popp
The transparency principle of the General Data Protection Regulation requires data-processing information to be clear, precise, and accessible. While Large Language Models (LLMs) show promise in this context, their probabilistic nature raises challenges for ensuring truthfulness and comprehensibility. This paper presents an exploratory evaluation of eight Privacy Question Answering (QA) systems – including LLMs, retrieval-augmented generation, and alignment-based approaches – on two datasets. We propose an evaluation framework that maps both traditional NLP and LLM-as-a-judge metrics to the legal requirements of comprehensibility and precision. Results show that no single system consistently excels across all metrics, and that system rankings can vary depending on the choice of metric and thresholding. We highlight open questions and emphasize the need to translate legal requirements into technical evaluation criteria. Our work provides a foundation for a more robust evaluation of Privacy QA systems.
LDS Contractual Framework: Principles, Status and Implementation
Penny Labropoulou | Kossay Talmoudi | Dimitrios Gkoumas | Katerina Gkirtzou | Miltos Deligiannis | Leon Voukoutis | Athanasia Kolovou | Khalid Choukri | Stelios Piperidis | Dimitrios Galanis
Penny Labropoulou | Kossay Talmoudi | Dimitrios Gkoumas | Katerina Gkirtzou | Miltos Deligiannis | Leon Voukoutis | Athanasia Kolovou | Khalid Choukri | Stelios Piperidis | Dimitrios Galanis
To strengthen competitiveness and digital sovereignty, the European Union has promoted the development of Common European Data Spaces to enable secure and interoperable data sharing between participants for various sectors. Data spaces combine technical infrastructure with governance mechanisms to ensure trust, transparency, data sovereignty and interoperability. Their operation must comply with the evolving European regulatory framework as well as contractual law. This paper presents the strategy adopted in the Language Data Space (LDS) to operationalise these requirements, focusing on its contractual framework and supporting instruments. It outlines the governing principles designed to ensure lawful, transparent, and fair data transactions while safeguarding the rights and obligations of data providers and consumers alike. It further describes the actual framework, and the recommended data sharing licences, with a particular emphasis on the LDS standard licence. Finally, it presents the automation tools designed and developed to support the relevant workflows while serving a wide range of users that have little or no knowledge of technical and legal complexities.
Authorship Attribution in the Times of LLMs within the Framework of the CRediT Taxonomy
Pawel Kamocki | Andreas Witt
Pawel Kamocki | Andreas Witt
This article examines the concept of authorship in the context of generative language models and other uses of Artificial Intelligence, and how this new ‘authorshipness’ can be represented in metadata. It analyses authorship under copyright law and proposes a metadata-based approach to disclosing the use of AI in publications, drawing on the widely adopted CRediT taxonomy developed by the National Information Standards Organization (NISO), and informed by guidance from the United States Copyright Office (USCO) and the International Association of Scientific, Technical and Medical Publishers (STM).
DeID-Clinic: A Risk-Aware Pseudonymization Framework for Clinical Text De-identification and Re-identification Risk Assessment
Angel Paul | Dhivin Shaji | Lifeng Han | Warren Del-Pinto | Goran Nenadic | Suzan Verberne
Angel Paul | Dhivin Shaji | Lifeng Han | Warren Del-Pinto | Goran Nenadic | Suzan Verberne
The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sharing while preserving downstream utility. This paper presents DeID-Clinic, a multi-layered framework for automated pseudonymization and re-identification risk assessment of clinical free-text data. Our approach integrates domain-adapted transformer models, including BioBERT and ClinicalBERT, into the MASK de-identification framework to improve the detection and masking of protected health information (PHI). Beyond entity recognition, we introduce a novel document-level risk assessment module that quantifies residual re-identification risk using a combination of k-anonymity, l-diversity, t-closeness, contextual similarity, and entity co-occurrence analysis. Experiments conducted on the i2b2 2014 de-identification dataset demonstrate strong performance, achieving macro-level F1 scores above 0.96 for several entity categories, while enabling quantitative prioritization of high-risk documents for further review. Our results highlight the effectiveness of combining neural de-identification with explicit risk modeling, supporting privacy-preserving data sharing in sensitive domains. Although evaluated on clinical text, the proposed framework is generalizable to other privacy-critical domains such as legal and administrative documents, where reliable pseudonymization and risk-aware anonymization are essential.
Distilling Human-Aligned Privacy Sensitivity Assessment from Large Language Models
Gabriel Loiseau | Damien Sileo | Damien Riquet | Maxime Meyer | Marc Tommasi
Gabriel Loiseau | Damien Sileo | Damien Riquet | Maxime Meyer | Marc Tommasi
Accurate privacy evaluation of textual data remains a critical challenge in privacy-preserving NLP. Recent work has shown that LLMs can serve as reliable privacy evaluators, achieving strong agreement with human judgments; however, their computational cost and impracticality for processing sensitive data at scale limit real-world deployment. We address this gap by distilling the privacy assessment capabilities of Mistral Large 3 (675B) into lightweight encoder models with as few as 150M parameters. Leveraging a large-scale dataset of privacy-annotated texts spanning 10 diverse domains, we train efficient classifiers that preserve strong agreement with human annotations while dramatically reducing computational requirements. We validate our approach on human-annotated test data and demonstrate its practical utility as an evaluation metric for de-identification systems.
Birds of a Feather: Do Embedding Representations of Personal Information Flock Together?
Maria Irena Szawerna | Simon Dobnik
Maria Irena Szawerna | Simon Dobnik
Personally identifiable information (PII or PI) can appear in a wide variety of linguistic data, posing both ethical and legal challenges for conducting research and developing applications involving such texts. In this paper, we investigate the alignment between automatic clustering of FastText and Transformer embedding representations of personal information spans sourced from essays written by adult learners of Swedish as a second language and the general and detailed personal information labels assigned to these spans by expert annotators. Our goals are to assess the extent of overlap between the semantic categories and evaluate the semantic coherence of the human-assigned classes, which may have implications for de-identification procedures. We observe that while contextual embeddings, especially ones from a specialized word-in-context model, produce relatively good clustering results, they only partly map to the human understanding of how to classify personal information.
Modelling Legal Compliance in a Consent Wizard Application as Part of a Research-Centered and User-Oriented Data Infrastructure
Aliena Strathmann | Marc-Levin Joppek | Maryam Mohammadi | Katja Politt | Paul T. Schrader | Annett B. Jorschick | Hendrik Buschmeier
Aliena Strathmann | Marc-Levin Joppek | Maryam Mohammadi | Katja Politt | Paul T. Schrader | Annett B. Jorschick | Hendrik Buschmeier
Recent research calls for data management infrastructures that explicitly operate within the bounds of ethical and legal constraints, and facilitate adherence to Open Science principles by integrating automated support for planning, collection, storage, use, reuse, and sharing of data within. Legal and ethical requirements of data processing have become increasingly complex, introducing administrative barriers to scientific research investigating data generated by human participants, which encompasses a vast majority of humanities research. In response to this, we present RUDI (“Research-centered User-oriented Data Infrastructure”), a modular framework grounded in an interdisciplinary approach informed by legal, computational and linguistic expertise. This paper introduces its first component; a configurable and dynamically adaptive consent form generator in the form of a “wizard” web application. We outline how legal aspects are modeled within, and highlight its concrete benefits for administrative aspects of research. Further, we discuss the contextualization of data within the research domain by leveraging the use of standardized ontology within the framework.
Balancing FAIR and GDPR: A Governance Framework for Oral Archives
Elvira Mercatanti | Monica Monachini | Giovanni Abete | Silvia Calamai | Sergio Canazza | Alessandro Casellato | Virginia Niri | Cesarina Vecchia | Giulia Zitelli Conti | Giada Zuccolo
Elvira Mercatanti | Monica Monachini | Giovanni Abete | Silvia Calamai | Sergio Canazza | Alessandro Casellato | Virginia Niri | Cesarina Vecchia | Giulia Zitelli Conti | Giada Zuccolo
This paper presents a governance framework developed within the research project ROADS to support thesustainable management of oral archives, which constitute essential linguistic resources for interdisciplinary research and cultural heritage preservation. Oral archives raise complex ethical and legal challenges due to the hybrid nature of voice data, which function simultaneously as historical documents, scientific sources and biometric identifiers, thereby creating tensions between open science principles and data protection regulations. The proposed framework integrates FAIR principles (Findable, Accessible, Interoperable, Reusable) with Privacy by Design and the GDPR accountability principle through a multilayered approach. It introduces an access model that distinguishes between publicly available metadata and controlled access to identifiable audio materials, following trusted repository standards. The framework also incorporates consent management procedures and safeguards for legacy collections, enabling responsible data sharing while preserving scientific usability. More broadly, ROADS provides a transferable model to guide the transition from project-based archives to FAIR, sustainable and reusable research resources, ensuring compliance with data protection requirements and respect for the sensitivity of the documented contexts.
Legal Considerations in the Use of Synthetic Data for AI Development and Finetuning: The Case of LLMs4EU
Kossay Talmoudi | Khalid Choukri | Amélie Gourgeot | Florine Astruc
Kossay Talmoudi | Khalid Choukri | Amélie Gourgeot | Florine Astruc
This paper examines the legal implications of using synthetic data to develop and fine-tune general-purpose AI models in the European Union, using the LLMs4EU project as a case study. It situates synthetic data within the Union’s broader data policy and highlights it as a candidate tool for reconciling data availability with regulatory constraints. From a data-protection perspective, it analyses whether and when synthetic data should be classified as “personal data” under the GDPR. From a copyright and contractual standpoint, the paper assesses the risks that synthetic datasets may embed infringing content or derive from unlawfully trained models, in light of the GEMA v. OpenAI ruling on memorised works and emerging analyses of liability for AI-generated outputs, and considers the constraints imposed by model licensing and acceptable-use policies on using models to generate training data for other models. The paper concludes that synthetic data can play a valuable role in mitigating legal risks and enabling compliant AI development in LLMs4EU, but only if its generation and use are embedded in robust governance frameworks that address data protection, copyright and contractual obligations across the entire data value chain.
Evaluating Encoder- and LLM-Based Approaches for Robust Indirect Personal Identifier Detection
Christoph Otto | Ibrahim Baroud | Akiko Aizawa | Sebastian Möller | Roland Roller | Lisa Raithel
Christoph Otto | Ibrahim Baroud | Akiko Aizawa | Sebastian Möller | Roland Roller | Lisa Raithel
Removing explicit protected health information does not fully eliminate re-identification risk in clinical text. Contextual attributes such as socio-economic status, institutional affiliations or detailed life circumstances may still enable linkage attacks. These heterogeneous and sparsely distributed elements, termed Indirect Personal Identifiers, extend de-identification beyond fixed identifier lists and pose new modeling challenges. Therefore, we present the first systematic comparison of encoder-only models, prompt-based LLMs and hybrid pipelines for span-level IPI detection in English discharge summaries. A fine-tuned RoBERTa-large model improves on an existing baseline and substantially outperforms ChatGPT-5.2, achieving 0.906 micro-F1 and 0.724 macro-F1, compared to 0.509 micro-F1 and 0.487 macro-F1. Our findings indicate that IPI detection constitutes a distinct modeling regime characterized by class imbalance and high intra-class variability, where scaling model capacity alone does not guarantee macro-level robustness. We show that supervised encoder models currently provide the most reliable foundation for extending anonymization guarantees and future research.
VEIL: A Benchmark for Value-Preserving Entity Identification Limitation
Darina Gold | Shadi Rastegar | Alina Liebel | Alessandra Zarcone
Darina Gold | Shadi Rastegar | Alina Liebel | Alessandra Zarcone
Large Language Models (LLMs) are linked to several issues regarding Personally Identifiable Information (PII). PII can occur in the training data and can thus be accidentally leaked or extracted with malicious intent, or it can be inputted in LLM-based technologies by users through their prompts. A viable strategy to limit the LLMs’ exposure to PII is to filter input and output data by de-identifying PII, including personal names. This however poses a challenge: a name could refer to a private person in a context containing sensitive information (e.g., Michelangelo is an atheist), or it could refer to a famous artist in another context (e.g., Michelangelo’s Sistine Chapel), and masking the latter may hinder the LLMs’ capabilities in general-knowledge tasks. We tackle the problem of personal name de-identification and focus on the decision of which personal names need to be removed (and which should be kept), based on context. We present VEIL, a challenging benchmark for Value-preserving Entity Identification Limitation, for context-aware de-identification decisions on LLM training data, and compare the performance of different state-of-the-art systems on the task.