Khalid Choukri
Papers on this page may belong to the following people: Khalid Choukri, Khalid Choukri
2026
Development of Speech Corpus for Low-Resource Language- a Case of Sanskrit
Devendr Kumar | Girish Nath Jha | Khalid Choukri
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
Devendr Kumar | Girish Nath Jha | Khalid Choukri
Proceedings of the 8th Workshop on Indian Language Data: Resources and Evaluation
This paper presents a comprehensive framework for the development of a speech corpus for Sanskrit, designed to facilitate advances in Automatic Speech Recognition (ASR) and AI/ML research. The proposed corpus comprises over 107 hours of transcribed speech data, collected from diverse Sanskrit sources through a systematic and scalable pipeline. We detail the end-to-end methodology adopted for corpus creation, encompassing web crawling, data sanitization, audio downloading, and transcription alignment. Particular emphasis is placed on the methodological rigor applied at each stage, including source selection, preprocessing for quality assurance, transcription protocols, and forced alignment techniques. The paper further addresses the unique complexities inherent to Sanskrit, spanning its phonetic richness, intricate morphological structure, and distinctive syntactic patterns. By systematically addressing these dimensions, the resulting 107-hour corpus aims to serve as a foundational resource for speech technology research in Sanskrit.
LDS Contractual Framework: Principles, Status and Implementation
Penny Labropoulou | Kossay Talmoudi | Dimitrios Gkoumas | Katerina Gkirtzou | Miltos Deligiannis | Leon Voukoutis | Athanasia Kolovou | Khalid Choukri | Stelios Piperidis | Dimitrios Galanis
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Penny Labropoulou | Kossay Talmoudi | Dimitrios Gkoumas | Katerina Gkirtzou | Miltos Deligiannis | Leon Voukoutis | Athanasia Kolovou | Khalid Choukri | Stelios Piperidis | Dimitrios Galanis
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
To strengthen competitiveness and digital sovereignty, the European Union has promoted the development of Common European Data Spaces to enable secure and interoperable data sharing between participants for various sectors. Data spaces combine technical infrastructure with governance mechanisms to ensure trust, transparency, data sovereignty and interoperability. Their operation must comply with the evolving European regulatory framework as well as contractual law. This paper presents the strategy adopted in the Language Data Space (LDS) to operationalise these requirements, focusing on its contractual framework and supporting instruments. It outlines the governing principles designed to ensure lawful, transparent, and fair data transactions while safeguarding the rights and obligations of data providers and consumers alike. It further describes the actual framework, and the recommended data sharing licences, with a particular emphasis on the LDS standard licence. Finally, it presents the automation tools designed and developed to support the relevant workflows while serving a wide range of users that have little or no knowledge of technical and legal complexities.
Legal Considerations in the Use of Synthetic Data for AI Development and Finetuning: The Case of LLMs4EU
Kossay Talmoudi | Khalid Choukri | Amélie Gourgeot | Florine Astruc
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Kossay Talmoudi | Khalid Choukri | Amélie Gourgeot | Florine Astruc
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
This paper examines the legal implications of using synthetic data to develop and fine-tune general-purpose AI models in the European Union, using the LLMs4EU project as a case study. It situates synthetic data within the Union’s broader data policy and highlights it as a candidate tool for reconciling data availability with regulatory constraints. From a data-protection perspective, it analyses whether and when synthetic data should be classified as “personal data” under the GDPR. From a copyright and contractual standpoint, the paper assesses the risks that synthetic datasets may embed infringing content or derive from unlawfully trained models, in light of the GEMA v. OpenAI ruling on memorised works and emerging analyses of liability for AI-generated outputs, and considers the constraints imposed by model licensing and acceptable-use policies on using models to generate training data for other models. The paper concludes that synthetic data can play a valuable role in mitigating legal risks and enabling compliant AI development in LLMs4EU, but only if its generation and use are embedded in robust governance frameworks that address data protection, copyright and contractual obligations across the entire data value chain.
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Ingo Siegert | Maria Irena Szawerna | Khalid Choukri | Simon Dobnik | Paweł Kamocki | Therese Lindström Tiedemann | Pierre Lison | Ricardo Muñoz Sánchez | Ildikó Pilán | Lisa Södergård | Kossay Talmoudi | Elena Volodina | Xuan-Son Vu
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Ingo Siegert | Maria Irena Szawerna | Khalid Choukri | Simon Dobnik | Paweł Kamocki | Therese Lindström Tiedemann | Pierre Lison | Ricardo Muñoz Sánchez | Ildikó Pilán | Lisa Södergård | Kossay Talmoudi | Elena Volodina | Xuan-Son Vu
Proceedings of the Joint Workshop on Legal and Ethical Issues in Human Language Technologies and Computational Approaches to Language Data Pseudonymization, Anonymization, De-identification, and Data Privacy (LEGAL2026 and CALD-pseudo 2026) @ LREC 2026
Search
Fix author
Co-authors
- Kossay Talmoudi 3
- Florine Astruc 1
- Miltos Deligiannis 1
- Simon Dobnik 1
- Dimitrios Galanis 1
- Katerina Gkirtzou 1
- Dimitrios Gkoumas 1
- Amélie Gourgeot 1
- Girish Nath Jha 1
- Paweł Kamocki 1
- Athanasia Kolovou 1
- Devendr Kumar 1
- Penny Labropoulou 1
- Pierre Lison 1
- Ricardo Muñoz Sánchez 1
- Ildikó Pilán 1
- Stelios Piperidis 1
- Ingo Siegert 1
- Maria Irena Szawerna 1
- Lisa Södergård 1
- Therese Lindström Tiedemann 1
- Elena Volodina 1
- Leon Voukoutis 1
- Xuan-Son Vu 1