Sofia Contreras
2026
Privacy VITA: a new multilingual and multimodal annotated video corpus to evaluate anonymization systems
Jorge Rico | Sofia Contreras | Enrique Manjavacas Arevalo | Maria Viana | Ruben Pérez-Ramón | Maria Luisa Izquierdo | María José Vilella | Jaime Corton | Silvia Rodriguez | Fernando Espinza | Rafael Ginard | César Pérez | Jesús Arias | Richard Cook | Pablo Regodón | Manuel Moyano | Ricardo Heredia | Lara De Santos | Pierre Plaza | Jacqueline González
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Jorge Rico | Sofia Contreras | Enrique Manjavacas Arevalo | Maria Viana | Ruben Pérez-Ramón | Maria Luisa Izquierdo | María José Vilella | Jaime Corton | Silvia Rodriguez | Fernando Espinza | Rafael Ginard | César Pérez | Jesús Arias | Richard Cook | Pablo Regodón | Manuel Moyano | Ricardo Heredia | Lara De Santos | Pierre Plaza | Jacqueline González
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
In this paper, we introduce a new multimodal and multilingual public resource for evaluating anonymization models and solutions, comprising a collection of 516 annotated videos. Unlike previously available resources, the Privacy VITA corpus is both multilingual and multimodal, covering Video, Image, Text and Audio modalities. The development of this dataset is motivated by the growing need to anonymize private data in an increasingly multimedia-driven society. Furthermore, emerging regulations pose significant challenges for organizations and companies that manage sensitive information while ensuring legal compliance. This paper details the processes of collection, preprocessing, annotation, and curation of the corpus, as well as its overall scope. We hope this resource will contribute to the advancement of multimodal anonymization research.