Sofia Contreras


2026

In this paper, we introduce a new multimodal and multilingual public resource for evaluating anonymization models and solutions, comprising a collection of 516 annotated videos. Unlike previously available resources, the Privacy VITA corpus is both multilingual and multimodal, covering Video, Image, Text and Audio modalities. The development of this dataset is motivated by the growing need to anonymize private data in an increasingly multimedia-driven society. Furthermore, emerging regulations pose significant challenges for organizations and companies that manage sensitive information while ensuring legal compliance. This paper details the processes of collection, preprocessing, annotation, and curation of the corpus, as well as its overall scope. We hope this resource will contribute to the advancement of multimodal anonymization research.