Rocío Alaiz-Rodríguez
2026
When Privacy Helps: Pseudonymisation as a Strategy for Improved Cyber Incident Classification
Loya C. Haughton | Eduardo Fidalgo | Rocío Alaiz-Rodríguez | Manuel Castejón-Limas | Laura Fernández-Robles
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Loya C. Haughton | Eduardo Fidalgo | Rocío Alaiz-Rodríguez | Manuel Castejón-Limas | Laura Fernández-Robles
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Organisations often rely on Cyber Threat Intelligence (CTI) for collective defence against evolving threats. Therefore, it is important that these shared reports are both compliant with data protection legislation and analytically useful. Pseudonymisation is a recognised data protection measure, yet its effect on downstream utility remains underexplored in cybersecurity. Unlike other domains where personal identifiers carry predictive value, personal data in cyber incident reports may act as noise rather than signal — suggesting pseudonymisation could improve classification accuracy. This hypothesis is evaluated in two steps. First, we derive three new datasets by applying Data Masking, Data Tokenisation and Data Substitution to a subset of the CECILIA-10C-900 dataset. We then evaluate 21 models — spanning traditional Machine Learning (ML) classifiers, encoder transformers and QLoRA fine-tuned Large Language Models (LLMs) — on these datasets, for a CTI classification task based on the Spanish National Cybersecurity Institute’s incident taxonomy. The RoBERTa-base model achieved the highest overall weighted F1-score of 87.35% when Data Tokenisation was applied, while Llama-3.1-8B demonstrated the largest gain (+12.63 pp) with Data Masking. These findings reframe pseudonymisation from only a compliance measure into a preprocessing step that may simultaneously protect privacy and improve classification in specific model-technique pairings.
2024
Comparative Analysis of Natural Language Processing Models for Malware Spam Email Identification
Francisco Jáñez-Martino | Eduardo Fidalgo | Rocío Alaiz-Rodríguez | Andrés Carofilis | Alicia Martínez-Mendoza
Proceedings of the First International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Francisco Jáñez-Martino | Eduardo Fidalgo | Rocío Alaiz-Rodríguez | Andrés Carofilis | Alicia Martínez-Mendoza
Proceedings of the First International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Spam email is one of the main vectors of cyberattacks containing scams and spreading malware. Spam emails can contain malicious and external links and attachments with hidden malicious code. Hence, cybersecurity experts seek to detect this type of email to provide earlier and more detailed warnings for organizations and users. This work is based on a binary classification system (with and without malware) and evaluates models that have achieved high performance in other natural language applications, such as fastText, BERT, RoBERTa, DistilBERT, XLM-RoBERTa, and Large Language Models such as LLaMA and Mistral. Using the Spam Email Malware Detection (SEMD-600) dataset, we compare these models regarding precision, recall, F1 score, accuracy, and runtime. DistilBERT emerges as the most suitable option, achieving a recall of 0.792 and a runtime of 1.612 ms per email.
WAVE-27K: Bringing together CTI sources to enhance threat intelligence models
Felipe Castaño | Amaia Gil-Lerchundi | Raul Orduna-Urrutia | Eduardo Fidalgo Fernandez | Rocío Alaiz-Rodríguez
Proceedings of the First International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Felipe Castaño | Amaia Gil-Lerchundi | Raul Orduna-Urrutia | Eduardo Fidalgo Fernandez | Rocío Alaiz-Rodríguez
Proceedings of the First International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Considering the growing flow of information on the internet, and the increased incident-related data from diverse sources, unstructured text processing gains importance. We have presented an automated approach to link several CTI sources through the mapping of external references. Our method facilitates the automatic construction of datasets, allowing for updates and the inclusion of new samples and labels. Following this method we built a new dataset of unstructured CTI descriptions called Weakness, Attack, Vulnerabilities, and Events 27k (WAVE-27k). Our dataset includes information about 27 different MITRE techniques, containing 22539 samples related one technique and 5262 related to two or more techniques simultaneously. We evaluated five BERT-based models into the WAVE-27K dataset concluding that SecRoBERTa reaches the highest performance with a 77.52% F1 score. Additionally, we compare the performance of the SecRoBERTa on the WAVE-27K dataset and other public datasets. The results show that the model using the WAVE-27K dataset outperforms the others. These results demonstrate that the data within WAVE-27K contains relevant information and that the proposed method effectively built a dataset with a level of quality sufficient to train a machine-learning model.