LLM-based Defense Against Adversarial Abstracts in ML/AI Conference Reviewer Assignments

Mohamed Omar Cherif, Christophe Cerisara, Julien Falgas


Abstract
Large Machine Learning Conferences guarantee quality reviews of submitted papers by assigning each submission to reviewers who are experts in the relevant topics. This is typically realized by matching reviewers’ expertise to the paper abstract with text semantic embeddings. In order to try and maximize their acceptance score, malevolent authors may attack this assignment process by modifying their abstract so that it matches the expertise of colluding accomplices registered as reviewers. The success of such attacks has been recently demonstrated for realistic conference reviewing datasets with SPECTER embeddings. We propose in this work a defense mechanism against such attacks that leverages Large Language Models (LLM) to rewrite the submitted abstracts and remove the targeted alteration of the original abstract that were aimed at the colluding reviewers. We demonstrate experimentally the effectiveness of our defense that prevents assigning most malevolent abstracts to their colluding reviewer, while preserving the topic-based assignment of normal abstracts to expert reviewers.
Anthology ID:
2026.nlpaics-1.12
Volume:
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Month:
June
Year:
2026
Address:
Alicante, Spain
Editors:
Ruslan Mitkov, Rafael Muñoz, Elena Lloret, Tharindu Ranasinghe, Ernesto L. Estevanell-Valladares, Salima Lamsiyah, Andrés Montoyo, Saad Ezzini
Venue:
NLPAICS
SIG:
Publisher:
Department of Languages and Information Systems, University of Alicante
Note:
Pages:
113–121
Language:
URL:
https://aclanthology.org/2026.nlpaics-1.12/
DOI:
Bibkey:
Cite (ACL):
Mohamed Omar Cherif, Christophe Cerisara, and Julien Falgas. 2026. LLM-based Defense Against Adversarial Abstracts in ML/AI Conference Reviewer Assignments. In Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security, pages 113–121, Alicante, Spain. Department of Languages and Information Systems, University of Alicante.
Cite (Informal):
LLM-based Defense Against Adversarial Abstracts in ML/AI Conference Reviewer Assignments (Cherif et al., NLPAICS 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.nlpaics-1.12.pdf