Mohamed Omar Cherif
2026
LLM-based Defense Against Adversarial Abstracts in ML/AI Conference Reviewer Assignments
Mohamed Omar Cherif | Christophe Cerisara | Julien Falgas
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Mohamed Omar Cherif | Christophe Cerisara | Julien Falgas
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Large Machine Learning Conferences guarantee quality reviews of submitted papers by assigning each submission to reviewers who are experts in the relevant topics. This is typically realized by matching reviewers’ expertise to the paper abstract with text semantic embeddings. In order to try and maximize their acceptance score, malevolent authors may attack this assignment process by modifying their abstract so that it matches the expertise of colluding accomplices registered as reviewers. The success of such attacks has been recently demonstrated for realistic conference reviewing datasets with SPECTER embeddings. We propose in this work a defense mechanism against such attacks that leverages Large Language Models (LLM) to rewrite the submitted abstracts and remove the targeted alteration of the original abstract that were aimed at the colluding reviewers. We demonstrate experimentally the effectiveness of our defense that prevents assigning most malevolent abstracts to their colluding reviewer, while preserving the topic-based assignment of normal abstracts to expert reviewers.