Jack W. Stokes
2026
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
Andrew Zhao | Reshmi Ghosh | Vitor R. Carvalho | Emily Lawton | Keegan Hines | Gao Huang | Jack W. Stokes
Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
Andrew Zhao | Reshmi Ghosh | Vitor R. Carvalho | Emily Lawton | Keegan Hines | Gao Huang | Jack W. Stokes
Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
Large language model (LLM) systems increasingly power everyday AI applications such as chatbots, computer-use assistants, and autonomous robots, where performance often depends on manually well-crafted prompts. LLM-based prompt optimizers reduce that effort by iteratively refining prompts from scored feedback, yet the security of this optimization stage remains underexamined. We present the first systematic analysis of poisoning risks in LLM-based prompt optimization. Using HarmBench, we find systems are substantially more vulnerable to manipulated feedback than to query poisoning alone: feedback-based attacks raise attack success rate (ASR) by up to ΔASR = 0.48. We introduce a simple fake reward attack that requires no access to the reward model and significantly increases vulnerability. We also propose a lightweight highlighting defense that reduces the fake reward ΔASR from 0.23 to 0.07 without degrading utility. These results establish prompt optimization pipelines as a first-class attack surface and motivate stronger safeguards for feedback channels and optimization frameworks.
2025
Group Preference Alignment: Customizing LLM Responses from In-Situ Conversations Only When Needed
Ishani Mondal | Jack W. Stokes | Sujay Kumar Jauhar | Longqi Yang | Mengting Wan | Xiaofeng Xu | Xia Song | Jordan Lee Boyd-Graber | Jennifer Neville
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track
Ishani Mondal | Jack W. Stokes | Sujay Kumar Jauhar | Longqi Yang | Mengting Wan | Xiaofeng Xu | Xia Song | Jordan Lee Boyd-Graber | Jennifer Neville
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track
LLMs often fail to meet specialized needs of distinct user groups due to their one-size-fits-all approach, and there is limited understanding of what personalization each group expects.To address this, we propose GPA a group-aware personalization framework that captures context-specific preference variations and steers LLMs accordingly.Our approach involves: (1) Group-Aware Preference Extraction, which distills divergent preferences from real-world conversation logs into interpretable rubrics, and (2) Tailored Response Generation, using (a) GPA-CT, which adapts responses using learnt rubrics, and (b) GPA-FT, which finetunes models using rubric-guided synthetic data.Automatic and Human evaluations confirm that GPA improves group alignment without compromising perfomance on standard instruction-following benchmarks.