Incomplete Prompt Jailbreaks in Large Language Models

Yeonjea Kim; Bumjin Park; Jaesik Choi

Incomplete Prompt Jailbreaks in Large Language Models

Abstract

Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion remains vulnerable to incomplete harmful prompts. In this work, we formalize this phenomenon as incomplete prompt jailbreaks (IPJ) and provide a systematic empirical characterization of when and how incomplete prompts elicit harmful continuations. We analyze diverse attractor types associated with incomplete sentence continuation and show that LLMs systematically delay refusal until sentence termination. We further demonstrate that training models to refuse incomplete harmful prompts via parameter tuning is insufficient, failing to generalize across both content domains and attractor types. To enable fine-grained control, we identify two functional neurons: termination and continuation neurons. By clarifying their roles in sentence completion, we highlight the potential of neuron-level interventions for more precise and robust IPJ defenses.

Anthology ID:: 2026.findings-acl.1567
Volume:: Findings of the Association for Computational Linguistics: ACL 2026
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 31352–31368
Language:
URL:: https://aclanthology.org/2026.findings-acl.1567/
DOI:
Bibkey:
Cite (ACL):: Yeonjea Kim, Bumjin Park, and Jaesik Choi. 2026. Incomplete Prompt Jailbreaks in Large Language Models. In Findings of the Association for Computational Linguistics: ACL 2026, pages 31352–31368, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Incomplete Prompt Jailbreaks in Large Language Models (Kim et al., Findings 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.findings-acl.1567.pdf
Checklist:: 2026.findings-acl.1567.checklist.pdf

PDF Cite Search Checklist Fix data