Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering

Purva Chiniya, Kevin Joseph Scaria, Sagar Chaturvedi


Abstract
Large language models (LLMs) remain susceptible to jailbreak and direct prompt-injection attacks, yet the strongest defensive filters frequently over- refuse benign queries and degrade user experience. Previous work on prompt injection detection such as, GradSafe, detects unsafe prompts with a single “accept all” anchor token, but its threshold is brittle and it offers no deterministic guarantee that harmful content will not be emitted once decoding begins. We introduce Gradient-Controlled Decoding (GCD), a training-free guardrail that combines with both an acceptance anchor (“Sure”) and refusal anchor (“Sorry”) tightening the decision boundary and lowering false positives. In the mitigation stage, if a prompt is flagged, GCD preset-injects one or two refusal tokens ("Sorry, I can’t . . . ") before autoregressive decoding resumes, guaranteeing first- token safety regardless of sampling strategy. On ToxicChat, XSTest-v2, and AdvBench, GCD reduces false positives by 52% vs. GradSafe at comparable recall, lowers attack success rate by up to 20% vs. the strongest decoding-only baseline, adds under 15-20 ms latency on an average on V100 instances, transfers to LLaMA-2-7B, Mixtral-8×7B, and Qwen-2-7B, and requires only 20 template prompts. GCD is a lightweight, scalable safety layer for real-time LLM deployment.
Anthology ID:
2026.lrec-1.775
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
9884–9892
Language:
External URL:
https://lrec.elra.info/lrec2026-main-775
DOI:
10.63317/5axtpujejwcx
Bibkey:
Cite (ACL):
Purva Chiniya, Kevin Joseph Scaria, and Sagar Chaturvedi. 2026. Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 9884–9892, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering (Chiniya et al., LREC 2026)
Copy Citation: