CSULoRA: Closest Safe Update Low-Rank Adaptation

Oleksandr Marchenko, Adelaide Danilov, Aria Nourbakhsh, Salima Lamsiyah


Abstract
Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adversarial fine-tuning data can substantially weaken the safety behavior of aligned models. Existing safety-preserving LoRA methods often rely on hard interventions such as projection, pruning, thresholding, or additional training objectives. While these methods can suppress unsafe update directions, they may also remove task-relevant information or require extra tuning. We introduce CSULoRA, a post-hoc method for correcting trained LoRA adapters through closest safe update estimation. CSULoRA estimates a safety-aligned subspace from the weight displacement between a safety-aligned model and its corresponding base checkpoint. It then decomposes each LoRA update into fully aligned, partially aligned, and off-subspace components. Instead of discarding components outside the estimated safety subspace, CSULoRA solves a closed-form penalized minimum-change problem that preserves the fully aligned component while smoothly attenuating potentially unsafe directions according to their relative energy. In adversarial fine-tuning experiments, CSULoRA substantially reduces attack success rate while preserving most of the utility gains obtained from standard LoRA fine-tuning[<https://github.com/Oleksandr-MB/NLPAICS2026_CSULoRA>].
Anthology ID:
2026.nlpaics-1.11
Volume:
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Month:
June
Year:
2026
Address:
Alicante, Spain
Editors:
Ruslan Mitkov, Rafael Muñoz, Elena Lloret, Tharindu Ranasinghe, Ernesto L. Estevanell-Valladares, Salima Lamsiyah, Andrés Montoyo, Saad Ezzini
Venue:
NLPAICS
SIG:
Publisher:
Department of Languages and Information Systems, University of Alicante
Note:
Pages:
103–112
Language:
URL:
https://aclanthology.org/2026.nlpaics-1.11/
DOI:
Bibkey:
Cite (ACL):
Oleksandr Marchenko, Adelaide Danilov, Aria Nourbakhsh, and Salima Lamsiyah. 2026. CSULoRA: Closest Safe Update Low-Rank Adaptation. In Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security, pages 103–112, Alicante, Spain. Department of Languages and Information Systems, University of Alicante.
Cite (Informal):
CSULoRA: Closest Safe Update Low-Rank Adaptation (Marchenko et al., NLPAICS 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.nlpaics-1.11.pdf