Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification

Zenghao Duan; Zhiyi Yin; Zhichao Shi; Liang Pang (庞亮); Shaoling Jing; Zihe Huang; Jiayi Wu; Yu Yan; Jingcheng Deng (邓竞成); Huawei Shen (沈华伟); Xueqi Cheng (程学旗)

Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification

Zenghao Duan, Zhiyi Yin, Zhichao Shi, Liang Pang, Shaoling Jing, Zihe Huang, Jiayi Wu, Yu Yan, Jingcheng Deng, Huawei Shen, Xueqi Cheng

Abstract

Large language models (LLMs) exhibit exceptional performance but pose inherent risks of generating toxic content, restricting their safe deployment. While traditional methods (e.g., alignment) adjust output preferences, they fail to eliminate underlying toxic regions in parameters, leaving models vulnerable to adversarial attacks. Prior mechanistic studies characterize toxic regions as "toxic vectors" or "layer-wise subspaces", yet our analysis identifies critical limitations: i) Removed toxic vectors can be reconstructed via linear combinations of non-toxic vectors, demanding targeting of entire toxic subspace; ii) Contrastive objective over limited samples inject noise into layer-wise subspaces, hindering stable extraction. These highlight the challenge of identifying robust toxic subspace and removing them. Therefore, we propose GLOSS (GLobal tOxic Subspace Suppression), a lightweight method that mitigates toxicity by identifying and eliminating this global subspace from FFN parameters. Experiments on LLMs (e.g., Qwen3) show GLOSS achieves SOTA detoxification while preserving general capabilities without requiring large-scale retraining.

Anthology ID:: 2026.acl-long.1652
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 35697–35719
Language:
URL:: https://aclanthology.org/2026.acl-long.1652/
DOI:
Bibkey:
Cite (ACL):: Zenghao Duan, Zhiyi Yin, Zhichao Shi, Liang Pang, Shaoling Jing, Zihe Huang, Jiayi Wu, Yu Yan, Jingcheng Deng, Huawei Shen, and Xueqi Cheng. 2026. Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 35697–35719, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification (Duan et al., ACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.acl-long.1652.pdf
Checklist:: 2026.acl-long.1652.checklist.pdf

PDF Cite Search Checklist Fix data