Yian Wang
Author directoryPapers on this page may belong to the following people: Yian Wang, Yian Wang
2026
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
Yian Wang | Yuen Chen | Agam Goyal | Hari Sundaram
Findings of the Association for Computational Linguistics: ACL 2026
Yian Wang | Yuen Chen | Agam Goyal | Hari Sundaram
Findings of the Association for Computational Linguistics: ACL 2026
Large language models (LLMs) frequently generate toxic content, posing significant risks for safe deployment. Current mitigation strategies often degrade generation quality or require costly human annotation. We propose CausalDetox, a framework that identifies and intervenes on the specific attention heads causally responsible for toxic generation. Using the Probability of Necessity and Sufficiency (PNS), we isolate a minimal set of heads that are necessary and sufficient for toxicity. We utilize these components via two complementary strategies: (1) Local Inference-Time Intervention, which constructs dynamic, input-specific steering vectors for context-aware detoxification, and (2) PNS-Guided Fine-Tuning, which permanently unlearns toxic representations. We also introduceParaTox, a novel benchmark of aligned toxic/non-toxic sentence pairs enabling controlled counterfactual evaluation. Experiments on ToxiGen, ImplicitHate, and ParaDetox show that CausalDetox achieves up to 5.34% greater toxicity reduction compared to baselines while preserving linguistic fluency, and offers a 7× speedup in head selection.
2025
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
Agam Goyal | Vedant Rathi | William Yeh | Yian Wang | Yuen Chen | Hari Sundaram
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Agam Goyal | Vedant Rathi | William Yeh | Yian Wang | Yuen Chen | Hari Sundaram
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Large language models (LLMs) are now ubiquitous in user-facing applications, yet they still generate undesirable toxic outputs, including profanity, vulgarity, and derogatory remarks. Although numerous detoxification methods exist, most apply broad, surface-level fixes and can therefore easily be circumvented by jailbreak attacks. In this paper we leverage sparse autoencoders (SAEs) to identify toxicity-related directions in the residual stream of models and perform targeted activation steering using the corresponding decoder vectors. We introduce three tiers of steering aggressiveness and evaluate them on GPT-2 Small and Gemma-2-2B, revealing trade-offs between toxicity reduction and language fluency. At stronger steering strengths, these causal interventions surpass competitive baselines in reducing toxicity by up to 20%, though fluency can degrade noticeably on GPT-2 Small depending on the aggressiveness. Crucially, standard NLP benchmark scores upon steering remain stable, indicating that the model’s knowledge and general abilities are preserved. We further show that feature-splitting in wider SAEs hampers safety interventions, underscoring the importance of disentangled feature learning. Our findings highlight both the promise and the current limitations of SAE-based causal interventions for LLM detoxification, further suggesting practical guidelines for safer language-model deployment.