SLM as Guardian: Pioneering AI Safety with Small Language Model

Ohjoon Kwon; Donghyeon Jeon; Nayoung Choi; Gyu-Hwung Cho; Hwiyeol Jo; Changbong Kim; Hyunwoo Lee; Inho Kang; Sun Kim; Taiwoo Park

doi:10.18653/v1/2024.emnlp-industry.99

SLM as Guardian: Pioneering AI Safety with Small Language Model

Ohjoon Kwon, Donghyeon Jeon, Nayoung Choi, Gyu-Hwung Cho, Hwiyeol Jo, Changbong Kim, Hyunwoo Lee, Inho Kang, Sun Kim, Taiwoo Park

Abstract

Most prior safety research of large language models (LLMs) has focused on enhancing the alignment of LLMs to better suit the safety requirements of their use cases. However, internalizing such safeguard features into larger models brought challenges of higher training cost and unintended degradation of helpfulness. In this paper, we leverage a smaller LLM for both harmful query detection and safeguard response generation. We introduce our safety requirements and the taxonomy of harmfulness categories, and then propose a multi-task learning mechanism fusing the two tasks into a single model. We demonstrate the effectiveness of our approach, providing on par or surpassing harmful query detection and safeguard response performance compared to the publicly available LLMs.

Anthology ID:: 2024.emnlp-industry.99
Volume:: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track
Month:: November
Year:: 2024
Address:: Miami, Florida, US
Editors:: Franck Dernoncourt, Daniel Preoţiuc-Pietro, Anastasia Shimorina
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1333–1350
Language:
URL:: https://aclanthology.org/2024.emnlp-industry.99/
DOI:: 10.18653/v1/2024.emnlp-industry.99
Bibkey:
Cite (ACL):: Ohjoon Kwon, Donghyeon Jeon, Nayoung Choi, Gyu-Hwung Cho, Hwiyeol Jo, Changbong Kim, Hyunwoo Lee, Inho Kang, Sun Kim, and Taiwoo Park. 2024. SLM as Guardian: Pioneering AI Safety with Small Language Model. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 1333–1350, Miami, Florida, US. Association for Computational Linguistics.
Cite (Informal):: SLM as Guardian: Pioneering AI Safety with Small Language Model (Kwon et al., EMNLP 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.emnlp-industry.99.pdf
Poster:: 2024.emnlp-industry.99.poster.pdf

PDF Cite Search Poster Fix data