Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

Luis Espinosa Anke, Carla Perez Almendros


Abstract
Self-harm content is particularly challenging to detect using NLP techniques, and is also a high-stakes task which requires the highest accuracy to enable timely intervention or flagging at-risk users. We therefore present an analysis of how LLMs represent such self-harm content, which has downstream applications in self-harm detection, LLM intervention and governance and policing. In this paper, we focus on two datasets and four models, and perform two main experiments: (1) We train and evaluate linear probes across all layers of each model on two self-harm datasets: X-Sensitive and SH-Detection. Across both corpora, self-harm information crystallizes in the final 3 - 7% of network layers (93 to 97% depth). (2) We extract contrastive self-harm directions and, after performing a normaliation step, we find that the most accurate probes are not necessarily the most linearly separable. In particular, we find Gemma-3-4B to represent this contrastive self-harm direction in a slightly different, more intricate way than the other LLMs.
Anthology ID:
2026.nlpaics-1.16
Volume:
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Month:
June
Year:
2026
Address:
Alicante, Spain
Editors:
Ruslan Mitkov, Rafael Muñoz, Elena Lloret, Tharindu Ranasinghe, Ernesto L. Estevanell-Valladares, Salima Lamsiyah, Andrés Montoyo, Saad Ezzini
Venue:
NLPAICS
SIG:
Publisher:
Department of Languages and Information Systems, University of Alicante
Note:
Pages:
156–162
Language:
URL:
https://aclanthology.org/2026.nlpaics-1.16/
DOI:
Bibkey:
Cite (ACL):
Luis Espinosa Anke and Carla Perez Almendros. 2026. Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study. In Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security, pages 156–162, Alicante, Spain. Department of Languages and Information Systems, University of Alicante.
Cite (Informal):
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study (Espinosa Anke & Perez Almendros, NLPAICS 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.nlpaics-1.16.pdf