From Detection to Attribution: Forensic Linguistics and Adversarial Red Teaming as Complementary Responses to LLM Misuse

Rui Sousa-Silva


Abstract
The proliferation of Large Language Models (LLMs) has enabled the automation of cyber-attacks (including phishing, social engineering, and impersonation) at unprecedented scale, while existing safeguards remain routinely circumvented. Current detection approaches, predominantly based on stylometric and machine learning methods, face fundamental limitations against adaptive adversaries and struggle with the implicit, contextual, and pragmatic dimensions of language. This article proposes forensic linguistic analysis grounded in the theory of idiolect as a complementary approach to LLM-generated text detection and attribution. We adopt a red teaming methodology to generate synthetic toxic texts that bypass model guardrails to create controlled conditions for testing whether qualitative forensic analysis can succeed where quantitative approaches falter. The findings of our stylometric, character n-gram, and cluster analysis converge to provide moderate evidence that stylometric approaches succeed in discriminating authorship. However, they are not conclusive and hence fall short of current admissibility criteria across diverse jurisdictions. The article thus concludes that idiolect-based forensic analysis can distinguish genuine authorship from LLM-generated impersonation, even when surface features are manipulated. We discuss implications for legal and investigative contexts, where interpretable, theoretically-grounded expert analysis is required over black-box classifier outputs.
Anthology ID:
2026.nlpaics-1.13
Volume:
Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security
Month:
June
Year:
2026
Address:
Alicante, Spain
Editors:
Ruslan Mitkov, Rafael Muñoz, Elena Lloret, Tharindu Ranasinghe, Ernesto L. Estevanell-Valladares, Salima Lamsiyah, Andrés Montoyo, Saad Ezzini
Venue:
NLPAICS
SIG:
Publisher:
Department of Languages and Information Systems, University of Alicante
Note:
Pages:
122–133
Language:
URL:
https://aclanthology.org/2026.nlpaics-1.13/
DOI:
Bibkey:
Cite (ACL):
Rui Sousa-Silva. 2026. From Detection to Attribution: Forensic Linguistics and Adversarial Red Teaming as Complementary Responses to LLM Misuse. In Proceedings of the Second International Conference on Natural Language Processing and Artificial Intelligence for Cyber Security, pages 122–133, Alicante, Spain. Department of Languages and Information Systems, University of Alicante.
Cite (Informal):
From Detection to Attribution: Forensic Linguistics and Adversarial Red Teaming as Complementary Responses to LLM Misuse (Sousa-Silva, NLPAICS 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.nlpaics-1.13.pdf