SimLLM: Detecting Sentences Generated by Large Language Models Using Similarity between the Generation and its Re-generation

Hoang-Quoc Nguyen-Son; Minh-Son Dao; Koji Zettsu

SimLLM: Detecting Sentences Generated by Large Language Models Using Similarity between the Generation and its Re-generation

Hoang-Quoc Nguyen-Son, Minh-Son Dao, Koji Zettsu

Abstract

Large language models have emerged as a significant phenomenon due to their ability to produce natural text across various applications. However, the proliferation of generated text raises concerns regarding its potential misuse in fraudulent activities such as academic dishonesty, spam dissemination, and misinformation propagation. Prior studies have detected the generation of non-analogous text, which manifests numerous differences between original and generated text. We have observed that the similarity between the original text and its generation is notably higher than that between the generated text and its subsequent regeneration. To address this, we propose a novel approach named SimLLM, aimed at estimating the similarity between an input sentence and its generated counterpart to detect analogous machine-generated sentences that closely mimic human-written ones. Our empirical analysis demonstrates SimLLM’s superior performance compared to existing methods.

Anthology ID:: 2024.emnlp-main.1246
Volume:: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2024
Address:: Miami, Florida, USA
Editors:: Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 22340–22352
Language:
URL:: https://aclanthology.org/2024.emnlp-main.1246
DOI:
Bibkey:
Cite (ACL):: Hoang-Quoc Nguyen-Son, Minh-Son Dao, and Koji Zettsu. 2024. SimLLM: Detecting Sentences Generated by Large Language Models Using Similarity between the Generation and its Re-generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 22340–22352, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):: SimLLM: Detecting Sentences Generated by Large Language Models Using Similarity between the Generation and its Re-generation (Nguyen-Son et al., EMNLP 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.emnlp-main.1246.pdf
Software:: 2024.emnlp-main.1246.software.zip

PDF Cite Search Software