Corpus and Evaluation Measures for Automatic Plagiarism Detection

Alberto Barrón-Cedeño; Martin Potthast; Paolo Rosso; Benno Stein

Corpus and Evaluation Measures for Automatic Plagiarism Detection

Alberto Barrón-Cedeño, Martin Potthast, Paolo Rosso, Benno Stein

Abstract

The simple access to texts on digital libraries and the World Wide Web has led to an increased number of plagiarism cases in recent years, which renders manual plagiarism detection infeasible at large. Various methods for automatic plagiarism detection have been developed whose objective is to assist human experts in the analysis of documents for plagiarism. The methods can be divided into two main approaches: intrinsic and external. Unlike other tasks in natural language processing and information retrieval, it is not possible to publish a collection of real plagiarism cases for evaluation purposes since they cannot be properly anonymized. Therefore, current evaluations found in the literature are incomparable and, very often not even reproducible. Our contribution in this respect is a newly developed large-scale corpus of artificial plagiarism useful for the evaluation of intrinsic as well as external plagiarism detection. Additionally, new detection performance measures tailored to the evaluation of plagiarism detection algorithms are proposed.

Anthology ID:: L10-1016
Volume:: Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)
Month:: May
Year:: 2010
Address:: Valletta, Malta
Editors:: Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Mike Rosner, Daniel Tapias
Venue:: LREC
SIG:
Publisher:: European Language Resources Association (ELRA)
Note:
Pages:
Language:
URL:: http://www.lrec-conf.org/proceedings/lrec2010/pdf/35_Paper.pdf
DOI:
Bibkey:
Cite (ACL):: Alberto Barrón-Cedeño, Martin Potthast, Paolo Rosso, and Benno Stein. 2010. Corpus and Evaluation Measures for Automatic Plagiarism Detection. In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10), Valletta, Malta. European Language Resources Association (ELRA).
Cite (Informal):: Corpus and Evaluation Measures for Automatic Plagiarism Detection (Barrón-Cedeño et al., LREC 2010)
Copy Citation:
PDF:: http://www.lrec-conf.org/proceedings/lrec2010/pdf/35_Paper.pdf

PDF Cite Search Fix data