Quality Focused Approach to a Learner Corpus Development

Roberts Darģis, Ilze Auziņa, Kristīne Levāne-Petrova, Inga Kaija


Abstract
The paper presents quality focused approach to a learner corpus development. The methodology was developed with multiple design considerations put in place to make the annotation process easier and at the same time reduce the amount of mistakes that could be introduced due to inconsistent text correction or carelessness. The approach suggested in this paper consists of multiple parts: comparison of digitized texts by several annotators, text correction, automated morphological analysis, and manual review of annotations. The described approach is used to create Latvian Language Learner corpus (LaVA) which is part of a currently ongoing project Development of Learner corpus of Latvian: methods, tools and applications.
Anthology ID:
2020.lrec-1.49
Volume:
Proceedings of the Twelfth Language Resources and Evaluation Conference
Month:
May
Year:
2020
Address:
Marseille, France
Editors:
Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association
Note:
Pages:
392–396
Language:
English
URL:
https://aclanthology.org/2020.lrec-1.49
DOI:
Bibkey:
Cite (ACL):
Roberts Darģis, Ilze Auziņa, Kristīne Levāne-Petrova, and Inga Kaija. 2020. Quality Focused Approach to a Learner Corpus Development. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 392–396, Marseille, France. European Language Resources Association.
Cite (Informal):
Quality Focused Approach to a Learner Corpus Development (Darģis et al., LREC 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.lrec-1.49.pdf