LaVA – Latvian Language Learner corpus
Roberts Darģis | Ilze Auziņa | Inga Kaija | Kristīne Levāne-Petrova | Kristīne Pokratniece
Proceedings of the Thirteenth Language Resources and Evaluation Conference
This paper presents the Latvian Language Learner Corpus (LaVA) developed at the Institute of Mathematics and Computer Science, University of Latvia. LaVA corpus contains 1015 essays (190k tokens and 790k characters excluding whitespaces) from foreigners studying at Latvian higher education institutions and who are learning Latvian as a foreign language in the first or second semester, reaching the A1 (possibly A2) Latvian language proficiency level. The corpus has morphological and error annotations. Error analysis and the statistics of the LaVA corpus are also provided in the paper. The corpus is publicly available at: http://www.korpuss.lv/id/LaVA.
Quality Focused Approach to a Learner Corpus Development
Roberts Darģis | Ilze Auziņa | Kristīne Levāne-Petrova | Inga Kaija
Proceedings of the Twelfth Language Resources and Evaluation Conference
The paper presents quality focused approach to a learner corpus development. The methodology was developed with multiple design considerations put in place to make the annotation process easier and at the same time reduce the amount of mistakes that could be introduced due to inconsistent text correction or carelessness. The approach suggested in this paper consists of multiple parts: comparison of digitized texts by several annotators, text correction, automated morphological analysis, and manual review of annotations. The described approach is used to create Latvian Language Learner corpus (LaVA) which is part of a currently ongoing project Development of Learner corpus of Latvian: methods, tools and applications.