Reusing a Multi-lingual Setup to Bootstrap a Grammar Checker for a Very Low Resource Language without Data

Inga Lill Sigga Mikkelsen; Linda Wiechetek; Flammie A. Pirinen

doi:10.18653/v1/2022.computel-1.19

Reusing a Multi-lingual Setup to Bootstrap a Grammar Checker for a Very Low Resource Language without Data

Inga Lill Sigga Mikkelsen, Linda Wiechetek, Flammie A Pirinen

Abstract

Grammar checkers (GEC) are needed for digital language survival. Very low resource languages like Lule Sámi with less than 3,000 speakers need to hurry to build these tools, but do not have the big corpus data that are required for the construction of machine learning tools. We present a rule-based tool and a workflow where the work done for a related language can speed up the process. We use an existing grammar to infer rules for the new language, and we do not need a large gold corpus of annotated grammar errors, but a smaller corpus of regression tests is built while developing the tool. We present a test case for Lule Sámi reusing resources from North Sámi, show how we achieve a categorisation of the most frequent errors, and present a preliminary evaluation of the system. We hope this serves as an inspiration for small languages that need advanced tools in a limited amount of time, but do not have big data.

Anthology ID:: 2022.computel-1.19
Volume:: Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of Endangered Languages
Month:: May
Year:: 2022
Address:: Dublin, Ireland
Editors:: Sarah Moeller, Antonios Anastasopoulos, Antti Arppe, Aditi Chaudhary, Atticus Harrigan, Josh Holden, Jordan Lachler, Alexis Palmer, Shruti Rijhwani, Lane Schwartz
Venue:: ComputEL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 149–158
Language:
URL:: https://aclanthology.org/2022.computel-1.19/
DOI:: 10.18653/v1/2022.computel-1.19
Bibkey:
Cite (ACL):: Inga Lill Sigga Mikkelsen, Linda Wiechetek, and Flammie A Pirinen. 2022. Reusing a Multi-lingual Setup to Bootstrap a Grammar Checker for a Very Low Resource Language without Data. In Proceedings of the Fifth Workshop on the Use of Computational Methods in the Study of Endangered Languages, pages 149–158, Dublin, Ireland. Association for Computational Linguistics.
Cite (Informal):: Reusing a Multi-lingual Setup to Bootstrap a Grammar Checker for a Very Low Resource Language without Data (Lill Sigga Mikkelsen et al., ComputEL 2022)
Copy Citation:
PDF:: https://aclanthology.org/2022.computel-1.19.pdf
Video:: https://aclanthology.org/2022.computel-1.19.mp4

PDF Cite Search Video Fix data