Exploring word embeddings and phonological similarity for the unsupervised correction of language learner errors

Ildikó Pilán; Elena Volodina

Exploring word embeddings and phonological similarity for the unsupervised correction of language learner errors

Abstract

The presence of misspellings and other errors or non-standard word forms poses a considerable challenge for NLP systems. Although several supervised approaches have been proposed previously to normalize these, annotated training data is scarce for many languages. We investigate, therefore, an unsupervised method where correction candidates for Swedish language learners’ errors are retrieved from word embeddings. Furthermore, we compare the usefulness of combining cosine similarity with orthographic and phonological similarity based on a neural grapheme-to-phoneme conversion system we train for this purpose. Although combinations of similarity measures have been explored for finding error correction candidates, it remains unclear how these measures relate to each other and how much they contribute individually to identifying the correct alternative. We experiment with different combinations of these and find that integrating phonological information is especially useful when the majority of learner errors are related to misspellings, but less so when errors are of a variety of types including, e.g. grammatical errors.

Anthology ID:: W18-4514
Volume:: Proceedings of the Second Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature
Month:: August
Year:: 2018
Address:: Santa Fe, New Mexico
Editors:: Beatrice Alex, Stefania Degaetano-Ortlieb, Anna Feldman, Anna Kazantseva, Nils Reiter, Stan Szpakowicz
Venue:: LaTeCH
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 119–128
Language:
URL:: https://aclanthology.org/W18-4514/
DOI:
Bibkey:
Cite (ACL):: Ildikó Pilán and Elena Volodina. 2018. Exploring word embeddings and phonological similarity for the unsupervised correction of language learner errors. In Proceedings of the Second Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, pages 119–128, Santa Fe, New Mexico. Association for Computational Linguistics.
Cite (Informal):: Exploring word embeddings and phonological similarity for the unsupervised correction of language learner errors (Pilán & Volodina, LaTeCH 2018)
Copy Citation:
PDF:: https://aclanthology.org/W18-4514.pdf

PDF Cite Search Fix data